Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
753b7cc
fix: Roll back the changes related to onnxslim in #308 and #305 (#310)
KakaruHayate Jun 28, 2026
36da953
fix: pool frame-level speaker mixes for Mix-LN (#317)
KakaruHayate Aug 2, 2026
ee7463f
some minor fixes (#318)
yxlllc Aug 3, 2026
8a07f76
fix: pad mel with spec_min instead of 0.0 in acoustic collater (#313)
KakaruHayate Aug 4, 2026
96bb14d
some minor fixes (#321)
yxlllc Aug 11, 2026
e2307b1
feat: Triton fused SoftSignGLU kernel for LYNXNet2 (#312)
KakaruHayate Aug 13, 2026
1802dfb
some minor fixes (#322)
yxlllc Aug 16, 2026
1b0a6da
some minor fixes (#325)
yxlllc Aug 19, 2026
1ee5eb8
some minor fixes (#329)
yxlllc Aug 25, 2026
d9f8d64
Expose RoPE theta via config API; cache fp32 cos/sin tables (#320)
KakaruHayate Aug 29, 2026
4975e19
some minor fixes (#330)
yxlllc Sep 3, 2026
1c18c59
probe and cap max_frames for the batch sampler (#327)
yxlllc Sep 3, 2026
336cf01
enable mixed-precision for layernorm (#328)
yxlllc Sep 3, 2026
8935f61
Rolling back the training precision of the last layernorm in LynxNet …
KakaruHayate Sep 15, 2026
86518e8
Support dual-timestep reflow (#323)
yxlllc Sep 15, 2026
d942b63
fix energy calculation (#332)
yxlllc Sep 15, 2026
bf78948
add missed configuration items (#334)
yxlllc Sep 18, 2026
8333dd6
Share manager across multiprocessing queues (#335)
yxlllc Sep 20, 2026
2169f65
Merge openvpi/DiffSinger main (up to 8333dd6) into LiuYunPlayer fork
KakaruHayate Sep 21, 2026
52e0256
feat(mixln): add config switch to disable speaker-shuffle during trai…
KakaruHayate Aug 23, 2026
be6a9f1
chore(configs): drop redundant comment on mixln_shuffle_speakers
KakaruHayate Aug 23, 2026
d9f2337
chore(common_layers): drop redundant shuffle_speakers comment
KakaruHayate Aug 23, 2026
89fc557
fix: address CodeRabbit review findings from upstream sync
KakaruHayate Sep 21, 2026
815a98c
docs: document config options missing from ConfigurationSchemas
KakaruHayate Sep 21, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,17 +1,17 @@
# DiffSinger (OpenVPI maintained version)

[![arXiv](https://img.shields.io/badge/arXiv-Paper-<COLOR>.svg)](https://arxiv.org/abs/2105.02446)
[![arXiv](https://img.shields.io/badge/arXiv-Paper-b31b1b.svg)](https://arxiv.org/abs/2105.02446)
[![downloads](https://img.shields.io/github/downloads/openvpi/DiffSinger/total.svg)](https://github.com/openvpi/DiffSinger/releases)
[![Bilibili](https://img.shields.io/badge/Bilibili-Demo-blue)](https://www.bilibili.com/video/BV1be411N7JA/)
[![license](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://github.com/openvpi/DiffSinger/blob/main/LICENSE)

This is a refactored and enhanced version of _DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism_ based on the original [paper](https://arxiv.org/abs/2105.02446) and [implementation](https://github.com/MoonInTheRiver/DiffSinger), which provides:

- Cleaner code structure: useless and redundant files are removed and the others are re-organized.
- Better sound quality: the sampling rate of synthesized audio are adapted to 44.1 kHz instead of the original 24 kHz.
- Better sound quality: the sampling rate of synthesized audio is adapted to 44.1 kHz instead of the original 24 kHz.
- Higher fidelity: improved acoustic models and diffusion sampling acceleration algorithms are integrated.
- More controllability: introduced variance models and parameters for prediction and control of pitch, energy, breathiness, etc.
- Production compatibility: functionalities are designed to match the requirements of production deployment and the SVS communities.
- Production compatibility: functionalities are designed to match the requirements of production deployment and the Singing Voice Synthesis (SVS) communities.

| Overview | Variance Model | Acoustic Model |
|:-------------------------------------------------------------------------------------:|:-------------------------------------------------------------------------------------:|:-------------------------------------------------------------------------------------:|
Expand Down Expand Up @@ -69,7 +69,7 @@ TBD

## Disclaimer

Any organization or individual is prohibited from using any functionalities included in this repository to generate someone's speech without his/her consent, including but not limited to government leaders, political figures, and celebrities. If you do not comply with this item, you could be in violation of copyright laws.
Any organization or individual is prohibited from using any functionalities included in this repository to generate someone's voice without his/her consent, including but not limited to government leaders, political figures, and celebrities. If you do not comply with this item, you could be in violation of copyright laws.

## License

Expand Down
22 changes: 0 additions & 22 deletions basics/base_binarizer.py
Original file line number Diff line number Diff line change
@@ -1,7 +1,6 @@
import json
import pathlib
import pickle
import random
import shutil
import warnings
from copy import deepcopy
Expand Down Expand Up @@ -121,15 +120,6 @@ def split_train_valid_set(self, prefixes: list):
if prefix in self.item_names:
valid_item_names[prefix] = 1
prefixes.pop(prefix)
# Add prefixes that exactly matches item name without speaker id to test set
for prefix in deepcopy(prefixes):
matched = False
for name in self.item_names:
if name.split(':')[-1] == prefix:
valid_item_names[name] = 1
matched = True
if matched:
prefixes.pop(prefix)
# Add names with one of the remaining prefixes to test set
for prefix in deepcopy(prefixes):
matched = False
Expand All @@ -139,15 +129,6 @@ def split_train_valid_set(self, prefixes: list):
matched = True
if matched:
prefixes.pop(prefix)
for prefix in deepcopy(prefixes):
matched = False
for name in self.item_names:
if name.split(':')[-1].startswith(prefix):
valid_item_names[name] = 1
matched = True
if matched:
prefixes.pop(prefix)

if len(prefixes) != 0:
warnings.warn(
f'The following rules in test_prefixes have no matching names in the dataset: {", ".join(prefixes.keys())}',
Expand Down Expand Up @@ -195,9 +176,6 @@ def process(self):
self.item_names = sorted(list(self.items.keys()))
self._train_item_names, self._valid_item_names = self.split_train_valid_set(test_prefixes)

if self.binarization_args['shuffle']:
random.shuffle(self.item_names)

self.binary_data_dir.mkdir(parents=True, exist_ok=True)

# Copy spk_map, lang_map and dictionary to binary data dir
Expand Down
3 changes: 2 additions & 1 deletion basics/base_task.py
Original file line number Diff line number Diff line change
Expand Up @@ -337,7 +337,8 @@ def train_dataloader(self):
size_reversed=True,
required_batch_count_multiple=hparams['accumulate_grad_batches'],
shuffle_sample=True,
shuffle_batch=True
shuffle_batch=True,
probe_and_cap_max_frames=True
)
return torch.utils.data.DataLoader(
self.train_dataset,
Expand Down
7 changes: 6 additions & 1 deletion configs/acoustic.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,9 @@ base_config:

task_cls: training.acoustic_task.AcousticTask

# Enable Triton-fused Linear+SoftSignGLU kernels for LYNXNet2 backbones.
use_fused_kernels: false

dictionaries: {}
extra_phonemes: []
merged_phoneme_groups: []
Expand All @@ -19,7 +22,6 @@ fmin: 40
fmax: 16000

binarization_args:
shuffle: true
num_workers: 0
augmentation_args:
random_pitch_shifting:
Expand Down Expand Up @@ -56,6 +58,7 @@ use_spk_id: false
num_spk: 1
use_mix_ln: false
mix_ln_layer: [0, 2]
mixln_shuffle_speakers: false
use_energy_embed: false
use_breathiness_embed: false
use_voicing_embed: false
Expand All @@ -75,12 +78,14 @@ shift_mouth_opening_args:
teacher_use_amp: true

diffusion_type: reflow
use_dual_timestep: true
time_scale_factor: 1000
timesteps: 1000
max_beta: 0.02
enc_ffn_kernel_size: 3
use_rope: true
rope_interleaved: false
rope_theta: 10000
use_stretch_embed: true
use_variance_scaling: true
rel_pos: true
Expand Down
1 change: 0 additions & 1 deletion configs/base.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,6 @@ datasets: []
binary_data_dir: null
binarizer_cls: null
binarization_args:
shuffle: false
num_workers: 0

audio_sample_rate: 44100
Expand Down
5 changes: 4 additions & 1 deletion configs/templates/config_acoustic.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,7 @@ pe_ckpt: 'checkpoints/rmvpe/model.pt'
hnsep: vr
hnsep_ckpt: 'checkpoints/vr/model.pt'
vocoder: NsfHifiGAN
vocoder_ckpt: checkpoints/nsf_hifigan_44.1k_hop512_128bin_2024.02/model.ckpt
vocoder_ckpt: checkpoints/pc_nsf_hifigan_44.1k_hop512_128bin_2025.02/model.ckpt

use_lang_id: false
num_lang: 1
Expand All @@ -45,6 +45,7 @@ num_spk: 1

use_mix_ln: false
mix_ln_layer: [0, 2]
mixln_shuffle_speakers: false

# NOTICE: before enabling variance embeddings, please read the docs at
# https://github.com/openvpi/DiffSinger/tree/main/docs/BestPractices.md#choosing-variance-parameters
Expand Down Expand Up @@ -86,9 +87,11 @@ augmentation_args:

# diffusion and shallow diffusion
diffusion_type: reflow
use_dual_timestep: true
enc_ffn_kernel_size: 3
use_rope: true
rope_interleaved: false
rope_theta: 10000
use_stretch_embed: true
use_variance_scaling: true
use_shallow_diffusion: true
Expand Down
2 changes: 2 additions & 0 deletions configs/templates/config_variance.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -72,6 +72,7 @@ mouth_opening_estimator_ckpt: 'checkpoints/r3moe/0508_s2k_noise_aug_0.15/ema_mod
enc_ffn_kernel_size: 3
use_rope: true
rope_interleaved: false
rope_theta: 10000
use_stretch_embed: false
use_variance_scaling: true
hidden_size: 384
Expand All @@ -96,6 +97,7 @@ glide_types: [up, down]
glide_embed_scale: 11.313708498984760 # sqrt(128)

diffusion_type: reflow
use_dual_timestep: true

pitch_prediction_args:
pitd_norm_min: -8.0
Expand Down
6 changes: 5 additions & 1 deletion configs/variance.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,9 @@ base_config:

task_cls: training.variance_task.VarianceTask

# Enable Triton-fused Linear+SoftSignGLU kernels for LYNXNet2 backbones.
use_fused_kernels: false

dictionaries: {}
extra_phonemes: []
merged_phoneme_groups: []
Expand All @@ -15,7 +18,6 @@ win_size: 2048 # FFT size.
midi_smooth_width: 0.06 # in seconds

binarization_args:
shuffle: true
num_workers: 0
prefer_ds: false

Expand All @@ -38,6 +40,7 @@ predict_mouth_opening: false
enc_ffn_kernel_size: 3
use_rope: true
rope_interleaved: false
rope_theta: 10000
use_stretch_embed: false
use_variance_scaling: true
rel_pos: true
Expand Down Expand Up @@ -113,6 +116,7 @@ lambda_pitch_loss: 1.0
lambda_var_loss: 1.0

diffusion_type: reflow # ddpm
use_dual_timestep: true
time_scale_factor: 1000
schedule_type: 'linear'
K_step: 1000
Expand Down
5 changes: 2 additions & 3 deletions deployment/exporters/variance_exporter.py
Original file line number Diff line number Diff line change
Expand Up @@ -759,7 +759,7 @@ def _optimize_merge_pitch_predictor_graph(
assert check, 'Simplified ONNX model could not be validated'

onnx_helper.model_override_io_shapes(
pitch_predictor, output_shapes={'pitch_pred': (1, 'n_frames')}
pitch_predictor, output_shapes={'x_pred': (1, 'n_frames')}
)
print(f'Running ONNX Simplifier #1 on {self.pitch_predictor_class_name}...')
pitch_predictor, check = onnxsim.simplify(pitch_predictor, include_subgraph=True)
Expand Down Expand Up @@ -814,8 +814,7 @@ def _optimize_merge_variance_predictor_graph(
else (1, len(self.model.variance_prediction_list), 'n_frames')
}
)
print(f'Running ONNX Simplifier #1 on'
f' {self.multi_var_predictor_class_name}...')
print(f'Running ONNX Simplifier #1 on {self.multi_var_predictor_class_name}...')
var_diffusion, check = onnxsim.simplify(var_diffusion, include_subgraph=True)
assert check, 'Simplified ONNX model could not be validated'
onnx_helper.graph_fold_back_to_squeeze(var_diffusion.graph)
Expand Down
16 changes: 1 addition & 15 deletions deployment/modules/fastspeech2.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
import torch.nn.functional as F

from modules.commons.common_layers import NormalInitEmbedding as Embedding
from modules.fastspeech.acoustic_encoder import FastSpeech2Acoustic
from modules.fastspeech.acoustic_encoder import FastSpeech2Acoustic, uniform_attention_pooling
from modules.fastspeech.variance_encoder import FastSpeech2Variance
from utils.hparams import hparams
from utils.phoneme_utils import PAD_INDEX
Expand All @@ -18,20 +18,6 @@
f0_mel_max = 1127 * np.log(1 + f0_max / 700)


def uniform_attention_pooling(spk_embed, durations):
_, T_mel, _ = spk_embed.shape
ph_starts = torch.cumsum(torch.cat([torch.zeros_like(durations[:, :1]), durations[:, :-1]], dim=1), dim=1)
ph_ends = ph_starts + durations
mel_indices = torch.arange(T_mel, device=spk_embed.device).view(1, 1, T_mel)
phoneme_to_mel_mask = (mel_indices >= ph_starts.unsqueeze(-1)) & (mel_indices < ph_ends.unsqueeze(-1))
uniform_scores = phoneme_to_mel_mask.float()
sum_scores = uniform_scores.sum(dim=2, keepdim=True)
attn_weights = uniform_scores / (sum_scores + (sum_scores == 0).float()) # [B, T_ph, T_mel]
ph_spk_embed = torch.bmm(attn_weights, spk_embed)

return ph_spk_embed


def f0_to_coarse(f0):
f0_mel = 1127 * (1 + f0 / 700).log()
a = (f0_bin - 2) / (f0_mel_max - f0_mel_min)
Expand Down
Loading