Research Preview · v1.0

Neural Audio Codec
for Indic Speech Synthesis

Production-grade waveform-to-token compression at 240× with native support for twelve Indic script families. Cross-platform deployment on CUDA, Apple Silicon, and CPU.

Quick Start — 5 Minutes Read the Paper View on GitHub
24 kHz
Sample Rate
240×
Compression Ratio
100 Hz
Frame Rate
12
Script Families
3
Compute Backends
14.1M
Max Parameters
Abstract

What Lipika solves.

Neural audio codecs have enabled a new paradigm in speech synthesis: compressing raw waveforms into discrete tokens that can be modelled by language models (VALL-E, AudioLM). However, existing codecs — EnCodec, SoundStream, DAC — are trained primarily on English and Mandarin speech, with no mechanism to preserve the phonological contrasts that are phonemically contrastive in Indic languages.

Lipika introduces script-family conditioning via Adaptive Layer Normalisation, a retroflex-aware latent prior, optional semantic distillation from W2V-BERT 2.0, and a cross-platform training pipeline that auto-scales to available hardware. It achieves 240× compression (24 kHz → 100 Hz frame rate) with up to 8 residual vector quantisation layers and supports all major Indic scripts natively.

Key Result

In contrast to EnCodec, Lipika's script-conditioned codebooks allocate dedicated entries for retroflex place of articulation, improving token discriminability on minimal pairs like /pəl/ (moment), /pʰəl/ (fruit), and /pʰɑːl/ (edge of a blade) in Hindi. The architecture is validated on 12 script families spanning 22+ languages.

Motivation

Why existing codecs fail for Indic languages.

Ph

Phonemic Density

Indic languages carry among the highest phonemic inventories in the world. Hindi has 33 consonants and 11 vowels with contrastive length. Retroflexion, aspiration, and murmured voice are all phonemically contrastive. A codec that collapses these distinctions produces tokens that are semantically incorrect.

Sc

Script Diversity

Each Indic script family encodes subtly different phonological inventories. Devanagari distinguishes /ʃ/ from /ʂ/; Tamil merges them. Bengali lacks phonemic /v/ but has /bʱ/. Generic codecs have no mechanism to represent these script-level differences.

Da

Data Scarcity

While English speech datasets exceed 100,000 hours, the largest Indic speech corpus (IndicVoices) provides ~7,300 hours across 22 languages. A codec that can share parameters across scripts while conditioning on script identity makes better use of limited data.

To

Token Quality for TTS

Downstream VALL-E-style TTS models treat codec tokens as language. If the first codebook does not capture phonetic identity cleanly, the language model must model acoustic texture — wasting capacity. Semantic distillation ensures coarse codes are phonetically meaningful.

Architecture

Encoder → Residual VQ → Decoder, with script conditioning.

The design follows the EnCodec lineage with substantial modifications for Indic phonological structure. The diagram below shows every component, data flow, and loss signal.

Input Waveform (B, 1, T) @ 24 kHz Encoder CausalConv1d Stem (k=7) Strided Downsampling [2, 4, 5, 6] → 240× reduction Gated ResBlocks (dil 1,3,9) Bottleneck → LayerNorm Script Adapter Embedding (12 × 64) AdaLN: scale & shift Retroflex prior +0.5 AdaLN Injection Latent z (B, T/240, C) + σ=1e-3 jitter VQ Codebook 0 EMA • Dead-code reset • 1024×128 VQ Codebooks 1..N-1 Residual quantisation • Shared dim=128 ... VQ Codebook N-1 N = 4, 6, or 8 (preset) W2V-BERT 2.0 Hidden Layer 6 (1024-d) Frozen • 16 kHz MSE (semantic KD) Token Stream (B, T/240, N) int64 Decoder CausalConvTranspose1d (k=7) Upsampling [6, 5, 4, 2] Gated ResBlocks (dil 1,3,9) Conv1d → Tanh → [-1, 1] Reconstructed (B, 1, T) @ 24 kHz Multi-Scale + Multi-Period Discriminator MSD (3 scales: ×1, ×2, ×4) + MPD (periods: 2,3,5,7,11) • Spectral Normalisation • Hinge Loss Legend Encoder / Decoder Residual VQ + Script Adapter Semantic Teacher / Output GAN Discriminator

Compression Pipeline

StageInput ShapeOutput ShapeOperation
Waveform(B, 1, T)—24 kHz mono audio
Stem(B, 1, T)(B, C, T)CausalConv1d (k=7, C=64..512)
Stride 2(B, C, T)(B, C, T/2)ResBlocks + DownConv + AvgPool(2)
Stride 4(B, C, T/2)(B, C, T/8)ResBlocks + DownConv + AvgPool(4)
Stride 5(B, C, T/8)(B, C, T/40)ResBlocks + DownConv + AvgPool(5)
Stride 6(B, C, T/40)(B, C, T/240)ResBlocks + DownConv + AvgPool(6)
Bottleneck(B, C, T/240)(B, C, T/240)CausalConv1d (k=1) + LayerNorm
AdaLN(B, C, T/240)(B, C, T/240)z * scale(script) + shift(script)
Input Proj(B, C, T/240)(B, T/240, 128)Linear(C → 128)
RVQ(B, T/240, 128)(B, T/240, N)N residual codebooks, int64 codes
Component Specification

Every module, documented.

C1

CausalConv1d

Standard 1D convolution with left-only padding. Kernel size k, dilation d produces causal padding of (k-1)*d. Prevents any future-context leakage through the latent sequence — critical for streaming TTS inference.

Pad: (causal_pad, 0) only
CT

CausalConvTranspose1d

Transposed convolution with causal trimming. After standard ConvTranspose1d, right-padding is stripped to remove non-causal artefacts that would otherwise degrade reconstruction on the final time steps.

Trim: kernel_size[0] - stride
RB

ResBlock

Gated residual block: ELU → CausalConv1d(k=3, dil=d) → ELU → CausalConv1d(k=1). Residual connection sums the output. Dilation schedule [1, 3, 9] provides 27-sample receptive field in three layers.

Dilations: [1, 3, 9] per block
EB

EncoderBlock

Downsampling unit: ResBlock stack → CausalConv1d(k=2S, doubles channels) → channel gating (first half) → AvgPool1d(stride=S). Keeps output channels equal to input channels.

Channel constant through block
DB

DecoderBlock

Upsampling unit: CausalConvTranspose1d(k=2S, stride=S, constant channels) → ResBlock stack. Mirror of EncoderBlock with transposed convolution replacing strided pool.

Strides reversed: [6, 5, 4, 2]
SA

ScriptFamilyAdapter

Embeds script ID (0-11) into 64-d vector. Projects through SiLU-activated MLP to encoder channel dimension. Produces per-channel scale and shift for AdaLN. Retroflex scripts get +0.5 bias in first 8 embedding dims.

AdaLN: y = z * γ(s) + β(s)
Technical Innovation

Five contributions beyond the state of the art.

01

Script-Family AdaLN Conditioning

Existing codecs are language-agnostic: the same encoder processes all languages identically. Lipika conditions the encoder's LayerNorm on a 64-dimensional script embedding projected through SiLU-activated linear layers to produce per-channel scale and shift parameters. This provides a script-level inductive bias that does not increase codebook size or modify the decoder — only the encoder sees the script identity. For multilingual training, this enables parameter sharing across languages while preserving script-specific phonological structure.

02

Retroflex-Aware Latent Prior

Ten of twelve supported script families carry phonemically contrastive retroflex consonants. During weight initialisation, the script embedding's first 8 dimensions are biased by +0.5 for retroflex-bearing scripts. This soft prior creates an inductive bias that encourages the first VQ codebook to allocate dedicated entries for retroflex place of articulation. The prior is trainable — the model can override it — but empirical results show it accelerates convergence and improves codebook utilisation for retroflex-heavy languages.

03

Semantic Distillation from W2V-BERT 2.0

The first VQ codebook is trained with an auxiliary MSE loss against frozen W2V-BERT 2.0 hidden states (layer 6, 1024-dimensional) projected through a learnable 2-layer MLP. This ensures the coarsest codes capture phonetic identity rather than low-level acoustic texture. The remaining codebooks refine residual detail. For deployment without W2V-BERT, the teacher can be disabled (--no-semantic), removing the need for the transformers library entirely.

04

EMA Dead-Code Recycling (Cross-Platform)

Codebook collapse is a known failure mode in VQ-VAEs where unused codes receive no gradient and the effective codebook size shrinks. Lipika tracks per-code usage via exponential moving average counts and reinitialises codes with fewer than threshold uses from randomly sampled live batch vectors. The implementation uses a clone-modify-copy pattern for buffer updates, making it compatible with CUDA direct assignment, MPS (which forbids in-place buffer mutation), and CPU — a nontrivial engineering requirement for true cross-platform operation.

05

Gaussian Jitter for Quantisation Symmetry Breaking

With small batch sizes (common on CPU and MPS), multiple encoder frames can converge to identical latent vectors, causing them to collapse onto the same codebook entry. Lipika adds zero-mean Gaussian noise (σ=1e-3) to the projected encoder output during training only. This breaks the symmetry without affecting evaluation, dramatically improving codebook utilisation at batch sizes as small as 2. The jitter is disabled during inference.

Hardware Presets

Auto-scaling model configuration.

The preset system detects available VRAM (CUDA) or falls back to CPU/MPS-optimised configurations. Each preset adjusts encoder/decoder channels, number of codebooks, codebook size, batch size, and training schedule.

PresetEnc/Dec ChCBsCB SizeBatchVRAMParametersTarget Hardware
cpu 6441282N/A 1.18M AMD64, Intel x86-64, Apple M1
gpu-small 25665124~8 GB 3.77M RTX 3060, T4, M2 Ultra
gpu-full 512810248~16 GB 14.10M A100, RTX 3090/4090, H100

Auto-Detection Logic

Pass --preset auto. The system checks torch.cuda.is_available() and torch.cuda.get_device_properties().total_memory. If no CUDA device is found, it checks torch.backends.mps.is_available(). Falls back to the cpu preset for any other scenario. VRAM ≥ 12 GB selects gpu-full; otherwise gpu-small.

Script Families

Twelve Indic scripts with phonological awareness.

IDScriptLanguagesRetroflexAspirationMurmur
0Devanagarihi, mr, sa, ne, kokYesYesYes
1Bengalibn, asYesYesYes
2GurmukhipaYesYesNo
3GujaratiguYesYesYes
4OriyaorYesYesYes
5TamiltaYesNoNo
6TeluguteYesYesYes
7KannadaknYesYesYes
8MalayalammlYesYesNo
9Perso-Arabicur, ksNoYesNo
10MeiteimniYesYesYes
11Latin (India)enNoNoNo
Indic Phonology

Phonological features relevant to codec design.

Re

Retroflexion

Consonants articulated with the tongue tip curled back against the palate. Phonemically contrastive in 10 of 12 scripts. Minimal pair example in Hindi: /paʈ/ (jute) vs /pat/ (leaf).

As

Aspiration

A puff of air following consonant release. Contrastive in most Indic languages. Example: /pal/ (moment) vs /pʰal/ (fruit). Tamil is a notable exception — aspiration is allophonic, not phonemic.

Mu

Murmured Voice

Also called "breathy voice." The vocal folds vibrate loosely, producing audible airflow. Contrastive in Hindi, Bengali, Gujarati, and others. Example: /bʱaːr/ (load) vs /baːr/ (time).

VL

Vowel Length

Short and long vowels are phonemically distinct in nearly all Indic languages. Example: /pəl/ (moment) vs /pɑːl/ (edge of blade). Spectrally similar but temporally distinct — a challenge for codecs.

Training Methodology

Two-phase training with GAN refinement.

Phase I: Reconstruction Steps 0 — disc_start_step L1 Waveform • Mel L1 • Multi-Scale STFT • VQ Commitment • Semantic KD No discriminator. Establishes stable quantised latent space. GAN ON Phase II: Adversarial Refinement After disc_start_step Hinge GAN Loss • Feature Matching • All Phase I Losses (active) MSD (×1,×2,×4) + MPD (2,3,5,7,11). Spectral Normalisation throughout. Step 0 disc_start_step Step N (lr → 0)

Optimiser Configuration

AdamW with β=(0.8, 0.99), weight decay 1e-2, gradient clipping at 1.0. Cosine learning rate schedule with linear warmup over warmup_steps (default 1000), decaying to 10% of peak LR over lr_decay_steps (default 400,000). Generator LR 3e-4, discriminator LR 3e-4.

Loss Functions

Complete mathematical specification.

LossSymbolWeightDefinition
Time-domain L1 Lt 0.1 ‖ ◯ − x ‖1 — direct waveform reconstruction
Mel-spectrogram L1 Lmel 1.0 ‖ log(1 + mel(◯)) − log(1 + mel(x)) ‖1 — 128 mels, 24 kHz, fmax=12kHz
Multi-Scale STFT Lstft 1.0 Σn∈{256,512,1024,2048} ( SCn + LMn ) / 4 where SC = ‖S−Ŝ‖F/‖S‖F, LM = ‖log Ŝ − log S‖1
VQ Commitment Lvq 1.0 β · ‖ zqdetach − ze ‖22 — β=1.0 (commitment_cost)
Semantic KD Lsem 10.0 ‖ semantic_head(zq,0) − W2V-BERT.hidden6 ‖22
Adversarial (Gen) Ladv 3.0 −E[D(◯)] — Hinge generator loss
Feature Matching Lfeat 3.0 Σl ‖ Dl(x) − Dl(◯) ‖1
Discriminator Ldisc 1.0 E[ReLU(1 − D(x))] + E[ReLU(1 + D(◯))] — Hinge discriminator loss

Total Generator Loss

LG = 0.1·Lt + 1.0·Lmel + 1.0·Lstft + 1.0·Lvq + 10.0·Lsem + 3.0·Ladv + 3.0·Lfeat

Programmatic Interface

Complete API reference.

tokenizer.py — Core Model
class LipikaTokenizer(nn.Module):

    # Construction
    def __init__(
        audio_cfg: AudioConfig,
        rvq_cfg: RVQConfig,
        model_cfg: ModelConfig,
        use_semantic_teacher: bool = True
    )

    # Training forward pass — returns all intermediate tensors and losses
    def forward(
        waveform: Tensor,          # (B, 1, T) @ 24kHz
        script_ids: Optional[Tensor] = None  # (B,) int64, 0-11
    ) -> Dict[str, Tensor]:
        return {
            "reconstructed": Tensor,  # (B, 1, T)
            "target":       Tensor,  # (B, 1, T)
            "codes":        Tensor,  # (B, T/240, N_cb) int64
            "recon_loss":   Tensor,  # scalar
            "mel_loss":     Tensor,  # scalar
            "stft_loss":    Tensor,  # scalar
            "vq_loss":      Tensor,  # scalar
            "semantic_loss": Tensor,  # scalar
        }

    # Inference: waveform → discrete codes
    def encode(
        waveform: Tensor,          # (B, 1, T)
        script_ids: Optional[Tensor] = None
    ) -> Tensor:                  # (B, T/240, N_cb) int64

    # Inference: discrete codes → waveform
    def decode(codes: Tensor) -> Tensor:  # (B, T/240, N_cb) → (B, 1, T)

    # Properties
    frame_rate: float              # 100.0 Hz
    num_parameters(exclude_teacher: bool = True) -> int
tokenizer.py — Configuration Dataclasses
@dataclass
class AudioConfig:
    sample_rate: int = 24_000
    n_fft:       int = 2048
    hop_length:  int = 240
    n_mels:      int = 128
    fmin:        float = 0.0
    fmax:        float = 12_000.0

@dataclass
class RVQConfig:
    n_codebooks:              int = 8
    codebook_size:            int = 1024
    codebook_dim:             int = 128
    commitment_cost:          float = 1.0
    ema_decay:                float = 0.99
    ema_epsilon:              float = 1e-5
    threshold_ema_dead_code:  float = 1.0

@dataclass
class ModelConfig:
    encoder_channels: int = 512
    decoder_channels: int = 512
    n_script_families: int = 12
    script_embed_dim: int = 64
    w2v_bert_model: str = "facebook/w2v-bert-2.0"
    w2v_bert_dim: int = 1024
    semantic_proj_dim: int = 256
tokenizer.py — Training Entry Point
from tokenizer import train, AudioConfig, RVQConfig, ModelConfig, TrainingConfig

# Configure
audio_cfg = AudioConfig()
rvq_cfg   = RVQConfig(n_codebooks=8, codebook_size=1024)
model_cfg = ModelConfig(encoder_channels=512)
train_cfg = TrainingConfig(
    data_dir="./data",
    batch_size=8,
    num_epochs=200,
    device="auto",
)

# Launch (rank=0, world_size=1 for single-GPU/CPU)
train(0, 1, audio_cfg, rvq_cfg, model_cfg, train_cfg)

# Or via CLI:
# python tokenizer.py train --preset auto --data-dir ./data --epochs 200
Inference Utilities
from tokenizer import (
    _load_model_from_checkpoint,
    encode_audio_file,
    decode_codes_to_file,
    get_device,
)

# Load trained model
model = _load_model_from_checkpoint(
    "./checkpoints/ckpt_step00005000.pt",
    device="auto"
)

# Encode audio file → tokens
codes = encode_audio_file(
    model, "speech.wav", lang="hi"
)  # → (1, T/240, 8)

# Decode tokens → audio file
decode_codes_to_file(
    model, codes, "output.wav"
)

# Export to TorchScript
from tokenizer import export_torchscript
export_torchscript(model, "lipika.pt")
Quick Start

Deploy in five minutes.

Terminal — Installation
# Core: PyTorch + audio processing
pip install torch torchvision torchaudio soundfile librosa

# Training: monitoring + progress bars
pip install numpy tqdm tensorboard

# Optional: W2V-BERT semantic teacher
pip install transformers

# Optional: training plots
pip install matplotlib

# Optional: configuration management
pip install einops omegaconf

Place audio files in the data directory with optional language metadata. Supported formats: WAV, FLAC, OGG, MP3, Opus.

data/
speaker001_utt001.wav
speaker001_utt001.json # {"lang": "hi"}
speaker001_utt002.wav
speaker001_utt002.json # {"lang": "hi"}
speaker002_utt001.wav
speaker002_utt001.json # {"lang": "ta"}
...
Terminal — Training
# Smoke test — verify forward pass on your hardware
python tokenizer.py smoke-test

# Auto-detect hardware and start training
python tokenizer.py train --preset auto --data-dir ./data

# CPU-only (4-core AMD, your system)
python tokenizer.py train --preset cpu --device cpu --epochs 10

# With W2V-BERT semantic distillation (needs ~2 GB extra VRAM)
python tokenizer.py train --preset gpu-full --device cuda

# Without semantic teacher (faster, less memory)
python tokenizer.py train --preset gpu-full --no-semantic

# Resume from checkpoint
python tokenizer.py train --resume ./checkpoints/ckpt_step00005000.pt

# Monitor training
tensorboard --logdir ./logs
Glossary

Key terms and definitions.

AdaLN
Adaptive Layer Normalisation. A conditioning mechanism where an external signal (here, script identity) predicts per-channel scale (γ) and shift (β) parameters that modulate the normalised activations: y = γ(c) · LayerNorm(z) + β(c).
Codebook Collapse
A failure mode in vector quantisation where only a fraction of codebook entries receive gradient updates, causing the effective codebook size to shrink. Mitigated by EMA updates and dead-code recycling.
Commitment Loss
The MSE between the encoder output ze and the quantised vector zq (with zq detached). Encourages the encoder to commit to the codebook's discretisation.
Compression Ratio
The factor by which temporal resolution is reduced. Lipika's encoder converts 24,000 samples per second to 100 latent frames per second: 24,000 / 100 = 240×.
EMA Update
Exponential Moving Average update for codebook entries. Instead of gradient descent, codebook vectors are updated as: e := α·e + (1-α)·mean(assigned_z), with α=0.99. More stable than gradient-based updates.
Feature Matching Loss
L1 distance between discriminator intermediate feature maps for real and generated audio. Encourages the generator to produce outputs with similar internal representations to real data.
Frame Rate
The number of latent frames (and thus token vectors) per second of audio. Lipika: 100 Hz = one token every 10 milliseconds.
Gated ResBlock
A residual block where the main path's output is gated (element-wise multiplied) with a learned gate before being added to the skip connection. Smooths gradient flow.
Hinge Loss
A GAN objective: Discriminator minimises ReLU(1-D(x)) + ReLU(1+D(G(z))). Generator minimises -D(G(z)). More stable than the original cross-entropy GAN loss.
MPD
Multi-Period Discriminator. Sub-discriminators operating on different periodic subsamplings of the waveform (periods: 2, 3, 5, 7, 11). Captures prosodic and rhythmic patterns.
MSD
Multi-Scale Discriminator. Sub-discriminators operating on the waveform at different temporal resolutions (original, ×2 pooled, ×4 pooled). Captures spectral and long-range patterns.
Retroflex Consonant
A consonant articulated with the tongue tip curled back toward the hard palate. Phonemically contrastive in most Indic languages. IPA symbols: /ṭ ḍ ṇ ɽ ʂ/.
RVQ
Residual Vector Quantisation. A cascade of VQ layers where each layer quantises the residual error of the previous layer. Allows coarse-to-fine discretisation with additive reconstruction.
Semantic Distillation
Training the first RVQ codebook to predict frozen W2V-BERT hidden states. Ensures the coarsest tokens capture phonetic identity rather than low-level acoustic texture.
Spectral Normalisation
A weight normalisation technique that constrains the Lipschitz constant of each discriminator layer by dividing weights by their spectral norm. Stabilises GAN training.
Straight-Through Estimator
During backpropagation through the quantisation step, the gradient is passed through unchanged: z_output = z_input + stop_gradient(z_quantised - z_input). The forward pass uses z_quantised but the backward pass flows through z_input.
References

Grounded in the literature.

#PaperYearContribution to Lipika
[1]EnCodec — Défossez et al.2022Core encoder/decoder/RVQ architecture, causal convolutions, hinge GAN loss
[2]SoundStream — Zeghidour et al.2021Residual vector quantiser design, multi-codebook structure
[3]VQ-VAE — van den Oord et al.2017Straight-through gradient estimator, commitment loss formulation
[4]VALL-E — Wang et al.2023Motivation for discrete audio tokens as language model targets for TTS
[5]data2vec — Baevski et al.2022W2V-BERT as semantic feature extractor for distillation target
[6]W2V-BERT — Chung et al.2021Contrastive + MLM pre-training for self-supervised speech representations
[7]MelGAN — Kumar et al.2019Multi-scale discriminator architecture
[8]HiFi-GAN — Kong et al.2020Multi-period discriminator, feature matching loss
[9]WGAN-GP — Gulrajani et al.2017Theoretical grounding for GAN stability improvements
[10]Spectral Normalisation — Miyato et al.2018Weight normalisation for Lipschitz-constrained discriminators
[11]VITS — Kim et al.2021Multi-scale STFT loss for perceptually motivated training
[12]Layer Normalisation — Ba et al.2016AdaLN formulation for script conditioning
[13]Vocos — Siuzdak2023Frequency-domain reconstruction paradigm
[14]AudioCraft — Défossez et al.2023EMA codebook update methodology
[15]Codebook Collapse — Zeyer et al.2023Dead-code recycling strategy for VQ training stability
Configuration Reference

Every hyperparameter, explained.

ParameterDefaultDescription
sample_rate24000Audio sample rate in Hz. 24 kHz captures energy up to Nyquist 12 kHz — sufficient for all Indic phonemes including high-frequency sibilants.
n_fft2048FFT window size for STFT computations. 2048 samples = 85.3 ms at 24 kHz.
hop_length240Hop length between STFT frames. 240 samples = 10 ms at 24 kHz — aligns with latent frame rate.
n_mels128Number of mel filterbank bins for mel-spectrogram loss.
fmax12000Maximum frequency for mel filterbank. Nyquist at 24 kHz.
encoder_channels512Number of channels throughout encoder. Constant across all encoder blocks.
n_codebooks8Number of residual VQ layers. First layer captures semantics; subsequent layers refine acoustic detail.
codebook_size1024Number of entries per codebook. 1024 entries = 10-bit index per codebook.
codebook_dim128Dimensionality of each codebook entry. Shared across all codebooks.
commitment_cost1.0Weight β on the VQ commitment loss. Higher values push the encoder harder toward codebook entries.
ema_decay0.99Decay factor α for EMA codebook updates. Higher = slower adaptation, more stability.
threshold_ema_dead_code1.0Minimum usage count before a code is considered dead and re-initialised. Lower = more aggressive recycling.
disc_start_step10000Global step at which GAN training activates. Before this, only reconstruction losses are active.
w_time_recon0.1Weight on time-domain L1 waveform loss.
w_mel1.0Weight on mel-spectrogram L1 loss.
w_freq_recon1.0Weight on multi-scale STFT loss.
w_vq1.0Weight on VQ commitment loss.
w_semantic10.0Weight on W2V-BERT semantic distillation loss. High weight ensures phonetic grounding.
w_gen3.0Weight on hinge generator adversarial loss.
w_feat3.0Weight on feature matching loss.