Pocket-TTS Korean 100M

Pocket-TTS Korean 100M is a Korean zero-shot text-to-speech model compatible with the official Pocket-TTS Python package and CLI. It is a 6-layer student depth-distilled from the Korean 24-layer teacher seastar105/pocket-tts-korean-300m.

This is a community model and is not an official Kyutai release.

Model details

  • Architecture: Pocket-TTS FlowLM with Lagrangian Self Distillation (LSD) and the Mimi neural audio codec
  • Exact parameter count: 109,502,146
    • FlowLM: 89,447,809
    • Mimi: 20,054,337
  • FlowLM: 6 transformer layers, model dimension 1024, 16 attention heads
  • Korean tokenizer: SentencePiece, 4,000 tokens
  • Audio: mono, 24 kHz, 12.5 latent frames per second
  • Published weight: step-6k EMA checkpoint.
  • Bundle contents: EMA FlowLM weights and the frozen Mimi codec in one Pocket-TTS-format safetensors bundle
  • Weight format: float32 safetensors
  • Student initialization: Kyutai English 6-layer checkpoint, languages/english_2026-04/model.safetensors, revision 19f95fe2df36e79fbd9f10008595cc4c977a0fcc
  • Distillation teacher: raw step-50k checkpoint of the Korean 24-layer teacher

The rounded 100M name includes both the 89.4M-parameter FlowLM and the 20.1M Mimi codec. The 6-layer student is substantially smaller than the Korean 300M teacher and is the intended Pocket-TTS deployment model.

Quick start

Install and run the official Pocket-TTS CLI:

uvx pocket-tts generate \
    --config hf://seastar105/pocket-tts-korean-100m/korean.yaml@d5418f368ed97bc982ab80c5b639fe28f280636b \
    --voice ./voice_prompt.wav \
    --text "안녕하세요. 한국어 음성 합성 모델입니다."

The config revision is pinned so an older cached release cannot silently load the previous weight.

The voice prompt should contain clean speech from a speaker who has consented to voice cloning.

Python usage:

from pocket_tts import TTSModel
import scipy.io.wavfile

model = TTSModel.load_model(
    config="hf://seastar105/pocket-tts-korean-100m/korean.yaml@d5418f368ed97bc982ab80c5b639fe28f280636b"
)
voice_state = model.get_state_for_audio_prompt("./voice_prompt.wav")
audio = model.generate_audio(
    voice_state,
    "안녕하세요. 한국어 음성 합성 모델입니다.",
)
scipy.io.wavfile.write(
    "korean_tts.wav",
    model.sample_rate,
    audio.detach().cpu().numpy(),
)

Pocket-TTS requires Python 3.10 or newer and PyTorch 2.5 or newer. Keep the loaded model and voice state in memory when synthesizing multiple utterances.

Training

The student was initialized from the released English 6-layer checkpoint. Its text embedding was reset for the Korean tokenizer, then the trained Korean embedding and frozen flow/EOS heads were transferred from the Korean teacher. The remaining student backbone was optimized by depth distillation against the raw final teacher checkpoint.

Note: In our experiments, from-scratch Korean training collapsed very quickly. The released teacher was therefore warm-started from Kyutai's English checkpoint, and this student uses pretrained initialization plus depth distillation rather than from-scratch training.

Note: All intermediate training checkpoints are archived in seastar105/pocket-tts-checkpoints.

  • Dataset: seastar105/emilia-yodas-ko-filtered-pocket-tts, with 918,609 training utterances (2,276.88 hours) and 9,472 validation utterances (23.00 hours)
  • Training: 50,000 steps, global batch size 64, on 4 NVIDIA RTX 5090 GPUs
  • Optimization: AdamW, peak learning rate 4e-4, 1,000-step warmup followed by cosine decay, EMA decay 0.999

The exact resolved training arguments are included in training_args.yaml. The training audio and transcripts are not redistributed in this repository.

Evaluation

All 25 EMA checkpoints were evaluated on the complete 500-item Korean zero-shot split of yuekai/CV3-Eval.

CV3 Korean student EMA checkpoint sweep

Lower CER is better; higher speaker similarity and UTMOS are better. The 2k checkpoint is an immature outlier, especially for CER. Automated scores are useful for checkpoint selection but do not replace listening tests.

The step-6k EMA checkpoint is published because it achieved the lowest CER in the sweep (7.755%) while retaining a UTMOS score of 2.3694. UTMOS continued to decline at later checkpoints. The 2k checkpoint had higher UTMOS but an unusable 48.981% CER, so it was not selected.

Best checkpoint by metric

Model Lowest CER ↓ Highest speaker sim. ↑ Highest UTMOS ↑
Student 6,000 (7.755%) 2,000 (0.9187) 2,000 (2.7888)

Protocol

  • Checkpoints: EMA model.safetensors, steps 2k–50k at 2k intervals, from archive revision fd4eadb8b1eac9181029f398564ccc8a767f5d8a.
  • Data: all 500 zero_shot_ko items, revision 6ea9d3650fffcbed7c6279e6d1546d01ef1d2796.
  • Generation: seed 0, temperature 0.3, CFG 2.0, one decode step, EOS threshold -1.0, maximum 30 seconds, and full prompt audio.
  • Intelligibility: openai/whisper-large-v3 with Korean forced, revision 06f233fe06e710322aca913c1bc4249a0d71fce1. CER removes punctuation and preserves spaces; no-space CER additionally removes spaces.
  • Voice and quality: microsoft/wavlm-base-plus-sv speaker similarity and UTMOS.
  • Storage: audio was deleted immediately after all metrics for each item were committed. No evaluation audio is stored in this repository.

Machine-readable aggregate results are available as CSV and JSON. The earlier five-checkpoint raw-training evaluations remain under eval/cv3_ko_raw_step*.json for provenance and are distinct from this EMA sweep.

Raw scores — 25 EMA checkpoints
Step CER ↓ No-space CER ↓ Speaker sim. ↑ UTMOS ↑ No EOS
2,000 48.981% 53.918% 0.9187 2.7888 9
4,000 10.062% 9.900% 0.8906 2.4276 0
6,000 7.755% 7.750% 0.8872 2.3694 0
8,000 8.723% 8.616% 0.8878 2.3245 0
10,000 8.032% 8.042% 0.8880 2.3162 0
12,000 8.133% 8.071% 0.8868 2.3076 0
14,000 8.265% 8.404% 0.8883 2.2930 0
16,000 9.322% 9.349% 0.8881 2.2945 0
18,000 8.358% 8.467% 0.8874 2.2735 1
20,000 7.984% 7.899% 0.8858 2.2858 0
22,000 9.115% 9.395% 0.8871 2.2884 0
24,000 9.833% 10.043% 0.8839 2.2649 0
26,000 8.816% 8.879% 0.8851 2.2513 0
28,000 9.380% 9.309% 0.8864 2.2607 0
30,000 9.309% 8.897% 0.8847 2.2567 0
32,000 8.904% 8.770% 0.8879 2.2421 0
34,000 9.648% 9.659% 0.8845 2.2504 0
36,000 9.639% 9.510% 0.8823 2.2466 1
38,000 8.723% 8.742% 0.8808 2.2280 4
40,000 9.807% 10.244% 0.8811 2.2423 2
42,000 9.771% 9.733% 0.8795 2.2355 4
44,000 9.089% 9.177% 0.8824 2.2310 4
46,000 10.991% 11.539% 0.8809 2.2113 2
48,000 9.353% 9.292% 0.8819 2.2253 2
50,000 9.974% 10.106% 0.8826 2.2265 3

Limitations

  • English reading quality is very poor; treat this model as Korean-only for practical use.
  • Code-switching, numbers, abbreviations, rare names, and unusual punctuation were not systematically evaluated.
  • Some later EMA checkpoints produced generations that reached the 30-second limit without EOS; the published step-6k checkpoint produced none.
  • UTMOS is an automated estimate and may be less reliable for Korean than for the data on which it was developed.
  • No human listening study, demographic fairness audit, or robustness audit has been performed.
  • Output quality and speaker identity depend strongly on prompt cleanliness, duration, recording conditions, and consented speaker coverage.
  • Pocket-TTS processes one request at a time and is not thread-safe.

Responsible use

Only clone or imitate a voice with the speaker's explicit and lawful consent. Do not use this model for impersonation, fraud, deception, misinformation, harassment, privacy invasion, or any unlawful or harmful activity. Clearly disclose synthesized audio where listeners could reasonably mistake it for a genuine recording.

License and attribution

The model is released under CC BY 4.0, following the license of the Kyutai Pocket-TTS base weights. Credit Kyutai for Pocket-TTS and cite the original project when using this derivative model:

Users remain responsible for complying with the licenses and terms applicable to their prompts, generated content, and downstream uses.

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for seastar105/pocket-tts-korean-100m

Finetuned
(23)
this model

Dataset used to train seastar105/pocket-tts-korean-100m

Paper for seastar105/pocket-tts-korean-100m