Instructions to use seastar105/pocket-tts-korean-100m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Pocket-TTS
How to use seastar105/pocket-tts-korean-100m with Pocket-TTS:
from pocket_tts import TTSModel import scipy.io.wavfile tts_model = TTSModel.load_model("seastar105/pocket-tts-korean-100m") voice_state = tts_model.get_state_for_audio_prompt( "hf://kyutai/tts-voices/alba-mackenna/casual.wav" ) audio = tts_model.generate_audio(voice_state, "Hello world, this is a test.") # Audio is a 1D torch tensor containing PCM data. scipy.io.wavfile.write("output.wav", tts_model.sample_rate, audio.numpy()) - Notebooks
- Google Colab
- Kaggle
Pocket-TTS Korean 100M
Pocket-TTS Korean 100M is a Korean zero-shot text-to-speech model compatible with the official Pocket-TTS Python package and CLI. It is a 6-layer student depth-distilled from the Korean 24-layer teacher seastar105/pocket-tts-korean-300m.
This is a community model and is not an official Kyutai release.
Model details
- Architecture: Pocket-TTS FlowLM with Lagrangian Self Distillation (LSD) and the Mimi neural audio codec
- Exact parameter count: 109,502,146
- FlowLM: 89,447,809
- Mimi: 20,054,337
- FlowLM: 6 transformer layers, model dimension 1024, 16 attention heads
- Korean tokenizer: SentencePiece, 4,000 tokens
- Audio: mono, 24 kHz, 12.5 latent frames per second
- Published weight: step-6k EMA checkpoint.
- Bundle contents: EMA FlowLM weights and the frozen Mimi codec in one Pocket-TTS-format safetensors bundle
- Weight format: float32 safetensors
- Student initialization: Kyutai English 6-layer checkpoint,
languages/english_2026-04/model.safetensors, revision19f95fe2df36e79fbd9f10008595cc4c977a0fcc - Distillation teacher: raw step-50k checkpoint of the Korean 24-layer teacher
The rounded 100M name includes both the 89.4M-parameter FlowLM and the 20.1M Mimi codec. The 6-layer student is substantially smaller than the Korean 300M teacher and is the intended Pocket-TTS deployment model.
Quick start
Install and run the official Pocket-TTS CLI:
uvx pocket-tts generate \
--config hf://seastar105/pocket-tts-korean-100m/korean.yaml@d5418f368ed97bc982ab80c5b639fe28f280636b \
--voice ./voice_prompt.wav \
--text "안녕하세요. 한국어 음성 합성 모델입니다."
The config revision is pinned so an older cached release cannot silently load the previous weight.
The voice prompt should contain clean speech from a speaker who has consented to voice cloning.
Python usage:
from pocket_tts import TTSModel
import scipy.io.wavfile
model = TTSModel.load_model(
config="hf://seastar105/pocket-tts-korean-100m/korean.yaml@d5418f368ed97bc982ab80c5b639fe28f280636b"
)
voice_state = model.get_state_for_audio_prompt("./voice_prompt.wav")
audio = model.generate_audio(
voice_state,
"안녕하세요. 한국어 음성 합성 모델입니다.",
)
scipy.io.wavfile.write(
"korean_tts.wav",
model.sample_rate,
audio.detach().cpu().numpy(),
)
Pocket-TTS requires Python 3.10 or newer and PyTorch 2.5 or newer. Keep the loaded model and voice state in memory when synthesizing multiple utterances.
Training
The student was initialized from the released English 6-layer checkpoint. Its text embedding was reset for the Korean tokenizer, then the trained Korean embedding and frozen flow/EOS heads were transferred from the Korean teacher. The remaining student backbone was optimized by depth distillation against the raw final teacher checkpoint.
Note: In our experiments, from-scratch Korean training collapsed very quickly. The released teacher was therefore warm-started from Kyutai's English checkpoint, and this student uses pretrained initialization plus depth distillation rather than from-scratch training.
Note: All intermediate training checkpoints are archived in seastar105/pocket-tts-checkpoints.
- Dataset: seastar105/emilia-yodas-ko-filtered-pocket-tts, with 918,609 training utterances (2,276.88 hours) and 9,472 validation utterances (23.00 hours)
- Training: 50,000 steps, global batch size 64, on 4 NVIDIA RTX 5090 GPUs
- Optimization: AdamW, peak learning rate 4e-4, 1,000-step warmup followed by cosine decay, EMA decay 0.999
The exact resolved training arguments are included in training_args.yaml.
The training audio and transcripts are not redistributed in this repository.
Evaluation
All 25 EMA checkpoints were evaluated on the complete 500-item Korean zero-shot split of yuekai/CV3-Eval.
Lower CER is better; higher speaker similarity and UTMOS are better. The 2k checkpoint is an immature outlier, especially for CER. Automated scores are useful for checkpoint selection but do not replace listening tests.
The step-6k EMA checkpoint is published because it achieved the lowest CER in the sweep (7.755%) while retaining a UTMOS score of 2.3694. UTMOS continued to decline at later checkpoints. The 2k checkpoint had higher UTMOS but an unusable 48.981% CER, so it was not selected.
Best checkpoint by metric
| Model | Lowest CER ↓ | Highest speaker sim. ↑ | Highest UTMOS ↑ |
|---|---|---|---|
| Student | 6,000 (7.755%) | 2,000 (0.9187) | 2,000 (2.7888) |
Protocol
- Checkpoints: EMA
model.safetensors, steps 2k–50k at 2k intervals, from archive revisionfd4eadb8b1eac9181029f398564ccc8a767f5d8a. - Data: all 500
zero_shot_koitems, revision6ea9d3650fffcbed7c6279e6d1546d01ef1d2796. - Generation: seed 0, temperature 0.3, CFG 2.0, one decode step, EOS threshold -1.0, maximum 30 seconds, and full prompt audio.
- Intelligibility:
openai/whisper-large-v3with Korean forced, revision06f233fe06e710322aca913c1bc4249a0d71fce1. CER removes punctuation and preserves spaces; no-space CER additionally removes spaces. - Voice and quality:
microsoft/wavlm-base-plus-svspeaker similarity and UTMOS. - Storage: audio was deleted immediately after all metrics for each item were committed. No evaluation audio is stored in this repository.
Machine-readable aggregate results are available as CSV and JSON. The earlier five-checkpoint raw-training evaluations remain under eval/cv3_ko_raw_step*.json for provenance and are distinct from this EMA sweep.
Raw scores — 25 EMA checkpoints
| Step | CER ↓ | No-space CER ↓ | Speaker sim. ↑ | UTMOS ↑ | No EOS |
|---|---|---|---|---|---|
| 2,000 | 48.981% | 53.918% | 0.9187 | 2.7888 | 9 |
| 4,000 | 10.062% | 9.900% | 0.8906 | 2.4276 | 0 |
| 6,000 | 7.755% | 7.750% | 0.8872 | 2.3694 | 0 |
| 8,000 | 8.723% | 8.616% | 0.8878 | 2.3245 | 0 |
| 10,000 | 8.032% | 8.042% | 0.8880 | 2.3162 | 0 |
| 12,000 | 8.133% | 8.071% | 0.8868 | 2.3076 | 0 |
| 14,000 | 8.265% | 8.404% | 0.8883 | 2.2930 | 0 |
| 16,000 | 9.322% | 9.349% | 0.8881 | 2.2945 | 0 |
| 18,000 | 8.358% | 8.467% | 0.8874 | 2.2735 | 1 |
| 20,000 | 7.984% | 7.899% | 0.8858 | 2.2858 | 0 |
| 22,000 | 9.115% | 9.395% | 0.8871 | 2.2884 | 0 |
| 24,000 | 9.833% | 10.043% | 0.8839 | 2.2649 | 0 |
| 26,000 | 8.816% | 8.879% | 0.8851 | 2.2513 | 0 |
| 28,000 | 9.380% | 9.309% | 0.8864 | 2.2607 | 0 |
| 30,000 | 9.309% | 8.897% | 0.8847 | 2.2567 | 0 |
| 32,000 | 8.904% | 8.770% | 0.8879 | 2.2421 | 0 |
| 34,000 | 9.648% | 9.659% | 0.8845 | 2.2504 | 0 |
| 36,000 | 9.639% | 9.510% | 0.8823 | 2.2466 | 1 |
| 38,000 | 8.723% | 8.742% | 0.8808 | 2.2280 | 4 |
| 40,000 | 9.807% | 10.244% | 0.8811 | 2.2423 | 2 |
| 42,000 | 9.771% | 9.733% | 0.8795 | 2.2355 | 4 |
| 44,000 | 9.089% | 9.177% | 0.8824 | 2.2310 | 4 |
| 46,000 | 10.991% | 11.539% | 0.8809 | 2.2113 | 2 |
| 48,000 | 9.353% | 9.292% | 0.8819 | 2.2253 | 2 |
| 50,000 | 9.974% | 10.106% | 0.8826 | 2.2265 | 3 |
Limitations
- English reading quality is very poor; treat this model as Korean-only for practical use.
- Code-switching, numbers, abbreviations, rare names, and unusual punctuation were not systematically evaluated.
- Some later EMA checkpoints produced generations that reached the 30-second limit without EOS; the published step-6k checkpoint produced none.
- UTMOS is an automated estimate and may be less reliable for Korean than for the data on which it was developed.
- No human listening study, demographic fairness audit, or robustness audit has been performed.
- Output quality and speaker identity depend strongly on prompt cleanliness, duration, recording conditions, and consented speaker coverage.
- Pocket-TTS processes one request at a time and is not thread-safe.
Responsible use
Only clone or imitate a voice with the speaker's explicit and lawful consent. Do not use this model for impersonation, fraud, deception, misinformation, harassment, privacy invasion, or any unlawful or harmful activity. Clearly disclose synthesized audio where listeners could reasonably mistake it for a genuine recording.
License and attribution
The model is released under CC BY 4.0, following the license of the Kyutai Pocket-TTS base weights. Credit Kyutai for Pocket-TTS and cite the original project when using this derivative model:
- Project: https://github.com/kyutai-labs/pocket-tts
- Base model: https://huggingface.co/kyutai/pocket-tts
- Paper: https://arxiv.org/abs/2509.06926
Users remain responsible for complying with the licenses and terms applicable to their prompts, generated content, and downstream uses.
- Downloads last month
- -
Model tree for seastar105/pocket-tts-korean-100m
Base model
kyutai/pocket-tts