KrenTransV0-1

MWire Labs | WMT 2026 Submission | Team: KrenTransV0-1

Overview

NLLB-200-1.3B fine-tuned on 11 Northeast Indian languages paired with English. Extended vocabulary: 256,212 tokens (8 custom lang tags added).

This Checkpoint

  • CP2 โ€” retrain after critical tokenization bug fix (run 1 used batch["src_lang"][0] instead of per-sample src_lang)
  • Trained: 60,000 steps | Effective batch: 96 | bf16
  • mni_Mtei excluded (Meitei Mayek script unsupported by NLLB subword vocab)
  • Best val loss: ~2.245 (plateaued at step ~26k)

Languages

Code Language Script Category
asm_Beng Assamese Bengali Cat 1
lus_Latn Mizo Latin Cat 1
kha_Latn Khasi Latin Cat 1
mni_Beng Meitei Bengali Cat 1
njz_Latn Nyishi Latin Cat 1
brx_Deva Bodo Devanagari Cat 2
nag_Latn Nagamese Latin Cat 2
trp_Latn Kokborok Latin Cat 2
tgj_Latn Tangsa Latin Cat 2
mjw_Latn Karbi Latin Cat 2

Known Limitations

  • Category 2 languages (low data) show repetition loops with greedy decode โ€” use num_beams=4
  • Karbi (1,745 train pairs) and Tangsa (4,875) severely data-starved

Training Data

  • Train: 1,366,790 pairs | Dev: 74,964 pairs
  • Sources: WMT 2026 official data + HuggingFace datasets
Downloads last month
15
Safetensors
Model size
1B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support