RACER IS OP — Odia Web Corpus
Odia NLP datasets curated by RACER IS OP — web corpora for low-resource Odia language modeling
Updated • 112Note v4: 1M+ rows. Merged v1-v3, added SFT/ChatML instruction-tuning data. Both pretrain & SFT splits.
saidutta69/Odia-Web-Corpus-v3
Viewer • Updated • 470k • 60Note v3: 1M rows. Enhanced dedup, quality filtering, standardized splits. Improved over v2.
saidutta69/Odia-Web-Corpus-v2
Viewer • Updated • 2.17M • 129Note v2: 1M rows (900K train). Parquet format with train/test/val splits. Expanded & refined from v1.
saidutta69/Odia-Web-Corpus-v1
Viewer • Updated • 650k • 137Note v1: 650K docs, ~0.9GB. First Odia web corpus — raw scraped+filtered JSONL from public web sources.
saidutta69/Odia-Web-Corpus-v5
Viewer • Updated • 4.16M • 293Note v5: 4.16M rows of raw Odia web-scraped text (news, instructions, general web). Feeds into the pretrain dataset.
saidutta69/odia_pretrain_dataset
Viewer • Updated • 10.3M • 38Note 9.75M rows. Rebuilt from scratch because access requests went unanswered. CC-BY-4.0. Ready for training.
-
saidutta69/odia_pretrain_dataset_v2
Viewer • Updated • 12.3M • 207
saidutta69/odia-eval-benchmark
Viewer • Updated • 125k • 37Note Odia eval suite — multiple-choice, QA, classification, translation, NER. Pairs with the pretrain corpora.