ram-18m Training Script
Training script for ram-18m: an 18,290,304-param LLaMA-style language model.
Architecture
| Parameter | Value |
|---|---|
| d_model | 384 |
| n_heads | 6 |
| n_kv_heads | 2 (GQA) |
| head_dim | 64 |
| n_layers | 7 |
| FFN | SwiGLU, 4x (1536) |
| Norm | RMSNorm |
| Positional | RoPE (θ=10000) |
| Vocab | 8192 (BPE) |
| Tied embed/head | yes |
| Total params | 18,290,304 |
Default Training Config
- Data: FineWeb-Edu L3 (sample-100BT), ~2B tokens
- Optimizer: AdamW (β=0.9/0.95, wd=0)
- LR: 2e-4, cosine decay to 2e-5, warmup 200 steps
- Batch: 32, seq_len 512
- Steps: 12,207 (~2B tokens)
- Grad clip: 1.0
Usage
# 1. Prepare data (trains BPE tokenizer + tokenizes 2B tokens)
python3 train_ram_18m.py --stage prepare
# 2. Train
python3 train_ram_18m.py --stage train
# 3. Or do both
python3 train_ram_18m.py --stage all
# 4. Eval a checkpoint
python3 train_ram_18m.py --stage eval --ckpt ckpt_step12207.pt
Requirements
pip install torch transformers datasets numpy tokenizers
Notes
- Requested by GGUFGuy in model-requests #30
- The model is NOT trained yet — this is the script only.
- GPU recommended (RTX 3090+ for reasonable speed); CPU works but is ~50x slower.
- Checkpoints saved every 500 steps to the script directory.