ram-18m Training Script

Training script for ram-18m: an 18,290,304-param LLaMA-style language model.

Architecture

Parameter Value
d_model 384
n_heads 6
n_kv_heads 2 (GQA)
head_dim 64
n_layers 7
FFN SwiGLU, 4x (1536)
Norm RMSNorm
Positional RoPE (θ=10000)
Vocab 8192 (BPE)
Tied embed/head yes
Total params 18,290,304

Default Training Config

  • Data: FineWeb-Edu L3 (sample-100BT), ~2B tokens
  • Optimizer: AdamW (β=0.9/0.95, wd=0)
  • LR: 2e-4, cosine decay to 2e-5, warmup 200 steps
  • Batch: 32, seq_len 512
  • Steps: 12,207 (~2B tokens)
  • Grad clip: 1.0

Usage

# 1. Prepare data (trains BPE tokenizer + tokenizes 2B tokens)
python3 train_ram_18m.py --stage prepare

# 2. Train
python3 train_ram_18m.py --stage train

# 3. Or do both
python3 train_ram_18m.py --stage all

# 4. Eval a checkpoint
python3 train_ram_18m.py --stage eval --ckpt ckpt_step12207.pt

Requirements

pip install torch transformers datasets numpy tokenizers

Notes

  • Requested by GGUFGuy in model-requests #30
  • The model is NOT trained yet — this is the script only.
  • GPU recommended (RTX 3090+ for reasonable speed); CPU works but is ~50x slower.
  • Checkpoints saved every 500 steps to the script directory.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support