TinyStories-40M

A 39.6M-parameter LLaMA-style transformer trained from scratch on the full TinyStories corpus. Published as a from-scratch training demonstration at the ~40M scale — it is not a coherent story generator (see Quality below).

Architecture

Parameter Value
Layers 12
Hidden size 512
Attention heads (Q) 8
Attention heads (KV) 4 (GQA 2:1)
Head dim 64
FFN (SwiGLU) 1408
Vocab size 8192 (BPE)
Context length 512
RoPE θ 10000
Norm RMSNorm (pre-norm)
Tied embeddings Yes
Precision FP32

Total parameters: 39,596,544 (86 tensors, head weight tied to the token embedding). Verified against the published model.safetensors.

Training

  • Data: roneneldan/TinyStories (~1.9B chars, ~490M tokens at seq 512)
  • Steps: 15,000
  • Batch size: 64 sequences × 512 tokens (32,704 tok/step)
  • Optimizer: AdamW, lr 3e-4, cosine decay (min-lr-frac 0.1), 300-step warmup
  • Hardware: RTX 5090 (32 GB), ~4 hours
  • Final train loss: 3.23 (step 15000)

Quality

Honest picture, from running the published weights (temp 0.7, top-k 50):

  • On the canonical prompt "Once upon a time" the model produces a story-like first sentence, then degrades:

    "Once upon a time, his mom and his clapped and cheered for him. They all enjoyed spending the rest of their special day in the park, his mom's light and a Stop being himself."

  • On other prompts it is weaker still — often a single short phrase or an immediate stop:

    "In the forest" → "In the forest things." "She opened the door" → "She opened the door."

  • Longer generations drift into incoherent, non-grammatical text.

This is what a 40M from-scratch model on a single 490M-token corpus demonstrably does: it learns the surface distribution of story text (word order, names, punctuation, the "Once upon a time" register) but does not sustain coherent narrative. It is published as a training demonstration, not as a usable story generator.

Usage

This is a custom transformers model — load with trust_remote_code=True:

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "Compactbot/tinystories-40m", trust_remote_code=True, torch_dtype=torch.float32
)
tok = AutoTokenizer.from_pretrained("Compactbot/tinystories-40m")

prompt = "Once upon a time"
ids = tok(prompt, return_tensors="pt").input_ids
out = model.generate(ids, max_new_tokens=100, temperature=0.7, top_k=50)
print(tok.decode(out[0], skip_special_tokens=True))

Note: the bundled generate() always samples (no do_sample flag); pass temperature and top_k to control it.

Notes

  • From-scratch training: initialized randomly, trained end-to-end on TinyStories. No pre-training from any other model.
  • GQA (Grouped Query Attention) 2:1, SwiGLU, RoPE — LLaMA-2 recipe at 40M scale.
  • ~151 MB in FP32 — under 200 MB.

Files

file what
model.safetensors 39,596,544 params, 86 tensors, FP32
config.json architecture + training metadata
modeling_tinystories.py the model class (loaded via trust_remote_code)
tokenizer.json BPE-8k tokenizer (HF tokenizers format)
Downloads last month
-
Safetensors
Model size
39.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Compactbot/tinystories-40m