CompactLM-5M
A from-scratch ~5M-parameter language model, trained from zero on a 300 MB slice of diverse real web text (chemistry, code, literature, general web). This is an independent small-model build in the "fits on a floppy disk" range β not a fine-tune of a bigger model.
Architecture
LLaMA-style decoder, built from scratch:
| Field | Value |
|---|---|
| Parameters | 4,912,992 |
| Layers | 6 |
| Hidden size | 224 |
| Attention | GQA β 7 query heads, 2 KV heads, head_dim 32 |
| FFN | SwiGLU, intermediate 576 |
| Norm | RMSNorm (eps 1e-6) |
| Positional | RoPE (theta 10000) |
| Embeddings | Tied (input = output head) |
| Vocab | 8192 (BPE, trained on the corpus) |
| Context | 512 |
| Dtype | float32 |
Training
- Data:
/corpus_slice300mβ 300 MB of diverse real web text, tokenized to ~76.06M tokens (BPE, 8192 vocab). - Steps: 10,000 @ batch 32 Γ seq 512
- Optimizer: AdamW, betas (0.9, 0.95), weight decay 0.1
- LR: 3e-4, cosine decay with 10% warmup, floor 10%
- Grad clip: 1.0
- Hardware: NVIDIA RTX 5090 (32 GB), CUDA
Measured results (independently recomputed)
- Val perplexity: 58.91 β computed on the 1.52M-token held-out tail (last 2% of the corpus), token-level, by the author.
- Unigram baseline: 1453.67 on the same held-out tail.
- The model beats the unigram floor by ~25Γ, i.e. it genuinely learned context, not just token frequencies.
Note: the in-training val_loss (β0.008) is not a reliable number β the
training loop's validation slice leaked from the training stream. The 58.91
above is the honest held-out figure.
What it is good at / not
At 5M parameters this model produces grammatical first sentences but is far from fluent: greedy decoding is coherent for roughly the first sentence and then collapses into repetition loops (e.g. "the world's largest city in the world is the world's largest cityβ¦"), and sampled decoding (top-p 0.9, temp 0.7) is more varied but still drifts into incoherence within a few sentences. Factual recall is weak. It is a demonstration of from-scratch small-model training, not a useful general assistant.
Sample outputs (greedy, temp 0.0)
The capital of France is β "The capital of France is a very important part of the world's economy."
In machine learning, a neural network β "In machine learning, a neural network is a very important part of the development of the system."
To make a cup of tea, you need β "To make a cup of tea, you need to be sure to use a cup of coffee."
Sample outputs (temp 0.7, top-p 0.9)
Once upon a time, there was a β "Once upon a time, there was a great chance to have a lot of life."
The sun rises in the β "The sun rises in the sun. Collecting a land, which is known for its brightness in a space, has been caused by the dirt and swords."
Files
Version
This is v2 of CompactLM-5M. It replaces an earlier 6,162,688-param build (d256/4L/4H, vocab 12288) whose card honestly noted it produced grammatical word-salad. This v2 retrain (d224/6L/GQA, vocab 8192) produces grammatical first sentences (better than the v1 word-salad) but still collapses into repetition loops under greedy decoding, so it is published as a from-scratch training demonstration, not as a coherent generator.
model.safetensorsβ weights (27 MB, F32).head.weightandtok.weightare tied (identical values); both keys are present for loaders that expect an untied head.tokenizer.jsonβ BPE tokenizer (8192 vocab)config.jsonβ architecture config
Reproduction
Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact
script, tokenizer and training log are not bundled here; the architecture is
fully specified in config.json and above.
- Downloads last month
- 1,058