CompactLM-5M

A from-scratch ~5M-parameter language model, trained from zero on a 300 MB slice of diverse real web text (chemistry, code, literature, general web). This is an independent small-model build in the "fits on a floppy disk" range β€” not a fine-tune of a bigger model.

Architecture

LLaMA-style decoder, built from scratch:

Field Value
Parameters 4,912,992
Layers 6
Hidden size 224
Attention GQA β€” 7 query heads, 2 KV heads, head_dim 32
FFN SwiGLU, intermediate 576
Norm RMSNorm (eps 1e-6)
Positional RoPE (theta 10000)
Embeddings Tied (input = output head)
Vocab 8192 (BPE, trained on the corpus)
Context 512
Dtype float32

Training

  • Data: /corpus_slice300m β€” 300 MB of diverse real web text, tokenized to ~76.06M tokens (BPE, 8192 vocab).
  • Steps: 10,000 @ batch 32 Γ— seq 512
  • Optimizer: AdamW, betas (0.9, 0.95), weight decay 0.1
  • LR: 3e-4, cosine decay with 10% warmup, floor 10%
  • Grad clip: 1.0
  • Hardware: NVIDIA RTX 5090 (32 GB), CUDA

Measured results (independently recomputed)

  • Val perplexity: 58.91 β€” computed on the 1.52M-token held-out tail (last 2% of the corpus), token-level, by the author.
  • Unigram baseline: 1453.67 on the same held-out tail.
  • The model beats the unigram floor by ~25Γ—, i.e. it genuinely learned context, not just token frequencies.

Note: the in-training val_loss (β‰ˆ0.008) is not a reliable number β€” the training loop's validation slice leaked from the training stream. The 58.91 above is the honest held-out figure.

What it is good at / not

At 5M parameters this model produces grammatical first sentences but is far from fluent: greedy decoding is coherent for roughly the first sentence and then collapses into repetition loops (e.g. "the world's largest city in the world is the world's largest city…"), and sampled decoding (top-p 0.9, temp 0.7) is more varied but still drifts into incoherence within a few sentences. Factual recall is weak. It is a demonstration of from-scratch small-model training, not a useful general assistant.

Sample outputs (greedy, temp 0.0)

The capital of France is β†’ "The capital of France is a very important part of the world's economy."

In machine learning, a neural network β†’ "In machine learning, a neural network is a very important part of the development of the system."

To make a cup of tea, you need β†’ "To make a cup of tea, you need to be sure to use a cup of coffee."

Sample outputs (temp 0.7, top-p 0.9)

Once upon a time, there was a β†’ "Once upon a time, there was a great chance to have a lot of life."

The sun rises in the β†’ "The sun rises in the sun. Collecting a land, which is known for its brightness in a space, has been caused by the dirt and swords."

Files

Version

This is v2 of CompactLM-5M. It replaces an earlier 6,162,688-param build (d256/4L/4H, vocab 12288) whose card honestly noted it produced grammatical word-salad. This v2 retrain (d224/6L/GQA, vocab 8192) produces grammatical first sentences (better than the v1 word-salad) but still collapses into repetition loops under greedy decoding, so it is published as a from-scratch training demonstration, not as a coherent generator.

  • model.safetensors β€” weights (27 MB, F32). head.weight and tok.weight are tied (identical values); both keys are present for loaders that expect an untied head.
  • tokenizer.json β€” BPE tokenizer (8192 vocab)
  • config.json β€” architecture config

Reproduction

Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact script, tokenizer and training log are not bundled here; the architecture is fully specified in config.json and above.

Downloads last month
1,058
Safetensors
Model size
6.75M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support