math0_mid
d125m (134.1M parameters) on OpenWebMath + DeepMind maths (70/30), trained on RTX 3060 12GB (Home GPU). Started 7 Oct 2026, ended 7 Oct 2026.
Hypothesis
250M more pre-training tokens for MATH0, 70% its maths web pages and 30% DeepMind arithmetic, algebra and number drills, teach it the arithmetic fine-tuning could not.
What we learned
Running: continued pre-training from MATH0; scores follow.
Scores
Lower is better on the headline; change against the parent run, MATH0.
| Test set | MATH0-MID | MATH0 (parent) | Change |
|---|---|---|---|
| OpenWebMath held-out, clean headline Mathematical web pages never used in training, minus every page with copied passages in the training text. |
1.07346 | 1.06595 | +0.7% |
| GSM8K test Grade-school maths word problems with worked solutions, scored as text. |
1.08254 | 1.07016 | +1.2% |
| MATH test, clean Competition problems with solutions, mostly in LaTeX, minus any with copied text in the training pages. |
0.86696 | 0.86027 | +0.8% |
| OpenWebMath held-out, all Every held-out OpenWebMath page, including the ones with passages copied into training pages. |
0.99899 | 0.99258 | +0.6% |
| MATH test, all All 5,000 MATH test problems, including the ones with copied text in the training pages. |
0.85298 | 0.84625 | +0.8% |
Bits per byte: how many bits the model needs, on average, to predict each byte of text it never saw in training. Lower is better, and it compares models with different tokenizers fairly, because every model is charged for the same bytes.
Curves
Training loss
Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.
Held-out bits per byte
The validation split, evaluated during training, by step. Lower is better.
Throughput
Tokens per second, by step.
Model
- Preset
- d125m
- Parameters
- 134,105,856
- Outside embeddings
- 84,953,856
- Layers · heads · width
- 12 · 12 · 768
- Tokenizer
- Shared unigram 32k
Data
- Slice
- OpenWebMath + DeepMind maths (70/30)
- Words
- —
- Training tokens
- 1,555,523,031
- Passes
- 0.16 passes
- Tokens seen
- 250,183,680
Compute
- GPU
- RTX 3060 12GB
- Where
- Home GPU
- Steps
- 5,090 / 5,090
- GPU hours
- 2.83
- Cost
- —
- Spot restarts
- —
Lineage
homed125m