math1_sft
d125m (134.1M parameters) on GSM8K, MATH, MetaMathQA and Orca-Math worked solutions, trained on RTX 3060 12GB (Home GPU). Started 6 Oct 2026, ended 6 Oct 2026.
Hypothesis
Fine-tuning MATH0 on about 35 times more worked solutions (GSM8K and MATH plus MetaMathQA and Orca-Math, 518,583 examples, one pass) teaches it the arithmetic the first fine-tune lacked.
What we learned
No. GSM8K with three examples 1.97% (26 of 1,319) vs 2.43% for MATH0-SFT, inside the noise. With no examples 4.09% vs 2.43%: the answer form, not better maths. DeepMind maths fell to 4.07% from 6.59%. At 125M, more fine-tuning data does not buy arithmetic.
Scores
Higher is better on the headline; change against the parent run, MATH0.
| Test set | MATH1-SFT | MATH0 (parent) | Change |
|---|---|---|---|
| GSM8K exact match headline The share of the 1,319 GSM8K test problems the model answers exactly right, with three worked examples in the prompt. |
1.97% | 0.91% | +1.06 pts |
| GSM8K exact match, no examples The same 1,319 problems with no worked examples in the prompt: only the question. Meaningful for fine-tuned models, which learned the answer format. |
4.09% | — | |
| DeepMind maths, interpolate Exact answers on 5,600 questions from the DeepMind mathematics test set (100 from each of 56 kinds: arithmetic, algebra, calculus, number facts), five worked examples in the prompt. |
4.07% | 7.61% | −3.54 pts |
| DeepMind maths, extrapolate The same on 1,500 harder questions (15 kinds) with bigger numbers or longer expressions than the training data has. |
2.2% | 2.4% | −0.20 pts |
Exact match: the share of problems answered exactly right, as a percentage. Higher is better.
Curves
Model
- Preset
- d125m
- Parameters
- 134,105,856
- Outside embeddings
- 84,953,856
- Layers · heads · width
- 12 · 12 · 768
- Tokenizer
- Shared unigram 32k
Data
- Slice
- GSM8K, MATH, MetaMathQA and Orca-Math worked solutions
- Words
- —
- Training tokens
- —
- Passes
- —
- Tokens seen
- —
Compute
- GPU
- RTX 3060 12GB
- Where
- Home GPU
- Steps
- 16,206 / 16,206
- GPU hours
- 2.81
- Cost
- —
- Spot restarts
- —
Lineage
homed125mfine-tune