MATH1-SFT

Math DoneNeutral

math1_sft

d125m (134.1M parameters) on GSM8K, MATH, MetaMathQA and Orca-Math worked solutions, trained on RTX 3060 12GB (Home GPU). Started 6 Oct 2026, ended 6 Oct 2026.

Hypothesis

Fine-tuning MATH0 on about 35 times more worked solutions (GSM8K and MATH plus MetaMathQA and Orca-Math, 518,583 examples, one pass) teaches it the arithmetic the first fine-tune lacked.

What we learned

No. GSM8K with three examples 1.97% (26 of 1,319) vs 2.43% for MATH0-SFT, inside the noise. With no examples 4.09% vs 2.43%: the answer form, not better maths. DeepMind maths fell to 4.07% from 6.59%. At 125M, more fine-tuning data does not buy arithmetic.

Scores

Higher is better on the headline; change against the parent run, MATH0.

GSM8K exact match
1.97%
exact match · headline
+1.06 pts
Every published score of MATH1-SFT, against its parent MATH0
Test setMATH1-SFTMATH0 (parent)Change
GSM8K exact match headline
The share of the 1,319 GSM8K test problems the model answers exactly right, with three worked examples in the prompt.
1.97% 0.91%+1.06 pts
GSM8K exact match, no examples
The same 1,319 problems with no worked examples in the prompt: only the question. Meaningful for fine-tuned models, which learned the answer format.
4.09% —
DeepMind maths, interpolate
Exact answers on 5,600 questions from the DeepMind mathematics test set (100 from each of 56 kinds: arithmetic, algebra, calculus, number facts), five worked examples in the prompt.
4.07% 7.61%−3.54 pts
DeepMind maths, extrapolate
The same on 1,500 harder questions (15 kinds) with bigger numbers or longer expressions than the training data has.
2.2% 2.4%−0.20 pts

Exact match: the share of problems answered exactly right, as a percentage. Higher is better.

Curves

No training curves were published for this run.Its scores above are what it published.

Model

Preset
d125m
Parameters
134,105,856
Outside embeddings
84,953,856
Layers · heads · width
12 · 12 · 768
Tokenizer
Shared unigram 32k

AdamW fine-tune, answer-only loss, 1 epoch, from MATH0

Data

Slice
GSM8K, MATH, MetaMathQA and Orca-Math worked solutions
Words
—
Training tokens
—
Passes
—
Tokens seen
—

Compute

GPU
RTX 3060 12GB
Where
Home GPU
Steps
16,206 / 16,206
GPU hours
2.81
Cost
—
Spot restarts
—

Lineage

homed125mfine-tune