math0_sft
d125m (134.1M parameters) on GSM8K + MATH worked solutions, trained on RTX 3060 12GB (Home GPU). Started 1 Oct 2026, ended 1 Oct 2026.
Hypothesis
Fine-tuning MATH0 on worked solutions (the GSM8K and MATH training sets) teaches it to solve grade-school word problems.
What we learned
Partly: GSM8K exact match 2.43% (32 of 1,319) vs 0.91% before. It learned the answer format (1,308 of 1,319 answers end in one) but not the arithmetic. At 125M the limit is reasoning, not format.
Scores
Higher is better on the headline; change against the parent run, MATH0.
| Test set | MATH0-SFT | MATH0 (parent) | Change |
|---|---|---|---|
| GSM8K exact match headline The share of the 1,319 GSM8K test problems the model answers exactly right, with three worked examples in the prompt. |
2.43% | 0.91% | +1.52 pts |
| GSM8K exact match, no examples The same 1,319 problems with no worked examples in the prompt: only the question. Meaningful for fine-tuned models, which learned the answer format. |
2.43% | — | |
| DeepMind maths, interpolate Exact answers on 5,600 questions from the DeepMind mathematics test set (100 from each of 56 kinds: arithmetic, algebra, calculus, number facts), five worked examples in the prompt. |
6.59% | 7.61% | −1.02 pts |
| DeepMind maths, extrapolate The same on 1,500 harder questions (15 kinds) with bigger numbers or longer expressions than the training data has. |
2.33% | 2.4% | −0.07 pts |
Exact match: the share of problems answered exactly right, as a percentage. Higher is better.
Curves
Model
- Preset
- d125m
- Parameters
- 134,105,856
- Outside embeddings
- 84,953,856
- Layers · heads · width
- 12 · 12 · 768
- Tokenizer
- Shared unigram 32k
Data
- Slice
- GSM8K + MATH worked solutions
- Words
- —
- Training tokens
- —
- Passes
- —
- Tokens seen
- —
Compute
- GPU
- RTX 3060 12GB
- Where
- Home GPU
- Steps
- 1,362 / 1,362
- GPU hours
- 0.22
- Cost
- —
- Spot restarts
- —
Lineage
homed125mfine-tune