math0_mid_sft
d125m (134.1M parameters) on GSM8K + MATH worked solutions, trained on RTX 3060 12GB (Home GPU). Started 7 Oct 2026, ended 7 Oct 2026.
Hypothesis
MATH0-SFT's fine-tune, run on MATH0-MID instead of MATH0, scores higher than MATH0-SFT on GSM8K and DeepMind maths.
What we learned
Queued after MATH0-MID; scores follow.
Scores
Higher is better on the headline; change against the parent run, MATH0-MID.
GSM8K exact match
2.43%
exact match · headline
| Test set | MATH0-MID-SFT | MATH0-MID (parent) | Change |
|---|---|---|---|
| GSM8K exact match headline The share of the 1,319 GSM8K test problems the model answers exactly right, with three worked examples in the prompt. |
2.43% | — | |
| GSM8K exact match, no examples The same 1,319 problems with no worked examples in the prompt: only the question. Meaningful for fine-tuned models, which learned the answer format. |
2.12% | — | |
| DeepMind maths, interpolate Exact answers on 5,600 questions from the DeepMind mathematics test set (100 from each of 56 kinds: arithmetic, algebra, calculus, number facts), five worked examples in the prompt. |
9.52% | 19.54% | −10.02 pts |
| DeepMind maths, extrapolate The same on 1,500 harder questions (15 kinds) with bigger numbers or longer expressions than the training data has. |
6.73% | 14.8% | −8.07 pts |
Exact match: the share of problems answered exactly right, as a percentage. Higher is better.
Curves
No training curves were published for this run.Its scores above are what it published.
Model
- Preset
- d125m
- Parameters
- 134,105,856
- Outside embeddings
- 84,953,856
- Layers · heads · width
- 12 · 12 · 768
- Tokenizer
- Shared unigram 32k
Data
- Slice
- GSM8K + MATH worked solutions
- Words
- —
- Training tokens
- —
- Passes
- —
- Tokens seen
- —
Compute
- GPU
- RTX 3060 12GB
- Where
- Home GPU
- Steps
- 1,362 / 1,362
- GPU hours
- 0.22
- Cost
- —
- Spot restarts
- —
Lineage
homed125mfine-tune