MATH0-MID-SFT

Math Done

math0_mid_sft

d125m (134.1M parameters) on GSM8K + MATH worked solutions, trained on RTX 3060 12GB (Home GPU). Started 7 Oct 2026, ended 7 Oct 2026.

Hypothesis

MATH0-SFT's fine-tune, run on MATH0-MID instead of MATH0, scores higher than MATH0-SFT on GSM8K and DeepMind maths.

What we learned

Queued after MATH0-MID; scores follow.

Scores

Higher is better on the headline; change against the parent run, MATH0-MID.

GSM8K exact match
2.43%
exact match · headline
Every published score of MATH0-MID-SFT, against its parent MATH0-MID
Test setMATH0-MID-SFTMATH0-MID (parent)Change
GSM8K exact match headline
The share of the 1,319 GSM8K test problems the model answers exactly right, with three worked examples in the prompt.
2.43% —
GSM8K exact match, no examples
The same 1,319 problems with no worked examples in the prompt: only the question. Meaningful for fine-tuned models, which learned the answer format.
2.12% —
DeepMind maths, interpolate
Exact answers on 5,600 questions from the DeepMind mathematics test set (100 from each of 56 kinds: arithmetic, algebra, calculus, number facts), five worked examples in the prompt.
9.52% 19.54%−10.02 pts
DeepMind maths, extrapolate
The same on 1,500 harder questions (15 kinds) with bigger numbers or longer expressions than the training data has.
6.73% 14.8%−8.07 pts

Exact match: the share of problems answered exactly right, as a percentage. Higher is better.

Curves

No training curves were published for this run.Its scores above are what it published.

Model

Preset
d125m
Parameters
134,105,856
Outside embeddings
84,953,856
Layers · heads · width
12 · 12 · 768
Tokenizer
Shared unigram 32k

AdamW fine-tune, answer-only loss, 3 epochs, from MATH0-MID

Data

Slice
GSM8K + MATH worked solutions
Words
—
Training tokens
—
Passes
—
Tokens seen
—

Compute

GPU
RTX 3060 12GB
Where
Home GPU
Steps
1,362 / 1,362
GPU hours
0.22
Cost
—
Spot restarts
—

Lineage

homed125mfine-tune