f11_slp1_uni8k_d700m_plus_all_v2_4x
d700m (704.1M parameters) on plus_all v2 slice (489.9M words, 4 passes (0.16 done)), trained on L40S 48GB (AWS). Started 8 Oct 2026.
Hypothesis
4 passes instead of 3 improves ex-Gita (~0.9% predicted by the passes proxy).
What we learned
F10 exactly (700M, filtered plus_all v2 slice, same seed, schedule shape and weight decay) with one change: four passes (129,096 steps) instead of three. On one L40S 48GB GPU on AWS, on-demand.
Scores
Curves
Training loss
Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.
Held-out bits per byte
The validation split, evaluated during training, by step. Lower is better.
Throughput
Tokens per second, by step.
Model
- Preset
- d700m
- Parameters
- 704,128,512
- Outside embeddings
- 679,552,512
- Layers · heads · width
- 24 · 12 · 1,536
- Tokenizer
- SLP1 unigram 8k
Data
- Slice
- plus_all v2 slice
- Words
- 489,850,000
- Training tokens
- 1,586,326,034
- Passes
- 4 passes (0.16 done)
- Tokens seen
- 260,505,600
Compute
- GPU
- L40S 48GB
- Where
- AWS
- Steps
- 5,300 / 129,096
- GPU hours
- 3.01
- Cost
- $5.64
- Spot restarts
- 0
Lineage
awsl40sd700m