F11

Sanskrit Running

f11_slp1_uni8k_d700m_plus_all_v2_4x

d700m (704.1M parameters) on plus_all v2 slice (489.9M words, 4 passes (0.16 done)), trained on L40S 48GB (AWS). Started 8 Oct 2026.

Hypothesis

4 passes instead of 3 improves ex-Gita (~0.9% predicted by the passes proxy).

What we learned

F10 exactly (700M, filtered plus_all v2 slice, same seed, schedule shape and weight decay) with one change: four passes (129,096 steps) instead of three. On one L40S 48GB GPU on AWS, on-demand.

Scores

No scores published yet. The held-out scores are published when the run finishes; the live card above shows the latest evaluation during training.

Curves

Training loss

Cross-entropy per token, by step. The first few percent of the run, far higher, run off the top; hover or the table has every value.

Held-out bits per byte

The validation split, evaluated during training, by step. Lower is better.

Throughput

Tokens per second, by step.

Model

Preset
d700m
Parameters
704,128,512
Outside embeddings
679,552,512
Layers · heads · width
24 · 12 · 1,536
Tokenizer
SLP1 unigram 8k

Muon + AdamW, rope, qk_norm, relu^2, untied head, weight decay

Data

Slice
plus_all v2 slice
Words
489,850,000
Training tokens
1,586,326,034
Passes
4 passes (0.16 done)
Tokens seen
260,505,600

Compute

GPU
L40S 48GB
Where
AWS
Steps
5,300 / 129,096
GPU hours
3.01
Cost
$5.64
Spot restarts
0

Lineage

awsl40sd700m