Own the training path
Train a tokenizer and a transformer from your own data, beginning with random weights. The toolkit supports single-GPU and distributed training, plus a desktop studio for exploring the process.
03 / Model research
TRAIN · MEASURE · REPEAT
A transformer training toolkit covering tokenization, training, exact checkpoint resumption, and inference, with an inspectable experimental record.
View repositoryFrom the public project record reviewed 08 October 2026. View source and measurement context ↗
Inside the project
Train a tokenizer and a transformer from your own data, beginning with random weights. The toolkit supports single-GPU and distributed training, plus a desktop studio for exploring the process.
Checkpoints restore optimizer and random-number state. Resumption checks reject changes to the schedule, tokenizer, or dataset so a resumed run does not silently become a different experiment.
Predictions are registered before experiments, and the findings record whether they held. The public history includes failed transfer experiments and a headline withdrawn after a data-contamination audit.
Training a model is not the same as demonstrating reliable reasoning. A previously reported tie with Phi-4-mini was withdrawn after contamination was found. The model’s clean program-synthesis results were much weaker. The linked research record retains the correction and the unsuccessful experiments.
Follow the evidence