Featured research Technical series / 01
Scaling a Diversified Model Civilization
We post-trained populations of models to develop complementary capabilities. This extends scaling beyond a single model: we can train more models, not just bigger ones.
2× the models trained
≈ 2.02× the parameters in one model
Proportional loss reduction: our evaluation loss versus Chinchilla’s reducible language-model loss, which also requires 2.35× the training data.
Read the full research