Model Rankings
Compare foundation models by benchmark scores when selecting a base model.
You can access the Model Rankings page to explore benchmark-based evaluation scores for a wide range of foundation models, including both open-source and proprietary options. This page helps you compare performance across standardized tasks to make informed decisions when selecting a base model for fine-tuning or inference.
Each model is evaluated on a variety of popular benchmarks such as:
- AGIEval, ARC, MMLU – General academic and reasoning tasks
- GSM8K, Math, DROP – Math and numerical reasoning
- BoolQ, PIQA, SIQA – Commonsense and logic
For a complete list of benchmark datasets used in evaluation, see the Benchmark Datasets section.
Scores range from 0 to 1, with higher values indicating stronger performance.
Back to top
Was this page helpful?