Skip to main content

Model Rankings

Compare foundation models by benchmark scores when selecting a base model.

You can access the Model Rankings page to explore benchmark-based evaluation scores for a wide range of foundation models, including both open-source and proprietary options. This page helps you compare performance across standardized tasks to make informed decisions when selecting a base model for fine-tuning or inference.

Each model is evaluated on a variety of popular benchmarks such as:

  • AGIEval, ARC, MMLU – General academic and reasoning tasks
  • GSM8K, Math, DROP – Math and numerical reasoning
  • BoolQ, PIQA, SIQA – Commonsense and logic

For a complete list of benchmark datasets used in evaluation, see the Benchmark Datasets section.

Scores range from 0 to 1, with higher values indicating stronger performance.


Back to top