Model Rankings
Compare foundation models by benchmark scores when selecting a model.
Feature no longer available
AI Studio no longer offers model rankings. It cannot be used from the console or the API, and this page is kept for reference only. See AI Studio for what is available.
You can access the Model Rankings page to explore benchmark-based evaluation scores for a wide range of foundation models, including both open-source and proprietary options. This page helps you compare performance across standardized tasks to make informed decisions when selecting a model for inference.
Each model is evaluated on a variety of popular benchmarks such as:
- AGIEval, ARC, MMLU – General academic and reasoning tasks
- GSM8K, Math, DROP – Math and numerical reasoning
- BoolQ, PIQA, SIQA – Commonsense and logic
Scores range from 0 to 1, with higher values indicating stronger performance.
Was this page helpful?