The frontier leaderboard is now a list of effort settings, not models
Of the top 15 entries on Artificial Analysis' intelligence index, most are the same handful of models at different reasoning-effort settings — Claude Opus 5 appears at max, xhigh, high and medium; GPT-5.6 Sol does the same. The spread within a single model is wide: Opus 5 ranges from 60.7 down to 56.3 across its effort tiers, which is larger than the gap between many distinct models. Choosing a model is increasingly choosing a latency and cost dial on a model you've already picked, and benchmark comparisons that don't state the effort setting are getting harder to read.
AI engineering practiceIndustry analysisreasoning-effortbenchmarksmodel-selectionartificial-analysis