Skip to content

Choosing the everyday model.

A method for choosing an agent’s default model, using knowledge honesty, benchmark efficacy, and cost per correct answer.

The research behind the Everyday Model Index.

Most work is short

Across our own agent telemetry, more than nine in ten prompts resolve inside fifteen minutes. These are the tasks that make up much of the observed workload. Duration motivates evaluating a less expensive default; it does not, by itself, establish which model can complete a task.

Short by count. Uneven by cost.

Green marks work completed in under 15 minutes. Violet marks the longer-running tail.

Coding agents

17,609 measured prompts
Prompt count
94.4%under 15 minutes
Token spend
71.2%under 15 minutes

5.6% of prompts account for 28.8% of spend.

Knowledge work

3,796 measured prompts
Prompt count
91.9%under 15 minutes
Token spend
49.7%under 15 minutes

8.1% of prompts account for 50.3% of spend.

Each square represents approximately one percentage point. Aggregate observations motivate examining a lower-cost default; they do not establish savings from rerouting these prompts.
Where the saving is, and where it is not

Those sub-15-minute prompts carry 71% of coding spend and 50% of knowledge-work spend. At the two ends of the qualifying set, a correct answer costs $0.14 on the cheapest approved configuration against $2.04 on Claude Fable 5.1 at max, a 15× difference at 86% efficacy against 98%. These are benchmark costs, not measured savings from rerouting those prompts.

The caveat is the tail. It is small by count and large by cost, larger in knowledge work than in coding, so an everyday model is not a licence to stop paying attention to it. The escalation rule below handles it. Route on predicted duration, not on prompt count, or the saving will disappoint.

Zaun internal data · August 2026 · a prompt is measured from submission to its last completion event

Three questions

Every configuration is asked the same three, in this order. Fail the first and cost never matters; fail the second and it does not reach the ranking.

Qualification comes before price.

GPT-6 Astra (low) clears both gates, takes its honesty grade, and enters the ranking.

Knowledge honesty+40.5Net correct per 100 questions

Above the +2.58 noise-band boundary. Answers when unsure 47%: grade A.

Benchmark efficacy87.9%Combined share of the best scores

Above the published 85% threshold.

Cost ranking$0.17Per correct answer

Compare with the other qualifying configurations.

Fails honestyGPT-5.6 Luna (low)

−14.7 net correct per 100 questions. Low cost cannot override failed honesty.

Below the efficacy barGPT-5.6 Sol (low)

77.8% efficacy. Honesty passes, but capability falls below the bar.

Honesty inconclusiveGPT-5.6 Sol (non-reasoning)

+1.1 net correct per 100 questions. The margin stays inside the noise band.

Recorded configurations from the index snapshot. The diagram explains a selection rule, not a running router.
How the three measures are defined
First

Is it honest?

Right answers minus confidently wrong ones. Below zero it asserts more than it knows. Below noise it is out. Then, among the configurations that pass, the same knowledge test yields a grade for what a model does when it does not know: of the questions it gets wrong, the share it answers anyway. A answers at most 50%, so it declines at least as often as it answers; B at most 75%; C more. The grade does not disqualify; it orders the ranking and says where each configuration may run.

Second

Is it capable enough?

How close it gets to the best score on each test, averaged across six benchmarks. The six are graduate and research-level science, frontier exams, scientific coding, terminal work and long-document reasoning. We use 85% of the best score across them as the bar to enter the ranking. This is a benchmark threshold, not a guarantee for every workload.

Third

What does it cost?

What one correct answer costs once you have paid for the attempts that failed. The ranking sorts on honesty grade first and this second, so a cheap configuration cannot leapfrog its grade.

What the data shows

GPT-5.6 Sol

Reasoning effort changes the configuration. The lowest settings here fail a gate before they can enter the ranking.

Cheapest qualifying effort: medium$0.1485.5% efficacy
Maximum effort$0.4193.6% efficacy
Cost and efficacy across GPT-5.6 Sol effort settingsHorizontal axis: cost per correct answer in dollars on a log scale. Vertical axis: benchmark efficacy. Highlighted medium: $0.14, 85.5% efficacy. The dashed line is the 85% threshold. The full configuration data is available in the Everyday Model Index.50%70%100%85%Efficacy$0.1$0.25$0.5$1$2Cost per correct answer (USD, log scale)Cost and efficacy across GPT-5.6 Sol effort settingsHorizontal axis: cost per correct answer in dollars on a log scale. Vertical axis: benchmark efficacy. Highlighted medium: $0.14, 85.5% efficacy. The dashed line is the 85% threshold. The full configuration data is available in the Everyday Model Index.50%70%100%85%Efficacy$0.1$0.5$2Cost / solve (USD, log scale)
3.0× the benchmark cost per correct answer at max effort.

Efficacy changes by 8.1 percentage points. The filled point marks the cheapest qualifying setting; the dashed line marks the 85% bar.

Recorded configurations connected in effort order. Cost uses a log scale. The plot does not measure latency or intermediate settings.

Claude Opus 5

The same comparison within another model family. Effort labels are specific to each vendor.

Cheapest qualifying effort: medium$0.3889.3% efficacy
Maximum effort$0.9493.0% efficacy
Cost and efficacy across Claude Opus 5 effort settingsHorizontal axis: cost per correct answer in dollars on a log scale. Vertical axis: benchmark efficacy. Highlighted medium: $0.38, 89.3% efficacy. The dashed line is the 85% threshold. The full configuration data is available in the Everyday Model Index.50%70%100%85%Efficacy$0.1$0.25$0.5$1$2Cost per correct answer (USD, log scale)Cost and efficacy across Claude Opus 5 effort settingsHorizontal axis: cost per correct answer in dollars on a log scale. Vertical axis: benchmark efficacy. Highlighted medium: $0.38, 89.3% efficacy. The dashed line is the 85% threshold. The full configuration data is available in the Everyday Model Index.50%70%100%85%Efficacy$0.1$0.5$2Cost / solve (USD, log scale)
2.5× the benchmark cost per correct answer at max effort.

Efficacy changes by 3.7 percentage points. The filled point marks the cheapest qualifying setting; the dashed line marks the 85% bar.

Recorded configurations connected in effort order. Cost uses a log scale. The plot does not measure latency or intermediate settings.

Claude Fable 5.1

The same comparison within another model family. Effort labels are specific to each vendor.

Cheapest qualifying effort: low$0.4089.7% efficacy
Maximum effort$2.0497.5% efficacy
Cost and efficacy across Claude Fable 5.1 effort settingsHorizontal axis: cost per correct answer in dollars on a log scale. Vertical axis: benchmark efficacy. Highlighted low: $0.40, 89.7% efficacy. The dashed line is the 85% threshold. The full configuration data is available in the Everyday Model Index.50%70%100%85%Efficacy$0.1$0.25$0.5$1$2Cost per correct answer (USD, log scale)Cost and efficacy across Claude Fable 5.1 effort settingsHorizontal axis: cost per correct answer in dollars on a log scale. Vertical axis: benchmark efficacy. Highlighted low: $0.40, 89.7% efficacy. The dashed line is the 85% threshold. The full configuration data is available in the Everyday Model Index.50%70%100%85%Efficacy$0.1$0.5$2Cost / solve (USD, log scale)
5.1× the benchmark cost per correct answer at max effort.

Efficacy changes by 7.9 percentage points. The filled point marks the cheapest qualifying setting; the dashed line marks the 85% bar.

Recorded configurations connected in effort order. Cost uses a log scale. The plot does not measure latency or intermediate settings.

Five findings from the same snapshot. Each includes the comparison and its supporting evidence.

3.3×Use the reasoning lever, but keep the big model.

Claude Sonnet 5 at max effort costs less per token than Claude Fable 5.1, and 3.3 times more per correct answer, because it uses more tokens and solves fewer problems per attempt. Fable 5.1 at low effort is cheaper and 9 points more capable. The sticker price told the opposite story.

Claude Sonnet 5 max$1.3381% efficacyClaude Opus 5 medium$0.3889% efficacyClaude Fable 5.1 low$0.4090% efficacycost per solve, same vendor, all pass the honesty gate
Sonnet 5 max $1.33 per solve at 81% efficacy · Fable 5.1 low $0.40 at 90% · same vendor, both pass the honesty gate
20 of 71The cheapest configurations are among the ones that make things up.

20 configurations are right less often than they are confidently wrong. GPT-5.6 Luna (low) would top the whole ranking at $0.013 per solve if the gate did not exist. It is excluded because every confident error is undone by a person, at a labor rate that swamps the saving. These models look fine on a standard test: their median GPQA Diamond score is 90%. Honesty is not something you can read off accuracy.

noise band−40−20+0+20+40$0.03$0.1$0.3$1every GPT-5.6 Luna settingGPT-5.6 Sol medium, cheapest to qualifybenchmark cost per solve, log scale · vertical: honesty margin per 100 questions
Green passes, amber is within measurement noise of zero, red fails · every GPT-5.6 Luna setting fails, best margin −10.3
The effort setting moves cost more than the model name does.

GPT-5.6 Sol alone spans $0.14 to $0.41 per solve and 86% to 94% efficacy across its effort settings. That is most of the useful range without changing model. The same is true of every frontier family here, which is why a configuration, model plus effort, is the unit ranked and never the model alone.

Sol at low (78%) falls under the bar; at non-reasoning its honesty is too close to call · the setting matters on both gates
5 → 2Capability is spread across labs. Declining when unsure is not.

21 configurations from 5 labs (Anthropic, Google, Meta, Moonshot and OpenAI) are approved, one has open weights among them. Grade A, the ones that decline at least as often as they answer when unsure, comes from 2 of those labs (Meta and OpenAI). Grade C is GPT-5.6 Sol alone. The field is wider than the three names most buying conversations start with on capability; on honesty when unsure, it is narrower, and it is moving fast: GPT-6 Astra, released 3 September 2026, was added on 4 September 2026 and holds 4 of the top 10 places.

1. GPT-6 Astra low$0.17 A · unsure 47%2. Muse Spark 1.3 xhigh$0.21 A · unsure 31%3. GPT-6 Astra medium$0.26 A · unsure 47%4. GPT-6 Astra high$0.36 A · unsure 45%5. GPT-6 Astra xhigh$0.55 A · unsure 48%6. Gemini 3.8 Flash high$0.22 B · unsure 55%7. Claude Opus 5 medium$0.38 B · unsure 61%8. Claude Fable 5.1 low$0.40 B · unsure 66%9. Kimi K3 max$0.50 B · unsure 53%10. Claude Fable 5.1 medium$0.51 B · unsure 69%11. Claude Opus 5 high$0.55 B · unsure 61%12. Claude Fable 5.1 high$0.73 B · unsure 69%13. Claude Opus 5 xhigh$0.73 B · unsure 60%14. GPT-6 Astra max$0.76 B · unsure 51%15. Claude Opus 5 max$0.94 B · unsure 61%16. Claude Fable 5.1 xhigh$1.32 B · unsure 71%17. Claude Fable 5.1 max$2.04 B · unsure 73%18. GPT-5.6 Sol medium$0.14 C · unsure 91%19. GPT-5.6 Sol high$0.19 C · unsure 91%20. GPT-5.6 Sol xhigh$0.26 C · unsure 92%21. GPT-5.6 Sol max$0.41 C · unsure 92%the 21 approved configurations in rank order: honesty grade, then cost per correct answerpurple: grade A · grey: grade B · red: grade C
Open-weight rows are marked in the index; their price is the creator's own rate, which third-party hosts routinely undercut
0.77The honesty gate measured knowledge. The trait it missed is stable, and it now grades the ranking.

The margin of correct answers over confident errors correlates 0.77 with accuracy on the same test. It is a knowledge score, and a model clears it by knowing a lot even if it bluffs every time it does not. Underneath sits the quantity the gate was meant to catch: of the questions a configuration gets wrong, the share it answers anyway. Among the 21 approved configurations it runs from 31% to 92%. It barely moves across effort settings within a family while accuracy does, and it is uncorrelated with capability (-0.02), so it is read as a trained disposition rather than a difficulty artefact. It is not a disqualifier, because its cost depends on whether the task verifies its outputs; it is a grade that orders the ranking and says where each configuration may run. What does not generalise is this test's error rate: the base rate of being wrong is task-specific, which is why the grade is the conditional propensity and not a price on confident errors.

margin vs accuracy r = 0.77 · answers-when-unsure vs capability r = -0.02 · among 21 approved the propensity runs 31 to 92%; 16 answer more than they decline, 4 more than three times in four

When to leave it

Three conditions to evaluate when deciding whether to leave the default route.

1The reasoning is hard.
CritPt, research-level physics, still separates tiers sharply. Two cheap configurations elsewhere on the board score exactly zero on it while still billing for the attempt. Passing GPQA Diamond does not predict passing CritPt, so a high headline score is not a licence to send it your hardest work.
2The model must state facts it cannot look up.
That is what the honesty column measures. Plus 31 means that across 100 questions the model finished 31 correct answers ahead of its confident mistakes. Minus 10.3 means it finished about 10 confident mistakes behind. Every GPT-5.6 Luna configuration sits below zero, so none qualifies at any effort. Put retrieval in front of the model and the gate relaxes, because it is no longer asserting from memory.
3The task will run long.
Fewer than one prompt in ten runs past 15 minutes, but those prompts carry 29% of coding spend and 50% of knowledge-work spend. The index scores bounded work only and does not adjudicate what to escalate to. Route on predicted duration, not on prompt count.

A stated default is a measurement baseline

Why the escalation rule is what makes spend legible, rather than a separate concern from it.

Two different things produce a large bill. A task that legitimately took the escalation path is expensive because the work was hard, and that is the rule working as intended. A task that ran long with no escalation decision behind it is something else: a retry loop, an agent re-reading its own context, tool calls driven by injected instructions, or a route around the default that nobody approved.

Without a declared default and a declared path up, those two are indistinguishable. Every large number looks equally plausible, and the shape of the spend guarantees there will be large numbers: fewer than one prompt in ten runs past fifteen minutes, and those prompts carry 29% of coding spend and 50% of knowledge-work spend.

With a stated default, the same concentration becomes tractable. Spend outside policy is a defined and small population, which is a set you can look at one row at a time.

What this is evidence for

Where a published model-selection standard is already evidence for a control the reader owes. Verified 2026-09-10. The last column is the part we cannot do for you.

NIST AI RMF 1.0 (NIST AI 100-1)

The closest fit of the four. Four of these subcategories ask for very nearly what the index already measures, and are marked as exact fits.

ControlWhat it requiresEvidence the index gives youWhat you still owe
MAP 2.2Information about the AI system’s knowledge limits and how system output may be utilized and overseen by humans is documented. Documentation provides sufficient information to assist relevant AI actors when making informed decisions and taking subsequent actions.The honesty gate measures knowledge limits directly: net correct against confident error, per configuration, plus the hallucination rate underneath it.Documenting oversight for your own workflow.
MAP 3.2Potential costs, including non-monetary costs, which result from expected or realized AI errors or system functionality and trustworthiness - as connected to organizational risk tolerance - are examined and documented.The labor calculator, parameterised to your analyst rate, converts confident errors into the non-monetary cost the subcategory asks for.Your own error-cost and detection-rate data.
MANAGE 3.2Pre-trained models which are used for development are monitored as part of AI system regular monitoring and maintenance.Release-triggered re-evaluation of every pre-trained model in scope, with a typed changelog.Monitoring your deployed configuration, not only the market.
MEASURE 1.1Approaches and metrics for measurement of AI risks enumerated during the Map function are selected for implementation starting with the most significant AI risks. The risks or trustworthiness characteristics that will not – or cannot – be measured are properly documented.The Limitations section and the published sensitivity sweep state what is not measured and how much the answer moves.Your unmeasured risks.
GOVERN 1.2The characteristics of trustworthy AI are integrated into organizational policies, processes, and procedures.Standard v0.1 is the policy object: a default, an escalation path, a disqualified set and a review trigger.Adoption, roles and enforcement.
GOVERN 1.3Processes and procedures are in place to determine the needed level of risk management activities based on the organization's risk tolerance.The 85% efficacy bar and the two-standard-error honesty band are explicit, movable risk-tolerance parameters, published with a sensitivity table.Choosing your own bar.
GOVERN 1.5Ongoing monitoring and periodic review of the risk management process and its outcomes are planned, organizational roles and responsibilities are clearly defined, including determining the frequency of periodic review.A stated review trigger tied to frontier releases, with a changelog entry on every check.Your review cadence and named owner.
GOVERN 6.1Policies and procedures are in place that address AI risks associated with third-party entities, including risks of infringement of a third party’s intellectual property or other rights.Every model in the index is a third party; the two gates are the third-party control.Contracts, IP terms and data terms.
MAP 2.3Scientific integrity and TEVV considerations are identified and documented, including those related to experimental design, data collection and selection (e.g., availability, representativeness, suitability), system trustworthiness, and construct validation.Methodology stated in advance, versioned scripts, a dated snapshot, and a rejected elbow-detection rule reported rather than hidden.Construct validation on your workload.
MAP 4.2Internal risk controls for components of the AI system including third-party AI technologies are identified and documented.The deny-list, with reason codes and per-platform API identifiers where confirmed.The enforcement point itself.
MANAGE 1.3Responses to the AI risks deemed high priority as identified by the Map function, are developed, planned, and documented. Risk response options can include mitigating, transferring, avoiding, or accepting.The deny-list is avoidance; the escalation rule is mitigation. Both are documented.Your accepted risks.
MANAGE 2.4Mechanisms are in place and applied, responsibilities are assigned and understood to supersede, disengage, or deactivate AI systems that demonstrate performance or outcomes inconsistent with intended use.Reason codes give a stated deactivation criterion per configuration.The mechanism that acts on it.

airc.nist.gov/docs/playbook.json, 72 subcategories, retrieved 2026-09-10

ISO/IEC 42001:2023

Annex A holds 38 controls across A.2 to A.10. The requirement column below is our paraphrase: ISO's own wording is copyright and is not reproduced here.

ControlWhat it requiresEvidence the index gives youWhat you still owe
A.6.2.2AI system requirements and specificationParaphrasedAI system requirements are specified before build.The configuration, model plus effort, is the specified unit; the two gates are stated acceptance criteria.Your functional and non-functional requirements.
A.6.2.4AI system verification and validationParaphrasedSystems are verified and validated before deployment.Per-configuration benchmark evidence, reproducible from a dated snapshot and versioned scripts.Validation in your own deployment conditions.
A.6.2.6AI system operation and monitoringParaphrasedOperation and monitoring are planned and performed.The release trigger, the changelog, and the deviation-detection section.Runtime monitoring of your own traffic.
A.9.4Intended use of the AI systemParaphrasedThe system is used only as intended.The bounded-work scope statement and the three escalation triggers define intended use and its edges.Your intended-use record.
A.10.3SuppliersParaphrasedSupplier-provided AI is assessed and managed.Per-vendor evidence, first-party price basis, open-weight and deprecation flags.Supplier security and commercial assessment.
A.5.2AI system impact assessment processParaphrasedA process exists for assessing AI system impacts.The labor arithmetic supplies a quantified error-impact input.The assessment itself.
A.2.2 / A.2.4AI policy / Review of the AI policyParaphrasedAn AI policy exists and is reviewed.Standard v0.1 and its stated review trigger.Approval and review records.

Annex A identifiers and titles cross-checked against two independent published listings, consistent. Clause structure and the Annex A title verified from the standard's own contents page in the ISO/IEC sample PDF.

EU AI Act (Regulation 2024/1689)

Not all of this is in force. Article 4 binds deployers today. The high-risk obligations apply from 2 December 2027 under Annex III and 2 August 2028 under Annex I, and Article 15 binds providers, reaching most deployers only through Article 25.

ControlWhat it requiresEvidence the index gives youWhat you still owe
Art. 4AI literacyProviders and deployersIn force since 2 February 2025Providers and deployers shall take measures to ensure a sufficient level of AI literacy among their staff and other persons operating and using AI systems on their behalf.A published, reproducible method for choosing and disqualifying models, including why a confident error is expensive, is usable directly as staff training material.Records of the measures you took.
Art. 26(1)Obligations of deployers: use per the instructions for useDeployersHigh-risk, from 2 December 2027ParaphrasedDeployers take appropriate technical and organisational measures to use high-risk systems in accordance with the provider's instructions for use.Configuration, not model name, is the approved unit. Instructions for use specify settings, so approving a bare model name cannot demonstrate conformity.Mapping each vendor's instructions onto your approved configuration.
Art. 26(5)Obligations of deployers: monitoring and suspensionDeployersHigh-risk, from 2 December 2027ParaphrasedDeployers monitor operation, inform the provider of suspected risks, and suspend use where needed.Reason codes give a stated suspension criterion; the dated deny-list is the record of what was suspended and why.The monitoring and suspension mechanism.
Art. 26(6)Obligations of deployers: log retentionDeployersHigh-risk, from 2 December 2027ParaphrasedAutomatically generated logs are retained for at least six months.Not covered by the index. Listed so the gap is visible rather than implied away.Log retention. All of it.
Art. 25Responsibilities along the AI value chainA deployer that becomes a providerTracks the high-risk datesParaphrasedA deployer becomes a provider if it puts its name or trademark on a high-risk system, substantially modifies it, or modifies its intended purpose, including by putting a general-purpose model to a high-risk use.Evidence that model selection was made on measured accuracy and honesty rather than sticker price, from a dated public snapshot.Determining whether a trigger applies to you. Building a product on a general-purpose model is squarely in the third one.
Art. 15Accuracy, robustness and cybersecurityProviders of high-risk systemsAnnex III 2 December 2027, Annex I 2 August 2028ParaphrasedHigh-risk systems achieve an appropriate level of accuracy and declare their accuracy metrics in the instructions for use.Per-configuration accuracy on seven public evaluations, from a dated snapshot with a stated costing method.Your own system's accuracy declaration. The index measures models, not your system.

Article texts retrieved 2026-09-10

SOC 2 (AICPA Trust Services Criteria)

The framework most readers already carry, and the one where model configuration lands inside an existing control. The operative word is “configures”.

ControlWhat it requiresEvidence the index gives youWhat you still owe
CC8.1Change managementThe entity authorizes, designs, develops or acquires, configures, documents, tests, approves, and implements changes to infrastructure, data, software, and procedures to meet its objectives.If the configuration is the approved unit, then moving effort from medium to low is a change to software configuration and falls inside CC8.1. The index supplies a documented approval basis and a dated snapshot per change.The change record and the approval itself.

TSC CC8.1, verified 2026-09-10

No control identifier appears here without verification against a named source. This page argues that confident wrongness is the expensive failure mode, so a clause reference we had not checked would undo the argument. Reproduction terms differ by framework: US Government work, freely published. Requirement text quoted verbatim. EU legislation, freely reproducible. Article text quoted or closely summarised. AICPA copyright. One short criterion quoted with attribution; do not reproduce further TSC text. ISO copyright, explicitly not reproducible. Identifiers and titles are cited; control text is PARAPHRASED and never quoted. The freely available sample (14 pages) stops at clause 4.4 and does not contain Annex A, so the wording was never available to us in any case.

How we measured

Public benchmarks, one costing method throughout, and a rule stated in advance rather than fitted to the answer.

Abstract

We rank 71 model configurations from 30 models and 11 labs on the cost of one correct answer, using seven public component evaluations of the Artificial Analysis Intelligence Index and first-party prices, retrieved 1 September 2026, with GPT-6 Astra from pages retrieved 4 September 2026. A configuration enters the ranking only if it clears a knowledge-honesty gate by more than measurement noise and holds at least 85% of the best score across six accuracy evaluations; 21 do. Cost per solve is cost per attempt divided by share solved, combined by geometric mean so no single benchmark dominates. Every approved configuration then carries an honesty grade from a stable, trained trait, the share of unsure questions it answers anyway: A at most 50% (it declines at least as often as it answers), B at most 75%, C above. The trait is nearly constant across effort settings within a family and uncorrelated with capability, while the knowledge margin alone correlates 0.77with accuracy. Ranking is lexicographic, grade then cost, so rank 1 is the default. 21 of 71 configurations are approved: 5 grade A, 12 grade B, 4 grade C. We publish the sensitivity of the qualifying set to the bar and the benchmark set, and of the grades to their boundary, and every figure traces to a dated snapshot and versioned scripts.

1

Scope and inclusion

Every model on the Artificial Analysis leaderboard carrying a measured Intelligence Index score and a score on all seven component evaluations used here. A configuration missing even one benchmark cannot be scored the way the others are. Imputing the gap would mean inventing a number, and averaging over what happens to be present would quietly reward models for the benchmarks they skipped.

The rule is mechanical on purpose. Anyone can run it against the leaderboard and get the same list, which is what makes 'what is missing' an answerable question rather than a matter of who we happened to think of.

  • In practice it is a Terminal-Bench rule. In practice the coverage rule is a Terminal-Bench v2.1 rule. It is the newest of the seven and Artificial Analysis has run it on 229 of 637 records, so it is the single missing evaluation for 395 of the 434 models that cannot be scored here. Whole labs are excluded by it alone, Amazon's Nova line and Microsoft's Phi among them. That is a property of benchmark coverage, not of those models.
  • This is not the Intelligence Index. The Artificial Analysis Intelligence Index weights nine evaluations. The two omitted here, GDPval-AA at 20% and tau3-Banking at 14%, are long-horizon agentic work and out of scope by design. They are also the expensive ones: GDPval alone can be around 59% of a model's published cost per index task. Cost per solve here is therefore a different and much smaller number than the cost per task Artificial Analysis publishes, and the two should never be quoted against each other.
  • Muse Spark 1.3 at max effort. Scored on all seven evaluations and would rank near the top, but Artificial Analysis publishes no price for it at all and lists no serving host. A configuration with no price cannot enter a cost ranking. Its xhigh sibling is priced and does appear.
  • Four non-reasoning variants of Kimi and DeepSeek. Missing Terminal-Bench v2.1 entirely, and their index is flagged estimated rather than measured.
  • Five vendor-deprecated configurations. Superseded by a newer release. This is a buying guide, so a model you cannot adopt going forward is out, though it is named here rather than quietly dropped.
2

Benchmarks

Seven component evaluations of the Artificial Analysis Intelligence Index v4.1.1. Six are scored for accuracy and enter cost per solve; the seventh, AA-Omniscience, is the honesty gate and is never averaged in.

Table 1. Benchmarks in scope, what each measures, and how much it separates the field.

BenchmarkWhat it measures, and why it is in scope
GPQA DiamondGraduate-level science, multiple choice. The floor check. Nearly every current model clears it, so it separates almost nothing.saturated
HLEHumanity's Last Exam: hard closed-ended questions across many fields. Still separates tiers sharply. One of the three that carry real signal.discriminates
CritPtResearch-level physics reasoning. The hardest test in scope. Two cheap configurations score exactly zero while still billing.discriminates
SciCodeScientific coding, graded on sub-problems. Compressed. Ranks three through forty span nine points.saturated
Terminal-Bench v2.1Short agentic coding in a terminal, minutes per task. The one borderline inclusion: agentic, but bounded. Removing it shifts costs about 15% and does not change the ranking.discriminates
AA-LCRReasoning over documents around 100,000 tokens. Compressed. The cheap tier is genuinely competitive here.saturated
AA-OmniscienceKnowledge, with a penalty for confident errors. Not scored for accuracy. It is the honesty gate.the gate
3

Metrics

3.1

Cost per solve

Take the price of one attempt on a benchmark and divide it by the share of problems the configuration got right. That is what a correct answer costs when you can spot a miss and retry. Do this on each of the six accuracy benchmarks, then combine them with a geometric mean so no single hard or expensive benchmark dominates.

cost per solvee = cost per attempte ÷ share solvede
index cost per solve = geometric mean over the six accuracy evaluations e
Worked example
GPT-5.6 Sol at medium solves 92.6% of GPQA Diamond and an attempt costs $0.0157, so one correct answer costs $0.0157 divided by 0.926, which is $0.017. Repeat across the other five benchmarks and take the geometric mean: $0.14 per solve. Claude Sonnet 5 at max costs less per token but solves fewer problems per attempt, so it lands at $1.33.
3.2

Efficacy

On each of the six accuracy benchmarks, divide this configuration's score by the highest score any configuration reached on that benchmark. Average those six shares. 100% means it matched the leader everywhere.

3.3

Honesty and the noise band

On the AA-Omniscience knowledge test a correct answer scores plus one, a confident wrong answer minus one, and declining to answer zero. The total is expressed per 100 questions, so the scale runs from minus 100 to plus 100. Passing means above zero by more than the noise band described next.

AA-Omniscience is 6,000 questions, each scoring plus one, minus one or zero. The net is a mean of bounded scores, so its standard error is at most about 1.3 points on the published per-100 scale. A configuration within two standard errors of zero is reported as too close to call rather than passed or failed, because the test cannot separate it from chance. Six configurations sit in that band, and one of them, GPT-5.6 Terra at max effort, nets plus 0.05, which is four hundredths of a standard error from zero.

3.4

The 85% bar

A stated target, not a fitted one. An elbow-detection rule was tried first and rejected: with four to eight frontier points it is decided by whichever two neighbours happen to tie.

3.5

Honesty grades

The knowledge margin is mostly a knowledge score: across all 71 configurations it correlates 0.77 with accuracy on the same test and only -0.33 with the behaviour the gate was meant to catch. A model clears it comfortably by knowing a great deal and bluffing whenever it does not: Claude Fable 5.1 at max posts the highest margin among approved configurations that answer more than they decline, +43.5, while answering 73% of the questions it gets wrong.

So the grade reads the conditional directly: of the questions a configuration does not get right, the share it answers anyway rather than declining. Three things in the snapshot say this is a trained trait rather than an artefact of difficulty or effort. Within a family it barely moves while accuracy does (GPT-5.6 Sol: 89 to 93% across 6 settings while accuracy runs 49 to 59%). Across all configurations it is uncorrelated with capability on the six accuracy benchmarks (-0.02). And it is slightly positively correlated with accuracy (0.33): more knowledgeable models guess a little more, not less, which is what training that rewards answering predicts.

answers when unsure = confident errors ÷ (100 − accuracy)  ·  A ≤ 50%  ·  B ≤ 75%  ·  C above  ·  rank = (grade, cost per solve)

A trait is a risk class, not a disqualifier; what it costs you depends on whether the task verifies its outputs. A grade A configuration may run unattended, including where it states facts nobody checks. Grade B is approved, and for unverified assertion the standard says prefer A or add verification. Grade C is approved only where outputs are verified before use. 50% is the one boundary with a meaning, declining at least as often as answering; 75% is a chosen value, and Table 3 shows that anywhere from 75% to 90% isolates the same family as grade C. What generalises is the conditional propensity, not this test's error rate: the base rate of being wrong is task-specific, so confident errors per 100 here are not carried over to other work.

4

Results

Of 71 configurations, 21 clear the knowledge-honesty gate beyond noise and the 85% bar and are approved, from 5 labs (Anthropic, Google, Meta, Moonshot and OpenAI). By honesty grade: 5 grade A (Meta, OpenAI), 12 grade B (Anthropic, Google, Moonshot, OpenAI) and 4 grade C (GPT-5.6 Sol). GPT-6 Astra at low ranks first, grade A at $0.17 per correct answer with 88% efficacy, followed by Muse Spark 1.3 at xhigh at $0.21 and GPT-6 Astra at medium at $0.26. The approved set spans a 15-fold cost range, from $0.14 to $2.04, for 12 points of efficacy. The cheapest approved configuration by cost alone, GPT-5.6 Sol at medium at $0.14, answers 91% of the time it is unsure, takes grade C and ranks 18th (Figure 1).

Of the rest, 20 pass the honesty gate but fall under the 85% bar, 6 sit inside the noise band on honesty and are not ranked, 20 are confidently wrong more often than right, and 4 score zero on at least one benchmark, which leaves cost per solve undefined. The cheapest configuration measured, GPT-5.6 Luna at low at $0.013 per solve, is in the failing tier; without the gate it would rank first.

Figure 1. The 21 approved configurations in rank order, honesty grade then cost per correct answer; bars are cost, colour is grade.

1. GPT-6 Astra low$0.17 A · unsure 47%2. Muse Spark 1.3 xhigh$0.21 A · unsure 31%3. GPT-6 Astra medium$0.26 A · unsure 47%4. GPT-6 Astra high$0.36 A · unsure 45%5. GPT-6 Astra xhigh$0.55 A · unsure 48%6. Gemini 3.8 Flash high$0.22 B · unsure 55%7. Claude Opus 5 medium$0.38 B · unsure 61%8. Claude Fable 5.1 low$0.40 B · unsure 66%9. Kimi K3 max$0.50 B · unsure 53%10. Claude Fable 5.1 medium$0.51 B · unsure 69%11. Claude Opus 5 high$0.55 B · unsure 61%12. Claude Fable 5.1 high$0.73 B · unsure 69%13. Claude Opus 5 xhigh$0.73 B · unsure 60%14. GPT-6 Astra max$0.76 B · unsure 51%15. Claude Opus 5 max$0.94 B · unsure 61%16. Claude Fable 5.1 xhigh$1.32 B · unsure 71%17. Claude Fable 5.1 max$2.04 B · unsure 73%18. GPT-5.6 Sol medium$0.14 C · unsure 91%19. GPT-5.6 Sol high$0.19 C · unsure 91%20. GPT-5.6 Sol xhigh$0.26 C · unsure 92%21. GPT-5.6 Sol max$0.41 C · unsure 92%
5

Sensitivity

A stricter bar narrows the field.

The honesty gate and six benchmarks stay fixed. Only the efficacy threshold changes.

Qualifying configurations decline from 37 at 75% efficacy to 14 at 90%. The published 85% bar admits 21 configurations.02040Qualifying configurations37332721191475%85%90%Efficacy threshold
Published rule / 85%21

qualifying configurations
from 5 labs

GPT-5.6 Sol (medium)

Cheapest approved by cost alone, $0.14 per correct answer, before the honesty grade orders the ranking.

The six points reproduce the published sensitivity sweep. The full threshold values and labs appear in Table 2 below.

The grade boundary orders the ranking, so it is swept too. The A boundary is fixed at 50%, the point at which a configuration declines at least as often as it answers; only the B/C boundary moves.

Table 3. Grade counts as the B/C boundary moves, A fixed at 50%. Grade C models named.
Grade B up toABCGrade C models
60%5412Claude Fable 5.1, Claude Opus 5, GPT-5.6 Sol
66%588Claude Fable 5.1, GPT-5.6 Sol
75%published5124GPT-5.6 Sol
80%5124GPT-5.6 Sol
90%5124GPT-5.6 Sol

Two judgement calls set who appears: the 85% bar and the benchmark set. Both move the answer, so both are published rather than footnoted.

Table 2. Configurations and labs that qualify as the efficacy bar moves; the honesty gate is held fixed.

Efficacy barQualifyLabs represented
90.0%14Anthropic, Meta, OpenAI
87.5%19Anthropic, Meta, Moonshot, OpenAI
85.0% (used here)21Anthropic, Google, Meta, Moonshot, OpenAI
82.5%27Anthropic, Google, Meta, Moonshot, OpenAI, Z.ai, xAI
80.0%33Alibaba, Anthropic, Google, Meta, Moonshot, OpenAI, Z.ai, xAI
75.0%37Alibaba, Anthropic, Google, Meta, Moonshot, OpenAI, Z.ai, xAI

One benchmark carries most of that. Remove CritPt and 13 more configurations clear the same bar, from 8 labs instead of 5. Research-level physics is where these models separate, and whether it belongs in your definition of everyday work is a judgement you should make rather than inherit.

6

Limitations

The index informs model selection; it does not evaluate an end-to-end router. Latency, tool access, context limits, retry behavior, and escalation policy still need to be tested on the workload being routed. Cost per solve is a benchmark estimate, not a measured production bill.

  • Data collected from benchmarks as of 1 September 2026; ranking published 13 September 2026. GPT-6 Astra was released 3 September 2026 and added from pages retrieved 4 September 2026. There are new releases and re-grades between index patches, so figures move. Zaun Research will update this index regularly.
  • The retry assumption. Cost per solve assumes a failed attempt can be detected and retried. That holds on graded, bounded work, which is the scope here. It does not hold on long agentic runs, which is one reason those are out of scope.
  • The grade boundaries. 50% has a meaning (declines at least as often as it answers); 75% is a chosen value, published with its sweep (Table 3). The grade is measured on one knowledge test; its stability across effort settings and its independence from capability are within-test evidence, and abstention cannot be measured on the six accuracy benchmarks because they force an answer. Tasks with verification turn a confident error into a detected one, which is why the grade sets where a configuration may run rather than whether it is approved.
  • Per-token price is not in the ranking. It does not predict cost per solve. One model here lists at 40% of another per token and costs more per task, because it uses more tokens and more turns.
  • Effort settings are not comparable across models. One vendor’s medium is not another’s. Read a configuration as a whole, never the effort label on its own. GPT-5.6 Sol at medium and Claude Opus 5 at medium list within 20% of each other per output token ($20 against $25 per million), yet one CritPt attempt costs $0.13 on the first and $1.40 on the second. The gap is tokens spent per task, not price per token.
  • Open-weight rows are priced at the creator’s rate, which is the expensive end. Marked open. Third-party hosts frequently undercut it, in one case by three to four times, so those configurations are likely cheaper in practice than shown here. The two NVIDIA rows are different again: NVIDIA offers no first-party API, so they carry a cross-provider median and are marked median price.
  • Platform premiums are not applied. Prices are first-party. Buying the same model through a cloud marketplace can add materially to it.
  • Long-running work is out of scope. Every benchmark here is bounded. The escalation triggers above point at the tail; this index does not rank models for it.
7

Data and reproducibility

Every figure on this page is computed from a dated snapshot of public sources, and the scripts that turn the snapshot into the ranking are versioned with it. Per-evaluation scores and cost per attempt were read from the Artificial Analysis evaluation pages [3-9] and checked against the published cost per task on each; index composition and weights from the Artificial Analysis methodology [2]; model-level cost per index task and per-token prices from the Artificial Analysis model pages [1]. First-party per-token prices for Anthropic and OpenAI were confirmed against the vendors’ own pricing pages [10, 11]; the cloud-platform multipliers discussed in the limitations come from the Bedrock, Azure and Vertex price lists [12-14]. Every other vendor’s first-party price is as recorded by Artificial Analysis on that model’s page [1]. The task-length figures are Zaun internal telemetry [15], August 2026, one prompt measured from submission to its last completion event, published in aggregate only. This is version 0.1 of the index, a preliminary snapshot; later versions will cite the snapshot date they supersede.

References

  1. Artificial Analysis, model pages and Intelligence Index leaderboard, v4.1.1. artificialanalysis.ai/models. Retrieved 1 September 2026; GPT-6 Astra pages 4 September 2026.
  2. Artificial Analysis, Intelligence benchmarking methodology. artificialanalysis.ai/methodology/intelligence-benchmarking.
  3. Rein, D. et al., GPQA: A Graduate-Level Google-Proof Q&A Benchmark, 2023, arXiv:2311.12022. Scored by Artificial Analysis, artificialanalysis.ai/evaluations/gpqa-diamond.
  4. Phan, L. et al., Humanity’s Last Exam, 2025, arXiv:2501.14249. Scored by Artificial Analysis, artificialanalysis.ai/evaluations/humanitys-last-exam.
  5. Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark, 2025, arXiv:2509.26574. Scored by Artificial Analysis, artificialanalysis.ai/evaluations/critpt.
  6. Tian, M. et al., SciCode: A Research Coding Benchmark Curated by Scientists, 2024, arXiv:2407.13168. Scored by Artificial Analysis, artificialanalysis.ai/evaluations/scicode.
  7. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces, 2026, arXiv:2601.11868; tbench.ai, artificialanalysis.ai/evaluations/terminalbench-v2-1. Version 2.1 scored by Artificial Analysis.
  8. Artificial Analysis, Long Context Reasoning (AA-LCR). artificialanalysis.ai/evaluations/artificial-analysis-long-context-reasoning.
  9. Artificial Analysis, AA-Omniscience: knowledge with a penalty for confident errors. artificialanalysis.ai/evaluations/omniscience.
  10. Anthropic, Claude pricing. platform.claude.com/docs/en/about-claude/pricing. Retrieved 1 September 2026.
  11. OpenAI, API pricing. developers.openai.com/api/docs/pricing. Retrieved 1 September 2026.
  12. Amazon Web Services, Amazon Bedrock pricing. aws.amazon.com/bedrock/pricing. Retrieved 1 September 2026.
  13. Microsoft, Azure Retail Prices API, Foundry meters, eastus2. prices.azure.com/api/retail/prices. Retrieved 1 September 2026.
  14. Google Cloud, Vertex AI generative AI pricing. cloud.google.com/vertex-ai/generative-ai/pricing. Retrieved 1 September 2026.
  15. Zaun Research, task-length distribution per prompt, coding agents and knowledge work, August 2026. Internal telemetry, aggregate figures only.