Skip to content

The Everyday
Model Index

A published standard for the models your agents run by default.

Security teams are now asked to approve AI model choices, and the spend behind them, without a method for doing it. This index is one. 71 configurations, each a model at one reasoning-effort setting, pass or fail a knowledge-honesty gate and a capability bar. Those that pass are graded by what they do when they do not know, from A (declines at least as often as it answers) to C (bluffs), and ranked by grade, then by what a correct answer costs. 21 are approved, 26 held for review and 24 denied. The result is a dated standard you can adopt, cite and enforce.

of 71 configurations from 30 modelsSnapshot 13 September 2026 · Latest source addition 4 September 2026

Find your default.

A configuration is a model at one reasoning-effort setting. Approved ones are ranked by honesty grade, then by what a correct answer costs.

How the index is calculated
Lab
Status

21 of 71 configurations · approved

Sort by

ApprovedClears the knowledge-honesty gate beyond measurement noise and holds 85% of the best capability score. Carries an honesty grade and a rank: grade first, then cost.

Configurations from the dated index snapshot. Approved configurations clear the knowledge-honesty gate beyond noise and hold 85% of the best capability score; they carry an honesty grade (A declines at least as often as it answers when unsure, B answers up to three times in four, C more) and are ranked by grade, then benchmark cost per correct answer.
RankModel & effortStatus
1$0.17A47%Approved
2$0.21A31%Approved
3$0.26A47%Approved
4$0.36A45%Approved
5$0.55A48%Approved
6$0.22B55%Approved
7$0.38B61%Approved
8$0.40B66%Approved
9$0.50B53%Approved
10$0.51B69%Approved
11$0.55B61%Approved
12$0.73B69%Approved
13$0.73B60%Approved
14$0.76B51%Approved
15$0.94B61%Approved
16$1.32B71%Approved
17$2.04B73%Approved
18$0.14C91%Approved
19$0.19C91%Approved
20$0.26C92%Approved
21$0.41C92%Approved

Rank is by honesty grade first, then benchmark cost per correct answer, among approved configurations, and stays fixed when you filter or sort. Grade A answers at most 50% of the questions it gets wrong, B at most 75%, C more. Prices are from the dated snapshot; each configuration records its price basis and retrieval date.

Adopt

Default model configuration standard

v0.4Snapshot 13 September 2026 · supersedes 11 September 2026

  1. Default

    Bounded work routes to rank 1: the approved configuration with the best honesty grade, and within that grade the lowest benchmark cost per correct answer. Approval means it knows more than it bluffs beyond measurement noise and holds 85% of the best capability score. The grade is how it behaves when it does not know: A declines at least as often as it answers (at most 50% answered when unsure), B answers up to three times in four (at most 75%), C more.

    GPT-6 Astra low · grade A · $0.17 per correct answer · answers when unsure 47% · +40.5 margin · 88% capability

    Cheapest by cost alone is GPT-5.6 Sol (medium) at $0.14, grade C: it answers 91% of the time it is unsure, so it ranks 18th and is approved only where outputs are verified before use. How the grade boundary moves the ranking

  2. Where each grade may run

    Grade A (5): When it does not know, it declines at least as often as it answers. Approved for all bounded work, including stating facts nobody checks. Grade B (12): When it does not know, it answers up to three times in four. Approved; for unverified factual assertion prefer grade A or add verification. Grade C (4, GPT-5.6 Sol): When it does not know, it answers more than three times in four. Approved only where outputs are verified before use.

  3. Escalation

    A task leaves the default when the reasoning is research-grade, when the model must state facts it cannot look up, or when predicted duration runs past fifteen minutes. It goes to the next configuration in rank order at least 5 capability points above the default.

    Claude Fable 5.1 high · grade B · $0.73 per correct answer · 94% capability · for unverified factual assertion, prefer a grade A configuration and accept the capability gap

  4. Disqualified

    24 configurations are denied on a measured finding. A further 26 need review against your own bar or returned an inconclusive honesty result. The two are not the same claim and are published separately.

    Deny-list, machine readable

  5. Approved unit

    Model plus reasoning effort. A model name on its own is not an approval, and effort labels are not comparable between vendors.

  6. Review

    The index is re-run when a model appears on the Artificial Analysis leaderboard with a measured Intelligence Index score and scores on all seven component evaluations used here. Sources are checked weekly. Most weeks there is nothing to add, and that is recorded. A change of default obligates re-approval.

    Changelog

Exceptions. Deviations from the default, the grade rules and the escalation rule are recorded with a stated reason and an owner.

Limits. Benchmark evidence, not a production measurement. The grade is measured on one knowledge test; the 50% boundary has a meaning, the 75% boundary is a chosen value with a published sweep. Latency, tool access, context limits and retry behaviour still need testing on the workload being routed.

Cite it

Dated snapshots are immutable. A correction publishes a new snapshot rather than rewriting a published one, so a citation keeps resolving to what it cited.

Configuration ids such as openai/gpt-6-astra@low are permanent. Where no vendor API identifier is confirmed, the row says unmapped rather than guessing. Which controls this is evidence for

Version history

What changed, and whether it obliges you.

The index is re-run when a model appears on the Artificial Analysis leaderboard with a measured Intelligence Index score and scores on all seven component evaluations used here. Sources are checked weekly. Most weeks there is nothing to add, and that is recorded.

  1. 13 September 2026

    METHODOLOGY

    Version 0.4. Honesty is graded, not gated. Approval is the knowledge-honesty gate beyond noise and the 85% capability bar, as in v0.1. Each approved configuration carries a grade from the share of unsure questions it answers anyway: A at most 50% (declines at least as often as it answers), B at most 75%, C above. Ranking is grade first, then benchmark cost, so rank 1 is the default and a cheap row cannot leapfrog its grade. The 75 boundary is a chosen value; the sweep is published.

  2. 13 September 2026

    DEFAULT CHANGED

    21 approved across 5 labs: 5 grade A, 12 grade B, 4 grade C. The default is GPT-6 Astra (low), rank 1, $0.17. Every GPT-5.6 Sol setting is grade C, approved only where outputs are verified before use; Sol (medium), cheapest by cost, ranks 18th.

  3. 13 September 2026

    METHODOLOGY

    Version 0.3 is superseded. Two hard gates on the trait approved four configurations from one lab, too few to be useful, and the +30 knowledge bar removed only the most honest configuration on the board. The 11 September files stay published.

  4. 11 September 2026

    METHODOLOGY

    Version 0.3. Honesty became two gates: a +30 knowledge margin and answering when unsure at most 50%. Superseded two days later.

  5. 11 September 2026

    METHODOLOGY

    Version 0.2 is withdrawn. It priced review time by charging this knowledge test's error rate against every everyday answer, overstating the term by roughly an order of magnitude.

  6. 10 September 2026

    METHODOLOGY

    Published the AA-Omniscience components behind the honesty gate: accuracy, the share of unsure answers given anyway, and confident errors per 100.

  7. 10 September 2026

    QUALIFYING SET

    The deny-list separates Deny, a measured integrity finding, from Review, which is our bar rather than a failure.

  8. 4 September 2026

    NEW MODEL

    GPT-6 Astra, released 3 September 2026, admitted from pages retrieved 4 September 2026.

  9. 1 September 2026

    NEW MODEL

    Version 0.1. First published snapshot: 71 configurations, 30 models, 21 approved, ranked by benchmark cost per correct answer.