Artificial Analysis Index v4.2 Explained: Private Tests Hit 40%

The ruler moved first. v4.2 doubles held-out weight and swaps in harder knowledge-work tests.

On September 4, 2026, Artificial Analysis did not just reshuffle names. It changed the ruler. Intelligence Index v4.2 doubles the weight of private tests.

The searchable event is Intelligence Index v4.2: held-out sets rise to 40 percent of the composite, a saturated public science exam leaves, and two harder knowledge-work suites enter. OfficeChai and AI Weekly repeated the same official changelog on September 5.

40%held-out / private weight
+2Briefcase and GDP.pdf
0GPQA Diamond left on the Index

What the changelog actually moved

Added

AA-Briefcase scores multi-week agentic knowledge work on a private set. GDP.pdf, from Surge AI, uses 100 PDFs, 4,592 pages, and 1,275 atomic criteria. All-pass means every criterion.

Removed

GPQA Diamond is marked saturated. The firm says it no longer splits the frontier. Keeping it as slide one in a vendor pack no longer explains an autumn 2026 gap.

ModelLabv4.2 Index
Claude Fable 5.1Anthropic57
GPT-6 AstraOpenAI55
Claude Opus 5Anthropic54
Claude Fable 5 / Muse Spark 1.3Anthropic / Meta53

What the new tasks reward

AA-Briefcase
Industry-built multi-week projects with thousands of source files. Rubrics check whether the work landed; pairwise grades cover analysis and presentation. The firm says Fable 5.1 and Opus 5 lead, then Astra and Muse Spark 1.3. Astra is about 85 Elo above GPT-5.6 Sol.
GDP.pdf
Single-turn professional document reasoning. Evidence sits in prose, tables, charts, footnotes, and exclusions. Astra 33.2%, Sol 28.2%, Fable 5.1 26.2%. The overall leader is not the document-slice leader.

Three ways to read the new number

  1. 01
    Read the weight first

    Forty percent now sits on held-out sets, including Briefcase, Omniscience, and CritPt solutions. The stated aim is less public-set gaming. v5 will raise that share again.

  2. 02
    Do not compare absolute scores across versions

    A v4.1 figure and a v4.2 figure are not the same exam. OfficeChai noted Astra barely moved on the old ruler; the new ruler shipped the next day.

  3. 03
    Split the slices from the headline

    Anthropic and OpenAI lead the composite. The cost-per-task frontier also includes Meta and Z.AI. Document work favors OpenAI. A single rank screenshot drops the slices.

Compare two elections after the districts changed and the totals still exist. The meaning does not.

The ruler is still moving

On September 7 the firm posted a v4.3 note: Terminal-Bench upgraded, AutomationBench-AA added with a private set. That is a rolling path to v5, not a frozen table.

If a group must pick an API to try next week, spreading the changelog on one page beats forwarding a rank image. Open a mic only when talk helps. A short oakmeet room is enough; no client install for an eval huddle. Limits are in eight people, about an hour.

Can a v4.2 score be compared with v4.1?

No. Weights, tasks, and part of the grading changed. You can say who sits higher on the new ruler. You cannot say a lab got N points smarter than last month.

Is AA-Briefcase a public download?

No. The firm describes it as a private held-out set, built to cut benchmaxxing. You can check the method write-up, not the item PDF.

Why drop GPQA Diamond?

Artificial Analysis says it saturated: frontier models crowd the ceiling, so gaps no longer mean capability gaps. The Index swapped the exam instead of citing a familiar, spent number.

Is v4.3 a different index?

No. The September 7 note is an increment on the same track: Terminal-Bench moves up, AutomationBench-AA joins. Treat it as a footnote that the ruler is not nailed down.

Start a room