Harbor-Index 2026 Explained: 82 Agentic Tasks, None Above 30%
Not another composite rank. Harbor first joins 80-plus agent suites, then keeps 82 tasks that still fail the frontier.
What landed on arXiv on September 3, 2026 was not a flagship weight file. It was a shared pipe. Harbor Adapters join 80-plus agent suites; Harbor-Index keeps 82 tasks that still fail the frontier.
The searchable event is Harbor-Index and arXiv:2609.04298. Lin Shi, Haowei Lin, and colleagues report eight models, 54 benchmarks, Terminus-2 plus one of three native harnesses. No pair clears 30%. GPT-5.5 with Codex leads at 28.0%. CCTest repeated the abstract.
The pipe and the item set are not the same object
Harbor Adapters
Heterogeneous suites become one format so any agent can share inference and grading. Parity experiments are the claim that a Harbor run still behaves like the source exam.
Harbor-Index
Eighty-two tasks from 29 suites. Breadth stays; routine re-runs get cheaper. tbench.ai says the seed campaign spent 226 billion tokens and more than $300,000; the distilled suite is a few hundred dollars.
| Funnel stage | Kept | Rule |
|---|---|---|
| Candidate pool | 6,627 | 54 adapters already wired |
| Difficulty filter | 1,311 | Frontier mix ≤33% of trials |
| Automated audit | 307 | Drop broken instruction/verifier pairs |
| Human + repair | 82 | 14 reviewers; re-run; drop the unfixable |
What the paper actually handed over
- Adapter layer
- More than 80 suites on Harbor. Parity work is there to show that a common format did not flatten the original tasks.
- Paired evaluation
- Eight models, 54 suites; Terminus-2 and one of three native harnesses each. Failures split into model, harness, environment, and tools.
- Compact meta-set
- 82 tasks, 29 sources. Difficulty filter, AI and human audit, repair loop. A ceiling under 30% is the design, not an accident.
Three ways to read the 82
- 01Read the harness before the model
Native CLIs often sit a little higher, without significance. Weaker models may keep about 7% of solves across a swap. A single composite screenshot drops the runtime.
- 02Split broken tasks from hard ones
Mismatched instructions and verifiers, vague success, missing dependencies all print a zero. The funnel exists so remaining fails look more like a capability cap.
- 03Treat it as a re-run kit, not an IQ total
CCTest notes that 82 tasks cannot stand in for every workflow. Adapters, results, and analysis are open so others can audit and extend — not freeze a forever table.
Share the pipe first. Keep the tasks that still fail. Rank is a side effect.
The ruler is left low on purpose
No tested pair clears 30%. That is the filter: frontier systems should still fail most of the time, or the set saturates.
If a group must pick an agent stack next week, the funnel and pass rates on one page beat a forwarded rank image. Open a mic only when talk helps. A short oakmeet room is enough; no client for an eval huddle. Limits are in eight people, about an hour.
Is Harbor-Index another composite leaderboard?
Not exactly. The paper ships adapters and a paired campaign, then an 82-task kit. tbench.ai treats it as a cheap high-signal suite, not a replacement for all 54 benchmarks.
Which number, 28.0% or 28.1%?
The abstract says GPT-5.5 with Codex is 28.0%, and no pair tops 30%. The tbench.ai page says 28.1%. Quote the source. Do not fuse them into one official score.
Why adapters first?
Agent suites need heavy environments and disagree on interfaces. Each new agent otherwise means a new integration. Harbor translates the mess into one format and checks the port with parity runs.
Is this the same ruler as the Intelligence Index?
No. Artificial Analysis publishes a weighted composite. Harbor-Index is an open 82-task agentic meta-set plus an adapter layer. Read them side by side; do not convert one into the other.