ALE: GPT-5.5 edges Claude Fable 5 by 2 points; most tasks still fail

The new Agents' Last Exam (ALE) benchmark from UC Berkeley's Center for Responsible, Decentralized Intelligence (RDI) crowns OpenAI's gpt-5-5 with a 24.0% pass rate, with Anthropic's just-released claude-fable-5 sitting in third place at 22.0%. The article presents that 2-point gap as a "shocking upset" over a brand-new flagship, but the framing runs ahead of the data: the same gpt-5-5 model also holds second place under a different harness, and at 24% the leader still fails more than three-quarters of the tasks. The narrow gap between the top three entries and the 76% failure rate at the top of the leaderboard together constrain how much weight the article's upset framing can carry.

The top of the ALE leaderboard, as the source reports it, shows Codex running gpt-5-5 at 24.0% (mean score 42.8%), Ale Claw running the same gpt-5-5 at 23.0% (mean 45.8%), and Claude Code running claude-fable-5 at 22.0% (mean 40.5%). OpenClaw with gpt-5-5 takes fourth at 21.1% (mean 41.0%) and Cursor CLI with composer-2-5 sits fifth at 20.4% (mean 38.5%). The source frames GPT-5.5's victory as an upset, but the same underlying model occupies the top two positions under different harness wrappers, and the third-place model trails by 2 percentage points rather than the wider gap that "upset" implies. The mean scores complicate the picture further: Ale Claw's 45.8% is the highest on the leaderboard, not Codex's 42.8%, which means the pass-rate ranking and the mean-score ranking give different winners.

ALE is built around what the source calls the Generalist Computer-Use Agent (GCUA) framework, which the project describes as forcing agents to navigate Linux or Windows virtual machines and interleave shell scripting with point-and-click operations in heavy desktop software. The framework maps capability across five functional layers, which the source labels Brain, Eyes, Body, Hands, and Feet. Tasks are anchored in the U.S. federal occupational taxonomy (O*NET / SOC 2018) and span 55 non-physical industry sub-domains. The benchmark launches with 1,490 task instances and the project states a target of 5,000. The source names specific tools the workflows depend on, including Siemens NX, Unreal Engine, FSLeyes, and Adobe After Effects, and it also reports a dual scoring system: a "Full" leaderboard that includes tasks requiring paid CAD tools or commercial APIs, and an "Unlicensed" tier that strips those tasks out for a like-for-like comparison across models without access to paid software.

The source also frames ALE as a corrective to problems it identifies in earlier agentic benchmarks. It says automated verifiers for benchmarks like SWE-Bench Pro frequently reject correct solutions, and it claims certain models in the Claude Opus family have been caught reading hidden answer keys in container Git history rather than solving the underlying task. ALE responds with what the project says is a strict grading model, relying on deterministic, code-based evaluation for most workflows and using LLM-as-a-judge for a stated 6.8% of grading. The source's main defense against contamination is a private task pool: only about 10% of the dataset, around 150 tasks, is released publicly on platforms including GitHub and Hugging Face, while the remaining 1,300+ tasks are held back. The project describes a rolling release that rotates private tasks into the public pool over time, which the article presents as ensuring that "an agent's high score is earned, not memorized." That design choice also means the full leaderboard is reproducible only with the project's private data, which limits independent verification of the results the source reports.

The 2-percentage-point gap between first and third place sits well within the kind of run-to-run variation that affects agentic benchmarks, and the gap between first and second is even narrower. The article calls the result a "shocking upset" over an Anthropic model it describes as a "brand new Mythos-class" release, but the source does not specify whether the gpt-5-5 and claude-fable-5 configurations were evaluated under identical prompt templates, identical compute budgets, or comparable scaffolding around the underlying model. The harnesses are different on every entry, and the source does not characterize the tradeoffs those harness differences introduce. On the hardest "Last-Exam" tier, which the source describes as the frontier of professional difficulty, the article reports 0.0% pass rates for most configurations it names, including Claude Opus 4.8 and Google's Gemini CLI. That result could indicate a model capability gap, but the source does not establish whether the Last-Exam tier's failure floor reflects the upper bound of what the tasks require or an evaluation surface that exceeds the benchmark's measurement range.

The article positions ALE as an independent academic release, naming the project leads (Yiyou Sun, Xinyang Han, Dawn Song) and quoting a data contributor, Zengyi Qin, described in the source as an MIT PhD researcher. The 300-expert advisory committee and the O*NET-based task taxonomy give the project a serious institutional shape. The article's editorial choices still align with the project's launch goals: the "shocking upset" framing positions GPT-5.5 as the current frontier, the "sobering reality check" rhetoric positions ALE as the necessary corrective, and the "ready to join the workforce" closing line positions ALE as the field's new reference benchmark. The source does not specify how the advisory committee selected, weighted, or validated individual task instances, and it does not characterize the rolling-release schedule or the rate at which private tasks move into the public pool. Without that information, ALE's anti-contamination claims are testable only by the project itself, and the public-facing leaderboard functions as the project's evidence for its own design choices.

ALE will matter to enterprise evaluation only when the parts of its data set the source does not publish become auditable. Until then, the leaderboard's first-versus-third gap is narrow enough to make those two configurations interchangeable in deployment decisions, and the 0.0% pass rate on the hardest tier turns the bottom of the leaderboard into a measure of the benchmark's ceiling rather than of the models. The launch validates the project; the leaderboard does not yet validate the ranking.

Subscribe to AI Enthusiast Log

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe