NeoCognition Lab in Palo Alto, known for the MMMU, Mind2Web, SWE-Bench Pro, and OSWorld 2 benchmarks, released ApprenticeBench — the first end-to-end test that hires AI for a real job: an agent sequentially completes 100 accounts payable clerk tasks in the ERP Odoo of a virtual construction company, Acme Home Builders, receiving a six-month invoice archive and a company directory. Claude Fable 5.1 solved 72% of tasks, GPT-6 Astra — 68%, while the best of two human testers managed 51%, and other models, including Gemini, Grok, Kimi, and Muse Spark, remained below 25%. The leaderboard with pricing is already open at apprenticebench.com.



What Happened
The NeoCognition team — creators of the MMMU, Mind2Web, SWE-Bench Pro, OSWorld 2 benchmarks and agent systems SeeAct, UGround, LLM-Planner, and HippoRAG — published ApprenticeBench, the first test combining computer use, continuous learning, and long agent runs in a real job. The agent works as an accounts payable clerk at the construction company Acme Home Builders within ERP Odoo: it receives a six-month invoice archive, from November 2025 to April 2026, a company directory, and a manual, and feedback is structured like onboarding a live employee — in the first month, every processed invoice is reviewed in detail, and later only rare comments remain at the end of the month. The realism of the scenarios and the solvability of each of the 100 tasks were independently confirmed by two professional accountants with an AP profile in the construction industry and a combined experience of over 30 years. Run results: Claude Fable 5.1 from Anthropic solved 72% of tasks, GPT-6 Astra from OpenAI — 68%, the best of two human testers — 51%, Gemini 3.8 Flash and Qwen3.8 Max — 24% each, Muse Spark 1.3 — 19%, Kimi K3 — 18%.
Context
Typical agent benchmarks measure individual skills: clicking a UI button, writing code, answering a question. ApprenticeBench for the first time tests the full cycle of 'employment' — learning on noisy data from a specific company, working in a GUI without backend APIs, and adapting to changing policies over a distance of 100 sequential tasks. Methodologically, the most valuable part is the 'smart novice' ablation: Fable 5 without an archive and feedback solves only 11 out of 100 tasks, with an archive — 31, with feedback — 35, with both channels at once — 43; thus, the measurement separates the gain from accumulated context from the model's basic capabilities. The authors link the record gap between closed and open models precisely to learning on a long run, not to the ability to use an interface: Kimi K3 barely referred to the archive and its own notes and degraded. The background is also important: previously, the transition of agent products from API to GUI noticeably reduced quality — this was called the 'computer use tax,' and its disappearance in the latest flagships became one of the main signals of the test.
Why This Matters for the Industry
For the industry, this is the first engineering-ready benchmark for GUI agents over the long distance, and it records a strict stratification of the market: the 'new employee' threshold was overcome only by two closed flagships released on September 1 and 3, while open models failed the onboarding stage on company data. Agent platform vendors now have a measurable metric that will likely migrate to marketing and corporate procurement requirements, and buyers have a pricing base: Fable 5.1's GUI mode costs $18.23 per task versus $6.95 via API and $7.21 for a human. There are three reproducible engineering steps: embed onboarding in an agent product according to the ablation recipe (company archive plus a feedback channel), evaluate GUI agents on real office tasks in Odoo-like systems based on the pair Fable 5.1 and GPT-6 Astra, and check unit economics against the open leaderboard. Kimi K3's failure additionally indicates where to focus development: agent memory management instead of growing notes and targeted model updates for long runs.
Why This Matters for Users
The leaderboard with pricing and links is open at apprenticebench.com, and independent snapshots are already being published by BenchLM — the latest dated September 15, 2026, so you can personally see how top models pass the same 100 invoices, where humans break down (fatigue, searching the archive), and where Kimi K3 degrades. The practical shift for the reader: the 'computer tax' has disappeared in the two latest flagships, and agents can now be evaluated on real office tasks in GUIs like Odoo, not just on synthetic clicks. In terms of money, however, API mode is still two to three times cheaper than GUI mode, and open models (Kimi K3 — 18%, Qwen3.8 Max — 24%) are not ready for such work: their main failure is precisely in learning on the company archive. Those who want to test agents on their own processes should reasonably launch a pilot on Fable 5.1 or GPT-6 Astra in an Odoo-like system, maintaining selective human control.
What Is Still Unknown / Limitations
The human comparison base is weak: 51% is the result of the best of only two testers, such a sample does not provide a population estimate, and the gap with 72% may be partially explained by individual differences, fatigue, and archive searching, not just model superiority. Cost figures are vendor data, with no independent verification yet. The benchmark was recently published: there have been no independent replications or methodological discussions yet, and the transfer of the format to other jobs has not been tested. Finally, even 72% of the best model means 28% errors — in accounts payable, agents cannot be launched without selective human verification.
Sources
- ApprenticeBench — NeoCognition Blog (primary source)
- ApprenticeBench — official benchmark site with leaderboard
- ApprenticeBench Leaderboard & Scores — BenchLM.ai (independent mirror, snapshot September 15, 2026)
Author
Look at AI, editorial team
