🛠 SWE-rebench Benchmark Update for AI Agents
A massive multilingual update to the SWE-rebench benchmark has been released to evaluate AI agents in programming. The test now covers 111 tasks across five languages: Go, Java, Python, Rust, and TypeScript. Anthropic Fable 5 emerged as the leader in Pass@1 (64.5%), while Anthropic Opus 5 was recognized as the most stable solution. The OpenAI GPT-5.6 Sol model showed high efficiency, completing tasks with a minimal number of actions.
🌍 The transition to multilingualism makes agent evaluation more realistic. The introduction of stability and cost metrics helps companies choose models for industrial deployment.
👤 It is now possible to evaluate AI agents not only by their capabilities but also by their predictability and the cost of solving real-world tasks in various languages.
Source 1: https://swe-rebench.com/ Source 2: https://teletype.in/@ibragim_bad/swe-rebench-2026-07
