🤖 DeepSWE: A New Benchmark for AI Programming Agents

DeepSWE has been introduced—a specialized benchmark for evaluating the capabilities of advanced AI agents in solving complex, long-term software engineering tasks. It utilizes 113 unique tasks from 91 repositories, protected against data leakage.

🌍 DeepSWE addresses the problem of "saturation" in current tests, where the gap between top models becomes statistically insignificant. This allows for better differentiation of agents based on their real-world behavior and planning.

👤 The industry is shifting from evaluating "the ability to write code" to evaluating "the ability to work as an engineer"—independently exploring repositories and making changes to systems.

Source 1: https://deepswe.datacurve.ai/ Source 2: https://github.com/datacurve-ai/deep-swe