Thursday, Oct 8 | --:--
Back to home

SWE-Bench Pro Audit: ~30% Broken Tasks, Recommendation Retracted

OpenAI’s coding-eval audit finds roughly 30% of SWE-Bench Pro tasks broken—hidden requirements, contradictory instructions, strict tests, incomplete grading—and retracts its prior recommendation that labs treat Pro as a leading coding eval. Score gaps may be noise as much as capability.

Times of AI Desk 5 min read San Francisco, CA View as Markdown
Cover illustration for SWE-Bench Pro Audit: ~30% Broken Tasks, Recommendation Retracted

Vendor leaderboards and marketing slides still hang on SWE-Bench-style resolve rates. If ~1 in 3 Pro tasks are broken, score gaps between GPT, Claude, Grok, and Gemini can be noise as much as capability—just as agentic coding models race on those same charts.

OpenAI on July 8, 2026 published “Separating signal from noise in coding evaluations,” reporting that a large share of SWE-Bench Pro tasks no longer reliably measure frontier coding ability. Model-based investigator agents plus five independent experienced software engineers supported the analysis. OpenAI estimates ~30% of Pro tasks are broken and retracts its earlier recommendation that labs treat SWE-Bench Pro as a primary coding eval—months after it had steered the field away from contaminated SWE-bench Verified toward Pro.

What the audit found

From OpenAI’s primary post and same-day X thread:

  • Headline: SWE-Bench Pro “no longer reliably measures frontier coding capability.”
  • Broken-task estimate: ~30% of tasks; automated pipeline flagged 200/27.4% broken; human annotation campaign flagged 249/34.1%.
  • Failure modes: Correct solutions fail due to hidden requirements, contradictory instructions, overly strict tests, or incomplete grading criteria.
  • Method: Model-based investigator agents at scale, plus independent review by five experienced software engineers.
  • Policy change: Retract prior recommendation that the research community use SWE-Bench Pro as a leading coding eval; call for harder, fairer, more trustworthy benchmarks as coding models improve.

Context: In February 2026, OpenAI had already stopped leaning on SWE-bench Verified over contamination and mismeasurement, pushing labs toward Pro. The July 8 audit closes that chapter: the successor benchmark is itself materially flawed under current scrutiny.

Limits

  • ~30% is OpenAI’s estimate (automated 27.4% vs human 34.1%); not an independent replication published with the post.
  • Retraction is of OpenAI’s recommendation, not deletion of the Pro dataset itself.
  • Other labs may keep citing Pro resolve rates until successor evals stabilize.

Sources

  • OpenAI: “Separating signal from noise in coding evaluations” (openai.com/index/separating-signal-from-noise-coding-evaluations/, July 8, 2026). Primary.
  • OpenAI X thread (July 8, 2026) summarizing 30% broken-task finding and retraction.
  • Prior context: OpenAI Feb 2026 move off SWE-bench Verified toward Pro.

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading