# SWE-Bench Pro Audit: ~30% Broken Tasks, Recommendation Retracted

Times of AI Desk · 2026-07-08 · Research

[https://timesof.ai/2026/07/openai-swe-bench-pro-audit-retracts-recommendation](https://timesof.ai/2026/07/openai-swe-bench-pro-audit-retracts-recommendation)

> OpenAI’s coding-eval audit finds roughly 30% of SWE-Bench Pro tasks broken—hidden requirements, contradictory instructions, strict tests, incomplete grading—and retracts its prior recommendation that labs treat Pro as a leading coding eval. Score gaps may be noise as much as capability.

Vendor leaderboards and marketing slides still hang on SWE-Bench-style resolve rates. If ~**1 in 3** Pro tasks are broken, **score gaps** between GPT, Claude, Grok, and Gemini can be noise as much as capability—just as agentic coding models race on those same charts.

**OpenAI** on **July 8, 2026** published **“Separating signal from noise in coding evaluations,”** reporting that a large share of **SWE-Bench Pro** tasks no longer reliably measure frontier coding ability. Model-based investigator agents plus five independent experienced software engineers supported the analysis. OpenAI estimates **~30%** of Pro tasks are **broken** and **retracts** its earlier recommendation that labs treat SWE-Bench Pro as a primary coding eval—months after it had steered the field away from contaminated **SWE-bench Verified** toward Pro.

## What the audit found

From OpenAI’s primary post and same-day X thread:

- **Headline**: SWE-Bench Pro “no longer reliably measures frontier coding capability.”
- **Broken-task estimate**: ~**30%** of tasks; automated pipeline flagged **200/27.4%** broken; human annotation campaign flagged **249/34.1%**.
- **Failure modes**: Correct solutions fail due to **hidden requirements**, **contradictory instructions**, **overly strict tests**, or **incomplete grading criteria**.
- **Method**: Model-based investigator agents at scale, plus independent review by five experienced software engineers.
- **Policy change**: **Retract** prior recommendation that the research community use SWE-Bench Pro as a leading coding eval; call for harder, fairer, more trustworthy benchmarks as coding models improve.

Context: In **February 2026**, OpenAI had already stopped leaning on **SWE-bench Verified** over contamination and mismeasurement, pushing labs toward **Pro**. The July 8 audit closes that chapter: the successor benchmark is itself materially flawed under current scrutiny.

## Limits

- ~30% is OpenAI’s estimate (automated 27.4% vs human 34.1%); not an independent replication published with the post.
- Retraction is of OpenAI’s **recommendation**, not deletion of the Pro dataset itself.
- Other labs may keep citing Pro resolve rates until successor evals stabilize.

## Sources

- OpenAI: [“Separating signal from noise in coding evaluations”](https://openai.com/index/separating-signal-from-noise-coding-evaluations) (openai.com/index/separating-signal-from-noise-coding-evaluations/, July 8, 2026). Primary.
- OpenAI X thread (July 8, 2026) summarizing 30% broken-task finding and retraction.
- Prior context: OpenAI Feb 2026 move off SWE-bench Verified toward Pro.
