Thursday, Oct 8 | --:--
Back to home

GPT-5.3-Codex — First ‘High’ Preparedness Coding Agent, Same Day as Opus 4.6

OpenAI released GPT-5.3-Codex (Feb 5): ~25% faster; vendor SOTA on SWE-Bench Pro and Terminal-Bench 2.0 (77.3%); multi-day autonomous app builds with real-time steering. First model classified High under Preparedness Framework with expanded cyber stack. Agentic coding as core infra—not autocomplete.

Times of AI Desk 5 min read San Francisco View as Markdown
Cover illustration for GPT-5.3-Codex — First ‘High’ Preparedness Coding Agent, Same Day as Opus 4.6

Two labs shipped major agentic coding upgrades the same day. OpenAI’s proprietary stake: coding agents as core software-team infrastructure, with the first High Preparedness designation—capability and cyber controls announced together.

OpenAI introduced GPT-5.3-Codex on February 5. Pitch: unify Codex-line coding advances with GPT-5.2-class professional reasoning; ~25% faster (company). Vendor benches: leads public SWE-Bench Pro; Terminal-Bench 2.0 77.3%; strong OSWorld-Verified computer use; matches/exceeds priors on GDPval knowledge-work tasks. Long-horizon: iterate on full apps/games over days and millions of tokens; skills for web dev, debugging, creative generation; real-time steering and progress updates in updated Codex app. Internal dogfood: early versions used to debug training runs, optimize infra, analyze evals, help build the release. Preparedness: first model classified High capability; expanded cyber safety stack (safety training, monitoring, trusted access pilots for defensive research, threat-intel enforcement); Trusted Access for Cyber pilot launched alongside.

Claims vs checks

Product features and High designation are OpenAI primary. SWE-Bench Pro / Terminal-Bench / OSWorld / GDPval numbers are vendor-cited—confirm on public boards. “25% faster” is company relative claim (baseline unspecified in brief). Dual-use cyber framing is OpenAI’s risk taxonomy, not a regulator’s.

Limits

  • Multi-day autonomy demos are curated; production reliability differs.
  • High preparedness ≠ public exploit capability scores.
  • Same-day Opus 4.6 comparisons require matched scaffolds—don’t mash leaderboards casually.

Sources

  • OpenAI: “Introducing GPT-5.3-Codex” (February 5, 2026). Primary.
  • OpenAI system card and Preparedness Framework references tied to the release.
  • Terminal-Bench, SWE-Bench Pro, OSWorld, GDPval evaluations cited in the announcement.

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading