Thursday, Oct 8 | --:--
Back to home

GPT-5.4: Agentic Computer Use Is the Product, Not Chat Elo

OpenAI GPT-5.4 (March 5) — Thinking/Pro with mini/nano following: native computer use, configurable reasoning effort, up to 1M context, Tool Search. Lab OSWorld ~75% (above human 72.4% baseline in OpenAI materials). Early Arena snapshots ~1502 Elo briefly over Opus 4.6 (~1494) — snapshot-sensitive; durable story is agents + computer use.

Times of AI Desk 4 min read San Francisco View as Markdown
Cover illustration for GPT-5.4: Agentic Computer Use Is the Product, Not Chat Elo

Frontier launches still get ranked on preference Elo. GPT-5.4’s proprietary frame: ship the agentic computer-use stack — configurable effort, Tool Search, long context — and treat early Arena leads as snapshot weather.

OpenAI unveiled GPT-5.4 (March 5) with Pro and (days later) mini/nano variants as its most capable model family for agent-style professional work: native computer-use APIs; configurable reasoning effort; context up to 1M tokens; coding/reasoning gains; Tool Search for dynamic tool calling; ChatGPT + API + Codex. Company materials: ~75% on OSWorld (slightly above human baseline 72.4%); strong SWE-bench-class coding. Resource shift toward developer and agent tools.

Source Snapshot Standing
LMArena / Arena Text March 2026 boards Secondary coverage: GPT-5.4 near ~1502 Elo, briefly above Claude Opus 4.6 (~1494) on some preference snapshots
Vendor OSWorld Launch materials ~75% vs human 72.4% baseline

Label early Arena leads as snapshot-sensitive; 5.4’s durable story is agentic computer use + pricing tiering, not permanent #1 preference.

Claims vs checks

Launch features are OpenAI primary. OSWorld/SWE figures are vendor (effort settings matter). Arena Elo is independent preference board but volatile week-to-week.

Limits

  • Computer-use success rates transfer unevenly outside eval harnesses.
  • 1M context availability/pricing varies by SKU.
  • Mini/nano economics belong in the March 17–18 files.

Sources

  • OpenAI: “Introducing GPT-5.4” (March 5, 2026).
  • API documentation and pricing for GPT-5.4 variants.
  • Benchmark reports (SWE-bench, OSWorld) and TechCrunch coverage (March 5, 2026).

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading