GPT-6.1 Sol Claims Near-Astra Work at a Fifth of Astra’s Token Price
OpenAI’s DevDay model is not the shelved Astra successor. GPT-6.1 Sol keeps the $2 / $10 sticker and claims near-GPT-6 Astra results on coding, computer use, and professional work at about one-fifth of Astra’s standard token prices. Those benches are the lab’s. Arena has not ranked it.

The model OpenAI shipped at DevDay is the cheap tier, tightened — not the frontier successor it had just pulled. GPT-6.1 Sol is an upgrade to GPT-6 Sol, priced the same, and sold as near GPT-6 Astra on agentic coding, computer use, and professional work at one-fifth of Astra’s standard input and output token prices. The hold on GPT-6.1 Astra, the day before, is what makes “near” do commercial work: buyers are offered Astra-class tasks without the SKU safety staff would not release.
API prices, from the OpenAI post (September 29): $2 per million input tokens, $0.10 cached input, $10 output. Cached input is, OpenAI says, 95% below this model’s own standard input price and 50% below GPT-6 Sol’s cached input price. Slug: gpt-6.1-sol. It is in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu. Not yet in Chat. Ultrafast, up to 8× token generation versus standard speed in Codex, is promised in the coming days, not on day one. The DevDay recap says Astra Ultrafast is already in the API and in Work and Codex on Pro 500 and Enterprise. That speed tier is not 6.1 Sol.
The sticker matches Claude Sonnet 5.5 and GPT-6 Sol. The claim that has to be true is cost per task against Astra, not the menu price.
Lab claims
OpenAI says competitor numbers are taken from public reports, and that its own runs were in a research environment or via API, which can differ from production ChatGPT. Effort level is doing a lot of work in every row.
- DeepSWE 1.1: matches GPT-6 Astra at roughly one-fifth the cost, and beats GPT-6 Sol’s best score by 6.4 percentage points at a lower effort and cost. The prose post does not print the absolute percent.
- GDP.pdf (professional questions over complex PDFs): higher than Opus 5.5 with fallbacks at less than half the cost per task across tested settings, and near Astra at roughly one-fifth the cost per task.
- AutomationBench 1.0.6 (47 tools): 2.2 percentage points above Opus 5.5 at medium effort, at roughly a third of the cost, and 4.8 points above GPT-6 Sol at that same setting. OpenAI again flags that a published Fable 5.1 point omits fallback cost, which hit on about 40% of tasks.
- OSWorld 2.0 offline, partial reward, release v2026.08.08: seven points above GPT-6 Sol at maximum effort and less than half the cost; within 2.1 points of Astra at maximum effort at roughly one-seventh the cost per task.
- Terminal-Bench Science 0.1: more than doubles GPT-6 Sol at maximum effort, at less than half the cost. OpenAI’s average cost per task at max: $5.47 for 6.1 Sol, $23.21 for Opus 5.5, $23.80 for Astra. Astra still leads the score, at 68.1%. OpenAI says to use Astra for the hardest scientific tasks.
- Factuality, an internal set of de-identified chats where a user had already flagged an error — not typical traffic. At low effort, the share of answers with a factual error falls from 11.4% to 7.7%, about a 32% reduction. Across tested settings, the error rate stays within 1.9 percentage points of Astra at less than one-fifth the cost per task.
On alignment, OpenAI says 6.1 Sol is closer to Astra than GPT-6 Sol was, more willing to state limits, and showed no attempts to bypass an automated safety reviewer — the same result it reports for Astra and GPT-6 Sol. One printed stress number: when the search tool is broken, 6.1 Sol fails to say so in 2.1% of cases, versus 4.9% for GPT-6 Sol, 1.5% for Astra, and 28.7% for Luna, all at maximum effort, on prompts chosen to elicit failures.
Independent checks
There is no matching public scorecard yet. No Arena Elo for GPT-6.1 Sol appears in the launch post or in TechCrunch’s rewrite. None is invented here.
Artificial Analysis had, the day before, put Claude Sonnet 5.5 max at Intelligence Index 56 and Opus 5.5 max at 58, and described GPT-6 Sol configurations as the cheaper comparison at lower Sonnet efforts. A settled max-effort Intelligence Index for GPT-6.1 Sol was not in those writeups. The model was hours old. TechCrunch restates OpenAI’s claims, including the shelved-Astra context. It is not a third-party bench.
Until AA or Arena prints a dated row, “near Astra” means OpenAI’s task-level cost curves, not a shared index. The Sonnet 5.5 piece is the caution: identical $2 / $10 stickers, and AA already showed max-effort Sonnet leaving the cost frontier by writing ~193,000 tokens a task.
Limits
- Not in consumer Chat. Work and Codex are the surfaces.
- Ultrafast for this model is announced, not available.
- Several headlines are “within N points” or “one-fifth the cost,” not a full score table. Absolute DeepSWE and AutomationBench percents are not in the prose we reviewed.
- Factuality and broken-tool tests are hard, non-representative sets. OpenAI says so.
- The Fable AutomationBench cost still omits fallbacks, by OpenAI’s own note.
- Arena: not yet ranked. Artificial Analysis: no settled max-effort Index number in the sources used here.