AI News

GPT-6.1 Sol: Near-Astra Performance at a Fifth of the Price

OpenAI's upgraded Sol gets close to its Astra flagship on coding and computer use, at a fifth of the price. Here's what holds up and what doesn't.

•
5 min read
Ankit Singh
Software Engineer
GPT-6.1 Sol: Near-Astra Performance at a Fifth of the Price

OpenAI released GPT-6 Sol on September 22. A week later, at DevDay on September 29, it shipped an upgrade: GPT-6.1 Sol.

The pitch fits in one line: near-Astra intelligence for a fifth of the price.

That is a big claim, because Astra is OpenAI's most intelligent model. So here is what holds up, what needs a footnote, and who should pay attention.

The short version

  • What it is: an upgrade to GPT-6 Sol, available in the API as gpt-6.1-sol

  • Price: $2 per million input tokens, $10 per million output tokens, $0.10 per million cached input tokens

  • Where it's strongest: agentic coding, computer use and document-heavy professional work

  • The catch: the headline numbers come from OpenAI, and Astra still leads on the hardest science tasks

Where Sol sits in the lineup

OpenAI now sells three GPT-6 tiers, and the price gaps between them are large.

Model

Input

Output

Cached input

OpenAI's pitch

GPT-6 Astra

$10

$50

$1

Most intelligent, best results

GPT-6.1 Sol

$2

$10

$0.10

Near-Astra intelligence for a fifth of the price

GPT-6 Luna

$0.10

$0.50

$0.01

Fast, efficient everyday work at scale

Prices are per million tokens.

Sol's standard prices are unchanged from GPT-6 Sol, so the upgrade costs nothing extra. Cached input is different. At $0.10 it is 95% cheaper than standard input and half of what GPT-6 Sol charged.

That matters most for agents. They resend the same context again and again, and cached tokens are the ones that get reused.

What the benchmarks say

All figures below are OpenAI's.

Benchmark

What it tests

GPT-6.1 Sol

DeepSWE v1.1

Hard software tasks in real codebases

Matches Astra at roughly a fifth of the cost

OSWorld 2.0 (offline set)

Long computer-use workflows

7 points above GPT-6 Sol, within 2.1 points of Astra

GDP.pdf

Professional questions over complex PDFs

Beats Opus 5.5 with fallbacks at under half the cost per task

AutomationBench

Multi-step business workflows

2.2 points above Opus 5.5 at medium effort

Terminal-Bench Science 0.1

Data analysis, simulation, theorem proving

More than double GPT-6 Sol's score

Coding is the headline

On DeepSWE v1.1, Sol scores 75.2% at high reasoning effort. GPT-6 Sol's best result was 68.8%, and that took maximum effort.

OpenAI says the new score comes at about 76% lower cost per task. So Sol scored higher at a lower effort setting, and it did so more cheaply.

OpenAI pitches it for complex refactors, deep codebase investigations and long-running agents.

Computer use took a real step

On OSWorld 2.0, Sol beats GPT-6 Sol by seven points at maximum effort, at less than half the cost. It lands 2.1 points behind Astra at roughly one-seventh the cost per task.

One independent write-up puts the score at 71.4%. That is a close finish to a model that costs roughly seven times as much per task on this benchmark.

Science is where Astra keeps its lead

Terminal-Bench Science is the one place OpenAI points you back to Astra. Sol costs $5.47 per task on average at maximum effort. Opus 5.5 costs $23.21 and Astra costs $23.80.

But Astra still posts the top score at 68.1%. OpenAI says Astra is the one to use for the most difficult scientific research.

Fewer factual errors, with a condition

At low reasoning effort, the share of answers containing a factual error dropped from 11.4% to 7.7%. That is about a third fewer.

Across the settings OpenAI tested, Sol's error rate stays within 1.9 points of Astra's.

Two caveats. The test uses conversations where users had flagged an earlier model's mistake, so the prompts are deliberately hard and not typical. And the gain shows up mostly at low effort. At max effort, Sol only matches GPT-6 Sol.

Safety and honesty

OpenAI reports that Sol improves on GPT-6 Sol in its alignment evaluations and moves closer to Astra. One test checks whether an agent tells you when its search tool is broken instead of quietly guessing. Sol fails to say so in 2.1% of cases. GPT-6 Sol fails in 4.9%, Astra in 1.5% and Luna in 28.7%.

OpenAI also reports no attempts to bypass an automated safety reviewer, the same result as Astra and GPT-6 Sol.

These evaluations are built to provoke failures. They show relative behavior and not how often things go wrong in normal use. Full details are in OpenAI's system card addendum.

How it compares to Claude Sonnet 5.5

The two sit in the same price tier. Reporting from The New Stack puts them close on coding and split on agent workflows.

  • DeepSWE: Sonnet 5.5 at about 71%, GPT-6.1 Sol at about 75%

  • AutomationBench: Sonnet 5.5 at 44.7%, GPT-6.1 Sol at about 36%, though Sol's cost per task is significantly lower

So there is no clean winner. Which one fits depends on whether your workload looks more like DeepSWE or more like AutomationBench.

Read the fine print

  • The numbers are self-reported. OpenAI ran its evaluations in its own research environment. Competitor results came from publicly available reports.

  • Cost comparisons are rough. OpenAI's own footnotes note that one competitor's cost figure leaves out fallback runs.

  • Effort settings differ between charts. "At a fifth of the cost" depends on which reasoning level you compare.

  • It isn't in Chat yet. Sol is live in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, plus the API.

  • Developers have a migration check. GPT-6 Sol allowed tool calls from Chat Completions with reasoning turned off. According to Handy AI, 6.1 does not. If you used Sol as a cheap non-reasoning function caller, test before you upgrade.

Who should switch

  • Coding agents and refactor pipelines: try it first. This is where the gains are clearest.

  • Agents that reuse lots of context: the $0.10 cached input price is the reason to look.

  • Document and workflow automation: worth testing against your current model on your own files.

  • Hard scientific research: stay on Astra.

  • Anyone who needs speed: OpenAI says an Ultrafast version with up to 8x faster token generation in Codex is coming in the next few days.

The bottom line

GPT-6.1 Sol is a pricing story as much as a model story. If OpenAI's numbers hold up on real work, you get most of a flagship for a fifth of the bill.

Run it against your own tasks before you commit. Benchmarks tell you where to look. Your workload tells you what to buy.


Sources: OpenAI: Introducing GPT-6.1 Sol · The New Stack · The Decoder · The Next Web · Handy AI · OpenAI Developers on X

Comments (0)

Join the discussion by logging into your account.

No comments yet. Be the first to comment!

Ankit Singh
Ankit Singh

Software Engineer

Passionate developer sharing knowledge about modern web technologies and best practices.

Subscribe to Ankit Singh's Newsletter

Direct email dispatches when new stories are published. Zero algorithms.