GPT-6.1 Sol: Pricing, Benchmarks, and What Changed
GPT-6.1 Sol brings stronger task performance at familiar token rates. Here's what changed in pricing, caching, benchmarks, and the case for switching.
GPT-6.1 Sol makes the everyday-model decision harder
OpenAI released GPT-6.1 Sol on September 29, just a week after GPT-6 Sol. If your reaction was “already?”, fair enough.
The useful question is whether it changes what you can afford to hand over to AI every day. A model that gets closer to flagship performance could make a difference to a team running coding agents or processing complicated documents. Especially if fewer tasks need an expensive second attempt.
OpenAI's September 29 API changelog confirms the release and the API identifier, gpt-6.1-sol.
The headline needs a little unpacking, though.
Cheaper than Astra doesn't mean 80% cheaper than the previous Sol. Confuse those comparisons and the whole buying argument goes sideways.
GPT-6.1 Sol pricing: the comparison that matters
GPT-6 Sol and GPT-6.1 Sol have the same standard input and output rates. The clear price cut is in cached input.
| Standard API rate, per million tokens | GPT-6 Sol | GPT-6.1 Sol |
|---|---|---|
| Input | $2.00 | $2.00 |
| Cached input | $0.20 | $0.10 |
| Output | $10.00 | $10.00 |
These rates apply to prompts with up to 272,000 input tokens. The API changelog records both releases and their prices.
So the upgrade case rests on better results at the same main token rates, plus cheaper cache reads. Any further savings depend on what happens during the task.
That's a meaningful distinction. A model can keep the same sticker price and become cheaper to use if it needs fewer retries. It can also look like a bargain and spend the saving on extra work that never produces an acceptable result.
Our AI Model Price Cuts and Efficiency guide covers that gap between a published rate and a useful outcome.
What the benchmarks actually say
OpenAI reports that GPT-6.1 Sol matches GPT-6 Astra on DeepSWE v1.1 at roughly one-fifth of the task cost. On AutomationBench, it improves on GPT-6 Sol by 4.8 percentage points at the same medium reasoning setting. Astra still leads the most difficult scientific research evaluation. These are vendor-reported results, not tests conducted by Deepapi. OpenAI's launch announcement
The tasks matter as much as the scores. Software engineering and business workflows require a model to keep track of intermediate results, use tools and recover when something fails.
A convincing answer is only the start.
Take a bug fix. The agent needs to find the cause, change the right code, run the relevant checks and leave the rest of the application working. A lovely explanation attached to a broken patch is still a broken patch.
Or take a document workflow. Reading a table correctly doesn't help much if the model then answers the wrong question or quietly drops a condition from the small print.
The launch evidence makes Sol worth testing for this kind of work. It doesn't establish that Sol can replace Astra across every task. Your repository, tools, instructions and acceptance criteria are part of the result.
A practical comparison should include failures. Run both models on representative tasks and record whether you accepted the output, how long it took, how often you intervened and what the entire run cost.
The demo that worked is easy to remember. The afternoon spent repairing the demo that didn't work belongs in the spreadsheet too.
Cheaper caching helps, but it isn't free context
Repeated context can become a large part of an agent's bill: repository background, tool definitions, project rules and earlier results all compete for space.
At GPT-6.1 Sol's standard cache-read rate, 200,000 cached input tokens cost $0.02. That's an arithmetic example, not a measured production run. It excludes cache writes, fresh input, output and tool charges.
Cache writes are listed at $2.50 per million tokens. The model supports a 1,050,000-token context window, but prompts above 272,000 input tokens trigger higher pricing for the full request: twice the input and cache rates, and 1.5 times the output rate. OpenAI's model reference
A million-token window gives you room. It doesn't give you a reason to fill it.
Before putting an entire codebase into every request, ask what the model actually needs. Relevant files and a clear task boundary can be more useful than a huge pile of background information.
Cache savings also depend on actual reuse. If your workflow keeps changing the reusable context, the bill may look very different from the estimate.
Our guide to reducing Codex token usage through context management is a useful next step if your agent spends more time rereading the project than changing it.
“Near Astra” still leaves a gap
A model can improve without becoming a universal replacement for the flagship.
OpenAI reports a decline in factual-error-containing answers from 11.4% to 7.7% at low reasoning effort. The prompts came from difficult conversations where users had flagged earlier errors; the result is not a general error rate for everyday usage. At launch, Sol was available through ChatGPT Work, Codex and the API, but not yet Chat. OpenAI's announcement
Those qualifications belong next to the claims. Otherwise, an encouraging evaluation becomes a promise the evidence never made.
For a team choosing a daily model, the question is which failures you can tolerate and catch. Routine edits with clear tests are easier to evaluate than ambiguous research or a cross-service bug with no reliable reproduction.
Start with work you understand well enough to judge. Then compare the harder cases. Don't let the name of the model do the reviewing for you.
Our earlier Astra versus GPT-5.6 Sol coding-cost comparison explains the same decision using a previous generation. Its measurements concern those older models, so they should not be treated as GPT-6.1 test results.
A capable agent still needs boundaries
OpenAI's GPT-6.1 Sol system card addendum discusses prompt injection, deceptive behavior and whether agents respect restrictions. It also notes that research and API evaluations can differ from production behavior because of prompts, tools and reasoning settings.
For anyone building an agent, that raises a practical design question: what is it allowed to do when it gets stuck?
A chatbot can leave a bad answer on the screen. An agent with write access can leave a bad change in a system. The difference matters even when the model is trying to help.
Define which actions it can take automatically, which require approval and how you'll inspect or undo the result. Keep a record of what actually happened. A stronger model doesn't configure those controls for you.
This is also part of cost. Work that needs to be unwound wasn't cheap just because the tokens were.
Should you switch your default model?
GPT-6.1 Sol is a reasonable candidate to test if you repeatedly run coding tasks, read complex documents or manage workflows with several tool steps.
If your current model already completes short, well-defined work reliably, the upgrade may have little value. If you regularly escalate to Astra, the interesting experiment is whether Sol can now finish some of those jobs.
Use a small set of real tasks and keep the setup comparable. Track accepted results, full-run cost, retries, elapsed time and human repair time. For subscription users, measure quota consumption separately from API dollars.
Deepapi's Price Reversal Phenomenon guide explores why a lower token rate can produce a higher task bill. Our broader API pricing guide covers how pricing conditions change the comparison.
GPT-6.1 Sol gives you a new option to put through that process. Whether it deserves the default slot is a decision your own work can settle.
Editorial note: Facts checked October 3, 2026. This is a source-based analysis, not a hands-on review. Calculated examples are identified as such. Pricing and access may change.
Continue exploring
More decisions worth reading
Follow the thread from this article to the next practical buying question.