GPT-6 Sol: OpenAI's Cost-Efficiency Mid-Tier Model, Scored 6.0 on the U365 CI-First Review at Half the Cost per Completed Task
Updated: 2 days ago
Status: Active | Last tested: 2026-09-23 (GPT-6 Sol, gpt-6-sol, released 2026-09-22) | Re-check: trigger-based (max 6 months)
Active: the tool is current and recommended.


Tool Snapshot
Tagline: "GPT-6 Sol is built for complex coding and agentic workflows." (OpenAI model reference, gpt-6-sol, read 2026-09-23.)
Category: Applied AI / Large Language Model (agentic coding and knowledge work)
Mid-tier of the GPT-6 family, below GPT-6 Astra. Primary use cases are recurring complex software work, agent runs that call tools, and business workflow automation.
Primary use cases:
Work through a multi-file change or migration in a repository, running the project's own tests, at roughly half the cost per task of the model it replaces.
Run an agent that calls tools, browses, and works in the background where the number that matters is cost per completed job rather than price per token.
Automate a business workflow across several applications and tools, where the published measurement is score against cost per task.
Operate a computer interface for long-running tasks at a fraction of the flagship rate.
Draft and check a technical document where the model is expected to say what it did and did not verify.
Produce a first pass over more material than you could read, at a price that makes the first pass cheap enough to throw away.
Pricing summary: Paid. API list price is $2 per million input tokens and $10 per million output tokens, down 50 percent from the GPT-5.6 Sol promotional rate of $4 and $20, and 60 percent below GPT-5.6 Sol's original list price of $5 and $30. Cached input is $0.20 per million, a 90 percent discount, and cache writes are billed at 1.25 times the uncached input rate, $2.50. Prompts above 272,000 input tokens pay 2 times the input and cache rates and 1.5 times the output rate for the whole request. Batch and Flex are 50 percent of Standard. Fast mode is 2 times the applicable rate. Regional processing adds 10 percent where available, and EU data residency on Sol and Luna is available only with Standard processing. In ChatGPT Work and Codex the models are metered in credits: Sol costs 50 credits per million input tokens and 250 per million output, shared across both surfaces. Consumer access needs a Plus, Pro, Business, Enterprise, or Edu plan; Enterprise and Edu administrators must enable the model for the workspace. Prices captured 2026-09-22 and 2026-09-23 from the OpenAI launch page, the model reference pages, and the pricing documentation.
Official links:
Launch page: https://openai.com/index/introducing-gpt-6-sol-and-luna/
Prompt caching announcement: https://openai.com/index/better-prompt-caching-for-gpt-6/
Model reference, GPT-6 Sol: https://developers.openai.com/api/docs/models/gpt-6-sol
Model reference, GPT-6 Luna: https://developers.openai.com/api/docs/models/gpt-6-luna
API pricing: https://developers.openai.com/api/docs/pricing
Codex local memories: https://learn.chatgpt.com/docs/customization/memories
Codex instruction files: https://learn.chatgpt.com/docs/agent-configuration/agents-md
Codex subagents: https://learn.chatgpt.com/docs/agent-configuration/subagents
GPT-6 Astra system card: https://deploymentsafety.openai.com/gpt-6-astra
Codex repository: https://github.com/openai/codex
Status page: https://status.openai.com
LLM-specific fields:
API model identifier: gpt-6-sol. No dated snapshot is published; the model ID itself is the snapshot.
Context window: 1,050,000 tokens, with 922,000 maximum input tokens and 128,000 maximum output tokens. Text and image input, text output. Audio and video are not supported.
Knowledge cutoff: 20 April 2026. (GPT-6 Astra is 30 April 2026; GPT-6 Luna is 18 May 2026.)
Effort levels: none, low, medium (default), high, xhigh, max. The parameter is reasoning.effort.
Parameters and architecture: not publicly disclosed. OpenAI publishes no parameter count, no architecture description, and no weights for any GPT-6 model.
Model variants: GPT-6 Sol only. There is no GPT-6 Terra, and OpenAI has not announced one. GPT-6 Astra sits above Sol at $10 and $50 per million tokens.
Available platforms: ChatGPT Work and Codex (web, CLI, IDE extension, iOS, and cloud tasks), the OpenAI API, and Amazon Bedrock. Enterprise and Edu workspaces need an administrator to turn the model on for the workspace. Not open weights and not self-hostable.
License: proprietary hosted service.
Comparison references: arena.ai (LMSYS Chatbot Arena) for head-to-head rankings and ollama.com/search for local deployment of other models. GPT-6 Sol cannot be run locally.
Delivery surface to check separately from the model: Codex CLI and the ChatGPT desktop app. That surface, not the API endpoint, is what the framework's clause 5.2.3-a and clause 7.5 findings in Section 7 rest on, because it writes memory files and runs parallel subagents.
At a Glance dashboard
Field | Value |
Category | Applied AI / Large Language Model (agentic coding and knowledge work) |
CI-First Benefit Score | 6.0 / 10 (CI-First Positive) |
Sub-scores | Time 7 / Quantity 7 / Quality 7 / Skill 3 |
CI-First Profile | Primary: Co-Worker and Assistant (level 2). Secondary: Analyst and Tester (level 4), Coach and Tutor (level 3) |
Collaboration Mode | Centaur. Cyborg is not available where more than one agent shares a view or channel (clause 7.5), and is permitted only for one supervised agent session with a stopping criterion set in advance |
Humics Protection | Humics-Neutral (-1 / +3) |
AI Imposture Risk | Medium overall, with Skill Illusion High |
Status | Active |
Last tested | 2026-09-23 |
Released | 2026-09-22 |
Access | ChatGPT Work and Codex on paid plans, the OpenAI API, Amazon Bedrock |
Price | $2 per million input tokens, $10 per million output, $0.20 cache read |
Context window | 1,050,000 tokens in, 128,000 tokens out |
Re-check triggers
Major release: OpenAI ships a GPT-6 Terra, a GPT-6.1 generation, or an updated checkpoint of Sol. There is no GPT-6 Terra and OpenAI has not announced one.
Price change: the published $2 per million input and $10 per million output rates move, or the $0.20 cache-read rate changes. OpenAI has told a reporter the reduction is permanent rather than promotional, which is itself worth re-testing rather than assuming.
Effort default change: the default effort is medium. The published benchmark rows are mostly at xhigh and max. If OpenAI raises the default, both the cost claim and this review's Time sub-score need re-measuring.
Independent re-measurement: any new like-for-like comparison of GPT-6 Sol against Claude Opus 5.5 on cost per completed task. The two launched about 90 minutes apart and no one has run both in one evaluation setup.
Regression watch: Artificial Analysis measured a roughly 100-Elo drop against GPT-5.6 Sol on its professional knowledge-work benchmark. If that gap closes on a re-run, or widens, the Quality sub-score needs revisiting.
API migration watch: function calling on Chat Completions works only when reasoning effort is none. If OpenAI restores it at other effort levels, the developer guidance in Section 5 changes.
The Problem
The price of a frontier model stopped being the thing that stops you from using it. What stops you now is the number of attempts a task needs. A cheap model that takes three tries, two of which you read before discarding, costs more than an expensive model that takes one. Vendors have moved the argument from price per token to cost per completed task for exactly that reason, and OpenAI makes that argument explicitly in this launch.
Three specific problems were live before this release.
The first is that the mid tier of OpenAI's own line-up had no home. GPT-6 Astra shipped on 3 September at $10 and $50 per million tokens, roughly double the GPT-5.6 flagship rate. Anyone who needed OpenAI's current generation for routine work was paying flagship prices for routine work, or staying on GPT-5.6 and accepting lower factuality.
The second is the bill for long agent runs. A coding agent that works for an hour carries the same instructions and context across many requests. If that shared context is billed at full input price every turn, the cost of an agent is dominated by re-reading material it has already read. OpenAI states its own internal token usage now exceeds $600 per day at API prices for the median researcher and $7,000 at the 90th percentile, which is a statement about the size of the bill, not about the value of the work.
The third is that reliability and cost were moving in opposite directions. The cheaper tier of the previous generation made more factual mistakes, so the cheap model was only safe on work where you could check the output mechanically. That constrained the cheap tier to formatting and extraction.
There is a fourth problem this release does not solve, and it belongs in the same paragraph. The cheaper and more capable the model gets, the less attention each individual output receives. Volume rises faster than the reading capacity of the person using it. Section 7 scores that as the High trap, and the reason it is High rather than Medium is the delivery surface rather than the model.
The Outcome
What changes for a reader who adopts this model:
Cost per completed task roughly halves. Artificial Analysis measured GPT-6 Sol at maximum effort at $1.06 to run its Intelligence Index, against $1.99 for GPT-5.6 Sol at the same setting. That is a measured number from a third party, not a vendor claim, and the price cut does most of the work: Sol uses slightly more output tokens per task than its predecessor, 31,000 against 29,000.
Factual errors drop sharply on the vendor's own test. OpenAI reports GPT-6 Sol makes about half as many mistakes as GPT-5.6 Sol, for example 5.1 percent against 10.8 percent at high effort. Independent measurement found the same direction with an important qualification: Artificial Analysis measured the hallucination rate falling from 92 percent to 60 percent, and traced much of that to the model attempting fewer questions, 83 percent against 99 percent, so accuracy on the questions it does answer fell from 59 percent to 54 percent. Declining to answer is a legitimate behaviour and it is not the same thing as knowing more.
Coding improves on independent measurement. Artificial Analysis scored the Coding Agent Index 57 for Sol at maximum effort, up 2 points from GPT-5.6 Sol, with Terminal-Bench 4.0 moving from 37 percent to 43 percent and SWE-Atlas-QnA from 54 percent to 58 percent.
The agent context bill falls for a second reason beyond the token price. OpenAI shipped higher default cache hit rates, discounts eligible shared prefixes reused within 30 minutes, and removed two things that used to break the cache: changing reasoning effort mid-conversation and enabling or disabling tools. It also added a caching dashboard and a diagnostics tool that explains a cache miss and names the cause. GitHub is quoted saying these changes reduced the share of prompt tokens needing fresh processing by more than 50 percent across billions of requests.
ChatGPT allowance rises. OpenAI's published estimates for local messages per five-hour window on Plus go from 10 to 100 for GPT-5.6 Sol to 15 to 150 for GPT-6 Sol, and on Pro 20x from 300 to 3,000. Those are the vendor's own estimates and cloud chats and image generation draw on the same credit pool.
The honest counterweight, stated once and then carried through this review:
The headline reduction is measured against a promotional rate, not a list rate. OpenAI's own sentence compares GPT-6 Sol to "GPT-5.6 promotional pricing", and GPT-5.6 Sol's list price was $5 and $30. Against the list rate the cut is 60 percent on input and 67 percent on output. Against the promotional rate it is 50 percent. Both numbers are true and they are not the same claim.
Sol does not lead OpenAI's own family. In OpenAI's own published tables, GPT-5.6 Sol still holds higher top scores on DeepSWE v1.1, 72.7 percent against 68.8 percent, and on OSWorld 2.0, 66.2 percent against 64.4 percent, at roughly double the cost per task.
Independent measurement found regressions on knowledge work. On GDPval-AA v2.1, which covers economically valuable tasks across 44 occupations, Artificial Analysis measured Sol losing about 100 Elo points against its predecessor, and traced the loss to lower presentation quality and incomplete results on manual inspection.
The migration is real work for existing API pipelines. Built-in tools and function calling should use the Responses API. On Chat Completions, function calling works only when reasoning effort is none, so a pipeline running GPT-5.6 Sol at medium or high cannot swap the model identifier and keep its behaviour. It moves endpoints or it loses reasoning on tool-calling turns.
The effort setting is now a large part of the bill and the choice is not obvious. On OpenAI's AutomationBench chart, Sol at max scores lower than at xhigh and costs 24 percent more. On DeepSWE, max adds 2.2 points over xhigh at 2.7 times the cost.
Who Should Use GPT-6 Sol
Learner type | Difficulty | Typical ROI | Career path |
Students (Bachelor, Master) | Intermediate | Usable through ChatGPT Work on a paid plan. Strong for working through a long document set, structuring an argument, and getting an answer you then reproduce yourself. The Skill Illusion is the live risk: a summary you cannot defend is not a skill. | UIT (Technology, AI, Data Science) tracks. UDA-confirmed anchor: AI Developer Specialist (18 days). |
Professionals (career upskilling) | Intermediate to Advanced | The strongest case is recurring complex engineering: feature work, code review passes, debugging, and data analysis, at roughly half the previous cost per completed task. Also strong for business workflow automation where the output is checkable. | |
Everyone (lifelong learners) | Beginner for chat in ChatGPT Work, Advanced for agent use | A 1M-token window and better factuality make it a good reading and reasoning partner. The agent surfaces this model is built around need a technical user to supervise, and supervising means naming a finish line and reading what came back. | SL-OS daily learning routine, LIPS Collect and Review phases. |
Skill level required: Intermediate for document and analysis work. Advanced for the agentic use this model is positioned for, because supervising a long run means setting the finish line, setting the stops, and reading the report.
Prerequisites: A paid ChatGPT plan (Plus, Pro, Business, Enterprise, or Edu) for ChatGPT Work and Codex, with an administrator turning the model on for the workspace on Enterprise and Edu. For the API, an OpenAI account with billing. Working knowledge of prompt structure and of verification practice. On the API, a decision about which endpoint you are on, because that decision has already changed behaviour for anyone using function calling through Chat Completions.
Typical time to first result: Under five minutes in ChatGPT Work. Under fifteen minutes in Codex for a scoped change with a stated finish line.
Typical time to competence: Ten to twenty hours of active use to learn effort selection, when to hand over a whole task, when to stop it, and how to read its report. Learning which effort level your own tasks actually need is the largest single saving available, and OpenAI's own charts show the relationship is not monotonic.
U365 Institutes Alignment
Institute | Relevance | Why |
UIT (Technology, AI, Data Science) | High (primary) | The clearest fit in the U365 curriculum, and the only institute that gains credential chains from this tool. Work through a multi-file change in a repository, run the project's own tests, and defend the result. The endpoint and effort configuration of an existing pipeline, including when function calling is available at all. Measured cost per completed task at roughly half the previous model, with the caveat that the vendor's headline compares against a promotional rate. Subagent governance under clause 7.5, because the vendor enables parallel subagents by default. |
UIB (Business Management, Entrepreneurship) | Medium | Two real competencies, and no more. The unit-economics question the vendor poses, which is not what a token costs but what a completed job costs, together with the baseline discipline of recording attempts, human minutes and error rate before adopting. And accountable automation governance: a defined check on the result, a named owner, a time and cost ceiling, and an accepted-risk decision. The tool teaches no management, leadership, finance or entrepreneurship content of its own. |
UIC (Digital Communication, Marketing) | Low to Medium | It drafts and restructures text and it supports content-operations automation as a supporting technical utility. It returns no sourced evidence, so verification stays with the Fellow. It builds no brand voice, audience judgment or campaign craft, and the independent knowledge-work regression was measured in presentation quality, which is the dimension this institute cares about most. |
UID (Digital Design, UX/UI) | Low to Medium | It reads images and turns an approved specification into interface code, and the vendor's own launch material shows it working through a site-layout change. It originates no design judgment, user research, information architecture or visual craft, so its role is implementation support inside a design process rather than design education. |
UDA confirms all four ratings as URC drafted them: [UIT](https://university-365.com/uit) High, [UIB](https://university-365.com/uib) Medium, [UIC](https://university-365.com/uic) Low to Medium, [UID](https://university-365.com/uid) Low to Medium. No row is moved. The reasons below are UDA's, and they are the reasons UDE should publish, because a rating without its reason reads as a label.
[UIT](https://university-365.com/uit) High, primary. Unchanged from URC. UDA adds that this is the only institute with credential chains attached, which is a stronger statement than the rating on its own.
[UIB](https://university-365.com/uib) Medium, confirmed rather than downgraded. UDA has twice moved a UIB rating down to Low to Medium on this series, for Jev AI and for Claude Code, because one narrow exercise is not a curriculum strand. It does not do so here, because two distinct UIB competencies are supported: the cost-per-completed-task analysis with a measured baseline, and accountable governance of an automated workflow with a defined check. UDA also states the limit that keeps the row at Medium and not higher, which is the measured knowledge-work regression in presentation quality and completeness. That regression lands on exactly the document-production work UIB programmes assess, so UIB Fellows should test the model against their own material before routing professional document production to it.
[UIC](https://university-365.com/uic) Low to Medium, confirmed. A general text-capable model supports drafting and restructuring, and content-operations automation is genuine supporting work. Neither is the institute's core disciplinary competency. No credential chain.
[UID](https://university-365.com/uid) Low to Medium, confirmed on the current record. Generating interface code from an approved specification is implementation. It builds no design competency and it produces no visual asset, because image generation is recorded as not supported on this model. One open question is flagged to URC in Section 7 of this document, because the URC review lists image generation both as a supported tool on the Responses API and as not supported on this model. If generation is genuinely reachable through this model, the visual-production statement changes and UID would move to Medium. UDA rates conservatively until URC rules.
Tool to Skill to Credential Chains
Tool skill | U365 competency | Credential anchor (verified published programme) | Institute | Stacks into |
Running a bounded agentic change in a real repository at a chosen reasoning effort: stating the finish line, reading the report before the diff, running the project's tests independently, and treating an unreadable diff as a task that was too large | AI-assisted software engineering at a controlled effort level | AI Developer Specialist (18 days, diploma) | UIT (Technology, AI, Data Science) | Bachelor of Science in IT (B.Sc.), then Master of Science in IT (M.Sc.) |
Configuring an existing AI pipeline for cost and correctness: choosing the endpoint for tool calls, setting reasoning effort deliberately, placing cache breakpoints, and checking the long-context billing threshold before a long run | Applied AI platform and cost engineering | Cloud Computing Specialist (30 days, diploma) | UIT (Technology, AI, Data Science) | Bachelor of Science in IT (B.Sc.), then Master of Science in IT (M.Sc.) |
Measuring cost per completed task and deciding on evidence: recording a baseline of attempts, human minutes and error rate, logging discarded runs, and comparing two effort levels or two models on the same task | AI cost and performance measurement | Data Scientist (60 days, diploma) | UIT (Technology, AI, Data Science) | Master of Science in IT (M.Sc.) |
Writing one scope line per agent before a run starts, keeping producer attribution, refusing self-approval, and settling dependencies before integration | Multi-agent workflow governance | Tech Leader (25 days, diploma) | UIT (Technology, AI, Data Science) | Bachelor of Science in IT (B.Sc.), then Master of Science in IT (M.Sc.) |
Building the business case for a delegated workflow: the baseline, the measured cost per completed task, the error rate after the change, and an explicit decision on whether the measured knowledge-work regression is acceptable for this process | AI unit economics and accountable automation governance | AI Business Specialist (18 days, diploma) | UIB (Business Management, Entrepreneurship) | Bachelor of Business Administration (B.B.A.), then Master of Business Administration (M.B.A.) |
The five anchors were verified live on 2026-09-23 against the Wix Online Programs catalogue for university-365.com using POST /online-programs/v3/programs/query. The catalogue returned 86 programmes, 79 published, unchanged from the 2026-09-22 measurement. Each anchor was read as a published programme with its duration and step count, and each public page was requested and returned HTTP 200 on the same date.
No micro-credential component title is asserted anywhere in this document. The UIT reconciliation of 2026-09-22 established that four component titles carried by earlier UDA alignment deliverables, including UIT AI Fundamentals MCC and UIT Applied AI Model Deployment MCC, were internal working names that never reached the catalogue. UDA does not use any of them here, and the chains anchor to programme names a Fellow can find and enrol in.
No access level is asserted for any individual programme. The catalogue read does not expose a per-programme access level, and this remains the rule UDA adopted on the Claude Opus 5.5 review. The rule is published; the per-programme level is not.
What UDA cannot verify and therefore does not claim. Whether a Fellow who completes a specialisation programme earns credit toward a named micro-credential inside it is not exposed by the catalogue. The two stacking columns above record the published programme hierarchy the catalogue does show: a specialisation diploma sits under the institute's degree path, and the degree programmes are open at the level the next subsection states. They are not a claim that one programme awards credit toward another.
University 365 has three academic access levels: DISCOVERY, INSIDER and SUPERHUMAN.
Micro-credential and specialised diploma programmes have Basic, Foundation and Expert levels. DISCOVERY Fellows can enrol in Basic-level programmes only. INSIDER Fellows can enrol in Basic and Foundation programmes. SUPERHUMAN Fellows can enrol in all of them.
University degree programmes have a single Expert level, and are open to SUPERHUMAN Fellows only. INSIDER and DISCOVERY Fellows cannot enrol in a degree programme without upgrading.
Because four of the five chains above stack into a Bachelor of Science in IT, a Master of Science in IT, a Bachelor of Business Administration or a Master of Business Administration, the degree outcome in four of the five chains is open to SUPERHUMAN Fellows only. UDE should state that consequence in one sentence and must not imply a degree pathway is available to every reader.
UDA maps no UIC or UID credential for GPT-6 Sol. Both institutes are rated Low to Medium and the tool builds neither institute's disciplinary competency. Creating a chain for a tool that does not build the competency would inflate the academic claim and mislead Fellows about where the skill is assessed. UIC and UID Fellows who need the underlying capability should take a relevant UIT programme as an elective or work with a UIT collaborator.
The chains apply only when the Fellow can state why the effort level they chose was the right one for that task, explain what the change does without reopening the conversation, reproduce the test evidence independently, and name what the run did not verify. Shipping a diff the model produced is not evidence of the Fellow's skill. The assessment artefact must include the Fellow's own finish line, the effort decision and its reason, the independent test record, and the rejected alternatives.
Deployment conditions on the Humics rating
UDA records these as deployment conditions rather than as score changes. Framework clause 4.2-a does not apply to this model as a generator, and the erosion condition becomes reachable the moment a deployment writes through a connected account in the person's own name.
Framework clause 4.2-a does not apply to the model as a generator, and erosion condition (a) becomes reachable the moment a deployment writes through a connected account in the person's own name. UDA restates its standing condition from the Claude Opus 5.5 review: no GPT-6 Sol deployment carries mail, messaging or Teams write access without a read-before-send gate. This is a U365 deployment condition, not a score change.
The recorded position of -1 is the honest position for supervised use. It is worse than -1 wherever an artefact produced by this model reaches a Fellow, a client, or an external reader without a named human reading it against the evidence. UDA states this as a condition on use rather than as a proposed score, because the framework rates the common case and the common case here does include supervision.
How GPT-6 Sol Works
Inputs: Text prompts, documents, images, code files, and structured API requests. Through the API, a conversation history plus tool definitions. Through Codex, a repository, its instruction files, and any skills you have written.
Outputs: Text, code, structured tool calls, and through Codex, edits to files. Up to 128,000 tokens per response.
Underlying technology:
Model: GPT-6 Sol, API identifier gpt-6-sol. Mid-tier of the GPT-6 family below GPT-6 Astra. OpenAI states Sol and Luna were trained with methods similar to Astra's.
Reasoning effort across six levels: none, low, medium (default), high, xhigh, max. Responses and Chat Completions are both supported as endpoints, but built-in tools and function calling should use the Responses API. Chat Completions supports function calling only when reasoning effort is none.
Supported tools on the Responses API: web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. Each carries its own per-call fee where applicable.
Prompt caching: eligible shared prefixes reused within a 30-minute window receive the discount. Cached input reads are billed at 10 percent of the uncached input rate. Explicit cache breakpoints let you choose which prompt prefixes are cached. Changing reasoning effort or tool availability mid-conversation no longer breaks reuse.
Long context: prompts above 272,000 input tokens are billed at 2 times the input and cache rates and 1.5 times the output rate for the whole request, rather than only for the tokens above the threshold.
Not supported on this model: fine-tuning, embeddings, image generation as an endpoint, video, speech, transcription, translation, moderation, and the Assistants API.
Delivery surface, checked separately from the model: Codex CLI and the ChatGPT desktop app. That surface writes local memory files and runs parallel subagents, which is the basis of both clause findings in Section 7. The model endpoint itself holds no state between calls.
Integrations: OpenAI API, ChatGPT Work, Codex (CLI, IDE extension, web, iOS, cloud tasks), Amazon Bedrock, MCP servers, and connectors available on paid plans.
Benchmark figures, OpenAI-published. OpenAI published selected comparisons rather than a full table this time. Its footnotes state that competitor scores come from public reports, that Claude Fable 5 stands in where Fable 5.1 was unavailable, and that evaluations ran in OpenAI's research environment rather than production ChatGPT.
Benchmark | What it measures | OpenAI's figure for GPT-6 Sol | Comparison OpenAI publishes |
AutomationBench 1.0.6 | Business workflows across 47 tools | 33.2% at xhigh, $0.27 per task | Astra at low 30.3% at 3.9 times the cost; Claude Opus 5 at max 26.9% at 11.1 times the cost; Fable 5.1 with Opus 5 fallback at max 31.4% at more than 8.9 times the cost |
Agents' Last Exam V1 | Long-horizon professional tasks across 55 sub-industries | 56.4% at max | Above Claude Opus 5's highest score for about 60 percent less per task. Opus 5's effort level is not stated. |
DeepSWE v1.1 | Software engineering in real codebases | 68.8% at max | Claude Fable 5 at xhigh 69.9 percent for about 80 percent more per task |
OSWorld 2.0 offline | Computer use, partial reward | 60.5% at xhigh | Claude Opus 5 at medium 60.3 percent for about 80 percent more per task |
FrontierCode 1.1 Main | Merge-ready code, graded on test quality, scope discipline, and codebase standards | 49.3% at max, $2.14 per task | Matches Fable 5.1. No score is published. |
Factuality, internal | De-identified conversations where users flagged a mistake | 5.1% error rate at high | 10.8% for GPT-5.6 Sol at the same setting |
Coding deception, internal | Answers containing detected deception under adversarial prompts | 1.3% | 10.4% for GPT-5.6 Sol |
Warning circumvention, internal | Attempts to work around an explicit access-denied message | 64.4% | 68.2% for GPT-5.6 Sol |
Read the method before the numbers, in five steps.
First, this is OpenAI's selection, not a full table. Analysts noted that GDPval and Terminal-Bench 4.0 are absent from this launch even though they are part of OpenAI's usual set. That absence matters because Artificial Analysis measured a regression on GDPval and an improvement on Terminal-Bench.
Second, the FrontierCode comparison is made against a competitor's weakest setting. OpenAI's sentence is that Sol matches Fable 5.1 at xhigh at much lower cost. On the chart OpenAI itself cites, xhigh is Fable 5.1's lowest score: Fable 5.1 at low scores 49.8 percent for $2.38, and Claude Opus 5 at medium holds the chart's top score at 53.4 percent. Sol's real advantage on that benchmark is cost against its own predecessor, not capability against the leader.
Third, the competitor set stops at Claude Opus 5 and Claude Fable 5.1. Claude Opus 5.5 was released by Anthropic about 90 minutes before this launch, at $4 and $20 per million, and it appears nowhere in OpenAI's comparisons. Two vendors published launch tables the same evening, each omitting the other's newest model.
Fourth, one comparison understates the competitor on purpose and OpenAI says so: the Fable 5.1 datapoint on AutomationBench omits the cost of the Opus 5 fallback runs that occurred on about 40 percent of tasks.
Fifth, the alignment rows come from deliberately adversarial setups and OpenAI states the tests do not measure failure rates in typical use. The warning-circumvention row is the one to read closely: Sol still attempted to work around an explicit access denial in 64.4 percent of unmitigated runs. That is an improvement, and it is not compliance.
Independent benchmark results:
Source | Method | Result |
Artificial Analysis | Intelligence Index v4.3.2, maximum effort | 48, against 47 for GPT-5.6 Sol, 53 for GPT-6 Astra, and 58 for Claude Opus 5.5. Sol gains 1 point on its predecessor and stays a clear step below both flagships. |
Artificial Analysis | Coding Agent Index, maximum effort, in OpenAI's Codex environment | 57, up 2 points from GPT-5.6 Sol. Terminal-Bench 4.0 moved 37 percent to 43 percent, and SWE-Atlas-QnA 54 percent to 58 percent. |
Artificial Analysis | Cost per Intelligence Index task | $1.06 against $1.99 for GPT-5.6 Sol, about 50 percent less, while using slightly more output tokens per task, 31,000 against 29,000. |
Artificial Analysis | AA-Omniscience, knowledge and hallucination | Hallucination rate 60 percent, down from 92 percent. The model attempted 83 percent of questions against 99 percent before, and accuracy on what it answered fell from 59 percent to 54 percent. The index moved from 22 to 27. |
Artificial Analysis | GDPval-AA v2.1, economically valuable tasks across 44 occupations, Elo | About 100 Elo points below GPT-5.6 Sol, 1,487 against 1,588. Its team reports the regressions were driven mainly by reduced presentation quality and incomplete results after manually inspecting hundreds of outputs. |
Artificial Analysis | AutomationBench-AA, its own adaptation | 62 percent against 60 percent for GPT-5.6 Sol. |
Artificial Analysis | Terminal-Bench 4.0, its own run | 44 percent against 40 percent for GPT-5.6 Sol. |

Available platforms: ChatGPT Work and Codex (web, CLI, IDE extension, iOS, cloud tasks), the OpenAI API through the Responses, Chat Completions, and Batch endpoints, and Amazon Bedrock. Not open weights, and not available for local deployment.
Getting Started
Required accounts: A paid ChatGPT plan for ChatGPT Work and Codex: Plus, Pro, Business, Enterprise, or Edu. Enterprise and Edu administrators must enable the model for the workspace before it appears to users. For the API, an OpenAI account with billing. No account is required to read the documentation. Free and Go users do not get Sol in the desktop app; they get Luna.
Installation: Nothing to install for ChatGPT Work in the browser or the app. Codex is available as a CLI, an IDE extension, a web surface, and cloud tasks. For API access, nothing beyond an HTTP client or one of the official SDKs.
First-time configuration:
In ChatGPT Work or Codex, open the model picker and select GPT-6 Sol. If it is not there, the rollout is gradual and OpenAI's own advice is to try again later.
Set the effort level deliberately and record what you chose. The default is medium. OpenAI's published rows are mostly at xhigh and max, so the default is not the setting the launch numbers describe.
On AutomationBench, OpenAI's chart shows max scoring lower than xhigh while costing 24 percent more. On DeepSWE, max adds 2.2 points over xhigh for 2.7 times the cost. Start at xhigh and raise it only where a retry would cost more than the extra effort.
Read the migration note before pointing existing API code at gpt-6-sol. Built-in tools and function calling should use the Responses API. On Chat Completions, function calling works only when reasoning effort is none. A pipeline running GPT-5.6 Sol at medium or high with tools cannot swap the model identifier and keep its behaviour.
If you run agents, turn on caching deliberately. Cached input reads cost 10 percent of the input rate, and OpenAI now discounts shared prefixes reused within 30 minutes. Explicit cache breakpoints let you choose where a cached prefix ends. Use the caching dashboard and the diagnostics tool rather than guessing which prefix was reused.
Check the 272,000-token threshold before a long run. Above it, the whole request is billed at 2 times the input and cache rates and 1.5 times the output rate, rather than only the excess.
In Codex, review your local memory settings. OpenAI's documentation states that local Codex memories are off by default and that turning them on is a deliberate step, while the same documentation set and the desktop app present the feature as something you manage from Settings. Decide where you stand rather than discovering it later. Section 7 explains why this matters.
First 15 minutes checklist:
☐ Give the model one real task from your own work in a single message, with a stated finish line and a stated stopping rule.
☐ Confirm which effort level you are on, then try xhigh and compare the answer and the wait against medium.
☐ Take one factual claim from the answer and verify it against a source outside the conversation. Then note whether the model said what it had and had not checked.
☐ Read the closing report and answer only the part waiting on you.
☐ If you use the API with tools, confirm which endpoint you are on and what effort level your tool-calling turns are running at.
Result: One real task completed end to end, a deliberate effort setting, one verified claim, and a known endpoint configuration. That is a working setup rather than an impression.
Real Workflows
Workflow 1: A recurring engineering task at a stated effort level
Learner type: Professional (career upskilling). Developer, technical lead, or platform engineer. CI-First benefit tags: Time, Quantity, Quality. Connects to: UIT engineering tracks. UDA to confirm the programme mapping. Time estimate: Thirty to ninety minutes of supervised work, most of it the model's runtime, for a scoped change with a stated finish line.
What you do vs what the tool does:
Step | You do | The tool does |
1 | State the finish line, the scope, and the stops. Decide the effort level before you start. | (Nothing yet) |
2 | Give it the repository and the instruction file. | Reads the instruction chain, then works through the code and runs the tests. |
3 | Read the report before reading the diff. | Returns what it changed, what it found, and what it needs from you. |
4 | Check the evidence behind each finding, not the claim. | Produces its evidence on request. |
5 | Run the test suite and static analysis yourself. | (Nothing. You verify.) |
6 | If the diff is too large to read, treat the task as too large and split it. | (Nothing. You decide.) |
Sample prompt (UP-Context method: context, task, constraints, output format):
Context: this repository is [project]. The work is [the change or the defect]. Done means [the tests pass / the behaviour is reproduced and fixed]. The relevant code is in [paths]. Task: complete the work above. Constraints: one module at a time. Do not change behaviour outside [scope]. Do not add dependencies. Stop and ask only when you cannot continue, or before anything destructive. Do not ask me to confirm steps that do not need a decision. Output format: a short table with file, change, and the evidence you used. Then three headings: Blocked on me, Changed, Found. State plainly anything you did not verify.
Verification checklist:
☐ Multi-Model Check: run the same task on GPT-6 Sol at a different effort level, or on Astra, and compare. A difference in what the two found is more informative than a difference in how the answer reads.
☐ External Source: run the project's own test suite and static analysis. Do not accept the model's statement that it works.
☐ Human Review: read the diff. If it is too large to read, the task was too large.
☐ CI-First Test: can you explain and defend the change without the model? If not, it is not ready.
Workflow 2: Business workflow automation with a cost per task you can measure
Learner type: Professional. Operations, finance, or support lead. CI-First benefit tags: Time, Quantity. Connects to: UIB (Business Management, Entrepreneurship). UDA to confirm the programme mapping. Time estimate: Half a day for the first workflow, including the measurement design.
What you do vs what the tool does:
Step | You do | The tool does |
1 | Pick one workflow that crosses at least three applications and has a checkable end state. | (Nothing yet) |
2 | Define the check. What does a correct result look like, and who confirms it? | (Nothing yet) |
3 | Record the baseline: attempts, human minutes, and error rate without the model. | (Nothing yet) |
4 | Run the workflow and log every attempt, including the ones you discard. | Calls the tools, completes the steps, and reports what it did. |
5 | Compute cost per completed task, not cost per token. | (Nothing. You measure.) |
6 | Keep the workflow only if the completed-task cost and the error rate both improved. | (Nothing. You decide.) |
Sample prompt:
Context: [workflow name] runs in [applications]. It currently takes [time] and about [error rate] of runs need a human fix. A correct result is [definition]. Task: complete this workflow end to end for the [N] items I have attached. Constraints: use only the tools listed. Do not invent missing fields, and leave them blank instead. Stop before anything that sends a message to a person outside the team. Output format: one row per item with status, the tools you called, and any field you left blank. Then a list of items you could not complete and why.
Verification checklist:
☐ Multi-Model Check: run the same batch on Luna and compare how many items each completed without a human fix.
☐ External Source: validate the output programmatically against the system of record. A well-formed result can still contain a wrong value.
☐ Human Review: have the workflow owner, not the person who built it, check a sample of accepted items.
☐ CI-First Test: can you describe the decision rules the workflow uses without reading the model's prompt? If not, the process is not documented.
Workflow 3: A first pass over more material than you could read
Learner type: Everyone (lifelong learners). Also students writing a literature review. CI-First benefit tags: Time, Quantity. Connects to: LIPS Collect and Review phases, MCC Research Methods. UDA to confirm the mapping. Time estimate: One hour for a first pass over 20 to 40 documents, plus the verification time you would have spent anyway.
What you do vs what the tool does:
Step | You do | The tool does |
1 | Define the question the first pass must answer, and what would make a document irrelevant. | (Nothing yet) |
2 | Supply only the documents you are willing to stand behind. | Reads the whole set in one window and works across it. |
3 | Ask for a claim-versus-source table rather than a summary. | Returns each claim with the document and passage it came from. |
4 | Spot-check at least three rows against the source documents. | (Nothing. You verify.) |
5 | Discard the rows that do not survive the spot check, and count how many did not. | (Nothing. You decide.) |
6 | Write the section yourself from the surviving rows. | (Nothing. This is the part that is yours.) |
Sample prompt:
Context: I am writing [length] on [topic]. The attached documents are the only sources. My expertise is [level]. Task: build a table of what these documents claim about [question]. Constraints: no information from outside the attachments. Do not state a number you cannot source. Where two documents disagree, show both rows. Output format: a table with claim, document, and the passage you relied on. Then a heading "Could not support" listing every claim you looked for and did not find.
Verification checklist:
☐ Multi-Model Check: run the same table request on Astra and compare the rows each produced. Rows that appear in one and not the other are where the uncertainty lives.
☐ External Source: read at least three cited passages in the original documents.
☐ Human Review: before you write, have someone read your argument, not the table.
☐ CI-First Test: can you defend any row you kept without reopening the model conversation? If not, you kept a row you do not own.
Strengths, Limits, and AI Imposture Risk
Strengths
CI-First Benefit | Strength | Evidence |
Time | Cost per completed task roughly halved without a capability loss on coding. | Artificial Analysis measured $1.06 per Intelligence Index task against $1.99, and the Coding Agent Index rose 2 points. |
Quantity | The same budget buys roughly twice the agent work, and the allowance in ChatGPT Work rose. | Independent cost-per-task halving, plus OpenAI's own published estimate moving from 10 to 100 local messages per five hours to 15 to 150 on Plus. |
Quality | Factual errors fell sharply, and the direction is confirmed independently. | OpenAI reports about half the mistakes; Artificial Analysis measured the hallucination rate falling from 92 percent to 60 percent. |
Skill | Marginal. The model is used for delegation rather than learning, and the delivery surface compounds that. | No teaching mode is claimed. Codex writes local memory files on the user's behalf, which Section 7 treats as a Skill Illusion vector. |
Limits
It is not OpenAI's best model, and OpenAI's own tables show it. GPT-5.6 Sol retains higher top scores on DeepSWE v1.1 and OSWorld 2.0, at roughly twice the cost per task.
Factual reliability improved partly by answering less. The model attempted 83 percent of questions on the knowledge benchmark against 99 percent for its predecessor, so its accuracy on what it did answer fell from 59 percent to 54 percent. A model that declines more is safer and not necessarily more knowledgeable.
Knowledge work regressed on independent measurement. Artificial Analysis measured about 100 Elo points lost against GPT-5.6 Sol on GDPval-AA v2.1, and its team attributes the loss to presentation quality and incomplete results.
Effort level is now the largest single variable in the bill, and the relationship is not monotonic. On one published chart max scores lower than xhigh and costs more.
The API migration is real work. Tools and function calling should move to the Responses API, and Chat Completions supports function calling only at effort none.
The vendor's benchmark selection omits benchmarks it publishes elsewhere. GDPval and Terminal-Bench 4.0 are missing from this launch's table, and both are places where independent testing found movement.
The comparison set stops short of the newest competitor. Claude Opus 5.5 shipped about 90 minutes before this release and appears in none of OpenAI's comparisons.
No open weights and no self-hosting. Nothing about this model is auditable from outside.
AI Imposture Risk
Trap | Rating | Evidence |
Time Illusion | Medium | The savings are real and measured, and the overhead is real too. Effort selection is now a decision the user has to make on every hard task, the migration to the Responses API is unpaid work, and the cheaper tier invites running the task more times rather than once properly. Cost per completed task fell at max; it is not guaranteed to fall at the setting you actually choose. |
Quantity Illusion | Medium | Cheaper output raises the volume a person can produce, and the reading capacity of the person does not rise with it. The published knowledge-work regression was in presentation quality and completeness, which are the two things that are easiest to miss when you are accepting more volume than you can read. |
Skill Illusion | High | Two mechanisms. First, the model's capability exceeds a beginner's ability to check it on exactly the work it is sold for, agent runs in real repositories, so a run that finishes looks like competence. Second, and this is the framework's clause 5.2.3-a: the Codex surface writes local memory files on the user's behalf. See the clause note below. |
Overall Imposture Risk: Medium, with Time Medium, Quantity Medium and Skill Illusion High. UDA requires all four values to be published together, because the point of the trap table is that one row is worse than the summary number. One trap is High with identifiable mitigations and two are Medium. Under framework Section 5.3, one High with clear mitigations is Medium rather than High.
Framework v1.2 clause note
Clause 5.2.3-a, agent-authored procedural memory: APPLIES. This is the second review in the series where the clause applies rather than returning a null, and it applies more broadly than it did on Claude Opus 5.5.
The mechanism is Codex local memory. OpenAI's documentation states that Codex turns useful context from eligible prior chats into local memory files, that the main memory files live under the Codex home directory and include summaries, durable entries, recent inputs, and supporting evidence from prior chats, that the code skips short-lived sessions and redacts secrets, and that memory updates happen in the background rather than at the end of every chat. The documentation tells the user to treat those files as generated state rather than as a control surface.
The no-lower-than-Medium floor is met, because the user holds a documented capability they did not write. The High threshold is met because the writes happen during use without a per-write human decision, and because the documentation itself discourages the routine reading practice that would catch it. One further point belongs in the record rather than being smoothed over: OpenAI's documentation states that local Codex memories are off by default, while the same documentation set and the desktop settings present memory as a feature the user manages. Where the user has no routine practice of reading what was written, the clause's High threshold is met, and that is where Skill Illusion High comes from.
Clause 7.5, team-level rooms: APPLIES to the delivery surface. OpenAI's subagent documentation states that current Codex releases enable subagent workflows by default, that the app surfaces each subagent thread so the user can inspect its work and the summary returned to the main chat, and that agents are given their own model and reasoning configuration. That is several agents working in one view alongside the human, which is the case clause 7.5 governs. It sets the consequence directly: each agent needs a written task boundary before it starts, the human reviews output per agent, and Cyborg is not available. The vendor's own guidance that read-heavy work suits parallel agents and parallel write-heavy work creates conflicts is the same point from the engineering side.
Clause 4.2-a, agent-mediated conversation: NULL, does not apply to the model as a generator. The live condition to watch on any deployment: erosion condition (a) would be met by a setup where an agent sends text in the person's own name through a connected account and the person does not read it first. That is a property of the connector, not of this model.
Recorded together: one clause applied to the model's delivery surface, one clause applied to its agent coordination, and one null. That mix is a finding, and it is a different mix from the Claude Opus 5.5 review, where 7.5 was recorded as applying only through the delivery surface rather than as an enabled-by-default behaviour.

U365 Co-Intelligence Rating
CI-First Profile
Primary profile: Co-Worker and Assistant (level 2). Secondary profiles: Analyst and Tester (level 4), Coach and Tutor (level 3).
Divergence from the sibling, stated rather than left to be noticed. The published GPT-5.6 Sol review records Co-Creator and Thought Partner (level 1) as the primary profile. This review moves the primary to level 2. The reason is the vendor's own positioning: OpenAI's model reference says GPT-6 Sol "is built for complex coding and agentic workflows", the launch page frames it as the model for recurring complex tasks, and Astra is explicitly named as the model to choose "when you want the best results and an uncompromising experience". That is delegation with review rather than co-creation. Analyst and Tester is recorded as secondary because the published coding and knowledge-work evaluations are first-pass and triage work.
Collaboration Mode
Recommended mode: Centaur. Alternative mode: Cyborg is not available where more than one agent shares a view or a channel, per clause 7.5. It is permitted for a single supervised agent session with a stopping criterion set before the run starts. Mode rationale: The framework's Section 7.2 rule applies directly: a tool whose Imposture Risk is Medium or High takes Centaur mode, because Centaur is safer. Here the case is stronger than the general rule. Codex runs parallel subagents by default and shows them in one view with the human, and Codex writes local memory files between sessions. Both of those move work out of a single supervised thread. Centaur mode is what keeps a written task boundary and a per-run review in place, and clause 7.5 removes Cyborg as an option once more than one agent is in the loop.
CI-First Benefit Score
Dimension | Score (0-10) | Rationale |
Time | 7 | Cost per completed task roughly halved on independent measurement, and coding improved. Offset by the effort-selection decision the user now has to make and by the API migration. |
Quantity | 7 | The same budget buys roughly twice the agent work and the ChatGPT allowance rose. The usable-output check holds it at 7 rather than higher: some of the extra volume arrives with lower presentation quality. |
Quality | 7 | Factual errors fell sharply and coding improved on independent measurement. Held back by the knowledge-work regression, by the accuracy drop on answered questions, and by the fact that capability is level with the model it replaces rather than above it. |
Skill | 3 | Delegation, not learning. The delivery surface writes memory the user did not author, and the framework requires that to be scored conservatively. |
CI-First Benefit Score: 6.0 / 10 (CI-First Positive)
The score does not move, and that is the finding
6.0 with sub-scores 7 / 7 / 7 / 3 is identical to the published GPT-5.6 Sol review at every dimension. That is not an oversight and it is not a failure to update.
The framework's Section 9.2 says to score the honest user, the net benefit rather than the gross benefit, the common case rather than the best case, and the user rather than the tool. GPT-6 Sol halves the price, raises factual reliability, improves the coding evaluation, and improves caching. None of that changes what the person using it has to do: state a finish line, choose an effort level, supervise, verify, and decide. A release that lowers the cost line does not move the human's capability line.
The framework's interpretation bands make the consequence explicit. GPT-6 Sol scores 6.0, which is CI-First Positive rather than CI-First Strong, while GPT-6 Astra scores 7.0 on the published review. A reader who wants a Strong band tool is choosing Astra, and the review should say so rather than implying that the cheaper model delivers the same result at a lower price.
Humics Protection Badge
Dimension | Rating | Rationale |
Creativity | Neutral (0) | It drafts, restructures, and proposes approaches, and OpenAI's own launch page shows it working through a design change. It does not originate the direction, and it neither trains nor replaces the user's ideation. Same rating as GPT-6 Astra and GPT-5.6 Sol. |
Critical Thinking | Erodes (-1) | The model is sold for work whose output is long and whose correctness is hard to check without doing the work. Its own reliability improvement partly takes the form of declining to answer, which a reader can mistake for confidence. The published warning-circumvention figure, 64.4 percent of runs attempting to work around an explicit access denial, is the clearest single signal that the model's compliance is not something a user should reason about by assumption. |
Social Authenticity | Neutral (0) | The model produces text and code rather than speaking in the user's name. OpenAI has deliberately changed its communication style toward shorter, less jargon-heavy answers, which affects the voice of the output and not the voice of the person. |
Humics Protection Score: -1 / +3 Badge: Humics-Neutral
Superhuman Usage Guidance
When to invite this tool:
Recurring complex engineering work with a checkable end state: feature work, debugging, review passes, migrations.
Agent runs that call tools and work in the background, where cost per completed task is the number that matters.
Business workflow automation across several applications, with a defined check on the result.
A first pass over more material than you could read, where every claim comes back tied to a source you supplied.
When to keep this tool out:
Work where the output cannot be checked by someone who did not do it. The model is strong enough to produce a plausible result and not reliable enough for you to skip the check.
The final judgement calls: what to publish, what to send to a client, what to tell a person. Those are the Humics this model does not supply.
Long unattended runs on a repository you cannot review afterwards. If the diff is too large to read, the task was too large, whatever the model's score says.
Environments where you need an auditable artifact. There are no weights and no way to inspect what the model does internally.
U365 method integration:
LIPS + CARE: route the model's outputs into the Collect phase, and keep the Action Plan and Review phases as your own work. The model can produce the collection; it cannot own the decision about what the collection means.
ULM + EVA: relevant to Career, through the cost-per-completed-task framing, which is a professional skill as much as a technical one. Weak fit for Body, Spirit, Social, and Quality of Life.
UP-Context: it responds well to explicit context, a stated role, a task, constraints, and a named output format. The workflows in Section 6 use that order.
SL-OS: usable as a workhorse inside an SL-OS automation layer, with the same caution as any delegated execution step: the check stays on your side of the boundary.
UNOP: no direct fit. The model is not built to teach, and its reliability gains come partly from declining to answer, which is not a pedagogy.
Over-delegation warning: the failure mode with this model is volume. At $2 and $10 per million tokens, running the task three times costs what running it once used to. That is useful when the runs are compared and dangerous when they are merely accumulated. The same arithmetic applies to agents: OpenAI enables parallel subagents by default, and more agents produce more output, not more verified output. The CI-First formula is the test. If your Human Intelligence input drops while the Artificial Intelligence term rises, the product falls. A reader who accepts more output without reading more of it has moved in that direction, and the published knowledge-work regression, which showed up as presentation quality and incomplete results, is exactly the kind of defect you find by reading and not by counting.

What Users Say
Aggregate Rating Table
Platform | Rating | Number of reviews | Link |
Hacker News | 1,337 points, 651 comments on the launch thread, read 2026-09-23 | Not a rating platform | |
Artificial Analysis | Intelligence Index 48 at maximum effort, Coding Agent Index 57, $1.06 per index task | Independent measurement, not user reviews | |
G2 | No model-level rating for GPT-6 Sol. G2 rates ChatGPT as a product and its pages return HTTP 403 to automated retrieval. No figure is reproduced here. | Not applicable | |
Capterra | No model-level rating found. | Not applicable | |
Trustpilot | No model-level rating. OpenAI is rated as a company, not per model. Page returns HTTP 403 to automated retrieval. | Not applicable | |
Product Hunt | ChatGPT, the product, has a listing. GPT-6 Sol has no separate listing. | Not applicable | |
Apple App Store | The ChatGPT app is rated in the region of 4.7 out of 5 from a very large number of ratings. This is the consumer product across all models and is not a rating of GPT-6 Sol. | Not model-specific | |
Mixed to positive. r/singularity and r/codex discussed the release on launch day. Reddit returns HTTP 403 to automated retrieval, so threads were read through search indexing rather than fetched. | Several threads |
Note on method, stated plainly because it affects how much this section is worth. No review platform rates an individual language model. Every aggregate score found covers the ChatGPT product or OpenAI the company, and third-party aggregators disagree with each other. Reproducing any of those numbers as a rating for GPT-6 Sol would be fabrication. The meaningful signals for a model released the previous day are the developer discussion, the independent benchmark platforms, and the practitioners who ran it themselves. Those are what this section reports.
Three sources cited in this review refuse automated retrieval: G2, Trustpilot, Capterra, the Reddit subreddit pages, and the openai.com/index/... marketing pages all return HTTP 403 to a scripted request. Their substance was reached through search indexing and through secondary reporting that quoted them. The Reddit and G2 entries above are named and linked to the community or platform root rather than to a specific thread, because a specific thread URL could not be verified from this environment. Nothing in this review states a number that came from a page it could not read.
What Users Praise
Three themes dominate the launch discussion.
The first is price against capability, and it is the theme practitioners reproduce themselves. The developer comment that carried the most weight on the launch thread was simply that half the price for the same or slightly better model is a large change, and a widely read comment argued that the cheaper tier now sits on the cost-efficiency frontier for most tasks, which makes routing routine work away from the flagship an easy decision. That point was made about Luna rather than Sol, and the same logic applies one tier up.
The second is token efficiency in real sessions. Several practitioners reported that Sol finishes the same work in a fraction of the time with fewer tokens than the previous flagship, including on agent runs. This is a distinct claim from the price cut: it is about how much work the model does rather than what each unit costs, and it is the claim that would matter most if it holds on your own tasks.
The third is the factual-reliability improvement, which shows up in the launch material as a headline and in the discussion as something users noticed without being told. Independent measurement supports the direction and qualifies the mechanism, as Section 4 records.
What Users Complain About
Four complaints recur.
The first is that the release moves the price and not the frontier. Reddit discussion of the independent benchmark results was split, with one widely read thread treating the same-tier scores as a disappointment and the prevailing reply being that cost reduction was the point of the release. Both positions are in the record, and the second is closer to what OpenAI published.
The second is that the model line-up is now hard to reason about. Sol at xhigh, Sol at max, Luna at max, and GPT-5.6 Sol at max land close enough on coding benchmarks that the choice is not obvious from the numbers, while their costs differ by a factor of five or more. Analysts made the same complaint in writing, and it is the complaint this review's Section 5 addresses.
The third is the API migration. Moving tools and function calling to the Responses API is required for anything beyond none effort on Chat Completions, and developers maintaining agent pipelines describe that as more work than a version bump.
The fourth is the vendor benchmark selection. Practitioners noticed that the launch table omits benchmarks OpenAI publishes elsewhere, and that the comparison set stops at Claude Opus 5 and Fable 5.1 while a newer Anthropic model shipped the same evening.
A fifth objection comes from the independent-lab side rather than from users: Artificial Analysis has previously been criticised for scoring GPT-6 Astra too low on its then-current benchmarks, and revised its index afterwards. The critique is fair to note and does not by itself invalidate this measurement; it does mean the knowledge-work regression should be re-checked on the next index revision rather than treated as settled.
Sentiment Summary
Overall sentiment: Predominantly positive on cost and reliability, with the same-tier capability result contested.
Key themes:
The price reduction is the substance of the release and it is confirmed independently.
Independent measurement shows the cost per task roughly halving and the capability staying level with the model it replaces.
Coding improved and knowledge work regressed, and both are in the independent data.
Factual reliability improved, partly by answering fewer questions.
The API migration to the Responses API is the most common practical complaint.
The choice between Sol, Luna, Astra, and GPT-5.6 Sol at various effort levels is harder than the launch material implies.
U365 Editorial Note
User sentiment and the CI-First evaluation agree on the two things that matter and diverge on the one that decides adoption.
They agree on cost. Users report the price change as the point of the release, and the framework scores Time at 7 because the cost per completed task roughly halved on independent measurement. They agree on the migration friction. Developers complain about moving to the Responses API, and the framework docks Time for exactly that reason rather than treating the price cut as free.
Where they diverge is the question a U365 reader brings. Developers evaluate this model as a component and mostly need it cheaper and equally reliable, and by that standard it is a good release that does what it says. The framework asks a different question: does this make you more capable? The measured answer here is no, and the arithmetic is blunt about it. Every dimension of the CI-First Benefit Score is identical to the published GPT-5.6 Sol review, and the Skill dimension stays at 3 while both clause 5.2.3-a and clause 7.5 apply to the delivery surface. GPT-6 Sol makes a capable person's engineering budget go further. It does not make a less capable person more capable, and the second half of that sentence is the one this review exists to say.
Comparison and Alternatives
Alternative | Choose the alternative if... | Choose GPT-6 Sol if... |
GPT-6 Astra (https://openai.com/index/gpt-6-astra/) | The outcome justifies five times the rate. Astra leads on OpenAI's own AutomationBench and OSWorld charts, and on the independent Intelligence Index at 53 against Sol's 48. It is the CI-First Strong tool in this family at 7.0. | You want most of the capability at a fifth of the price for recurring work. OpenAI's own framing is that Astra is for the uncompromising case and Sol is for the work that happens every day. |
The task is narrow, repeated at volume, and checkable by a validator: extraction, classification, summarization, routing, first-pass triage. Luna at maximum effort matches Sol at xhigh on the published DeepSWE chart for about a fifth of the cost. | The task is multi-step and the cost of an incomplete result is high, or the workflow needs a tool loop that runs longer than a few steps. Luna's published Terminal-Bench 4.0 result is 13 percent against Sol's 43 percent. | |
Claude Opus 5.5 (https://www.anthropic.com/claude-opus-5-5) | You want the top of the independent composite. Artificial Analysis scores it 58 against Sol's 48, released about 90 minutes apart on the same evening. | Price per token is the binding constraint and Sol's $2 and $10 fits where Opus 5.5's $4 and $20 does not. Do not treat this as a controlled comparison: no one has run both in one evaluation setup on cost per completed task. |
You are already on it and it meets your bar. In OpenAI's own tables it still holds higher top scores on DeepSWE v1.1 and OSWorld 2.0 than GPT-6 Sol. Its status is Active and its review is unchanged. | You want the price cut and the factuality improvement, and you accept that two of the vendor's own benchmark rows go the other way. | |
Claude Fable 5.1 (https://www.anthropic.com/claude-opus-5-5) | You need the tier Anthropic recommends for the most demanding long-horizon reasoning, and you accept $10 and $50 per million with fallback runs on some tasks. | The published coding comparison is within about one point and the cost difference is roughly five to six times. Read that comparison with its footnote: OpenAI matched against Fable 5.1's weakest effort setting on the chart it cites. |
Gemini 3.8 Flash (https://ai.google.dev) | Throughput dominates and the introductory input rate fits the budget. Note that the rate doubles on 2027-01-01. | The work needs the top of the quality range rather than the top of the volume range. |
Grok 4.7 (https://docs.x.ai) | You want a low list price under 200,000 prompt tokens and your work fits that window. | You need a 1M-token context at flat pricing, which Grok charges more for above 200K. |
A smaller or open-weights model, for example Kimi K3 (https://openrouter.ai/moonshotai/kimi-k3) or a local model via ollama.com/search | Data cannot leave your infrastructure, or you need an auditable artifact. GPT-6 Sol cannot be self-hosted or audited. | The task needs frontier-adjacent capability at volume, and the cost per completed task is lower than the cheap model's retries. |
Where GPT-6 Sol is clearly better: on cost per completed task in the mid tier, with the strongest same-day independent measurement in this release cycle behind it. Artificial Analysis measured $1.06 per Intelligence Index task against $1.99 for its predecessor while the coding index rose 2 points, and it measured the same halving for Luna. If your constraint is budget for recurring complex work and your verification is a test suite or a validator, this is the strongest position in OpenAI's line-up.
Where GPT-6 Sol is clearly worse: on the top of the capability range, where Claude Opus 5.5 scores 58 against 48 on the independent composite and GPT-6 Astra scores 53; on two of OpenAI's own published benchmark rows, where GPT-5.6 Sol still leads; on knowledge work, where independent measurement found about a 100-Elo regression against its predecessor; on open auditability, since there are no weights and no self-hosting; and on pipeline stability, because tool-calling pipelines on Chat Completions must move endpoints to keep reasoning. If a claim will be relied on without your own verification, this is not the component to rely on.
Verdict and Next Steps
Who should adopt it: Individuals and teams doing recurring, checkable work at volume: engineering, code review, debugging, migrations, workflow automation, and research synthesis from sources you control. Also anyone already paying for a ChatGPT plan whose allowance was the binding constraint, because the published estimates for the same plan roughly doubled.
When: Now, with three conditions. In ChatGPT Work and Codex this is a straightforward improvement and the effort setting is the main thing to set deliberately. If you call the API from production code, read the endpoint requirement first, because a tool-calling pipeline on Chat Completions at medium or higher effort cannot keep its behaviour by changing the model identifier. If your work is professional document production rather than engineering, test it against your own material before routing to it, because the independent knowledge-work regression is real and it showed up as presentation quality and completeness.
For what: Recurring complex engineering with a stated finish line, agent runs where cost per completed task matters, workflow automation with a defined check, and a first pass over more material than you could read.
The honest caveat, stated once: the price reduction is real and independently confirmed, the coding improvement is real and independently confirmed, and the CI-First score does not move from the model it replaces at any dimension. That is not a contradiction. The framework measures what the human gains, and GPT-6 Sol changes the cost line rather than the capability line. It also carries two structural findings that its predecessor's review recorded as null: the Codex surface writes memory files on the user's behalf, and it runs parallel agents by default. The tool is designed to make more work affordable. It is not designed to make verification unnecessary.
Related U365 content
INSIDE Tools Review: GPT-6 Astra (frontier tier of the same generation, scored 7.0)
INSIDE Tools Review: GPT-5.6 Sol (the model GPT-6 Sol replaces, scored 6.0)
INSIDE Tools Review: GPT-5.6 Luna (cheap tier of the previous generation, scored 4.8)
INSIDE Tools Review: GPT-5.6 Terra (the balanced tier, which has no GPT-6 successor)
INSIDE Tools Review: Claude Opus 5.5 (the same-evening competitor, scored 6.5)
[UDA to supply the U365 course and programme links for the UIT and UIB rows]
Related U365 content
INSIDE Tools Review: GPT-6 Astra, the frontier tier of the same generation, scored 7.0. https://www.university-365.com/post/gpt-6-astra-openai-s-frontier-model-for-computer-use-coding-and-science
INSIDE Tools Review: GPT-5.6 Sol, the model GPT-6 Sol replaces, scored 6.0. https://www.university-365.com/post/gpt-5-6-sol-openai-s-balanced-frontier-model-scored-6-0-on-the-u365-ci-first-review
INSIDE Tools Review: Claude Opus 5.5, the same-evening competitor, scored 6.5. https://www.university-365.com/post/claude-opus-5-5-anthropic-s-fable-5-1-class-model-at-40-percent-lower-cost-scored-6-5-on-the-u365
U365 Institutes overview: https://university-365.com/uit
UP-Context prompt packs
URC drafted three prompts for this section and UDA, as the academic owner, replaces all three. The URC drafts kept the UP-Context order and were sound in substance, but each had the same three gaps: no Role line, no verification close, and no memory-write disclosure. The third follows directly from framework clause 5.2.3-a, because the Codex surface writes local memory files on the user's behalf. Each pack below follows the UP-Context Method in full: context, role, task, constraints, output format, and a verification step.
Prompt pack 1: a bounded agentic change with the memory write declared
Context: [repository or project]. The work is [the change]. Done means [the check that passes]. The relevant code is in [paths]. I am working at reasoning effort [level], and I chose that level because [reason]. The files you may change are [paths]. The files and interfaces you must not change are [boundaries]. Role: AI as Co-Worker and Assistant (Profile 2). You execute. I define the intent, choose the effort level, approve the plan, read the report before the diff, and decide whether the work ships. Task: Inspect the repository, name the affected files and error paths, and propose a plan. Do not edit until I approve the plan. Then work one bounded increment at a time and report after each one. Constraints: Do not change behaviour outside [scope]. Do not add dependencies. Do not weaken, remove or skip a test to make a change pass. Never report a test result you did not run. Stop and ask only when you cannot continue, or before anything destructive. Do not ask me to confirm steps that do not need a decision. Before you finish, list every file you wrote outside this conversation, including any local memory file, with its path, and state what it now says. Output format: a table with file, change and the evidence you used. Then four headings: Blocked on me, Changed, Found, Wrote outside this conversation. State plainly what you did not verify. UP-Context verification: I run the full test suite and the linters outside this session, I read the final diff end to end, and I open every file listed under Wrote outside this conversation before I accept the run. If the diff is too large for me to read, the task was too large and I split it. I ship only code I can explain and defend without you.
Prompt pack 2: a cost-per-completed-task comparison on your own work
Context: my recurring task is [task]. Today it takes [human minutes] and about [error rate] of runs need a human fix. A correct result is [definition]. The two settings I am comparing are [setting A] and [setting B]. Role: AI as Analyst and Tester (Profile 4) for the measurement, and Co-Worker and Assistant (Profile 2) for the execution. I design the baseline, keep the ledger, and decide. You run the task at each setting and report the numbers. Task: run [N] real items of this task at [setting A], then the same [N] items at [setting B]. For each item, report whether it completed, how many attempts it took, and what it cost. Constraints: report the true attempt count, including the attempts I discarded. Report input and output tokens separately, because the cheaper setting may use more output tokens. Do not average away an outlier, and do not report a per-token price when what I asked for is the cost of a completed item. Output format: one row per item per setting with completed, attempts and cost. Then a summary table with, per setting, completed items, total attempts, cost per completed item, and cost per item that needed a human fix. Then one heading: "Where this does not generalise". UP-Context verification: I keep the ledger and compute the totals myself from your rows. I run one setting a second time on a fresh set to see whether the result holds. I write the chosen setting down with the date and the reason, and I re-check it whenever the model version or my task changes. If cost per completed item did not fall at the setting I actually use, I say so and I do not adopt it.
Prompt pack 3: a first pass where every claim keeps its source and the losses are counted
Context: I am writing [length] on [topic]. The attached documents are the only sources allowed. My expertise is [level]. The question the first pass must answer is [question], and a document is irrelevant if [irrelevance test]. Role: AI as Co-Worker and Assistant (Profile 2) for the reading pass, and Analyst and Tester (Profile 4) for the evidence. You extract. I decide what the collection means, and the writing is mine. Task: build a claim-versus-source table answering [question] from the attached documents. Constraints: no information from outside the attachments. Do not state a number you cannot source to a passage. Where two documents disagree, show both rows instead of resolving them. Do not summarise: give the passage you relied on. Do not tell me a document was relevant when you found nothing in it. Output format: the table, with claim, document, and the passage relied on. Then a heading "Could not support" listing every claim I looked for and you did not find. Then a heading "Documents with no usable passage", listing each attachment with nothing relevant. UP-Context verification: I spot-check at least three rows in the original documents, I count how many rows did not survive the check, and I record that count. I write the section myself from the surviving rows. I could not have written it from your summary, because you did not give me one. If I cannot defend a row without reopening this conversation, I drop the row.
U.Copilot integration
When U.Copilot should route a Fellow to GPT-6 Sol.
Route to GPT-6 Sol when the Fellow needs to:
Work through a multi-file change, migration or defect in a real repository with the project's own tests as the check, at roughly half the previous cost per completed task
Run an agent that calls tools and works in the background, where cost per completed job is the number that matters rather than price per token
Configure an existing pipeline for cost and correctness: choose the endpoint for tool calls, set reasoning effort deliberately, place cache breakpoints, and check the long-context billing threshold before a long run
Measure cost per completed task against a recorded baseline, then decide on evidence
Run a first pass over more material than the Fellow could read, where every claim comes back tied to a passage they supplied
Draft and check a technical document where the model is expected to say what it did and did not verify
When U.Copilot should route a Fellow away.
Route away when the Fellow:
Cannot read the code or the diff well enough to verify the work. A finished run looks like competence, and that is the Skill Illusion this model carries at High.
Needs professional document production. Independent measurement found about 100 Elo lost on GDPval-AA v2.1 against the previous model, in presentation quality and completeness. Test it against your own material before routing that work.
Wants a final judgement call delegated: what to publish, what to send to a client, what to tell a person
Needs an auditable artefact. There are no weights and no way to inspect the model internally.
Is running tools or function calling through Chat Completions at medium or higher. Function calling works only at reasoning effort none there, so the pipeline must move endpoint or lose reasoning on tool-calling turns.
Has not decided what the run may write. Codex writes local memory files, the vendor documents them as generated state, and clause 5.2.3-a rates that High.
State the score beside the risk. 6.0/10 CI-First Positive with Medium AI Imposture Risk, and name Skill Illusion High rather than only the overall Medium. Do not present the benefit without the risk.
Never quote the 50 percent price cut without its base. The vendor's sentence compares against the previous model's promotional rate of $4 and $20. Against that model's list rate of $5 and $30, the cut is 60 percent on input and 67 percent on output. Both numbers are true.
Never present the factual-reliability improvement as a knowledge gain. The hallucination rate fell from 92 percent to 60 percent while the model attempted 83 percent of questions against 99 percent before, so accuracy on the questions it did answer fell from 59 percent to 54 percent. Declining to answer is safer; it is not knowing more.
Never state the score without saying it is unchanged. Every dimension matches the published GPT-5.6 Sol review. That is a deliberate finding, and the post must explain it rather than let a reader assume the review was not updated.
State the clause mix. Clause 5.2.3-a applies at High, clause 7.5 applies to the delivery surface, clause 4.2-a is a null with a live condition. All three, not one.
Require a read of the memory files. If the Fellow cannot say what the surface wrote on their behalf, the workflow fails the CI-First test.
Require a written task boundary per agent wherever subagents run, which is the default in current Codex releases.
Never describe generated code as the Fellow's acquired skill. Skill evidence is the Fellow's finish line, effort decision, test design and defence of the result.
Name the volume risk. At $2 and $10 per million tokens, running the task three times costs what running it once used to. That is useful when the runs are compared and it is dangerous when they are merely accumulated.
SL-OS integration
For every substantive GPT-6 Sol engagement, store under the relevant LIPS Project, or under Career and Finance for skill development, one record containing:
The human-written objective, the finish line, and the acceptance criteria
The effort level chosen, the date, and the reason it was chosen
The CI-First Profile and the Collaboration Mode for the session
Instructions in force: the instruction file, the skills, the hooks, the scripts, and any subagent scopes
Model identifier gpt-6-sol, the access route (ChatGPT Work, Codex, or the API), and the endpoint where the API was used
The plan the human approved
Artifact links, diffs, and the real command and test output, kept as output rather than as a claim
The memory-write record: every local memory file the surface wrote or revised during the session, its path, and the date the human read it back. This is the clause 5.2.3-a record, and it is the field most likely to be omitted and the one that matters most.
Cost and effort record, including rework and every discarded run
The verification result: the test suite or validator run by the human, the external source checked, and the spot-check outcome
The human's own explanation of what the change does
Rejected suggestions and defects found
Collect: save the objective, the instruction set, the plan, the diffs, the real command results, the effort level, the cost, and the raw review findings.
Action Plan: before the run, write the permitted scope, the finish line and stops, the effort level, the reviewer, the test evidence, the time and cost ceiling, and the memory boundary.
Review: read the report before the diff, run the tests independently, open every file the run wrote outside the conversation, and read the memory files back. Compare against a second effort level or a second model where the risk warrants it.
Execute: accept only the verified change, and record the final commit, the decision, the unresolved risk, the accepted risk on the measured knowledge-work regression, and the learning outcome.
Career and Finance is the primary domain. The cost-per-completed-task framing is a professional skill as much as a technical one, and the ability to read a vendor claim against its own footnote base transfers directly to a Fellow's own budget decisions.
Quality of Life is conditional. It improves where the tool removes a repetitive technical burden the Fellow already understands. It does not improve Quality of Life where it adds a supervision load the Fellow did not previously carry, which is the failure mode at volume.
UDA does not recommend this tool for Body and Health, Spirit and Mind, Character and Emotions, or Social and Love Relationships. Nothing in the release addresses those domains, and the review's own method integration records the same weak fit.
Within EVA:
Explore: use a first pass over supplied material to find what is there, with every claim tied to a passage and the losses counted.
Visualize: map the proposed change, the effort choice, the cost, the verification evidence, and the human decision points.
Action Plan: choose what the run may execute, what stays human, what the run may write, and when it stops.
[UIT](https://university-365.com/uit) software-engineering Fellows: two to three sessions per week of 45 to 90 minutes, each closing with a real test record and a short explanation written without the tool.
[UIT](https://university-365.com/uit) agent-governance Fellows: one weekly review of the local memory files and the subagent scopes actually used, reading what was written rather than assuming what was written. This is the cadence that keeps clause 5.2.3-a from being a standing High.
[UIB](https://university-365.com/uib) Fellows: one measured workflow case, then a monthly review of cost per completed task, rework, accepted risk, and owner accountability.
[UIC](https://university-365.com/uic) and [UID](https://university-365.com/uid) Fellows: on demand for approved implementation or content-operations work. No dedicated routine is warranted.
All Fellows: monthly review of standing instructions, hooks, scripts and memory files. Remove or rewrite anything the Fellow cannot explain.
GPT-6 Sol has no required native Microsoft 365 integration. Use this manual evidence workflow:
Export the objective, the effort decision, the diff summary, the real test output, and the memory-write record.
Store them in the relevant OneDrive or SharePoint Project folder with a link to the commit, the pull request, or the run.
Record the workflow configuration and the learning reflection in OneNote.
Request review through the relevant Teams channel.
Keep the cost and rejection log where the next technical decision will see it.
Never store repository credentials, API keys, access tokens, or unredacted secret-bearing logs in LIPS. Note that the vendor documentation states the memory code redacts secrets; a redaction claim is not a reason to place a secret in a prompt.
GPT-6 Sol amplifies the Execute phase of CARE and supports Review when it produces an inspectable diff, real test output, and a declared memory record. Its strongest SL-OS contribution is affordable, verified execution at volume. Its failure mode is unowned volume: more output than the Fellow reads, memory written on their behalf, and no stable evidence trail. The SL-OS rule is therefore the same as for any delegated execution step: no artefact without its objective, its effort decision, its producer, its reviewer, its test evidence, its memory record, and the human's decision.
U365's Recommendations to Learn More
These resources were curated to help you go deeper on GPT-6 Sol. Every link and every video below was resolved on 2026-09-23. We prioritise material that teaches something this review does not cover.
Official learning resources
Introducing GPT-6 Sol and Luna: the vendor's own claims, the cost-per-task charts, the alignment results, and the footnotes that qualify the competitor comparison. https://openai.com/index/introducing-gpt-6-sol-and-luna/
Better prompt caching for GPT-6: the caching changes that reduce the agent bill, including cache-safe effort and tool toggles and the diagnostics tool. Read this before optimising an agent pipeline. https://openai.com/index/better-prompt-caching-for-gpt-6/
GPT-6 Sol model reference: the endpoint support matrix, the supported tools, the rate limits, and the exact token and cache prices. https://developers.openai.com/api/docs/models/gpt-6-sol
GPT-6 Luna model reference: the same details for the cheaper tier, including the knowledge cutoff and the endpoint matrix. https://developers.openai.com/api/docs/models/gpt-6-luna
Codex local memories: how Codex generates memory files from prior chats, where they are stored, and how to control them per chat. This is the basis of this review's clause 5.2.3-a finding. https://learn.chatgpt.com/docs/customization/memories
Codex subagents: how parallel agents are spawned, how custom agents are configured, and the vendor's own warning about parallel write-heavy workflows. https://learn.chatgpt.com/docs/agent-configuration/subagents
Custom instructions with AGENTS.md: the instruction chain and its precedence order, which is what makes a Codex run reproducible. https://learn.chatgpt.com/docs/agent-configuration/agents-md
GPT-6 Astra system card: the alignment work Sol and Luna build on, including the metagaming and alignment-generalisation sections. https://deploymentsafety.openai.com/gpt-6-astra
Video tutorials and channels
GPT-6 Sol and Luna, Pricing and Benchmarks, by United Top Tech (Published Sep 22, 2026)
GPT-6 Sol and Luna, Pricing and Benchmarks, by United Top Tech. A straight walk through the published rate card and the benchmark rows against GPT-6 Astra, useful for the numbers without the marketing framing. https://www.youtube.com/watch?v=rZ2myXdJYdI
GPT-6 Sol, First impressions, by Arena AI. Runs GPT-6 Sol against GPT-5.6 Sol and compares total tokens used and wall-clock time rather than only the answer quality. https://www.youtube.com/watch?v=jVFz7cTeQxk
GPT-6 Luna vs Sol, I Tested Both on Real Work, by AI Words Explained. Gives both models the same four workplace tasks and reports where the cheap tier held up and where it did not. Directly relevant to the routing decision in Section 10. https://www.youtube.com/watch?v=0wyl2G1WU5Q
GPT-6 Sol vs Claude Opus 5.5 LIVE, Which AI Model Is Better?, by The Neuron. A same-session comparison of the two models that launched on the same evening, which is the head-to-head this review could not find in writing. https://www.youtube.com/watch?v=X0ERFFbjEug
I reviewed Opus 5.5 and GPT-6 Sol live, and the results surprised me, by How I AI. A blind-comparison format, useful as a check on how much benchmark differences are visible in practice. https://www.youtube.com/watch?v=LMT-bknLmNo
Claude Opus 5.5 vs GPT-6 Sol, Everything You Need to Know, by Universe of AI. A structured side-by-side covering pricing, benchmarks, and the practical differences. https://www.youtube.com/watch?v=vG2rNycYdQQ
GPT-5.6 Luna First Test, Hands-On With OpenAI's CHEAPEST Model, by Bijan Bowen. Covers the previous-generation cheap tier on browser workflows, C++ game creation, 3D CAD modelling, and frontend design, which is a reasonable preview of what the Luna tier is asked to do. https://www.youtube.com/watch?v=1nf7VqduM3Y
Evaluating the GPT-5.6 family, by Braintrust. An independent evaluation design across 225 tasks and six models, useful for the method as much as the result. https://www.youtube.com/watch?v=5jzUVjno-OQ
Written tutorials and deep-dive articles
Artificial Analysis, GPT-6 Sol and Luna push the cost efficiency frontier: the same-day independent measurement, including the per-task costs, the coding index movement in both directions, and the hallucination and knowledge-work findings this review relies on. https://artificialanalysis.ai/articles/gpt-6-sol-and-luna-push-the-cost-efficiency-frontier
Artificial Analysis, GPT-6 Sol release page: the intelligence and cost figures at every effort level. https://artificialanalysis.ai/models/releases/gpt-6-sol
The New Stack, OpenAI releases GPT-6 Sol and Luna and cuts token prices in half: a same-day read that correctly notes nobody had run Sol and Opus 5.5 head-to-head. https://thenewstack.io/openai-gpt-6-sol-luna-release/
The Decoder, OpenAI's GPT-6 Sol and Luna cut prices in half but barely move the needle on performance: the independent regressions, the missing benchmarks, and the reasoning-level sprawl. https://the-decoder.com/openais-gpt-6-sol-and-luna-cut-prices-in-half-but-barely-move-the-needle-on-performance/
VentureBeat, OpenAI releases GPT-6 Sol and Luna models, slashing API costs 50 percent or more: includes the vendor statement that the pricing is permanent rather than promotional. https://venturebeat.com/technology/openai-releases-gpt-6-sol-and-luna-models-slashing-api-costs-50-or-more
Coursiv, GPT-6 Sol and Luna, pricing, benchmarks and availability: a clearly labelled reading of the launch numbers that states what is not independently confirmed. https://coursiv.io/blog/gpt-6-sol-luna
Codersera, GPT-6 Sol and Luna complete guide: per-effort transcriptions of OpenAI's charts, which is where the non-monotonic effort behaviour becomes visible. https://codersera.com/blog/gpt-6-sol-luna-complete-guide-2026/
Apidog, What is GPT-6 Luna: a careful reading of the pricing fine print, including that the 50 percent comparison runs against a promotional rate and the list-price reduction is larger. https://apidog.com/blog/what-is-gpt-6-luna/
Community and social
Hacker News launch thread: 1,337 points and 651 comments, including developer reports on token efficiency and the cost-per-task argument. https://news.ycombinator.com/item?id=49805509
r/singularity discussion of the independent benchmark results: the argument over whether same-tier performance at half the price is the point or a disappointment. Read via search indexing because Reddit returns HTTP 403 to automated retrieval. https://www.reddit.com/r/singularity/
r/OpenAI: the largest general OpenAI community, where the model-routing question plays out. Same retrieval limitation. https://www.reddit.com/r/OpenAI/
r/codex: the agent surface this model is built for, where effort levels and Codex configuration are discussed. Same retrieval limitation. https://www.reddit.com/r/codex/
OpenAI status page: check this before concluding the model is behaving badly. https://status.openai.com
Resources on X
Dedicated X channels:
@OpenAI: the official account, carrying the launch announcement. https://x.com/OpenAI
@OpenAIDevs: the developer-facing account, carrying the per-benchmark developer breakdown. https://x.com/OpenAIDevs
@ArtificialAnlys: the independent evaluation account, carrying the cost-efficiency frontier result and the token-efficiency caveat. https://x.com/ArtificialAnlys
X posts with video content:
The Artificial Analysis thread on the cost-efficiency frontier result: the cost-per-task halving, the coding index movement in both directions, and the 31,000 against 29,000 output-token comparison. https://x.com/ArtificialAnlys/status/2102462962758033624
The OpenAI Developers thread giving the per-benchmark developer summary, including the 88 percent lower reported cost per task claim on AutomationBench and the caching changes. https://x.com/OpenAIDevs/status/2102461432684282061
Tibor Blaho's summary thread covering both same-evening launches side by side, OpenAI's price cut and Anthropic's Opus 5.5, which is the clearest single-post comparison of the two. https://x.com/btibor91/status/2102511469556609105
These channels and posts were resolved on 2026-09-23. Thumbnail images must be captured from the posts themselves at assembly time, and each image must match its own post.
Glossary
CI-First Benefit Score
The average of four dimensions, each scored 0 to 10: Time, Quantity, Quality, and Knowledge and Skill. It answers whether using the tool makes Co-Intelligence more profitable than Human Intelligence alone. Bands: 0 to 2.0 CI-First Negative, 2.1 to 4.0 CI-First Neutral, 4.1 to 6.0 CI-First Positive, 6.1 to 8.0 CI-First Strong, 8.1 to 10.0 CI-First Transformative. The score accounts for the overhead of prompting and verification, not just the benefit the tool produces. GPT-6 Sol scores 6.0.
CI-First Profile
The role the AI plays in your working relationship. (level 1) Co-Creator and Thought Partner, (level 2) Co-Worker and Assistant, (level 3) Coach and Tutor, (level 4) Analyst and Tester, (level 5) Challenger and Devil's Advocate. Lower level numbers indicate higher AI autonomy. Assigning a profile before giving the AI a task is a core CI-First discipline. GPT-6 Sol is primarily a Co-Worker and Assistant (level 2).
Humics Protection Badge
A rating of whether a tool protects, leaves neutral, or erodes three human capabilities: Creativity, Critical Thinking, and Social Authenticity. Each is scored +1, 0, or -1, and the sum gives the badge. +2 to +3 is Humics-Friendly, -1 to +1 is Humics-Neutral, -2 to -3 is Humics-Risky. It measures whether the tool strengthens the human or contributes to AI Obesity. GPT-6 Sol is Humics-Neutral at -1 / +3.
AI Imposture Risk
The likelihood that a tool traps you in one of three illusions. The Time Illusion is the appearance of saving time when net time is lost. The Quantity Illusion is high volume that looks good but does not survive inspection. The Skill Illusion is the appearance of competence in you while the underlying skill is absent or eroding. Each trap is rated Low, Medium, or High with cited evidence, and the overall level is Low when all three are Low, High when two or more are High. GPT-6 Sol is Medium overall, with Skill Illusion High.
User Sentiment
The aggregated public opinion from review platforms, community forums, and repository activity. It is reported separately from the CI-First score because crowd sentiment can contradict a rigorous evaluation. Where the two agree, the finding is stronger. Where they diverge, the divergence is worth explaining.
Review Status
Review Status records the current standing of the tool at the time of the last test. Active: the tool is current and recommended. Active (updated): recently re-checked and the content was refreshed. Changed: a re-check trigger fired and an update is pending, so read the review with that in mind. Risky: the tool has significant unresolved issues, or it has been clearly surpassed by newer alternatives. Use it with caution and read the Limits section. Retired: the tool still works but is no longer recommended. Deprecated: the tool has been shut down or fundamentally changed. Retired and Deprecated posts include a Migration Path section.
Sources
Vendor primary sources
OpenAI, Introducing GPT-6 Sol and Luna, 2026-09-22: https://openai.com/index/introducing-gpt-6-sol-and-luna/
OpenAI, Better prompt caching for GPT-6, 2026-09-22: https://openai.com/index/better-prompt-caching-for-gpt-6/
OpenAI developer documentation, GPT-6 Sol model reference: https://developers.openai.com/api/docs/models/gpt-6-sol
OpenAI developer documentation, GPT-6 Luna model reference: https://developers.openai.com/api/docs/models/gpt-6-luna
OpenAI developer documentation, API pricing: https://developers.openai.com/api/docs/pricing
OpenAI developer documentation, prompt caching guide: https://platform.openai.com/docs/guides/prompt-caching
ChatGPT Learn, local Codex memories, covering generated memory files and their storage: https://learn.chatgpt.com/docs/customization/memories
ChatGPT Learn, custom instructions with AGENTS.md, covering the instruction chain and its precedence: https://learn.chatgpt.com/docs/agent-configuration/agents-md
ChatGPT Learn, subagents, covering parallel agents, custom agents, and the [agents] settings: https://learn.chatgpt.com/docs/agent-configuration/subagents
OpenAI Deployment Safety Hub, GPT-6 Astra system card: https://deploymentsafety.openai.com/gpt-6-astra
OpenAI, GPT-6 Astra launch page, for the family positioning: https://openai.com/index/gpt-6-astra/
OpenAI, GPT-5.6 launch page, for the previous-generation rate card: https://openai.com/index/gpt-5-6
OpenAI Codex repository: https://github.com/openai/codex
OpenAI status page: https://status.openai.com
Independent sources
Artificial Analysis, GPT-6 Sol and Luna push the cost efficiency frontier, 2026-09-22: https://artificialanalysis.ai/articles/gpt-6-sol-and-luna-push-the-cost-efficiency-frontier
Artificial Analysis, GPT-6 Sol release intelligence, performance and price: https://artificialanalysis.ai/models/releases/gpt-6-sol
Artificial Analysis, GPT-6 Luna release intelligence, performance and price: https://artificialanalysis.ai/models/releases/gpt-6-luna
The New Stack, OpenAI releases GPT-6 Sol and Luna and cuts token prices in half, including the absence of any Sol versus Opus 5.5 head-to-head: https://thenewstack.io/openai-gpt-6-sol-luna-release/
The Decoder, OpenAI's GPT-6 Sol and Luna cut prices in half but barely move the needle on performance, including the independent GDPval regression and the missing benchmarks: https://the-decoder.com/openais-gpt-6-sol-and-luna-cut-prices-in-half-but-barely-move-the-needle-on-performance/
VentureBeat, OpenAI releases GPT-6 Sol and Luna models, slashing API costs 50 percent or more, including the vendor statement that the pricing is permanent: https://venturebeat.com/technology/openai-releases-gpt-6-sol-and-luna-models-slashing-api-costs-50-or-more
MacRumors, OpenAI's new GPT-6 Sol and Luna models bring Astra improvements to cheaper tiers: https://www.macrumors.com/2026/09/22/openai-gpt-6-sol-luna/
9to5Google, Claude Opus 5.5 and OpenAI GPT-6 Sol and Luna lower costs, on the same-evening timing: https://9to5google.com/2026/09/22/claude-opus-5-5-and-openai-gpt-6-sol-luna-both-launch-today-with-lower-costs/
AWS, OpenAI GPT-6 Sol and GPT-6 Luna now generally available on Amazon Bedrock: https://aws.amazon.com/about-aws/whats-new/2026/09/openai-gpt-6-sol-luna-on-amazon-bedrock
Coursiv, GPT-6 Sol and Luna, pricing, benchmarks and availability: https://coursiv.io/blog/gpt-6-sol-luna
Codersera, GPT-6 Sol and Luna complete guide, with per-effort transcriptions of OpenAI's charts: https://codersera.com/blog/gpt-6-sol-luna-complete-guide-2026/
Apidog, What is GPT-6 Luna, on the pricing fine print: https://apidog.com/blog/what-is-gpt-6-luna/
OpenRouter, GPT-6 Sol pricing and providers, including the weighted average price actually paid and the measured cache-hit rate: https://openrouter.ai/openai/gpt-6-sol
Benchmark sources cited by the vendor: AutomationBench (https://zapier.com/benchmarks), Agents' Last Exam (https://agents-last-exam.org/), DeepSWE (https://deepswe.datacurve.ai/), FrontierCode (https://cognition.com/frontiercode), OSWorld 2.0 (https://osworld-v2.xlang.ai/)
Community and community-reported evidence
Hacker News, GPT-6 Sol and Luna launch thread, 1,337 points and 651 comments as of 2026-09-23: https://news.ycombinator.com/item?id=49805509
Reddit r/singularity, discussion of the independent benchmark results, read via search indexing: https://www.reddit.com/r/singularity/
Reddit r/OpenAI, general community discussion, read via search indexing: https://www.reddit.com/r/OpenAI/
Reddit r/codex, Codex configuration and effort-level discussion, read via search indexing: https://www.reddit.com/r/codex/
Internal sources
CI-First Evaluation Framework v1.2, sections 3 (benefit rubric), 4 (Humics, including clause 4.2-a), 5 (Imposture risk, including clause 5.2.3-a), 6 (profiles), 7 (collaboration modes, including clause 7.5), 9 (scoring procedure and principles). The framework is applied across the U365 INSIDE Tools library: https://www.university-365.com/tools
INSIDE Tools Post Template, including the LLM variant and the agent-surface fields. Every published review built on it is listed at: https://www.university-365.com/tools
Published INSIDE Tools Reviews used as internal comparisons: GPT-6 Astra, scored 7.0 (https://www.university-365.com/post/gpt-6-astra-openai-s-frontier-model-for-computer-use-coding-and-science), GPT-5.6 Sol, scored 6.0, GPT-5.6 Luna, scored 4.8, GPT-5.6 Terra, and Claude Opus 5.5, scored 6.5 (https://www.university-365.com/post/claude-opus-5-5-anthropic-s-fable-5-1-class-model-at-40-percent-lower-cost-scored-6-5-on-the-u365).
Faculty Note on Evidence Quality
Four claims from this release did not survive checking against primary sources, and the difference matters for anyone deciding what to build on.
First, the headline price reduction is measured against a promotional rate. OpenAI's own sentence compares GPT-6 Sol to "GPT-5.6 promotional pricing". GPT-5.6 Sol's list price was $5 and $30 per million, with a promotional rate of $4 and $20 running at least through 21 November 2026. Against the promotional rate the reduction is 50 percent. Against the list rate it is 60 percent on input and 67 percent on output. Both numbers are true, they lead to different budget models, and a reader quoting "50 percent cheaper" is quoting the smaller of the two. The stronger finding is the vendor's own: a spokesperson told a reporter the pricing is permanent, which means the comparison base will not recur and the reduction for anyone currently on promotional pricing is the 50 percent.
Second, the benchmark table is a selection and the selection omits the vendor's own benchmarks. GDPval and Terminal-Bench 4.0, both part of OpenAI's usual published set, are absent here. Artificial Analysis found a regression on the first and an improvement on the second, so the omission is not neutral in either direction. The competitor set also stops at Claude Opus 5 and Claude Fable 5.1, while Claude Opus 5.5 shipped about 90 minutes earlier and appears nowhere. Two vendors published launch tables the same evening, each leaving out the other's newest model. That is a pattern rather than an accident.
Third, the FrontierCode comparison is against the competitor's weakest setting. OpenAI's sentence is that Sol matches Claude Fable 5.1 at xhigh effort at much lower cost. On the chart OpenAI itself cites, xhigh is Fable 5.1's lowest score of the five effort levels, and Claude Opus 5 at medium holds the chart's top score. The comparison is fair as a cost-per-result argument and it is not a capability claim, and the sentence does not by itself distinguish the two.
Fourth, the factual-reliability improvement is partly a change in behaviour rather than in knowledge. OpenAI reports about half as many mistakes on an evaluation built from conversations where users had already flagged an error, and it states those conversations are not representative of typical use. Independent measurement found the same direction with a mechanism worth naming: Artificial Analysis measured the hallucination rate falling from 92 percent to 60 percent while the model attempted 83 percent of questions against 99 percent before, so accuracy on the questions it did attempt fell from 59 percent to 54 percent. A model that says "I don't know" more often is safer and it is not the same as a model that knows more. For a high-volume workflow the distinction matters, because a declined answer is a field you must fill another way.
All four cases teach the same lesson, which is the one this review is built on. When a vendor's headline sentence and the vendor's own footnotes say different things, read the footnotes, and say so when they differ. OpenAI published several of these footnotes itself, including the fallback-cost omission on its competitor's datapoint and the statement that its adversarial tests do not measure failure rates in typical use. That makes the footnotes the useful part.
Review conducted by URC under the CI-First Evaluation Framework, version 1.2. Scoring date 2026-09-23. Tool version reviewed: GPT-6 Sol (`gpt-6-sol`, released 2026-09-22). Framework version applied: 1.2.










Comments