Know what fits. Know why.
Change one constraint. See the current evidence that changes the call.
This field shows the exact reviewed scenario. Change the comparison type or constraints to change the decision brief.
The decision is held until every selected constraint has current support.
Decision scope: Route a production workflow with a high-volume routine lane and a smaller decision-critical lane without treating one model as the universal answer. The candidate evidence below comes from the fixed official-source record.
GPT-5.6 Sol
- Complex professional workSol clears the provider-positioning gate for the high-stakes evaluation lane.Official docs
Inspect evidence
Source recordGPT-5.6 Sol model card What the source saysThe official model card describes Sol as the frontier model for complex professional work, and current model guidance names it as the flagship-capability route.
What it does not proveProvider positioning is not independent evidence that Sol will satisfy the quality, reliability, or latency thresholds of this exact workload.
- Cost-sensitive volumeSol does not clear the selected routine-lane token-price cap.Not supported
Inspect evidence
Source recordGPT-5.6 Sol model card What the source saysThe official model card lists promotional pricing of $4.00 per million input tokens, $0.40 per million cached input tokens, and $20.00 per million output tokens, available at least through November 21, 2026.
What it does not proveThis compares published token rates only. It excludes discounts, Batch pricing, caching mix, tool-call fees, and the documented long-context uplift above 272,000 input tokens.
- Production surfaceSol clears the documented interface and context requirements for this production test.Official docs
Inspect evidence
Source recordGPT-5.6 Sol model card What the source saysThe official model card lists a 1,050,000-token context window, 128,000 maximum output tokens, Responses and Batch endpoints, streaming, function calling, structured outputs, and image input.
What it does not proveFeature availability does not establish application-level reliability, regional availability, account capacity, or acceptable tool behavior.
GPT-5.6 Luna
- Complex professional workLuna needs a representative workload evaluation before it can clear the high-stakes lane.Needs validation
Inspect evidence
Source recordGPT-5.6 Luna model card What the source saysCurrent official guidance positions Luna for efficient, cost-sensitive, high-volume workloads. The reviewed official sources do not position it as the complex professional-work tier.
What it does not proveThe absence of provider positioning is not evidence that Luna cannot perform the work. No workload-specific output was tested for this record.
- Cost-sensitive volumeLuna clears the selected published token-price cap for standard-length routine requests.Official docs
Inspect evidence
Source recordGPT-5.6 Luna model card What the source saysThe official model card lists $0.20 per million input tokens, $0.02 per million cached input tokens, and $1.20 per million output tokens.
What it does not provePrompts above 272,000 input tokens carry a documented uplift. Tool fees, discounts, Batch pricing, and the measured input-output mix still change effective cost.
- Production surfaceLuna clears the documented interface and context requirements for this production test.Official docs
Inspect evidence
Source recordGPT-5.6 Luna model card What the source saysThe official model card lists a 1,050,000-token context window, 128,000 maximum output tokens, Responses and Batch endpoints, streaming, function calling, structured outputs, and image input.
What it does not proveFeature availability does not establish application-level reliability, regional availability, account capacity, or acceptable tool behavior.
Evidence record
6 decision-critical claims| Evidence | Candidate | Status | Source | Checked |
|---|---|---|---|---|
| Complex professional work | GPT-5.6 Sol | Official docs | OpenAI | |
| Cost-sensitive volume | GPT-5.6 Sol | Not supported | OpenAI | |
| Production surface | GPT-5.6 Sol | Official docs | OpenAI | |
| Complex professional work | GPT-5.6 Luna | Needs validation | OpenAI | |
| Cost-sensitive volume | GPT-5.6 Luna | Official docs | OpenAI | |
| Production surface | GPT-5.6 Luna | Official docs | OpenAI |
Smallest test before committing
Run one blinded routing test on your real work. Turn provider eligibility facts into workload evidence before either model becomes a default route.
- Create a redacted set of at least 30 representative tasks, including routine cases, difficult cases, and known failure modes.
- Pre-register the quality floor, required schema, prohibited actions, review rubric, token budget, and latency ceiling.
- Run both model IDs with the same prompt, tools, reasoning setting, and retry policy, then blind model identity during human review.
- Record rubric results, structured-output validity, tool errors, tokens, end-to-end latency, retries, and reviewer corrections for every case.
- Route only the lane that clears every hard requirement, and keep an escalation path for tasks that remain uncertain.
Compare the tools people are choosing now.
Every published comparison starts with the job, shows the useful differences, preserves the main limitations, and links back to the product sources behind the record.