On this page

The short answer

Use GPT-6 Sol as a candidate for everyday application work, Astra for testing whether difficult tasks justify higher cost, and Luna for bounded volume workloads. These are evaluation starting points, not universal winners. Compare accepted outcomes and complete task cost on your own examples before changing production.

What to take away

  • ChatGPT plans and API model pricing are different.
  • Keep reasoning effort attached to every benchmark score.
  • A cheaper token rate does not guarantee a cheaper accepted result.

Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.

A dated model snapshot. Published metrics describe the named model, settings, and test. They are not Rubrex experiments or a universal ranking. Prices and availability can change after the research date.

Compare models and check the data date →

Which OpenAI models are covered

As checked on September 25, 2026, OpenAI’s model catalog lists GPT-6 Astra, Sol, and Luna as its flagship choices. It positions Astra for demanding reasoning, Sol for a balance of capability and cost, and Luna for focused, high-volume work. This briefing concerns API models. A ChatGPT subscription, its tools, and its limits are a separate product experience; an API token price should not be presented as the cost of a ChatGPT conversation.

Evidence: OpenAI model catalog [1]

Read the price and capacity together

The table records base catalog rates, not a quote for a completed task. Long-context tiers and generated reasoning can affect the bill, and a larger context window does not establish reliable use of every supplied document. Our recommendation is to test Sol as a practical baseline, then compare Astra on difficult failures and Luna on bounded, verifiable tasks. Keep the prompt, available tools, and acceptance rubric stable so the comparison measures the candidate model rather than a changing application.

Evidence: OpenAI model catalog [1]OpenRouter public model catalog [2]

API catalog snapshot checked September 25, 2026. USD per million tokens, base input/output rates; cached input, tools, long-context tiers, taxes, and routing may change the bill.
Model or propertyPublished valueHow to interpret it
GPT-6 Astra$10 input / $50 output1,050,000-token catalog context; tiered pricing
GPT-6 Sol$2 input / $10 output1,050,000-token catalog context; tiered pricing
GPT-6 Luna$0.10 input / $0.50 output1,050,000-token catalog context; tiered pricing

What independent results actually show

The saved Epoch AI EBR-bench results measure performance in the Earthborne Rangers game environment. They provide evidence about planning under that harness, not a general percentage of business tasks that each model can complete. The displayed configurations use max reasoning effort, and their run dates differ. Do not infer a reliable production ranking from the point estimates alone: the standard errors are material, and your tool permissions, task budget, and definition of success may produce a different ordering.

Evidence: Epoch AI: EBR-bench [3]

Epoch AI internal evaluations, CC BY 4.0. Snapshot checked September 25, 2026. Standard error is not a confidence interval; scores across different benchmarks are not interchangeable.
BenchmarkExact configurationScoreStandard errorRun date
EBR-benchGPT-6 Astra (max)76.19%8.38 percentage points2026-08-30
EBR-benchGPT-6 Sol (max)53.33%4.62 percentage points2026-09-22

A useful selection experiment

Build three task groups from your own product: routine work with objective checks, difficult cases requiring several decisions, and failures where a human must intervene. Run the same examples through candidate configurations and review outputs without model names when practical. Record accepted completion, severe failures, total token and tool cost, and p95 duration. Escalate to a more expensive configuration only when it improves an outcome your team actually values. A strong benchmark score is a reason to test a model, not permission to skip this step.

Limits of the evidence

This briefing synthesizes provider documentation and Epoch AI records; Rubrex has not reproduced these benchmark runs. The snapshot is dated, not live. Missing Luna results in the selected EBR table mean insufficient coverage here, not zero capability. Verify rates and availability before adoption.

Common questions

Which GPT-6 model is best for our SaaS?

Start with a task-specific baseline and compare accepted quality, cost, and latency. The right choice can differ between extraction, coding, and agent workflows.

Do these prices apply to ChatGPT?

No. They are API token rates. ChatGPT subscriptions and product usage limits are separate.

Sources & further reading

Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.

  1. OpenAI model catalog OpenAI · Living reference · Maintained documentationReviewed: current model names, context limits, and base API prices. Accessed September 25, 2026.
  2. OpenRouter public model catalog OpenRouter · Living reference · Maintained documentationReviewed: catalog fields; numerical prices taken from the saved September 25 API snapshot. Accessed September 25, 2026.
  3. Epoch AI: EBR-bench Epoch AI · 2026 · Independent benchmarkReviewed: benchmark definition and saved internal evaluation records; CC BY 4.0. Accessed September 25, 2026.

Questions or corrections? Write to Rubrex. Read our editorial approach.

From reading to a repeatable evaluation

Put a quality bar around your use case.

The AI Reliability Sprint covers one use case, a baseline, failure analysis, and one improvement experiment in 10 business days.

Explore the Sprint