A million output tokens cost $0.50 with GPT-6 Luna and $20 with Claude Opus 5.5. At forty times the token price, the expensive model has something to prove.
Consider two jobs: pulling an order number out of an email, and finding a bug in an unfamiliar repository. For the first, you can check the answer against a few lines of text. For the second, working out what to check may be much of the job. Buying the same model for both deserves a closer look.
Opus leads the chart. Luna is there for a reason.
In the Artificial Analysis Intelligence Index snapshot of 22 September 2026, Opus 5.5 scores 58. Fable 5.1 and GPT-6 Astra have 53, GPT-6 Sol has 48, and Luna has 37. These are index points, not the percentage of customer requests a model will resolve.
Our chart shows the nine highest-scoring distinct models, plus GPT-6 Luna. It is not a global top ten: we added Luna to keep all three models in this comparison visible. Each model appears once, using its highest listed score.
There is a detail worth reading under the bars. The selected Opus 5.5 and Fable 5.1 results use maximum reasoning effort with fallback enabled. The other configurations are labelled too. We have not matched them for cost or response time, so the tallest bar does not also promise the shortest wait or smallest bill.
Three candidates for different jobs
Anthropic describes Opus 5.5 as close to Fable 5.1 on most work, with improvements in coding, computer use and knowledge work. It also cautions that benchmark margins have become a less reliable guide to differences in everyday use. Even the company selling the model is asking readers to look beyond the leaderboard.
GPT-6 Sol targets complex coding and agentic workflows, making it another candidate for the repository problem. Record its adjustable reasoning effort when comparing results; two Sol runs need not use the same settings.
GPT-6 Luna is aimed at focused tasks in high volume. Try it on extraction or classification. Whether it finds the right order number is easier to check against the email itself than to infer from its 37 index points.
The rate card does not include the rework
The official standard text-token prices, checked on 22 September 2026, are in US dollars per million tokens:
- Opus 5.5: $4 input, $20 output; $0.20 for cache reads.
- GPT-6 Sol: $2 input, $10 output; $0.20 for cached input.
- GPT-6 Luna: $0.10 input, $0.50 output; $0.01 for cached input.
OpenAI has additional pricing conditions for long prompts, processing modes and regional processing. And forty times the token price need not mean forty times the task cost: token use and repeated attempts change the bill.
Anthropic's own figures show why the distinction is useful. In its launch announcement, Opus 5.5 input and output token prices are 20% below Opus 5. Its own tests estimate 40% lower cost on typical workloads at default settings. One figure describes the tariff; the other estimates a bill for doing the work.
If an employee has to rewrite the answer, the API invoice misses part of the expense. Include that work, retries and tool calls in the budget; our guide to AI assistant costs covers the wider calculation.
A model swap can break a working agent
The Opus 5.5 migration notes contain a less glamorous detail than the benchmark scores: thinking is always on, and forced tool use returns errors. An agent built around a forced tool call can therefore run into trouble before you get to judge the quality of its answer.
Try a reversible workflow first: a support-ticket routing suggestion that a person confirms, or a code change that still needs tests and review. Compare against the current model under the same instructions, tools and permissions. Include awkward requests and failed services, not just clean examples. Record response time and human corrections without widening the agent's access.
Buy the capability the job needs
If the order number is already in your system, a database lookup may settle the task. A fixed extraction rule is another option. Check automation without AI before paying any model to do it.
For the repository task, paying more per token could make sense if it cuts the work needed to finish correctly. Compare the cost of that finished task, including human time.
If you're weighing an upgrade, discuss the workflow with LindenTech. Start with a task that already costs your team time; keep the existing model unless the replacement earns its place.