What changed exactly?
Model rates, per million tokens, and by a lot. According to the table published by Simon Willison on September 22, 2026, GPT-6 Sol is half the price of GPT-5.6 Sol, for both input and output, and GPT-6 Luna does the same compared with the previous generation. Claude Opus 5.5, released the same day, is one fifth below Opus 5.
Two details matter to a business owner. First, an OpenAI spokesperson told The New Stack that this is the default price, not a launch promotion. Second, on the same day, OpenAI revised prompt caching: according to The New Stack, the discount on previously seen input tokens now covers most of their price, and the cache is no longer lost when changing tools or reasoning level. Also according to The New Stack, Anthropic reduced the cost of reading the cache by more than half. In other words, the listed price is falling, and the price actually paid for a talkative agent is falling even further.
Why can your bill still rise?
Because an agent does not ask one question: it asks ten to finish a task, and it rereads itself each time. That is the difference between a conversation in ChatGPT and an agent working in your tools.
The Blent team detailed the calculation in June 2026 for a research request handled by an agent in four cycles: 12,800 input tokens for 1,650 output tokens. Context accumulates, the agent must remember what it has already done, and the input grows on every turn. They note that actual consumption frequently exceeds the initial estimate by a factor of 5 to 10 when a system goes into production.
Remember the mechanism, not the figures: you mainly pay to send context back to the model, and the number of round trips depends on the difficulty of the task, not the advertised rate. A poorly bounded task is expensive on an inexpensive model.
Where does the money go in an agent that works every day?
Four areas, only one of which is the model price.
| Area | What makes it grow | What keeps it under control |
|---|---|---|
| Context reread on every turn | Long instructions, a history that is carried along, whole documents sent “just in case” | Caching, and above all organized context: the agent receives what this task needs |
| Number of turns | A vague task, an agent that searches because it lacks the right information | A bounded task, and a turn limit after which it hands control back |
| Retries | A call that fails and is tried again, an unusable result that is redone | Counting failures, not only successes |
| Model price | A top-tier model used to sort mail | The right model for each step: sorting and classifying do not require the same model as writing |
This is what we mean in a discovery audit: the real question is “how many times will this agent reread what, and for what result,” not “which model is best.”
So how much does an AI agent cost for a small business?
No one can tell you before looking at your task, and French pages that promise a numerical answer demonstrate this well: from one page to another, the published ranges vary tenfold. These pages are honest; they simply do not describe the same thing. Stema Partners prices the development of a simple case, RedArrow the implementation of an agent in a company, and Smartpoint a production agent at a large corporation. Three different scopes, three different answers, and none concerns your company.
An online range is therefore useless for making a decision. A useful estimate is built on your volume: how many inquiries arrive each week, how many quotes come from them, and how much time a person on your team spends today. That is the first figure to establish, and it is yours. Everything else follows from it, including the answer “there is nothing useful to build here.”
How do you budget a task rather than a subscription?
Four things to request in writing before signing anything, from your provider and from yourself.
- The measured cost of one complete execution, not an estimate: run the agent fifty times on real cases and review the bill, including failures.
- The expected monthly volume, in tasks rather than “users”: thirty inquiries per week and three hundred do not create the same project.
- A cap per task and per day, which refuses the call when the cap is exceeded instead of revealing the overrun on the monthly bill.
- What remains human: each approval removes a loop from the agent, and therefore cost, in addition to protecting your customer relationship.
We operate this way ourselves. Équipage IA is a company of one human and agents: every paid call to an external service (research, data, image generation) is logged with its cost, roles that use them have a cap, and the program refuses the call before making it when the cap is exceeded. This is what makes it possible to tell a customer what their system really costs, and to tell them before they discover it. This is the same logic we install in our prospecting agents and inquiry and quote agents.
Should you change models now?
Not urgently, and not without measurement. A less expensive model does not make an agent less expensive if it makes more mistakes: every retry is one more turn. To our knowledge, no independent test yet compares GPT-6 Sol and Claude Opus 5.5, and the results published on September 22 are those of the providers.
What can be decided immediately, however: organize what the agent needs to know so it stops rereading everything, and verify that your long prompts are being cached. That work applies to every model, including those released next month. This is where a durable automation begins.
FAQ
Will the price of AI continue to fall?
The two announcements of September 22, 2026 point in that direction. But the reduction applies to the price per token, not the cost of a task: as prices fall, agents are assigned longer work that consumes more. Do not build your calculation on a future price reduction.
Is a less expensive model enough for a small business?
Often yes, for sorting, classifying, extracting, or preparing. Tasks that require a top-tier model are rarer than people think, and nothing requires using the same model for every step. The choice is made step by step, backed by measurement.
How can we avoid unpleasant billing surprises?
Three requirements: a cap that blocks before spending, a log that shows what each execution cost, and measurement on real cases before connecting the agent to the entire flow. Without those three, you will discover the cost of your system at the same time as your accountant.
We review one task, your tools, and the decisions that must remain human. The written conclusion is a simple automation, an AI agent, or nothing useful to build.
Sources: model pricing table and cache reduction, “Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war,” Simon Willison, September 22, 2026 (simonwillison.net) · default pricing confirmed by an OpenAI spokesperson and caching details, “OpenAI releases GPT-6 Sol and Luna, and cuts token prices in half,” Frederic Lardinois, The New Stack, September 22, 2026 (thenewstack.io) · multi-cycle agent token consumption, “Coût des agents IA : optimiser le budget tokens,” Blent, June 12, 2026 (blent.ai) · cited price ranges, Stema Partners, June 2, 2026 (stemapartners.com), RedArrow, April 6, 2026 (redarrow.fr), Smartpoint, January 28, 2026 (smartpoint.fr). The openai.com announcement pages could not be opened from our servers (403 error): facts attributed to OpenAI come from the two independent sources above. No result figure obtained by us or a customer is cited in this article.
