Skip to content

How much it costs to put AI to work inside a company system, and what makes the bill go up or down

Pedro Cunha
Pedro Cunha
CTO at Epicora

Published on
updated on · 16 min read

In short

When AI runs inside a company system, you pay per token, and a customer service conversation costs cents: on Balcão, our WhatsApp customer service platform, triage comes to about R$ 0.15 (roughly US$ 0.03) per conversation. What weighs most is the text the AI rereads on every reply, and prompt caching bills that part at 10% of the price or less; model size comes next, with up to a 100x gap between the largest and smallest model from the same company. The bill also changes when you measure the result: in the simulation of a sales agent, the model that cost four times more brought more leads to the call, and the math by result came out in its favor.

How much a conversation handled by AI costs#

A conversation handled by AI inside a system costs cents. In the triage we set up on Balcão, our WhatsApp customer service platform, for the Fitness Academia gym chain and the OKSE agency, the estimate goes from R$ 0.10 for four turns to R$ 0.47 for sixteen on OpenAI's gpt-5.1, about US$ 0.02 to US$ 0.09. On the smaller model in the same line, R$ 0.02 to R$ 0.09 (under US$ 0.02).

These are estimates, calculated with the real token count of the triage prompt running at Fitness, OpenAI's price table and an exchange rate of R$ 5.50 to the dollar. A turn is one message from the contact and the AI's reply. The typical six-turn triage conversation comes to R$ 0.15. Both clients are in southern Brazil, where WhatsApp is the main channel between companies and customers.

Conversationgpt-5.1gpt-5-mini
4 turnsR$ 0.10R$ 0.02
6 turnsR$ 0.15R$ 0.03
10 turnsR$ 0.26R$ 0.05
16 turnsR$ 0.47R$ 0.09

The bill grows faster than the number of turns because, on every reply, the AI rereads the conversation from the beginning. OKSE stays on gpt-5.1 on purpose: the smaller model is a saving we are still validating, and we only switch the model of a service that already works when we can measure the difference without any other change happening at the same time.

This math applies when the AI runs inside a company system through the API, the door a program uses to talk to the model, as in our AI and automation projects. A monthly ChatGPT or Claude subscription has a fixed price per person and is a different bill.

What a token is and why writing costs more than reading#

A token is the piece of text the AI charges for, somewhere between a syllable and a word. The bill separates what goes in, which is everything the AI reads before answering, from what comes out, which is what it writes. On OpenAI's GPT-6 line and Anthropic's current models, each token written costs five times a token read, and on gpt-5.1 it costs eight times.

To get a sense of size, OpenAI says that in English a token is about four characters, or three quarters of a word, and that other languages have a different ratio. Rounding up, a page of text is close to a thousand tokens. Prices are given per million tokens, so US$ 2 per million means roughly two dollars for every thousand pages read.

What goes in is bigger than it looks. In customer service, on every reply the AI reads the company's instructions, which make up the prompt, the list of tools it can use, the conversation history and the new message. The bill pays for all of it.

What comes out is also bigger than what shows. Models that reason before answering spend tokens on that reasoning, and OpenAI explains that those tokens do not appear in the answer but are billed as output. A short answer can cost more than its length suggests, which is why the reasoning effort setting is a cost lever within the same model.

Why model size changes the bill by up to 100x#

OpenAI and Anthropic sell their models in sizes, and the price follows the size. On October 8, 2026, the large model from each company cost 100 times the small one: US$ 10 against US$ 0.10 per million tokens read and US$ 50 against US$ 0.50 per million tokens written. For 10,000 short customer service conversations a month, that is the difference between US$ 3 and US$ 300.

CompanySizeModelInput (US$ per million)Input reread from cacheOutput10,000 conversations
OpenAIsmallGPT-6 Luna0.100.010.50US$ 3
OpenAImediumGPT-6.1 Sol20.1010US$ 60
OpenAIlargeGPT-6 Astra10150US$ 300
AnthropicsmallClaude Haiku 5.50.100.010.50US$ 3
AnthropicmediumClaude Sonnet 5.520.1010US$ 60
Anthropicbetween medium and largeClaude Opus 5.540.2020US$ 120
AnthropiclargeClaude Fable 5.1100.2550US$ 300

The last column assumes 1,500 input tokens and 300 output tokens per conversation, without caching, which is a short conversation. Read that way, the table shows the middle model costs a fifth of the large one at both companies, and the small one costs a twentieth of the middle one. The comparison between the companies is per token: Anthropic notes that its models from Claude 4.7 on count about 30% more tokens for the same text, so the same conversation costs somewhat more on Claude. The Haiku 5.5 price applies to requests of up to 100,000 tokens, and the model came out on October 7, 2026 costing, according to Anthropic, about 75% less than Haiku 4.5.

The task picks the size. The small model handles short, repeated decisions, like sorting email, classifying an invoice or answering opening hours. The middle one handles everyday work with rules: customer service that follows a script, meeting summaries, proposal drafts. The large one is for what is rare and hard, like reading a long contract or cross-checking data from several systems.

Starting with the middle model is usually the safest path: move up when the answer does not work and move down what is repetitive. How much the AI decides on its own is a separate choice, explained in chat, assistant or AI agent.

What weighs most: the text the AI rereads on every reply#

What weighs most in the bill for AI customer service is the text the AI rereads on every reply: the company's instructions, the tools and the conversation history. Prompt caching holds that cost down by billing the part that repeats unchanged from the start at 10% of the input price on most OpenAI and Anthropic models.

In the simulated conversations of a sales agent we built on Balcão for an online training and diet coaching business, the AI rereads close to 13,000 tokens on every reply on gpt-5.1, and close to 90% of that comes from the cache. That is more than eight times the 1,500 tokens that example calculations usually assume. By our math on the measured tokens, without caching each conversation would cost about R$ 0.55 (US$ 0.10), not the R$ 0.15 it cost.

Caching works on the start of the text. The provider keeps the opening stretch it has already processed and, when the next call starts exactly the same way, bills that stretch at the lower price. Any change at the start makes the rest billed in full. That is why, on Balcão, the fixed prompt always goes first, and whatever changes from one contact to another (the date, the name, what the person answered on the form) goes at the end. Writing a stretch to the cache costs 1.25 times the input price on the newest models of both companies, and the discount comes on the reads that follow.

This leads to a conclusion that surprises people: shortening the prompt saves little while caching works, and shortening it in a way that changes the start on every call costs more. With the Balcão triage prompt, of about 4,000 tokens, the fixed part costs R$ 0.05 in a six-turn conversation. Half of that prompt, picked again for every message, would cost R$ 0.13. What belongs in that fixed prefix is the company knowledge the AI needs to answer, and how to organize it is covered in how to organize company knowledge for AI.

On October 7, 2026, Anthropic halved the cache read price of Sonnet 5.5, from US$ 0.20 to US$ 0.10 per million, and said that since those reads are a large share of consumption, the model became about 20% cheaper on most agent work.

Measuring without separating the cache is misleading. Until September, Balcão counted all the text read at full price, as a safety margin, and the dashboard showed a cost up to twice the OpenAI invoice. Today each call records how much came from the cache and stores the cost in Brazilian cents, rounded up. Anyone putting AI into production should ask for cost to be measured this way before using the number to set a price.

Triage: each request goes to a model of its size#

Triage means sending each request to the model size it needs. A small model reads every message, solves the simple ones and passes on only those that need more. In a worked example with 10,000 short conversations a month and the prices of October 8, 2026, triage takes the bill from US$ 300, all on the large model, down to US$ 17.40.

Here is the worked example. The small model reads all 10,000 messages and solves 8,000 on its own, for US$ 3. The middle model answers 1,900, for US$ 11.40. The large model answers the 100 that are rare and hard, for US$ 3. With the same number of tokens, the figures are the same with GPT-6 Luna, GPT-6.1 Sol and GPT-6 Astra and with Claude Haiku 5.5, Sonnet 5.5 and Fable 5.1, because the per-token prices of the three sizes match; with the same text, the Claude bill comes out about 30% higher. The split between 8,000, 1,900 and 100 is an example; each company's split comes from its real service history.

For the step that only decides, there are even models built for it. Jev, from TypeSafe AI, in early access, does not write text: it returns a yes or no, a score or a choice among given options. According to the company, input costs US$ 0.042 per million tokens and output is not billed.

The cheapest triage happens before the AI. At that coaching business, the ad leads to a form before WhatsApp, and almost everyone who talks to the sales agent has already filled it in. Someone who was only curious rarely reaches the AI, and every paid conversation starts with someone interested. In the Fitness and OKSE triage, which runs on a single model, the AI handles the first contact, fills in the lead record and hands the team whatever needs a person, which gives time back to the team without putting a large model on opening-hours questions.

When the more expensive model ends up cheaper#

The more expensive model ends up cheaper when you count by result instead of by conversation. In the simulation of that business's sales agent, Claude Opus 5.5 cost R$ 0.65 per conversation (about US$ 0.12), four times gpt-5.1, but 44% of the simulated leads accepted the sales call, against 19% to 31%. By our math, if that gap holds with real leads, each extra call costs R$ 2 to R$ 4 of AI.

The agent leads the WhatsApp conversation up to an invitation to a call with the business's staff. To measure the difference, we ran 16 simulated conversations, with eight types of lead (the suspicious one, the one who goes straight to the price, the one who asks for a person), each twice, and compared them with the last three rounds on gpt-5.1, the last of them with the same configuration.

On Opus, 94% of conversations ended without a form error, against 31% to 56% on gpt-5.1. A form error is a message that breaks the service rules, like asking two questions at once or using a tone outside what was agreed. The cost was R$ 0.65 per conversation against R$ 0.14 to R$ 0.16, with 91% of the reread text coming from the cache, against 88% to 89%. This is a simulation, not production: few conversations, with simulated leads, run before the agent went live.

By the math of results, the difference pays for itself. If the acceptance rate from the simulation holds in production, for every thousand conversations Opus costs about R$ 500 (around US$ 90) more and brings between 130 and 250 more calls, which comes to R$ 2 to R$ 4 (US$ 0.36 to US$ 0.70) of AI per additional call. The business stayed on Opus 5.5, and cost reduction was left for later, with the agent running. What makes a model follow the sales script closely is how the method is written, and that is covered in how to make AI follow your company's method.

How to cap spending before going live#

The spending cap is set before turning the AI on, in three layers: a daily ceiling for each client or operation, a global ceiling for the whole system and an alert whenever any ceiling is hit. On Balcão, each client's default is R$ 15 (about US$ 2.70) of AI per day and 500 calls to the model, and no limit is hit silently.

Each client's ceilings change from the dashboard, without releasing a new version of the system. There are also per-contact limits, 40 turns and 10 media files per day, which hold down the bill and anyone trying to use the service for something else, a security risk whose best-known case is prompt injection. The global ceiling, for the whole platform, stays with our administration.

When any ceiling is hit, the system logs the event and sends an alert to the Epicora on-call team and, when the ceiling is about the client's own cost or volume, also to the people the client named. If it is the client's cost or call ceiling, the AI stops answering for that client until the day turns or someone raises the ceiling; if it is a per-contact limit, only that contact is affected.

Before turning on AI customer service, this is the order we follow:

  1. Count the tokens of the real prompt, not of an example, and simulate conversations as long as the ones the company gets.
  2. Multiply by the prices of the chosen model, separating what comes from the cache.
  3. Set the daily and global ceilings, with an alert to someone who can act.
  4. Compare the measured cost with the provider invoice in the first week.

Frequently asked questions#

Can a ChatGPT or Claude subscription power AI customer service for my company?#

Not inside a system. A subscription is a fixed price per person to use the chat app: ChatGPT Plus costs US$ 20 a month, according to OpenAI. OpenAI and Anthropic both say that API usage, the door a system uses to talk to the model, is billed separately and by token. Anthropic now includes a monthly API credit in the Max and Team plans, but the bill for customer service at volume is the API bill.

How much does it cost per month to handle a thousand conversations with AI?#

In our estimate for Balcão triage, with six-turn conversations on gpt-5.1 at about R$ 0.15 each, a thousand conversations cost around R$ 150 (roughly US$ 27) of AI per month, and 10,000 cost around R$ 1,500 (roughly US$ 270). On gpt-5-mini, the smaller model in the same line, a thousand conversations come to about R$ 30. Longer conversations and larger models raise the bill. The messaging channel, the server and development are not in this number.

Can I always use the cheapest model?#

For simple, repeated decisions, like sorting email or answering opening hours, almost always. For conversations that follow rules, sales or long documents, the small model makes more mistakes, and a mistake costs more than the price difference. The safe path is to test the same service on two models with simulated conversations and compare the outcome, not only the cost.

Is it worth shortening the prompt to spend less?#

Not by much while caching works, because the text repeated from the start is already billed at 10% of the price or less. Shortening it in a way that changes the start of the prompt on every call costs more: by our math with the Balcão triage prompt, a fixed 4,000-token prompt costs R$ 0.05 per conversation, and half of it, changing every time, would cost R$ 0.13.

How can I know what the AI will cost before going live?#

Count the tokens of the real prompt, simulate conversations as long as the ones your company gets and multiply by the prices of the chosen model, separating what comes from the cache. After going live, compare the measured cost with the provider invoice in the first week. That comparison is how we found out the Balcão dashboard was showing up to twice the invoice.

Sources#

Next step#

If you want AI in customer service, triage or document reading and need to know what it will cost before turning it on, that is the work we do in AI and automation: the math with your real prompt, the model chosen by the task and spending ceilings in place, as in the triage for the OKSE agency.

Share
Pedro Cunha
Who writes here
Pedro Cunha
CTO at Epicora

Pedro Cunha leads engineering at Epicora, in Chapecó, Brazil. He writes about the technical decisions behind the systems the team puts into production — architecture, scope, estimation and applied AI.

Articles by Pedro Cunha

Contact

Let's talk about your project

Tell us what you need to solve. We reply fast, with people who understand both technology and business.

Prefer to talk directly?

Pick the channel you prefer. We reply fast, during business hours.

From the first conversation to go-live: efficiency, security and innovation.