The number on an AI vendor homepage is an entry price. The number that stops a rollout is usually a limit on another plan: seats, messages, pages, or tokens. A B2B buyer who compares those two numbers as if they were the same purchase will take a tool the team cannot run at the volume finance already approved.
Compare two options that can do the job. Read the limit that matches the rollout. Then, if either option bills a language model by the call, write the spend assumptions on the same page as the price. Someone in procurement should be able to rebuild the month without a meeting.
Start From the Rollout, Not the Category
Write the job in one paragraph before you open two pricing pages. Name the user, the finished work, the volume, and the input.
A 40-person engineering team wants a coding assistant inside the editor. Each engineer will use it on about 20 working days a month. A typical request sends a file excerpt and a short instruction. The answer should be a patch the engineer can reject. That is a different purchase from a support bot, an image generator, or a chatbot on the marketing site.
Four lines are enough to throw out most of a category:
- Who uses it, and how many seats that implies.
- What finished looks like. A suggestion in a pull request is not an autonomous commit.
- How often it runs. A pilot week is not 40 people for a month.
- What goes in, and how wrong the answer is allowed to be.
With that paragraph, two products are enough. A third vendor can wait until one of the two fails a limit you actually have.
Read the Limit That Will Bind First
On the pages we check for AI Compare, the first price is often a different product from the plan a team of 40 will land on. The AI Pricing Index 2026 tracks 727 AI tools and takes prices, plans, and free offers from the vendors’ own pricing pages.
In that catalog, the median tool lists four plans. 56.7 percent list four or more. 12.1 percent list exactly one. An entry row with a low cap is a real product for a lighter job. It is the wrong row for this rollout.
Ask which limit stops the work:
- Seats. A per-seat plan for 40 engineers is a different line from a plan built for one developer.
- A usage cap. Messages, completions, hours, or pages. Write down the cap and what the vendor does at the cap: a hard stop, a throttle, or an overage.
- The size of one request. A model context window and a plan’s monthly message cap are different limits. One says how much text fits in a call. The other says how many calls the plan paid for.
- Rate and fair use. A page that says unlimited still needs the sentence that defines normal use. If that sentence is missing, the cap is unpublished.
Development and code tools in this catalog are unusually likely to show a lasting free plan: 70.9 percent, against 8.3 percent in industry, construction, and logistics, and 54.7 percent across all 727 tools. Another 20.2 percent of the catalog offer a trial only, and 15.5 percent are paid from the first day.
A free coding tier you have used yourself is a weak predictor for a tool in another category, and it does not tell you whether the free tier survives 40 daily users.
Check Whether the Entry Price Is Ordinary, Then Leave It
Across the 216 tools in the index with a published US monthly entry price, the median is $20 a month. The mean is $46.88, because a smaller group starts much higher. 46.3 percent of those priced tools start below $20. These are US list prices for the entry paid plan, and only plans billed monthly are included. They describe this catalog. They do not describe every AI product on the market.
A $20 entry price is ordinary in that sample. Ordinary is not the same as affordable for 40 seats, and it is not the API bill.
Use the index to see whether a published entry price is odd. Then put the plan that covers 40 people, at the cap you calculated, into the comparison.
Keep the Seat Plan and the Meter in Different Columns
Some tools sell access. You pay for seats or a workspace, and the contract names a cap. The buyer question is whether that cap covers the month, what overage costs, and whether the cap is the one your engineers will hit.
Some tools sell model calls. You pay for tokens. Input tokens are the text you send. Output tokens are the text you get back.
A packaged assistant can also hide an API behind a seat price. If the order form does not say how usage is counted, you have a headline. The month shows up when the meter does.
OpenAI’s API pricing lists input, cached input, and output as separate rates per million tokens, prices long context on its own band, and bills web search as a per-call fee plus the tokens that search adds.
Anthropic’s model pricing states that Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer. This tokenizer produces approximately 30 percent more tokens for the same text. The exact increase depends on the content and workload shape.
Claude Sonnet 4.6 and earlier models use the previous tokenizer.
The same source file can be a different token count after you change models. The limit and the price both move. Write the model name next to the token assumption.
Write the LLM Budget So Finance Can Rerun It
When one option is an API, five lines decide the month:
- Requests.
- Average input tokens.
- Average output tokens.
- How much of the input repeats, and whether the price list you are using publishes a cache rate for that repeated part.
- The price list itself. The model maker, a cloud host, or a router. Those are different offers. A router price can differ from the maker’s direct price.
For this engineering team the sheet can start like this. Forty engineers, 20 working days, 25 requests a day. That is 20,000 requests in a month.
Say 1,200 input tokens, because a real request includes a file excerpt, and 400 output tokens, because you want a short patch.
If the same internal style guide is pasted into every call, test the cached-input rate for that repeated prefix only. The rest of the input stays on the normal input rate, unless the price list says otherwise.
The LLM API cost calculator is built for that sheet. It takes monthly requests, average input tokens, and average output tokens, and applies a published unit price.
The prices come from OpenRouter’s public list, which can differ from the model maker’s direct price, and the page shows when the data was loaded. Where a cache price is published, it uses it. Where it is not, it uses the normal input price rather than inventing a discount.
It also checks whether the prompt and the reply fit the model, and it withholds a figure when the call cannot run.
Treat the result as a budget for the model calls. Taxes, the contract, minimum commitments, and unpublished discounts stay outside it. So does the seat plan in the other column.
If input grows from 300 tokens in a demo to 1,200 tokens on a real file, the input portion of the bill is four times larger. The month rises by less than that when output stays at 400 tokens, because output has its own rate.
Hand Procurement Two Columns
Column One: The Packaged Assistant
The plan whose seat count and cap cover 20,000 suggestions. If the entry plan stops at a few hundred messages, its price refers to a smaller job. Leave it off the row.
Note overage and any commitment that is actually on the quote.
Column Two: The API
The five lines times the named rate card, plus anything that card prices separately, such as search or a regional endpoint. If a line is blank, the total stays blank.
Under both, one line for what price and limits did not measure: whether the suggestion is in the stack the team uses, and how long a bad patch takes to catch in review.
You can send the page when a colleague can rebuild both months from it.
The job names 40 engineers and 20,000 requests. Each option sits on the plan or the rate card that matches that job.
Every API figure has:
- Requests
- Input tokens
- Output tokens
- A cache decision
- A price list with a date
The entry price told you what is ordinary in a sourced catalog. The limit told you which plan is even a candidate. The assumptions told you what the LLM month costs.
About the Author
Martin Müller is the founder of AI Compare in Berlin. AI Compare publishes sourced prices, limits, and comparisons for AI tools.