Updated 2026-10-02: Sonnet has been upgraded to Claude Sonnet 5.5, with upstream and API model ID claude-sonnet-5-5. Input / output list prices remain $2 / $10 per million tokens, $1 / $5 on the shared pool, and $1.8 / $9 on the official line. Cache read lists at $0.20; 5-minute / 1-hour cache writes list at $2.50 / $4. Other model prices and catalogue descriptions below preserve the historical 2026-09-12 snapshot. For the current catalogue, routes and prices, see models and prices; the Sonnet source is the Anthropic model documentation.
Update 2026-09-30: claude-fable-5, claude-opus-5, gpt-5.6-sol and grok-4.6, which appear below, stopped working on 24 September; their successors are claude-fable-5-1, claude-opus-5-5, gpt-6.1-sol and grok-4.7. gpt-6-sol, named here on 28 September, stopped working on 30 September too. The sums use the prices of 12 September. The method still holds; for the numbers, use the current prices at ai.topxea.com/pricing.
$99. That is the month's bill for a coding agent sitting on claude-sonnet-5-5 over the shared pool, 22 working days. The usage is assumed: 3 million input tokens and 300,000 output tokens a day. Sonnet 5.5 unit prices were checked against the Anthropic model page on 2026-10-02. TopxAI bills 50% of list on the shared pool and 90% on the official line. Other model comparisons below remain the historical 2026-09-12 snapshot.
The arithmetic fits on one line: 66 × $1 + 6.6 × $5 = $66 + $33. The 66 is the 66 million input tokens written in millions, the 6.6 is 6.6 million output, and $1 and $5 are the shared pool's input and output prices. With usage like this, input is two thirds of the bill.
The same usage at the vendor's list price is $198. Through the official line at TopxAI, the API service the TopXEA team runs, it is $178.2. The shared pool halves the list price, $99 against $198; between TopxAI's two routes the most you can save is 44%, $99 against $178.2. The official line's $178.2 sits $79.2 above $99, and that $79.2 is the cost of picking the official line before a dedicated line is open.
Where 3 million input tokens a day come from: 60 requests averaging 50,000 tokens of repo context
All seven text models are priced in dollars per million tokens. A million tokens is roughly 750,000 English words, or somewhere between 500,000 and a million Chinese characters. Under the assumption in the opening, a coding agent gets through three of those a day.
The reason is how multi-turn conversations are billed. Every request resends the whole history so far, and the model's own answer from the last turn is charged as input on this one. Three million input a day can be 60 requests carrying 50,000 tokens of repository context on average; 300,000 output is each of those 60 turns writing 5,000 tokens of code and explanation. The longer the context and the more turns, the further the bill tilts toward input. If you want to spend less, cut the context first. Telling the model to be brief saves very little.
You do not have to guess your own usage. Each response comes back with a usage field holding the token counts for that call: input_tokens / output_tokens in the Anthropic shape, prompt_tokens / completion_tokens in the OpenAI shape. The console's usage log has one line per request with the model, the token counts, the cost, the latency and the group that actually served it. It does not record content.
The three vendors set the list price, and TopxAI only multiplies it by a factor
Each model in the endpoint carries a reference_price field, which is the price the vendor publishes. The endpoint also names where each number comes from. For Anthropic that is the pricing page; for OpenAI, the model pages for gpt-5.6-sol and gpt-6-astra; for xAI, the pricing page. The original endpoint stamped them as verified on 2026-09-09. The Sonnet row was rechecked against the Sonnet 5.5 model page on 2026-10-02: $2 input, $10 output, $0.2 cache read.
The official line multiplies by 0.9 and the shared pool by 0.5, and the factor is the same for input, output and cache read. All seven text models were checked for the original 2026-09-12 comparison; Sonnet 5.5 was rechecked on 2026-10-02. So Sonnet 5.5 on the shared pool is $2 × 0.5 = $1 for input and $10 × 0.5 = $5 for output, and putting those back into the $99 line gives exactly $66 plus $33.
The models and prices page reads straight from this endpoint, so a price change shows up there within five minutes. The Sonnet 5.5 row is dated 2026-10-02; all other rows preserve the 2026-09-12 snapshot. The live page governs the current catalogue, long-context tiers and prices.
Official line at 90%, shared pool at 50%, and one key runs on one route
The endpoint describes the routes in its own words as "Claude · Official key · 90% of list price" and "Claude · Shared pool · 50% of list price", with the same pair for OpenAI and Grok. Here we call them the official line and the shared pool. You pick one when you create a key. A key cannot run on both; if you need both, create two. Below is the original seven-model comparison, with the Sonnet row updated to 5.5. Units are USD per million tokens, and each cell is input / output / cache read.
| Model | List price | Official line (90%) | Shared pool (50%) |
|---|---|---|---|
| claude-fable-5-1 | 10 / 50 / 0.25 | 9 / 45 / 0.225 | 5 / 25 / 0.125 |
| claude-fable-5 | 10 / 50 / 1 | 9 / 45 / 0.9 | 5 / 25 / 0.5 |
| claude-opus-5 | 5 / 25 / 0.5 | 4.5 / 22.5 / 0.45 | 2.5 / 12.5 / 0.25 |
| claude-sonnet-5-5 (2026-10-02) | 2 / 10 / 0.2 | 1.8 / 9 / 0.18 | 1 / 5 / 0.1 |
| gpt-6-astra | 10 / 50 / 1 | 9 / 45 / 0.9 | 5 / 25 / 0.5 |
| gpt-5.6-sol | 4 / 20 / 0.4 | 3.6 / 18 / 0.36 | 2 / 10 / 0.2 |
| grok-4.6 | 2 / 6 / 0.5 | 1.8 / 5.4 / 0.45 | 1 / 3 / 0.25 |
Now walk the $99 usage across the table.
List price 66 × 2 + 6.6 × 10 = 132 + 66 = $198
Official line 66 × 1.8 + 6.6 × 9 = 118.8 + 59.4 = $178.2
Shared pool 66 × 1 + 6.6 × 5 = 66 + 33 = $99
The shared pool is five ninths of the official line, so 44% cheaper, and half of list. The old gpt-image-2 was not in this table: it had one route at $0.05 per request against a $0.10 reference price. It has since been removed. Current image models are billed by size, so that historical price no longer applies.
Behind the official line is a dedicated line, configured per account, and you get one by emailing support@topxea.com. The key itself you can create at any time, and the form says what happens if you do: while no dedicated line is configured, official-line requests fall back to the shared pool. Those requests are still billed at the official-line price, and the usage log records which group actually served them. Pick the official line without ever sending that email and you pay $79.2 extra for the same shared pool traffic, the $1.8 rate on the $1 road. That is the $79.2 from the opening. So do it in order: email first, wait for the line, then create the official-line key. Once the key exists, send one request and check that the group column in the log says official line before you put an agent on it.
Requests over the shared pool, in the privacy policy's words, "may include an intermediary API provider". Content on that leg is handled under that provider's terms; TopxAI has no control over it. On TopxAI's own side nothing is kept. Prompts and completions are relayed, never stored. The usage log holds counts and identifiers only. If your data must not pass through a third party, email first and ask what path the dedicated line takes, then read the prices. On both routes the upstream model is the same, served from the official upstream API, with no distilled copies and no third-party mirrors. Latency and success rates are not published for either route, and there is no estimate of them here. Every line of the log carries a latency figure. Run for a few days and read your own.
A 50% cache hit rate takes 30% off the bill, not half
Every text model in the price sheet also has a cache read column. When the prefix of a request matches the previous turn's, the vendor can read that stretch from cache and bills it at that column's price. For the Claude and OpenAI models the cache read price is 10% of input, except claude-fable-5-1 at 2.5% ($0.25 against $10); for grok-4.6 it is 25%. On the shared pool, sonnet reads cache at $0.1.
The formula changes in one place: the input term splits in two. Cost = input × (1 - hit rate) × input price + input × hit rate × cache read price + output × output price. Assume a 50% hit rate and $99 becomes 33 × 1 + 33 × 0.1 + 6.6 × 5 = 33 + 3.3 + 33 = $69.3. That is 30% less rather than 50%, because the $33 of output has not moved and the half of the input that missed the cache is still full price. The official line at the same hit rate comes to $124.74, also 30% less.
The hit rate is an assumption. What you actually get depends on whether each request's prefix matches the previous one, and the rules are in each vendor's API documentation. Put a timestamp or a random ID at the top of the system prompt and the prefix changes every turn, so whatever gets written to the cache is never read by anyone.
For the four Claude models, writing to the cache costs money too. The endpoint lists a write price 25% above plain input: sonnet on the shared pool writes at $1.25 against $1 for plain input, and the one-hour tier is $2; on the official line the same two are $2.25 and $3.6. One write plus one read is 1.25 + 0.1 = $1.35, under the $2 of sending the prefix twice, so a single hit pays it back. A one-turn request writes and nobody reads, and that 25% is wasted. The one-hour tier needs two hits to break even; leave it off for short conversations. For the OpenAI and xAI models this post works out cache reads only.
The same usage costs $198 on gpt-5.6-sol and $247.5 on claude-opus-5
With the route settled, the only variable left in the bill is the model name. Same 66 million input and 6.6 million output, all on the shared pool. gpt-5.6-sol: 66 × 2 + 6.6 × 10 = $198, exactly twice sonnet, and the same as sonnet at list. claude-opus-5: 66 × 2.5 + 6.6 × 12.5 = $247.5. grok-4.6: 66 × 1 + 6.6 × 3 = $85.8; the $13.2 it saves over sonnet is all in the output price, since both cost $1 for input on the shared pool.
Light usage, the occasional call, say a million input and 200,000 output a month: sonnet on the shared pool is 1 × 1 + 0.2 × 5 = $2, the official line $3.6, list $4. At that scale the route makes no difference worth thinking about; you are arguing over a dollar or two.
Leaving cache reads aside, the same phrase "a million tokens" runs from $1, input on sonnet or grok over the shared pool, to $50, output on either fable or on gpt-6-astra at list. A factor of 50. A quote that names neither the model nor the direction is a number not worth remembering.
Your own coding goes on the shared pool; if the request path matters, email for a dedicated line first
On the public pages, the two routes differ in price, and in that the shared pool may pass through an intermediary API provider. If you are running Claude Code on your own code, or batch jobs over your own data on your own machine, use the shared pool. The $99 bill is that case. If customer documentation or compliance material has requirements on where requests travel, email support@topxea.com first, get the line configured and the path answered, then create the official-line key, at 10% under list.
If you have both kinds of work, the gap decides whether to split into two keys. At light usage the routes are $1.6 a month apart, and one official-line key for everything is the simpler life. At a $79.2 gap, split: the path-sensitive work on the official line, the rest on the shared pool. Base URLs and environment variables are in the setup post.
When you create a key you can set the most it is allowed to spend, in dollars. For a key handed to CI or a colleague, set the cap at a month's budget; for sonnet on the shared pool that is around $99, and if the thing runs away that is all you lose. The balance alert is a webhook and nothing else: when the balance drops below your threshold it sends one POST to the URL you gave, default threshold $5, no email. The shared pool burns $4.5 a day, so $5 is a single day of headroom. Set it to five working days, somewhere around $22.5.
The 3 million and 300,000 in this post are assumptions. The balance is prepaid dollars, and top-ups are not refunded, except a verified duplicate charge, so keep the first top-up small. Create one shared pool key, let your own agent run for a week, add up that week's cost from the usage log and scale it to a month. To estimate a change of model or route, add up the input and output tokens from the usage fields separately, put them through the formula above, and the number that comes out is your actual bill.