Updated 2026-10-02: Sonnet has been upgraded to Claude Sonnet 5.5, with upstream and API model ID claude-sonnet-5-5. Input / output list prices remain $2 / $10 per million tokens, $1 / $5 on the shared pool, and $1.8 / $9 on the official line. Cache read lists at $0.20; 5-minute / 1-hour cache writes list at $2.50 / $4. Other model prices and catalogue descriptions below preserve the historical 2026-09-12 snapshot. For the current catalogue, routes and prices, see models and prices; the Sonnet source is the Anthropic model documentation.
Update 2026-09-30: claude-fable-5, claude-opus-5, gpt-5.6-sol and grok-4.6 stopped working on 24 September, and gpt-6-sol, which this note recommended on 28 September, stopped on 30 September; requests for them return HTTP 403. Use claude-fable-5-1, claude-opus-5-5, gpt-6.1-sol and grok-4.7. The two commands below that run as written now use gpt-6.1-sol. The price table is a record of 12 September; current prices are at ai.topxea.com/pricing. The site has since added setup guides for Cursor, Cline and other clients to its documentation at ai.topxea.com/docs.
Point the OpenAI SDK at https://ai.topxea.com/v1 and the Anthropic SDK at https://ai.topxea.com. Whether that is worth doing comes down to two numbers. The shared pool charges half the list price; the official line charges 90% of it. Output on claude-sonnet-5-5 lists at $10 per million tokens, so the shared pool bills $5 and the official line $9. If you run Claude Code on your own code, or batch jobs over your own data, use the shared pool. One line changes and the bill halves.
Two cases call for holding off. If there are rules about where your data may go, wait; the reason is in the section on keys. If you want the official line and have not yet emailed support@topxea.com, also wait. The official line is a dedicated line, configured per account. Until it is configured, requests on an official-line key are served by the shared pool and billed at the official-line price. That is the 90% rate for the 50% route.
One more fact sets the order of operations: top-ups are not refunded, except for a verified duplicate charge. Keep the first top-up small. Create a shared-pool key, get one curl through, then point the SDK at it. If you want the official line, send the email first and create that key only after support replies that the line is configured.
The SDK checks the path and the JSON shape, nothing else
An official SDK does little per request. It joins the base URL to the endpoint path and puts the key in a header. The body is serialized and the reply parsed, each against one fixed schema. Whose domain the request goes to never comes up. A gateway that serves the same paths with the same request and response shapes is, as far as the SDK can tell, the vendor's own server. Switching gateways changes which machine gets the request. The code that assembles messages and reads usage stays untouched.
The two SDKs disagree about what a base URL contains. The OpenAI SDK defaults to https://api.openai.com/v1: /v1 is part of the base, and the path is just /chat/completions. The Anthropic SDK defaults to https://api.anthropic.com, and the path is the full /v1/messages. The classic mistake when switching gateways is writing one SDK's base in the other's format. For TopxAI, one gets /v1 and the other does not. The two snippets below are independent of each other and use different keys.
import os
from openai import OpenAI
openai_client = OpenAI(base_url="https://ai.topxea.com/v1", api_key=os.environ["TOPXAI_KEY_OPENAI"])
from anthropic import Anthropic
claude_client = Anthropic(base_url="https://ai.topxea.com", api_key=os.environ["TOPXAI_KEY_CLAUDE"])
The two styles differ in headers and field names, and Claude answers to both
OpenAI style: the key sits in an Authorization: Bearer header, the endpoint is /v1/chat/completions, and the body carries a messages array with the system prompt as one of its entries. The reply is at choices[0].message.content; usage comes back as prompt_tokens and completion_tokens. /v1/responses is a second endpoint in the same style, with input in place of messages. Codex CLI uses it by default.
The Anthropic style parts ways at the headers. The key goes in x-api-key, you also need anthropic-version: 2023-06-01, and the endpoint is /v1/messages. max_tokens is required and the request errors without it; it is the field most often dropped when a body is ported over from the OpenAI side. The system prompt is a top-level field, not a message. Read the reply from blocks whose type is text in the content array; the first block may contain thinking. Usage is input_tokens and output_tokens. Both styles stream with stream: true, but the SSE events are shaped differently, so the parsing code does not carry over. Both replies carry usage. The token count for the call you just made is right there, so there is no need to wait for the usage log.
Model and style are not welded together. The four Claude models accept both styles, so a codebase written only against the OpenAI SDK can call claude-opus-5 without adding the Anthropic SDK. gpt-6-astra, gpt-5.6-sol and grok-4.6 take OpenAI style only, plus /v1/responses. The two GPT Image 2.5 models (gpt-image-2.5-sunburst and gpt-image-2.5-flare) live on /v1/images/generations and /v1/images/edits. The result is that one OpenAI SDK codebase reaches all seven text models. To switch vendor you change the model field and swap in a key created on that vendor's route. Anthropic-style clients such as Claude Code can only reach the four Claude models.
curl https://ai.topxea.com/v1/chat/completions \
-H "Authorization: Bearer $TOPXAI_KEY_OPENAI" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-6.1-sol","messages":[{"role":"user","content":"Say hi"}]}'
curl https://ai.topxea.com/v1/messages \
-H "x-api-key: $TOPXAI_KEY_CLAUDE" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-5-5","max_tokens":256,"messages":[{"role":"user","content":"Say hi"}]}'
You can send a request before you have a key. Hit any endpoint without one and you get a 401: type authentication_error, a message that starts with Invalid token, then a request id. The error body follows the endpoint. This is what came back on 2026-09-12, with the request id replaced by a placeholder:
{"error":{"message":"Invalid token (request id: <request-id>)","type":"authentication_error","code":"authentication_error","param":null}}
{"error":{"type":"authentication_error","message":"Invalid token (request id: <request-id>)"},"type":"error"}
The first one is what /v1/chat/completions and /v1/models return, with code and param inside the error. The second is from /v1/messages, with an extra type on the outer object. Each path answers in its own style. A 401 only says the hostname is right and the gateway is answering. Whether the path is right has to wait until a key exists, and the 404 section covers it.
The route is fixed when you create the key, and the official line needs an email to support first
A key's route is set at creation. One key cannot run on both routes; if you need both, create two. Routes are split by vendor. Each text vendor has an official line and a shared pool, named claude-official, claude-shared and so on in the API, with OpenAI and xAI following the same pattern. The GPT Image 2.5 models have their own per-image route. A key for Claude Code has to be created on the Claude side and a key for Codex on the OpenAI side, which is why the snippets above keep the keys in two variables.
Both routes run the same upstream model. The site's own wording: "Every model is served from its official upstream API. No distilled copies, no third-party mirrors." What differs is the price and the path the request takes. The price is a fixed ratio: input, output and cache reads all come in at 90% and 50% of list. The original seven text models were checked on 2026-09-12. The table preserves that historical comparison, with the Sonnet row updated to 5.5 on 2026-10-02. It shows input and output only, in USD per million tokens:
| Model | List price | Official line (90%) | Shared pool (50%) |
|---|---|---|---|
| claude-fable-5-1 | $10 / $50 | $9 / $45 | $5 / $25 |
| claude-fable-5 | $10 / $50 | $9 / $45 | $5 / $25 |
| claude-opus-5 | $5 / $25 | $4.5 / $22.5 | $2.5 / $12.5 |
| claude-sonnet-5-5 (2026-10-02) | $2 / $10 | $1.8 / $9 | $1 / $5 |
| gpt-6-astra | $10 / $50 | $9 / $45 | $5 / $25 |
| gpt-5.6-sol | $4 / $20 | $3.6 / $18 | $2 / $10 |
| grok-4.6 | $2 / $6 | $1.8 / $5.4 | $1 / $3 |
Captured 2026-09-12. List prices were checked against what Anthropic, OpenAI (gpt-6-astra, gpt-5.6-sol) and xAI publish, verified 2026-09-09. At the time, gpt-image-2 was billed per request at $0.05 against a $0.10 reference price. That model has since been removed; this price does not apply to the current GPT Image 2.5 models. Live numbers are on /ai-api, which reads straight from the API; a price change shows up there within 5 minutes.
What rules out the shared pool is where your data goes. The privacy policy's wording is that, depending on the route, a request may pass through an intermediary API provider, and the shared pool is such a route. Content that reaches that provider is handled under its own terms, and TopxAI has no control over that part. TopxAI itself keeps no content. The usage log has one line per request with the model, token counts, cost, latency and the group that actually served it, and no content. For your own code on your own data, that is enough. If a customer contract names which companies may receive the data, the shared pool is the wrong route.
The official line is selectable in the console, but selecting it is not the same as having the line configured. Until it is, the option changes only the billing, not the path the request takes. If a compliance file has to say "via the official key", the email is not optional. The extra 40% of list buys the route itself, and the site publishes no latency or success-rate figures for either line, so do not pay it for a quality difference you are imagining.
Balance, alerts and what a key can lock
The balance is prepaid US dollars. Stripe takes cards and NOWPayments takes USDT / USDC. A top-up counts as received once TopxAI has verified it with the payment provider. The checkout page sending you back does not count. The only balance alert is a webhook: when the balance drops below your threshold, one POST goes to the URL you gave, with an optional secret. The default threshold is $5. There is no email alert. If you want to be told, wire up a URL, even if all it does is forward to a group chat.
Four things can be set on a key at creation: a spend cap (a dollar amount, or unlimited), an expiry, a model allowlist (blank means all) and an IP allowlist (one per line, CIDR accepted). There is no per-key rate limit; rate limits are per account and per model, visible in the console. Do not set the cap to unlimited, even on your own key. Set it low and allow only the model you will use, so a leaked key can cost at most what is left in the cap. Same for keys handed to colleagues or CI.
Environment variables for Claude Code and Codex CLI
export ANTHROPIC_BASE_URL=https://ai.topxea.com
export ANTHROPIC_AUTH_TOKEN=$TOPXAI_KEY_CLAUDE
export ANTHROPIC_MODEL=claude-sonnet-5-5
claude
The third variable is optional. Set it and the default model is pinned to Claude Sonnet 5.5 (claude-sonnet-5-5). If you would rather not export every time, the same values work in the env field of ~/.claude/settings.json.
export OPENAI_BASE_URL=https://ai.topxea.com/v1
export OPENAI_API_KEY=$TOPXAI_KEY_OPENAI
codex -m gpt-6.1-sol
Codex hits /v1/responses by default, and TopxAI serves it. Cursor, Cline and the other clients have no published setup on the site, so this post does not cover them.
404, 401 and model not found each have one place to look
For a 404, look for a doubled /v1. Add a /v1 to the Anthropic SDK's base and the path becomes /v1/v1/messages. Build the URL yourself with the requests library, starting from a base that already ends in /v1 and appending the full /v1/chat/completions, and you land on /v1/v1/chat/completions. Both are 404s.
A 401 means the gateway did not recognize the key. The no-key response is shown above; start by echoing the environment variable, since it may be empty. The two styles also use different headers, Authorization: Bearer for OpenAI style and x-api-key for Anthropic style, so check that while you are there.
For model not found, check the name first. TopxAI recognizes only the names in its current catalogue. Sonnet 5.5 uses claude-sonnet-5-5. Check the models and prices page for the current catalogue; the historical comparison table in this post is not a complete model list. If the name is right and it still fails, check two more things: whether the key's model allowlist includes the model, and whether the key's route matches the request (going by route name, a key on claude-shared carries Claude models only). Whichever it is, call GET /v1/models. What comes back is everything this key can call right now, and it is where to look whenever a model will not go through.
Once everything works, open the usage log in the console. Whatever the group column says is the line this key went through, and whether an official-line key fell back to the shared pool is in the same column. How much a month costs, and how cache hits and multi-turn conversations inflate the bill, is worked through in the cost post.