TCO calculator
How much could you save?
Three questions about your AI usage, and an estimate of your yearly bill with and without TokenSaver.
How cautious should the estimate be?
CentralWith this setting, we assume that TokenSaver:
- •
- •
- •
Your estimate, per year
EstimateYou could save
—/ year
·
Today, without TokenSaver
—
With TokenSaver
—
Yearly cost
Estimate. Results are for illustrative purposes only. Actual costs vary based on usage patterns, use case, number of users, infrastructure costs, and applicable discounts. This calculator is not a substitute for a formal cost analysis.
How it is computed
- AI bill = requests × (tokens sent × input price + tokens received × output price), at each provider’s official list prices per 1M tokens.
- Cache: requests answered from the cache do not reach the model.
- Compression: the text sent with the remaining requests is shorter.
- Routing: part of the remaining requests go to the cheaper model (Pro and Enterprise).
- Subscription: Pro at €79 excl. VAT per person per month, converted to USD. Enterprise on quote.
Official list prices of each provider, standard tier, in USD per 1M tokens, as of October 2, 2026. Excludes negotiated discounts, batch pricing, provider prompt caching and taxes. Gemini 3.8 Flash: introductory price until 31 Dec 2026. DeepSeek: off-peak rates, peak hours cost double. € converted at 1 € = 1.15 $ (provisional rate).
AnthropicOpenAIGoogleMistralDeepSeek
Sources for the request sizes
- OpenRouter and a16z, State of AI, empirical study of 100 trillion tokens (arXiv): about 6K input and 400 output tokens per request on average in late 2025; programming requests routinely above 20K input tokens. arxiv.org
- Microsoft Learn, Azure AI Search: recommended RAG chunk size of 512 tokens, the basis of the Medium estimate (about 8 retrieved chunks plus the prompt). learn.microsoft.com
- OpenAI Help Center, What are tokens: 1 paragraph is about 100 tokens, 100 tokens about 75 words, the basis of the Short estimate. help.openai.com
- Anthropic Engineering, multi-agent research system: agents use about 4 times the tokens of chat interactions, multi-agent systems about 15 times. www.anthropic.com
Measure it on your own traffic.
The 30-day hosted trial measures cache, compression and routing on your real requests, run by run.