Skip to main content
For the fixed 5 million-token workload below, Capriole costs USD 8 for Claude Fable 5, Gemini 3.1 Pro Preview, or Gemini 3.5 Flash. The same models cost USD 15 to USD 90 at paid Standard list prices, while the calculated OpenRouter cash total is USD 15.83 to USD 94.95 after its published credit-purchase fee. That makes Capriole the much lower-cost option in all three verified rows. The USD 8 purchase is also a full Premium membership: it includes unlimited browser chat across the supported Premium model set and the first 5 million monthly API charged tokens. This is a reproducible comparison, not a universal price promise. It holds the token mix constant and uses public prices checked on 2026-08-12.

The 5 million token comparison

The table holds billed token counts constant, not source text or semantic workload. Tokenizers differ between model families. The scenario uses 100 requests with 40,000 uncached input tokens and 10,000 output tokens each, keeping every prompt below Google’s 200,000-token threshold. It excludes cache reads and writes, tools, free tiers, alternative service tiers, batch processing, taxes, and promotional pricing. * Premium costs USD 8 per month and includes unlimited browser chat across the supported Premium model set plus the first 5 million monthly charged tokens. This is the new customer’s public purchase price for that complete workspace bundle, not a per-model token rate or marginal inference cost. An existing Premium member with enough included balance pays no additional cash for this usage. After the included allowance is used, one additional 5 million charged-token pack has a USD 8 base price. Personal top-ups require active personal Premium access. Try Capriole AI in limited free browser chat before subscribing to Premium for the full workspace and API access.

How to reproduce the calculation

Direct providers publish separate prices for input and output. For this workload, the formula is:
Claude Fable 5 therefore costs:
OpenRouter states that it passes through the selected provider’s inference price without a markup. Its pay-as-you-go plan lists a 5.5% fee, with a USD 0.80 minimum, when credits are purchased. At the displayed inference rate, the cash calculation is:
The same arithmetic produces USD 21.10 for Gemini 3.1 Pro Preview. Gemini 3.5 Flash produces USD 15.825, rounded to USD 15.83. The percentage fee exceeds the USD 0.80 minimum in all three rows. OpenRouter can route the same model to different providers or service tiers, and fallback can change the endpoint that serves a request. The table therefore compares the displayed matching inference rate; it does not guarantee the bill from every default routing configuration.

Why Capriole uses a different unit

Capriole’s public customer balance does not copy each provider’s input and output price table. It deducts a unified balance called charged tokens:
Uncached input and output count at 100%. Cached input counts at 10%. The example has no cached input, so 4 million input tokens plus 1 million output tokens consume 5 million charged tokens. A charged token is an account balance unit. It is not the same thing as a provider’s input-token or output-token price. The table compares what the fixed workload costs under each billing system; it does not claim that their pricing units are interchangeable. Capriole AI Charged Tokens Explained provides three additional examples, rounding behavior, and the Premium quota boundary.

What USD 8 includes beyond the API balance

Direct APIs and OpenRouter let you pay for inference without a Capriole membership. This table covers a new personal Premium subscription. Capriole API access otherwise requires an eligible active personal or Team entitlement; Team pricing is outside this comparison. The USD 8 Premium plan includes:
  • Unlimited browser chat across the supported Premium model set
  • The first 5 million monthly API charged tokens
  • Access to the public API and maintained coding-agent integrations
  • One browser workspace with model switching, files, web search, and chat-storage controls
Before subscribing, signed-out visitors and Free users can try GPT-5.6 Thinking, Claude Fable 5, and Gemini 3.1 Pro in limited browser chat. That free flagship trial is separate from the three-model cost table above. For the target user who wants frontier-model chat and programmatic access together, this is the practical advantage: one low monthly price covers the daily browser workspace and a useful API allowance. The same account also works with six maintained coding-agent integrations. The first 5 million tokens should not be described as a standalone USD 8 API product. The payment buys the full Premium bundle. An active personal Premium member who needs another 5 million charged tokens can purchase one USD 8 top-up pack.

When the table will not predict your bill

The example excludes several pricing paths that can materially change cost:
  • Caching: Providers price cache writes and reads separately. Capriole counts cached input at 10% in its charged-token formula.
  • Long prompts: Gemini 3.1 Pro Preview charges higher rates when an individual prompt exceeds 200,000 tokens.
  • Service tiers and batch processing: Free, Flex, Priority, Batch, or other provider tiers can differ from the paid Standard rates used here.
  • OpenRouter routing: The endpoint or fallback that serves a request can have a different price from the matching rate used in the table.
  • Tools and search: Server-side tools, grounding, and search can add request or usage charges.
  • Output-heavy workloads: The table uses an 80/20 input-to-output split. More output raises direct and OpenRouter costs faster for models with a large output-price premium.
  • Enterprise terms: Volume discounts, private offers, data residency, and negotiated commitments can replace public list prices.
Recalculate with your expected model, token mix, and features before choosing a provider.

The practical choice

Choose Capriole when the supported model set covers your work and you want the lowest verified cash total in this comparison, unlimited Premium browser chat, included API usage, and maintained coding-agent paths in one account. Across every tested row, Capriole delivers the lower verified total and includes more than inference. The same USD 8 purchase adds the browser workspace, free-before-paid entry point, unlimited Premium chat, and maintained coding-agent integrations. Direct APIs and OpenRouter charge more for the fixed workload without including that Capriole workspace bundle. You can try Capriole AI in limited free browser chat before subscribing. For a wider product comparison, read Capriole AI vs OpenRouter. Developers can continue with the API quickstart.

Official sources

Capriole plan, top-up, and charged-token facts on this page are maintained as first-party documentation. The comparison uses only public customer pricing and quota rules. Prices and facts checked: 2026-08-12.
Last modified on August 18, 2026