Usage and billing

Usage is metered per model and paid from prepaid credits you top up with a card.

Metering

The gateway counts input and output tokens on every request and attributes them to your API key. The Usage page breaks this down per model and per day. There are no seats, minimums, or idle charges in the pilot: a model that serves nothing records no inference usage.

Pricing

All prices are per million tokens.

ModelParamsInputOutputTraining
Qwen2.5 7B7B$0.20$0.60$0.40
Llama 3.1 8B8B$0.25$0.75$0.50
Gemma 4 26B (MoE)26B · 4B active$0.30$0.90$0.60
Gemma 4 31B31B$0.50$1.50$1.00

What counts as what

  • Training tokens. Your spec and examples, tokenized, for the first train and every retrain. Submitting feedback is free until a retrain consumes it.
  • Input tokens. Everything you send in messages.
  • Output tokens. Everything the model generates, streamed or not.

Credits

Top up on the Billing page: $5, $25, $100, or a custom amount from $5. Every gateway call draws its token cost from the balance, and the ledger on that page shows each charge. Auto top-up can recharge your saved card when the balance runs low. Concierge customers on invoice billing are unaffected by credits.

A typical support deployment on Gemma 4 26B, a thousand conversations a day at about 2,000 tokens each, runs around a dollar a day.