Usage and billing
Usage is metered per model and paid from prepaid credits you top up with a card.
Metering
The gateway counts input and output tokens on every request and attributes them to your API key. The Usage page breaks this down per model and per day. There are no seats, minimums, or idle charges in the pilot: a model that serves nothing records no inference usage.
Pricing
All prices are per million tokens.
| Model | Params | Input | Output | Training |
|---|---|---|---|---|
| Qwen2.5 7B | 7B | $0.20 | $0.60 | $0.40 |
| Llama 3.1 8B | 8B | $0.25 | $0.75 | $0.50 |
| Gemma 4 26B (MoE) | 26B · 4B active | $0.30 | $0.90 | $0.60 |
| Gemma 4 31B | 31B | $0.50 | $1.50 | $1.00 |
What counts as what
- Training tokens. Your spec and examples, tokenized, for the first train and every retrain. Submitting feedback is free until a retrain consumes it.
- Input tokens. Everything you send in
messages. - Output tokens. Everything the model generates, streamed or not.
Credits
Top up on the Billing page: $5, $25, $100, or a custom amount from $5. Every gateway call draws its token cost from the balance, and the ledger on that page shows each charge. Auto top-up can recharge your saved card when the balance runs low. Concierge customers on invoice billing are unaffected by credits.
A typical support deployment on Gemma 4 26B, a thousand conversations a day at about 2,000 tokens each, runs around a dollar a day.