Legible
Every response carries the watt-hours it burned, the dollars that came to, and the share of your prompt served from cache. Cost stops being something you reverse-engineer from an invoice.
Reported per requestEnergy-priced inference API
Traversaal Energy is an inference API that bills the energy your requests consume. Every response comes back with what it used, in watt-hours and in dollars.
vs token priceSame request
The idea
Per-token prices bake each model’s markup into the unit itself, so a bigger model costs more per token by decree. A kilowatt-hour costs what a kilowatt-hour costs. Price on that and an efficient model becomes genuinely cheaper to call, not just cheaper on the rate card.
See it on a live request ↓Industry first
Every response carries the watt-hours it burned, the dollars that came to, and the share of your prompt served from cache. Cost stops being something you reverse-engineer from an invoice.
Reported per requestOne rate across the whole catalogue. Moving from a 31B to a 400B changes how much energy you consume. It never changes what you are charged per unit of it.
No per-model markupTrim a prompt, raise your cache rate, switch to a sparser model, defer batch work to the flex tier. Under energy pricing every one of those lands on your bill instead of our margin.
Optimisation pays youTry it
Switch models and watch both numbers move: what the request cost in energy, and what the identical call would have cost per token.
Start a conversation. Every response comes back with the watt-hours it consumed and what that cost.
Preview only. Completions are canned and served from the browser, not a live endpoint. Token prices are the posted rates for these models; energy figures are modelled from measured consumption under batched serving, at a 31% cache hit rate. Real requests meter their own energy.
Estimate
Put your volume and your current token rate side by side. The energy line is metered kilowatt-hours times the flat rate, and nothing else.
Energy pricing, same usage
$108 / mometered per monthat 0.09 Wh per 1K tokens
An estimate, not a quote. Energy per token moves with model architecture, request size, cache hit rate, and cluster load. Sparse models land well under this line, long-context reasoning runs above it. Your real bill meters each request.
Why Traversaal Energy
Metering you can read, serving that earns its watts, and a rate card that rewards efficiency instead of taxing it.
Watt-hours and cost on each completion, per-key trends in the dashboard, and side-by-side efficiency across models. Included even if you stay on per-token billing.
Continuous batching, tensor parallelism across GPUs, and prefix caching that skips recomputation on repeated context. 0.22s median time to first token on the fast tier.
Sparse mixture-of-experts models activate a fraction of their weights per token, so they meter far below a dense model of the same size, and under a flat rate you keep that difference.
“Every token you buy is a quantity of energy someone else measured, marked up, and rounded off. We would rather just show you the measurement.”
Nothing else changes
Change one line
Tell us the models and volume you run today and we will come back with what it meters at.