Energy-priced inference API

One rate for every model.
Billed in kilowatt-hours.

Traversaal Energy is an inference API that bills the energy your requests consume. Every response comes back with what it used, in watt-hours and in dollars.

$10/kWhFlat, every model
88%Under token rates
0.22sMedian time to first token

Same rate for a 31B
and a 400B.

Per-token prices bake each model’s markup into the unit itself, so a bigger model costs more per token by decree. A kilowatt-hour costs what a kilowatt-hour costs. Price on that and an efficient model becomes genuinely cheaper to call, not just cheaper on the rate card.

See it on a live request

Inference priced by
the kilowatt-hour.

01

Legible

Every response carries the watt-hours it burned, the dollars that came to, and the share of your prompt served from cache. Cost stops being something you reverse-engineer from an invoice.

Reported per request
02

Flat

One rate across the whole catalogue. Moving from a 31B to a 400B changes how much energy you consume. It never changes what you are charged per unit of it.

No per-model markup
03

Rewarding

Trim a prompt, raise your cache rate, switch to a sparser model, defer batch work to the flex tier. Under energy pricing every one of those lands on your bill instead of our margin.

Optimisation pays you

Send a prompt. Read the meter.

Switch models and watch both numbers move: what the request cost in energy, and what the identical call would have cost per token.

Demo / GLM-5.2
Preview

Start a conversation. Every response comes back with the watt-hours it consumed and what that cost.

Preview only. Completions are canned and served from the browser, not a live endpoint. Token prices are the posted rates for these models; energy figures are modelled from measured consumption under batched serving, at a 31% cache hit rate. Real requests meter their own energy.

What your month looks like, metered.

Put your volume and your current token rate side by side. The energy line is metered kilowatt-hours times the flat rate, and nothing else.

  • 01 Set your monthly token volume
  • 02 Enter the token rate you pay now
  • 03 Compare against energy pricing
Estimate your bill
$10.00 / KWH

Energy pricing, same usage

$108 / mo
$70.20 on flex80% cheaper$540 on tokens
10.8 kWh

metered per monthat 0.09 Wh per 1K tokens

An estimate, not a quote. Energy per token moves with model architecture, request size, cache hit rate, and cluster load. Sparse models land well under this line, long-context reasoning runs above it. Your real bill meters each request.

Three things every layer is built around.

Metering you can read, serving that earns its watts, and a rate card that rewards efficiency instead of taxing it.

01Reporting

Energy reporting, on every plan.

Watt-hours and cost on each completion, per-key trends in the dashboard, and side-by-side efficiency across models. Included even if you stay on per-token billing.

02Performance

Serving built for throughput.

Continuous batching, tensor parallelism across GPUs, and prefix caching that skips recomputation on repeated context. 0.22s median time to first token on the fast tier.

03Efficiency

More output per kilowatt-hour.

Sparse mixture-of-experts models activate a fraction of their weights per token, so they meter far below a dense model of the same size, and under a flat rate you keep that difference.

“Every token you buy is a quantity of energy someone else measured, marked up, and rounded off. We would rather just show you the measurement.”
OpenAI-compatibleOne endpoint, many modelsStreaming, tools, JSONReporting on every plan

Point your base URL at us.

Tell us the models and volume you run today and we will come back with what it meters at.