Skip to content

Pricing

Priced against productive GPU capacity, not seats.

Enterprise AI infrastructure software already prices per GPU per year. Metrale sits inside that range, ships with the control plane and the economics layer, and shows its payback on this page. Proposed list prices, September 2026.

Proposed sheet · September 2026

A small desktop compute box on a dark desk with a lavender status light.

Waitlist open
Community Edition

$0AGPL-3.0, forever

The engine and every recipe, free. For developers, labs and anyone running open models on hardware they own. Not released yet.

  • Metrale Engine, full source
  • Every model recipe in atlas-recipes
  • OpenAI, Anthropic and Responses APIs
  • LAN fleet manager, early access
  • Community support in Discord
Join the waitlist

Workstation and edgePROPOSED

$50per box per month, billed annually

A DGX Spark or Strix Halo class box serving an office, a branch or a field team. Commercial license, signed update channel, managed from the console.

  • Commercial license per box
  • Signed stable and LTS channels
  • Console access for every licensed box
  • Email support, next business day
  • Volume pricing from 25 boxes
Price a fleet of boxes

Proof of value

Fixed feefour weeks, credited on conversion

One model, one hardware target, one workload. A side by side ladder in week one and a receipt in dollars per workload at the end.

  • Scoped success criteria, yours or ours
  • Side by side against your current engine
  • Economics baseline of the target cluster
  • Forward deployed engineer for the four weeks
  • Fee credited against the first year on conversion
Scope a pilot

Market anchor

Where it sits.

Established enterprise AI infrastructure software already prices against GPU capacity. Metrale lists inside the range and includes the layers the others sell separately.

ProductListBasis
Red Hat AI Inference Server
Hardened vLLM, published list price
~$2,500per accelerator per year
NVIDIA AI Enterprise
Broad platform, OEM backed
$4,500per GPU per year
Metrale Enterprise
Engine, control plane and economics, realized $1,800 to $2,400 at scale
~$3,000per GPU per year, list

Third party prices are public list prices at the time of writing and belong to their owners. Metrale prices are proposed and subject to contract.

Illustrative contract economics

What a fleet costs to license.

At realized fleet scale pricing. Illustrative, not a forecast.

64 GPUs≈ $175K
256 GPUs≈ $550K
1,000 GPUs≈ $2.0M

What the platform meters

Priced on a number you can watch.

Proposed. The platform is being built to meter what it serves and show it live, so a renewal is read off the same number the console shows.

GPU hours and GPU count

A license is a cluster or a number of GPUs, discovered by the platform, not declared on a form.

Tokens per GPU second

The throughput the fleet actually produced, per GPU, per second, next to the throughput it could have.

Cost per million tokens

GPU, storage, network and platform cost over the tokens delivered, for your fleet and for the baseline you ran before.

Nothing through the gateway without a license

Every served request is entitled and counted, so the bill and the telemetry are the same record.

Payback

Find your payback period.

If a thing costs three thousand dollars and makes you a thousand a month, it pays for itself in three months, and everything after is upside. That is the number to walk to the CFO with. Three scenarios, every input editable, evidence class on every field.

The uplift frees GPUs. Freed GPUs are deferred purchases or rentals plus the power they burned. The license is what the uplift costs.

Uplift defaults to 1.20x, below the measured ratio on the GB10 ladder at C=128, because a datacenter part is not a Spark until we publish the receipt.

Payback period 4 months then $1,133,081 a year is upside
GPUs freed by the uplift
42.7
Deferred purchase or rental
$1,706,667
Power no longer burned
$40,815
Gross savings per year
$1,747,481
Metrale license per year
$614,400
Net per year
$1,133,081
Three year net
$3,399,244

A model, not a quote. Savings depend on your workload, your utilization and the uplift measured on your hardware during the pilot. Evidence classes: MEASURED from ladder.generated.json, the published concurrency ladder. PROPOSED a proposed list price from this page, the team can change it. USER yours to edit, the model recomputes as you type.

Questions

The questions we actually get asked.

Short answers. Each one is backed by something on this site or in the repository.

How is it priced?

Per GPU per year for the Enterprise Edition, with volume tiers as the fleet grows, and a per box license for workstation and edge deployments. Support and forward deployed engineering are priced separately. The Community Edition is free under AGPL-3.0. The pricing page lists the proposed sheet and a payback model with editable inputs.

How do you measure savings?

Against your own baseline. Economics records what each cluster cost per workload before Metrale takes traffic, then reports the delta as traffic moves. The pilot ends with a receipt in dollars per million tokens and dollars per successful workload, not a slide.

What is the payback period?

It depends on your fleet, your utilization and the uplift we measure on your workload. The model on the pricing page computes it from inputs you control. Every month after payback is upside, which is why we talk about payback rather than a percentage.

What about SOC 2 and compliance?

The architecture is built for regulated buyers, and SOC 2 readiness documentation, model risk documentation and pinned recipe governance packs are part of the first SLA engagements. Ask for the current state of the audit program when you book. We will tell you exactly where it is.

What support comes with it?

Community support in Discord for the open source engine. Enterprise includes a named engineer, a response SLA and a shared channel. Forward deployed engineering for the pilot and the cutover is scoped per engagement and credited against the first year on conversion.

Ask the rest in a working session, or read the deployment guide ↗.

Next step

Get the sheet, or get the receipt.

Email [email protected] for the full price sheet, or book a working session and we run the ladder on your workload.