Platform · Deployment
Hosted, your cloud account, on premises or air gapped. Same binary, same recipes.
Every deployment model runs the same signed artifacts through the same control plane. The difference is where the GPUs are and who holds the keys, and that is a policy decision you make once.
Who this is forInfrastructure leads choosing where inference should live, and procurement teams that need the options on one page.

In the console
Enrolling a machine with one command. The signature is verified, kernels are chosen for the silicon, the recipe warms, the readiness gate passes and the node starts serving.
Metrale Console. Demo data, recorded from the product mockup.
Deployment models
Hosted with private connectivity
Metrale operates the GPUs. Your applications reach them over private connectivity from your VPC with no public endpoint. Lowest friction, strongest protection for the kernel binaries.
Fits · Fast start, no GPU estate of your own
Dedicated
A dedicated GPU pool, cluster or VPC operated by Metrale for one customer. Promote from shared to dedicated without changing the API.
Fits · Enterprise isolation without an ops team
Bring your own cloud
The engine and router deploy into your AWS, Azure or GCP account on your GPU node pools through Terraform or Helm. No inference request leaves your account.
Fits · Regulated data, existing cloud commitments
On premises
Your datacenter, your racks, your network. One binary per node, signed recipes, the control plane inside your perimeter.
Fits · Owned GPU fleets, sovereign requirements
Air gapped
Signed artifacts installed from local media. Telemetry stays inside the network and exports on your schedule, or never.
Fits · Defense, classified and isolated networks
Workstation and edge
A DGX Spark or Strix Halo class box under a desk or in a branch office, licensed per box, managed by the same control plane.
Fits · SMB, branch offices, field deployments
Onboarding
What onboarding looks like
- 01
Choose hosted, your cloud, on premises or air gapped
- 02
Select region, GPU profile, models and capacity limits
- 03
Metrale validates the environment and generates the deployment
- 04
GPUs provision, the runtime selects the kernels for the silicon it finds
- 05
Models warm, health checks pass, you receive an endpoint
- 06
Routing, observation, repair, scaling, canary and rollback run from then on
You never have to understand CUDA compute capability, driver compatibility, Kubernetes GPU plugins or kernel rollout. Those are our problems.
Isolation
Isolation is a setting
| Tier | Isolation |
|---|---|
| Shared | Namespace and logical tenant |
| Enhanced | Dedicated GPU node pool |
| Enterprise | Dedicated cluster |
| Regulated | Dedicated VPC and cluster |
| Bring your own cloud | Your account, your keys |
Three owners, one platform
What changes between the three is who owns the ground, not what runs on it.
Proposed. From the platform architecture brief, September 2026. Every column runs the same operator, the same signed recipes and the same engine.
| Capability | Metrale Cloud | Your cloud | On premises and air gapped |
|---|---|---|---|
| Operator and signed recipes | Yes | Yes | Yes |
| Gateway and quotas | Metrale managed | Your package, or Metrale managed | Local |
| Central control plane | Native | An agent that only calls out | Connected, or a local subset |
| Where the data lives | The Metrale region you pick | Your cloud account | Your datacenter |
| Internet needed to serve | Service dependent | No | No |
| Serving through a control plane outage | Not applicable | Continues | Continues, fully air gapped if you choose |
Questions
The questions we actually get asked.
Short answers. Each one is backed by something on this site or in the repository.
Do I have to replace vLLM, SGLang or llama.cpp to use it?
No. Metrale deploys beside your current engine and takes traffic one model family at a time. Many teams keep the old engine for a family we do not ship a recipe for yet. You decide what moves, on your own side by side numbers.
How long does a deployment take?
One binary per node and one signed recipe per model. A single box runs in minutes from one install command. A fleet pilot has a side by side ladder against your current engine in week one. Production cutover is workload by workload over the following weeks, at your pace.
Can it run air gapped?
Yes. The engine is one binary with no runtime download and no Python environment to resolve. Recipes, models and kernels are delivered as signed artifacts and installed from local media. Telemetry can stay entirely inside the network and export on your schedule, or never.
Does it run in my cloud account?
Yes. Bring your own cloud deploys the engine and router into your AWS, Azure or GCP account, on your GPU node pools, through Terraform or Helm. The control plane sees configuration, licensing, versions and aggregate telemetry, and nothing else. Regulated buyers can run the customer pull GitOps mode, where Metrale never holds credentials to your account.
Which APIs does it expose?
OpenAI compatible chat and completions, the Anthropic Messages API and the Responses API, from the same binary, so existing SDKs, agents and gateways point at Metrale without code changes.
What support comes with it?
Community support in Discord for the open source engine. Enterprise includes a named engineer, a response SLA and a shared channel. Forward deployed engineering for the pilot and the cutover is scoped per engagement and credited against the first year on conversion.
Ask the rest in a working session, or read the deployment guide ↗.
Next step
See it against your own workload.
A side by side ladder on your hardware in week one. Your models, your criteria, your receipt.