Skip to content

Platform · Deployment

Hosted, your cloud account, on premises or air gapped. Same binary, same recipes.

Every deployment model runs the same signed artifacts through the same control plane. The difference is where the GPUs are and who holds the keys, and that is a policy decision you make once.

Who this is forInfrastructure leads choosing where inference should live, and procurement teams that need the options on one page.

A dark datacenter aisle with rows of racks receding to a vanishing point, status lights in green and cyan.

In the console

Enrolling a machine with one command. The signature is verified, kernels are chosen for the silicon, the recipe warms, the readiness gate passes and the node starts serving.

Enrolling a machine with one command. The signature is verified, kernels are chosen for the silicon, the recipe warms, the readiness gate passes and the node starts serving.

Metrale Console. Demo data, recorded from the product mockup.

Deployment models

Hosted with private connectivity

Metrale operates the GPUs. Your applications reach them over private connectivity from your VPC with no public endpoint. Lowest friction, strongest protection for the kernel binaries.

Fits · Fast start, no GPU estate of your own

Dedicated

A dedicated GPU pool, cluster or VPC operated by Metrale for one customer. Promote from shared to dedicated without changing the API.

Fits · Enterprise isolation without an ops team

Bring your own cloud

The engine and router deploy into your AWS, Azure or GCP account on your GPU node pools through Terraform or Helm. No inference request leaves your account.

Fits · Regulated data, existing cloud commitments

On premises

Your datacenter, your racks, your network. One binary per node, signed recipes, the control plane inside your perimeter.

Fits · Owned GPU fleets, sovereign requirements

Air gapped

Signed artifacts installed from local media. Telemetry stays inside the network and exports on your schedule, or never.

Fits · Defense, classified and isolated networks

Workstation and edge

A DGX Spark or Strix Halo class box under a desk or in a branch office, licensed per box, managed by the same control plane.

Fits · SMB, branch offices, field deployments

Onboarding

What onboarding looks like

  1. 01

    Choose hosted, your cloud, on premises or air gapped

  2. 02

    Select region, GPU profile, models and capacity limits

  3. 03

    Metrale validates the environment and generates the deployment

  4. 04

    GPUs provision, the runtime selects the kernels for the silicon it finds

  5. 05

    Models warm, health checks pass, you receive an endpoint

  6. 06

    Routing, observation, repair, scaling, canary and rollback run from then on

You never have to understand CUDA compute capability, driver compatibility, Kubernetes GPU plugins or kernel rollout. Those are our problems.

Isolation

Isolation is a setting

TierIsolation
SharedNamespace and logical tenant
EnhancedDedicated GPU node pool
EnterpriseDedicated cluster
RegulatedDedicated VPC and cluster
Bring your own cloudYour account, your keys

Three owners, one platform

What changes between the three is who owns the ground, not what runs on it.

Proposed. From the platform architecture brief, September 2026. Every column runs the same operator, the same signed recipes and the same engine.

CapabilityMetrale CloudYour cloudOn premises and air gapped
Operator and signed recipesYesYesYes
Gateway and quotasMetrale managedYour package, or Metrale managedLocal
Central control planeNativeAn agent that only calls outConnected, or a local subset
Where the data livesThe Metrale region you pickYour cloud accountYour datacenter
Internet needed to serveService dependentNoNo
Serving through a control plane outageNot applicableContinuesContinues, fully air gapped if you choose

Questions

The questions we actually get asked.

Short answers. Each one is backed by something on this site or in the repository.

Do I have to replace vLLM, SGLang or llama.cpp to use it?

No. Metrale deploys beside your current engine and takes traffic one model family at a time. Many teams keep the old engine for a family we do not ship a recipe for yet. You decide what moves, on your own side by side numbers.

How long does a deployment take?

One binary per node and one signed recipe per model. A single box runs in minutes from one install command. A fleet pilot has a side by side ladder against your current engine in week one. Production cutover is workload by workload over the following weeks, at your pace.

Can it run air gapped?

Yes. The engine is one binary with no runtime download and no Python environment to resolve. Recipes, models and kernels are delivered as signed artifacts and installed from local media. Telemetry can stay entirely inside the network and export on your schedule, or never.

Does it run in my cloud account?

Yes. Bring your own cloud deploys the engine and router into your AWS, Azure or GCP account, on your GPU node pools, through Terraform or Helm. The control plane sees configuration, licensing, versions and aggregate telemetry, and nothing else. Regulated buyers can run the customer pull GitOps mode, where Metrale never holds credentials to your account.

Which APIs does it expose?

OpenAI compatible chat and completions, the Anthropic Messages API and the Responses API, from the same binary, so existing SDKs, agents and gateways point at Metrale without code changes.

What support comes with it?

Community support in Discord for the open source engine. Enterprise includes a named engineer, a response SLA and a shared channel. Forward deployed engineering for the pilot and the cutover is scoped per engagement and credited against the first year on conversion.

Ask the rest in a working session, or read the deployment guide ↗.

Next step

See it against your own workload.

A side by side ladder on your hardware in week one. Your models, your criteria, your receipt.