Skip to content

Metrale Labs

The research arm.

Labs is where the engine gets its next order of magnitude. Kernels and compression, speculative decoding, memory and context, compilers and languages, protocols, agentic benchmarks, and the day zero bring up of every model that matters. The work lands in the open repository as pull requests with certified benchmarks.

A bright open plan research office with shared desks, whiteboards of diagrams and people working at the far end.

Research tracks

Kernels and quantization

Hand tuned attention, MoE, GDN and quantized GEMM per hardware target. NVFP4, FP8 and K quant expert kernels on raw blocks. TurboQuant+ KV cache compression.

Speculative decoding

MTP draft heads, DFlash block diffusion, lookup drafts into a wide verify, and a resolver that picks the verify width itself.

Memory and context

Tiered KV and SSM state across host RAM, NVMe and RDMA peers. Prefix caches that are prefilled once per fleet, not once per node. The context bus between agents.

Compilers and languages

One CUDA source compiled for NVIDIA and AMD through SCALE. Rust as the systems language. Interest in massively parallel functional runtimes, Bend and HVM among them, for the next abstraction.

Protocols and transport

Node to node links designed for a hostile network, one sided RDMA primitives shared by every tier, and the Citadel protocol lineage the founder brought to the company.

Agentic benchmarks

Contributor to the MLPerf edge agentic benchmark. BFCL and replayed agentic trajectories as gates, because a benchmark should look like the work.

Day zero model bring up

DeepSeek V4.1 Flash, Kimi K3, GLM 5.3, Qwen 3.8 Flash Next and Gemma 4 in open pull requests. The goal is that model vendors check Metrale the same week they check vLLM.

Inference economics

The correlation layer between operator telemetry and the ledger. Living benchmarks on the latest hardware and the latest models, because costs for equivalent quality keep falling and the measurement has to keep up.

Direction, not a shipped product. Everything below is labeled as where the work is going.

What we are building toward

Labeled as direction, not as a shipped product. A fleet where every idle machine on the network joins the mesh overnight as cache and compute. A repository that merges, corrects and versions itself around the clock with certified gates. A tuner that takes any hardware in any configuration, any model in any version, and has it running fast and stable on day zero. Each of these is a step we are taking now because the things we have to build today are on the way there.

Next step

Working on something adjacent?

Kernels, compilers, protocols, benchmarks, safety evaluation. If it advances inference on hardware people own, we want the conversation.