Skip to content

About Metrale

It started with two words.

In January 2026 a working improvement to llama.cpp was closed because it had been written with AI. The author wrote a short appeal to common sense, and when it went over everyone’s heads, answered the room with a question. Then he went and built the engine from scratch.

How it started

The exchange

Two comments on GitHub. XMR13 asks: why are you so aggressive? Your PR looks completely 100% AI generated. tbraun96, the author, replies: Your point? Two comments on GitHub. XMR13 asks: why are you so aggressive? Your PR looks completely 100% AI generated. tbraun96, the author, replies: Your point?
What came before it Two earlier comments. A maintainer writes that the pull request appears to contain substantial AI generated code without disclosure and cannot be accepted in its current form. The author replies that it works on the DGX Spark, that the requirement is like the 1960s when people thought compilers were sketchy, and that whether AI or a compiler, both translate one language to another. Two earlier comments. A maintainer writes that the pull request appears to contain substantial AI generated code without disclosure and cannot be accepted in its current form. The author replies that it works on the DGX Spark, that the requirement is like the 1960s when people thought compilers were sketchy, and that whether AI or a compiler, both translate one language to another.
llama.cpp pull request 18680, January 7 and 8, 2026 (UTC). Captured from the public thread. The whole thread ↗

697 stars later, the repository those two words started runs on hardware from NVIDIA and AMD, merged a kernel into Hugging Face Transformers, and out serves the incumbent on the published ladder. The answer to the question is this website.

January 7, 2026

The pull request

A loop attention model ran on a DGX Spark. The pull request adding it to llama.cpp was closed as containing AI generated code without disclosure. The reply argued that whether AI or a compiler, both translate one language to another, and that a community building AI tooling should not hold contempt for AI written code.

Read the thread ↗

January 8, 2026

“Your point?”

Asked why the pull request looked entirely AI generated, the author answered with two words. They became the first principle of the repository that followed. AI authored is the default. A human who writes code by hand explains why they were better than the machine.

Winter 2026

From scratch, in Rust

Months of trying to improve vLLM on the Spark had shown that the feedback loop from a kernel change to a number was too slow to learn from. The serving stack was rewritten in Rust with hand tuned CUDA, no Python, and a build that takes a minute instead of forty.

May 2026

One Reddit post

A stable 102 tokens per second on a DGX Spark, posted to r/LocalLLaMA. The star count went from a few dozen to a few hundred in a week and the Discord became the test fleet.

July 2026

Receipts

The fused Qwen Gated DeltaNet kernel merged into Hugging Face Transformers. MLCommons named the project a contributor to the new MLPerf edge agentic benchmark. AMD provided a Strix Halo desktop and the MLPerf submission went in from the same CUDA source on both vendors.

August 2026

The ladder

The concurrency ladder against the matched vLLM configuration was published with every rung lost on the way. Eight rungs, eight wins, and the margin widest at C=128.

September 2026

Metrale

The company took a new name, built on metron, the Greek word for measure, brought in commercial leadership that had scaled Anaconda, and set out to sell what the engine had proved. The industry measures inference in tokens per second. Metrale measures what that performance is worth.

What the name means

Measure first.

The name Metrale is built on metron, the Greek word for measure. It is the root of meter, metric, geometry and symmetry, and further back, of moon and month, the first measures of time. Protagoras used it when he called man the measure of all things. It is the oldest word for knowing how much.

Economies were run on instinct until they were measured. In 1934 Simon Kuznets gave the United States Congress its first national income accounts, and within a decade those accounts were how nations were compared and managed. Today the number is called GDP. Kuznets warned in the same report that the welfare of a nation can scarcely be inferred from its income. The number was a beginning, not a verdict.

Inference is where economies were then. Fleets report tokens per second the way a mill once reported spindle speed, and few can say what a GPU hour produced, what a workload cost, or when the fleet paid for itself. Metrale keeps those accounts: the gross product of inference, by model, cluster and business unit, against the baseline you ran before. Measure first, then manage.

Mission

Same silicon. Smarter inference. Stronger scalability.

AI worth having should run on hardware you own, whether that is an accelerator at the edge, the workstation under your desk, or a rack you operate. We build one engine for the whole range, verify it on the silicon we can put our hands on, and make what it produces accountable to the people paying for it.

Receipts, not adjectives

Every performance number on this site is generated from a record in the repository. If it is not in the repo, it is not on the page.

AI first, human accountable

AI authored is the default in the repository. Certified benchmarks gate every kernel change. People decide what ships.

Own the request path

Security, governance and economics are only exact when the engine is under the workload. Everything we build follows from that.

Open at the core

The Community Edition is AGPL-3.0 and always will be. The enterprise platform pays for the people who keep it that way.

Next step

Come build with us, or come buy from us.

Both conversations start the same way. Tell us what you run.