Why it exists
The same input stopped producing the same output.
Software testing rests on an assumption so old it is rarely stated: run the same input twice and you get the same output twice. Assert on the output, and you have a test.
An agent breaks that assumption on purpose. Ask it the same question twice and it will answer well both times, in different words, sometimes by a different route, occasionally in a different language. Traditional testing does not get harder here. It stops measuring anything.
So agent products get shipped on judgement instead. Somebody tries the thing a few times by hand, it behaves, and that becomes the evidence. It is not evidence, and everyone using it knows that.
Arena is the proving ground that replaces it: a population of synthetic situations — a call, an email thread, a meeting — run against the real agent and graded on what it did, not on the words it chose. Run enough times, that is a number, and a number is something you can hold a release to.
You cannot assert on the words. You have to grade the behaviour. The shift Arena is built around
The rest of the engine
One engine, five parts.
You are looking at one of them. They are not a stack you buy in order: each part runs our own companies first, and each one stands on its own. Here are the other four.
-
Agent harness
Airmond Live
Proactive agents that live where you do: Slack, iMessage, email, and on the phone.
-
Agent orchestration
Tartare Runs our own fleet
Every agent we run, ours or a customer’s: runners, credentials, guardrails, dashboards.
-
Agent fleet management
Latch Runs our own fleet
A list of features becomes a sequenced roadmap, and one person runs a fleet of dozens of agents against it.
-
Bare-metal deployments
Metal Alpha
Local models and local inference, orchestrating work on bare metal we rent.