Agent testing Alpha

Test an agent before a customer meets one.

The proving ground. Synthetic calls, emails and meetings that put an agent through the day before anyone real does.

Arena is alpha. It is a specification and a set of arguments, not software. Nothing described on this page is running, and there is nothing here to sign up for. We are publishing the thinking while it is still thinking.

Why it exists

The same input stopped producing the same output.

Software testing rests on an assumption so old it is rarely stated: run the same input twice and you get the same output twice. Assert on the output, and you have a test.

An agent breaks that assumption on purpose. Ask it the same question twice and it will answer well both times, in different words, sometimes by a different route, occasionally in a different language. Traditional testing does not get harder here. It stops measuring anything.

So agent products get shipped on judgement instead. Somebody tries the thing a few times by hand, it behaves, and that becomes the evidence. It is not evidence, and everyone using it knows that.

Arena is the proving ground that replaces it: a population of synthetic situations — a call, an email thread, a meeting — run against the real agent and graded on what it did, not on the words it chose. Run enough times, that is a number, and a number is something you can hold a release to.

You cannot assert on the words. You have to grade the behaviour. The shift Arena is built around