We run the comparison we could lose.
Most AI vendors show you a demo. We put our workflow on a cheap model, put the best frontier model on earth beside it with a hundred times the budget, drop your expert's real work into the pool unlabelled, and let a blind panel sort them out — with dollars and seconds on every single row.
The five-arm bench
Five machine arms, identical prompts, plus your expert's own withheld work sitting blind in the same judging pool. Quality, cost and time are recorded for every arm, every run.
| Arm | What it is given | What it is for | Measured |
|---|---|---|---|
| A0Frontier model, cold | The best model available, a competent prompt, and nothing else. No source material. | Establishes what you already get for free from the tools you have bought. If this arm already makes the unusual calls, the subject fails our gate and we tell you so instead of selling you a trial. |
|
| A1Frontier model, fully briefed | The same model, handed the entire source corpus in context — the books, the manuals, the SOPs, the archive. | Removes “well, it never saw the material” as an explanation for any gap. It saw all of it. |
|
| A2Frontier model, every advantage money buys | The same model plus retrieval, tools, live web access and full agentic scaffolding — run up to a ceiling of a hundred times what our arm costs, with the whole spend curve reported. | The arm that has to be beaten. A ceiling with a published curve is a benchmark; “unlimited tokens” is a shrug, so we set the ceiling and show what each score band cost to reach. |
|
| BCheap model, raw | The same inexpensive model our workflow runs on, with no workflow at all. | The floor. This is the arm that proves the result came from the captured method and not from the model underneath it — without it, any win is unattributable. |
|
| COur workflow on the cheap model | The captured method — the checklists, the decision forks, the severity weights, the negative space — executed by the same cheap model as the floor arm. | The product. Scored on the same rubric, against the same prompts, with its own cost and latency on the row. |
|
| GTYour expert's real work | Work your expert actually produced, withheld from every arm, dropped into the judging pool unlabelled. | Ground truth — and the trial's own integrity check. It is judged blind alongside the machines. | Judged, not timed |
The cheap-model floor arm is the one most vendors skip, and it is the one that makes a win mean anything. Without it, a good result might simply be the model doing well on its own — and nobody could tell the difference, including us.
If your expert does not win the blind panel, the trial is void.
Their real work is in the pool, unlabelled, judged by the same panel on the same rubric as every machine. If the panel does not put the human first, the panel is wrong — or the rubric is measuring the wrong thing, or the cases were badly chosen. Either way the run is thrown out and nothing from it is ever quoted to you. A benchmark whose judge cannot recognise excellence when it is sitting right there is not measuring excellence.
This rule exists to stop us. It is the single easiest way for a vendor to fool a customer, and the single easiest way for a team to fool itself.
What makes it a benchmark and not a demo
Identical prompts on every arm
Every arm is asked for exactly the same thing, in exactly the same words. The separation between arms is made by removing tools and material, and every tool call is logged so the air gap can be audited afterwards rather than taken on trust.
Dollars on every row
Every arm reports what it cost to run. The heavily-resourced arm reports its whole spend curve, so you can see what it had to spend to reach each score band — which is a far more useful number than any single headline.
Seconds on every row
Every arm reports its wall-clock latency too. A result that is better but takes four hours is a different product from one that is better in nine seconds, and we will not let a report hide which one you are being sold.
Blind judging, with the human in the pool
The judges do not know which output came from which arm. Your expert's own withheld work is in the pool as one more unlabelled entry. Before a judge is trusted on two thousand cases, it has to reproduce your expert's calls on twenty.
The result that actually settles it is one sentence long.
“Yes — that's mine.”
Your expert reads the output and recognises their own judgment in it. Including the strange call — especially the strange call, the one their peers would argue with, the one that is the reason you employ them and not somebody cheaper.
Every rule in the captured method cites where it came from: the interview turn, the worked example, the correction. So they can say it rule by rule rather than taking the whole thing on faith — and so can anyone who has to defend a decision later.
What else we track, and hand you
A score on its own is easy to dress up. These are the numbers that make a score honest.
- Cost to parity
- What the heavily-resourced arm has to spend before it matches ours.
- Capture cost
- How many hours of your expert's time the artifact consumed. Their time is the scarcest thing in the engagement and we report it like a cost, because it is one.
- Walls hit
- Every place a real person got stuck inside the capture product. A first-class metric, because the capture experience is the product, not the packaging.
- Stage-by-stage contribution
- Every step of our own pipeline is run with and without itself, quarterly. A step that moves nothing gets removed. Every corner has to show its number.
What you will not find on this page
A score, a multiple, or a sentence about beating a frontier model. We publish those one engagement at a time, from a logged run, and we name the arm and the budget every time we do. Marketing that outruns the bench destroys the bench — so this page describes the method, and the numbers arrive with the trial they came from.
We also will not tell you AI replaces your expert. It does part of their job, their way, so they get their time back. That is what is actually true, it is what people actually want, and it is the only version the expert themselves will cooperate with — and without their cooperation there is no product at all.
Use this on us.
We publish the seven questions we think every buyer should put to an AI vendor — and we expect you to put all seven to us first.