The landscape

A vendor who will not name competitors is hiding something.

So here is every serious alternative to us, by name, with an honest account of where each one is genuinely better than we are — and exactly where it stops. Most of them, you should buy.

Why the coding-agent playbook does not transfer

Coding tools are extraordinary because code has a free oracle: the tests pass or they do not. The agent can try, check and retry all night without a human in the loop.

The rest of the enterprise has no oracle. There is no unit test for “was that the right claim to deny”, “was that the right bid”, or “was that the right candidate”.

We build the oracle. That is what the captured expert method and the bench actually are — and it is why the companies who tried to copy the coding-agent playbook into the back office have been disappointed.

Every category, named

Product facts below were checked against current public documentation on 14 September 2026.

Agentic coding tools

Claude Code, OpenAI Codex, Cursor, GitHub Copilot, Cognition's Devin Desktop (formerly Windsurf)

Where they are genuinely better

Genuinely excellent, and we say so. Repo-wide reasoning, terminal-native execution, tight feedback loops, and productivity gains that show up in shipped software. If you have engineers, buy these. We use them ourselves.

Where they stop

They are aimed at work that has an oracle — the tests pass or they don't — and a shared professional idiom. Their steering layer has grown up: rules files, session memory, organisation-level custom instructions. But that captures a team's conventions, not how your specific person decides. And none of them, in their current public documentation, grades output against a named human expert's withheld judgment; what they ship is rubric-based model-as-judge evaluation and automated code review.

Complement, never competitor. Keep them. They don't do this and they aren't trying to.

Seats for everyone, plus enablement

Claude, ChatGPT and Gemini rolled out to all staff

Where they are genuinely better

This is often the highest-ROI first move a company can make, and we will say that to your face. Cheap, immediate, broad, no integration project. If you have not done it, do it before you talk to us.

Where they stop

Output quality tracks the median employee's prompting skill — your best expert gets a little faster, your weakest one gets confidently wrong faster. Shared projects and org-level instructions capture some house style, but not how your best person actually decides, so nothing is consistent and nothing compounds: month twenty-four looks like month one plus a better model. And the auditability on offer is usage auditability — every major enterprise tier now exports who asked what and when, into your compliance tooling. What none of them provides is reasoning provenance: a record of why a call was made, against whose standard, that you could defend later.

Seats make your people faster. They don't make your company keep anything. When your best adjuster retires, the seat stays and the judgment leaves.

Horizontal enterprise assistants

Microsoft 365 Copilot, Gemini Enterprise, Glean

Where they are genuinely better

Genuinely strong at the thing that is genuinely hard: secure, governed access to your own data across mail, files, chat and tickets. Enterprise search over a real corpus is a serious engineering achievement, procurement is easy, and it is often already in the contract. All three now let you assemble agents on top of that corpus, not merely search it.

Where they stop

The agents they build automate what your company has already written down. They answer “what did we say?” and “what is the documented process?” — not “how would our best person have judged this?” They are also tuned to serve everyone, which means tuned to the median.

We are a consumer of their plumbing, not a replacement for it. Their retrieval feeds our capture.

Platform agents where the data lives

Salesforce Agentforce 360, ServiceNow AI Agent Studio and Orchestrator, SAP Joule Studio, Palantir AIP

Where they are genuinely better

Real strengths. Agents inside the system of record with governed actions is the correct architecture for structured business process. Palantir in particular brings genuine ontology and operational depth, plus a services model that gets things deployed in hard environments.

Where they stop

They encode the process, which the company already knows. The workflow was never the hard part. The value we chase is the deviation from the process that the good ones make.

Vertical AI products

Harvey, Abridge, Sierra, Decagon, EvenUp, and the rest

Where they are genuinely better

The most honest section here. These are strong companies with deep domain investment and, in several cases, serious public eval harnesses — Harvey publishes BigLaw Bench, Sierra publishes τ-bench. The customer-service agent companies in particular have built much of what we are describing, for one domain. That is a validation of the model, not a refutation of it. They also customise per account: firm playbooks, agent operating procedures, per-clinician note style.

Where they stop

The customisation is organisational, not personal. You configure a firm's policy, its tone, its procedures — but the starting point is still the field's consensus, because a vertical product has to serve every customer in the field. For eighty percent of the work, industry-standard is fine and you should buy theirs. For the twenty percent where your expert's deviation is your margin, it cannot help — and the vendor cannot build it for you without building it for your competitor. You can also only buy one for a field somebody has already productised.

The question to ask is simple: does it do it your way, or the industry's way?

Build it yourself

LangChain and LangGraph, LlamaIndex, DSPy, Temporal, plus APIs

Where they are genuinely better

Genuinely viable for a company with a strong engineering team. The frameworks are real and mature — LangChain and LangGraph reached 1.0, DSPy-style automatic prompt optimisation is a real and actively advancing technique, and Temporal now markets durable execution for agents directly. You own everything.

Where they stop

They give you the machinery, not the method. The framework was never the hard part — capture and evaluation are, and every in-house team we have seen underestimates both by roughly an order of magnitude. They ship a working pipeline in six weeks and then spend eighteen months unable to answer “is it actually as good as Maria?”, because nobody built the oracle. Nothing in these frameworks tells you how to elicit one person's judgment, or how to know when you have captured it.

If you have the team and the patience, consider it. We would rather sell you the bench than lose you entirely.

Fine-tuning and custom training

Supervised tuning and reinforcement fine-tuning

Where they are genuinely better

Sometimes the right answer, and our own doctrine says so out loud: if you could fine-tune it, fine-tune it. With clean labelled pairs, training can beat a workflow — and the data bar has come down. OpenAI's own reinforcement fine-tuning guidance tells you to start with several dozen to a few hundred examples, though Google's Vertex tuning still asks for a hundred as a floor. We will tell a prospect this and disqualify ourselves from that use case.

Where they stop

A tuning set needs pairs that already exist. There is no corpus of how one particular person decides, and no vendor today markets tuning as a way to capture one named individual's judgment. The nearest thing on the market learns a person's writing voice from their files — that is style, not judgment.

The consultant who says they will build it better

A good independent, or a boutique shop

Where they are genuinely better

Sometimes true. A good independent building one narrow workflow for a motivated client can absolutely ship something great, faster and cheaper than an enterprise vendor.

Where they stop

What almost never comes with it: an evaluation harness, a repeatable capture methodology, provenance on every rule, maintenance after the invoice clears, and any way to prove the result beyond “looks good to me.”

Take this with you

Seven questions to ask any AI vendor

Use these on everyone you are considering. Use them on us first — that is the point of publishing them.

Download the checklist
  1. 1

    How will you capture what my expert knows but cannot articulate? Name the technique.

  2. 2

    How will you prove the output matches my expert, rather than a competent generalist?

  3. 3

    Show me the comparison against a top frontier model given the same material and a hundred times the budget.

  4. 4

    Who grades it, and how did you prove the grader agrees with my expert?

  5. 5

    What happens when my expert changes their mind?

  6. 6

    Who owns the captured method — me, or you?

  7. 7

    What does it cost per run, and what is the latency?

A vendor who cannot answer three and four has not built a system. They have built a demo.

What you get, depending on the road you take

These are paths, not enemies. Most companies should be on several of them at once — and we are the last column, not a replacement for the first four.

Comparison of five paths to enterprise AI across nine dimensions
DimensionSeats + trainingHorizontal assistantVertical productBuild in-houseAI Matrx Masterwork
Time to first valueDaysWeeksWeeks6–18 monthsWeeks per workflow
Captures tacit judgmentNoNoNoRarelyThe entire point
Consistent across staffNoPartialYesYesYes
Matches your expert, not the industry normNoNoFirm-level at bestMaybeYes
Compounds from your correctionsNoNoThe vendor's benefitIf you build itYes
Provable against withheld ground truthNoNoSometimesRarelyFive-arm bench
Survives the expert leavingNoNoNot applicablePartialYes
Who owns the captured methodNobodyThe vendorThe vendorYouYou
Run costPer seatPer seatPer seat or actionYour tokensCheap model, measured per run

This table describes the typical shape of each path, not a claim about any single product at a single moment. Vendor capabilities in this market move monthly; if you believe a cell is out of date, tell us and we will check it and change it.

How we engage

Four tiers. The first one exists to tell you whether you need the other three.

Audit

We run your candidate workflows through our gate and tell you which are worth capturing, which you should just buy seats for, and which you should fine-tune instead. We will disqualify workflows in writing. This tier sells the next one by being honest in this one.

Single Masterwork

One expert, one workflow, the full capture stack, the five-arm proof, and the expert's own sign-off at the end of it.

Vertical programme

Several experts across one function — including the ones who disagree with each other, kept apart and routed rather than averaged into a bland middle.

Capture platform

You run your own captures on our stack, with our protocols and our bench.

Our posture, in three sentences

We compete with almost none of these. We sit on top of the plumbing you have already bought, aimed at the twenty percent of work where your people's judgment — not the industry's — is the whole value.

Buy the seats. Keep the Copilot. Then point us at the six people whose judgment you can't afford to lose.

Start with the audit.

We will tell you which of your workflows are worth capturing and which you should simply buy seats for — in writing, including the ones we disqualify.