Nothing wrong settles. Here's how we know.
Models are sometimes wrong, ours included. Every workflow is built so that a wrong answer has to pass a measured cutoff and a check before it can settle, and every one is tested before it's published.
Wrong answers settled without a person, in every test we've run.
| Test | Settled wrong |
|---|---|
| Bank reconciliation, nine held-out monthsWording tuning never saw · 108 real exceptions, all surfaced | 0 |
| The same months, three different judgement modelsThe Judgement Engine and two frontier models in the same seat | 0 |
| Against a 1,154-line program an AI wrote for the jobThree unseen months · 36 of 36 exceptions | 0 |
| Custody statements from a live fund42 real lines first seen after tuning | 0 |
| UK pub VAT returns, a pub held out of tuning105 judgements · 15 of 15 problems reached the preparer | 0 |
| Deliberately wrong picks at 100% confidenceFed to a changed check before it went live · 36 of 36 rejected | 0 |
| Total | 0 |
How a workflow is tested
An answer key, and data the tuning never saw
Every workflow is tested before it's published. Its questions and cutoffs are tuned on one set of periods or entities and measured on another it has never seen, with the mess real files have and problems planted that a person must catch.
What we measure
- Settled wrong: answers that settled without a person and were wrong. This must be zero.
- Problems reaching a person: every one must.
- Items for people, before and after tuning: how much of the work is left for your team.
- Judgements right, at any confidence: how close a wrong answer came to settling.
Every change, tested the same way
When your team's rulings suggest a better question or cutoff, the change is measured on data it hasn't seen and against deliberately wrong answers before your admin can publish it.
What we claim
The Judgement Engine can be wrong, at high confidence. The claim is that nothing wrong settles: each wrong answer meets a check that refuses it, or a cutoff that sends it to a person. A pilot shows you the same on your own data.
Don't take our word for it. Test it on your data.
Bring one process; we'll run it alongside yours and show you every call.