Sumarity
How we test

Nothing wrong settles. Here's how we know.

Models are sometimes wrong, ours included. Every workflow is built so that a wrong answer has to pass a measured cutoff and a check before it can settle, and every one is tested before it's published.

The record

Wrong answers settled without a person, in every test we've run.

TestSettled wrong
Bank reconciliation, nine held-out monthsWording tuning never saw · 108 real exceptions, all surfaced0
The same months, three different judgement modelsThe Judgement Engine and two frontier models in the same seat0
Against a 1,154-line program an AI wrote for the jobThree unseen months · 36 of 36 exceptions0
Custody statements from a live fund42 real lines first seen after tuning0
UK pub VAT returns, a pub held out of tuning105 judgements · 15 of 15 problems reached the preparer0
Deliberately wrong picks at 100% confidenceFed to a changed check before it went live · 36 of 36 rejected0
Total0
The method

How a workflow is tested

An answer key, and data the tuning never saw

Every workflow is tested before it's published. Its questions and cutoffs are tuned on one set of periods or entities and measured on another it has never seen, with the mess real files have and problems planted that a person must catch.

What we measure

  • Settled wrong: answers that settled without a person and were wrong. This must be zero.
  • Problems reaching a person: every one must.
  • Items for people, before and after tuning: how much of the work is left for your team.
  • Judgements right, at any confidence: how close a wrong answer came to settling.

Every change, tested the same way

When your team's rulings suggest a better question or cutoff, the change is measured on data it hasn't seen and against deliberately wrong answers before your admin can publish it.

What we claim

The Judgement Engine can be wrong, at high confidence. The claim is that nothing wrong settles: each wrong answer meets a check that refuses it, or a cutoff that sends it to a person. A pilot shows you the same on your own data.

Don't take our word for it. Test it on your data.

Bring one process; we'll run it alongside yours and show you every call.