# Results

> Sumarity's measured results: every figure from a run, nothing wrong settled.

Source: https://www.sumarity.ai/results/

Results 

# Every figure from a run. Measured, not promised.

Each workflow below was designed through the product from an expert's description, then scored against a known answer on data its tuning never saw.

How we testThe full track record  

Tested workflows

## Measured, not promised.

Tested

### Bank reconciliation

0settled wrong
108 of 108exceptions surfaced
176 → 7judgement calls sent to a person

Tested

### UK VAT return for a small business (pub VAT)

0settled wrong
15 of 15problems reached the preparer
67 → 25items for people after the run

Tested

### Custody reconciliation

0settled wrong
77% → 97%lines typed right
47% → 58%lines settled without a person  

Everything we've run

## Settled wrong: zero.

| Test | Settled wrong
| Bank reconciliation, nine held-out monthsWording tuning never saw · 108 real exceptions, all surfaced | 0
| The same months, three different judgement modelsThe Judgement Engine and two frontier models in the same seat | 0
| Against a 1,154-line program an AI wrote for the jobThree unseen months · 36 of 36 exceptions | 0
| Custody statements from a live fund42 real lines first seen after tuning | 0
| UK pub VAT returns, a pub held out of tuning105 judgements · 15 of 15 problems reached the preparer | 0
| Deliberately wrong picks at 100% confidenceFed to a changed check before it went live · 36 of 36 rejected | 0
| Total | 0
