Till, card, cash, bank. Tied out, day by day.
Sumarity checks each day's till report against the card settlements, the cash-up and the bank, reads every supplier invoice, and works out the VAT return from what settled. A missing till report, a chargeback or £20 short reaches a person, with the day it happened.
Code proves the numbers. People settle the doubts.
| Date | Line | Amount | Decided by |
|---|---|---|---|
| 12.09 | Till report · card 2,418.60Equals the card settlement for the day | 2,418.60 | Code · tied |
| 13.09 | Cash banked 1,860.00Cash-ups 11.09 to 13.09 add up to the deposit | 1,860.00 | Code · tied |
| 13.09 | Paid-out · window cleanerA service, receipt kept · 0.93 | −38.50 | Engine · proven |
| 14.09 | Direct debit · food supplierPays the 10.09 delivery invoice · 0.95 | −1,206.35 | Engine · proven |
| 14.09 | No till reportCard and cash taken, but no till report for the day | – | To a person |
| 15.09 | Card chargeback −95.00A guest's card payment taken back | −95.00 | To a person |
|
Settled 4
To a person 2
Settled wrong 0
| |||
Your Z-reports, card settlements and cash-ups, checked against each other
One workflow for every site, each site's runs independent, with problems on an issues list until they're explained. Your accountant gets a quarter that already ties out.
Restaurant groups, pubs, bars and hotels, and the firms that keep their books
Is this deposit those days' cash, or something else?15 of 15 planted problems reached a person: a missing till report, a chargeback, cash short, a card payment rung as cash, takings from the month before.
What was this purchase for, and is any VAT reclaimable?103 of 105 judgements right on a pub held out of tuning; the nine boxes as truth once one mis-keyed invoice was corrected.
Who is owed what from the card tips?
Do the hours paid match the hours recorded?
Is every till export complete for the tax office?
Sumarity for other work: Asset managers Accounting and tax firms Finance teams · everything
wrong answers settled without a person, in every test we have run.
Models are sometimes wrong, ours included. The workflow is built so that a wrong answer has to pass a measured cutoff and a check in code before it can settle. So far, none has.
| Test | Settled wrong |
|---|---|
| Bank reconciliation, nine held-out monthsWording tuning never saw · 108 real exceptions, all surfaced | 0 |
| The same months, three different judgement modelsJudgement Engine and two frontier models in the same seat | 0 |
| Against a 1,154-line program an AI wrote for the jobThree unseen months · 36 of 36 exceptions | 0 |
| Custody statements from a live fund42 real lines first seen after tuning | 0 |
| UK pub VAT returns, a pub held out of tuning105 judgements · 15 of 15 problems reached the preparer | 0 |
| Deliberately wrong picks at 100% confidenceFed to a changed check before it went live · 36 of 36 rejected | 0 |
| Total | 0 |
Operations work is mostly rules. The rest is judgement.
Sumarity gives each part to whatever does it best: code for arithmetic, the Judgement Engine for lightning fast decisions, and people for what stays uncertain.
Describe
The expert writes a paragraph. Sumarity asks the questions that change the design, then writes the workflow.
Expert revisionsValidate
Numbers only from code, every doubt to a person, every path handled. What fails is repaired before anyone sees it.
43 errors → 0 in one repairReview
Every step on a canvas in plain words, editable and versioned. What runs is what was reviewed.
draft → validated → publishedRun
Code matches and ties out. The Judgement Engine decides, and learns when it cannot. People get an inbox.
2,112 lines in 4.5 sTune
Rulings become an answer key. Better questions are proposed, tested on data they haven't seen, and lets you decide.
77% → 97% on unseen linesIn this work, mistakes are expensive. So the rules aren't optional.
These aren't settings a busy team can switch off. The validator refuses a workflow that breaks them.
Wherever the work is mostly arithmetic, with judgement calls in between.
Every one of these has the same shape. Code does the matching and the sums, the Judgement Engine answers one narrow question per item, and whatever stays uncertain goes to a person.
Restaurants and hotels
Restaurant groups, pubs, bars, hotelsEvery day's till, card, cash and bank to tie out, every supplier invoice to read, and a return at the end of the quarter.
Is this deposit those days' cash, or something else?
Is this a valid VAT invoice, and for what?
What was this paid-out for?
Who is owed what from the card tips?
Accounting and tax services
Outsourced bookkeepers, tax preparers, practicesMany clients, the same process each period, and a filing at the end that has to be right.
What was this purchase used for, and is any VAT reclaimable?
Which account does this belong to?
Within policy?
Has the client's answer settled the question?
Finance and accounting
Controllers, shared services, treasuryMatching and tying out, where most lines are easy and the rest need someone who knows the business.
Is this receipt that invoice, paid days apart?
Which invoices does this remittance pay?
A price change, or overbilling?
Does it tie out, and is the variance explained?
Risk and compliance
Compliance, financial crime, onboarding teamsHigh-volume alerts where most are noise, every decision needs a record, and the real ones must never be missed.
The same party, or a namesake?
Worth escalating?
Complete and consistent?
Has anything material changed?
A workflow for every process. In every industry.
Standard workflows, each tested against an answer key before it's published. Start from the closest one and describe how yours differs.
Finance and accounting
Accounting, tax and Treuhand firms
Banking and payments
Wealth, asset management and fund administration
Insurance
Healthcare
Manufacturing and supply chain
Retail, hospitality and consumer
Logistics and transport
Energy and utilities
Telecoms
Real estate
Public sector
Legal and compliance
HR and payroll
Three processes, three domains. Every figure from a run.
Each workflow was authored through the product from an expert's description, then scored against a known answer.
Why not just hand it to an AI? We tried that too.
Same brief, same data, same scoring. These are the three things a capable team would try instead.
Each model got the brief and both exports and returned the finished reconciliation, five times on each of two data sets.
| Sumarity | Opus 5.5 | GPT-6 Sol | |
|---|---|---|---|
| Exceptions (of 108) | 108 | 108 | 95 |
| Same result, 10 runs | 10/10 | 10/10 | 2/10 |
| 2,112-line account | 4.5 s | 262 s | – |
| Record per line | Yes | No | No |
Opus wrote a careful 1,154-line program from the same brief and sample files. Both then ran on three months worded in ways neither had seen.
| Sumarity | AI-written code | |
|---|---|---|
| Exceptions (of 36) | 36 | 36 |
| Settled wrongly | 0 | 0 |
| Judgement items to people | 3 | 14 |
| False alarms | 3 | 14 |
The same workflow and checks, with only the judgement model swapped. 180 decisions with known answers.
| Engine | Opus 5.5 | GPT-6 Sol | |
|---|---|---|---|
| Answered right | 100% | 99% | 100% |
| Median per call | 0.21 s | 1.76 s | 1.47 s |
| Per 1,000 calls | Included | $4.67 | $1.36 |
| Settled wrong | 0 | 0 | 0 |
Reasoning, audit trails and human approval are now common. Checks in code, measured confidence and tested changes are not.
Checked, not just explained
Every call declares how it is checked. What code can't check, people audit by sample, and every doubt goes to a person.
Confidence measured
Cutoffs come from backtests against your team's rulings, not from what a model says about itself.
Changes tested first
Each improvement runs on data it hasn't seen, and against deliberately wrong answers, before an admin can publish it.
Bring one process. Run it alongside yours.
A few weeks in parallel on your own data, at our cost, then a line-by-line comparison: what it settled, what it sent to people, what it caught, and the time it took.
- 01Start from a tuned template, or describe itA paragraph from the person who runs the process today.
- 02Review the workflow with usOn the canvas, step by step, in plain words.
- 03Run it in parallelYour process stays the process of record.
- 04Your rulings tune its callsEvery inbox decision joins the answer key.
- 05Compare, line by lineThen decide, on your own evidence.
Tell us about the process. We'll come back to you.
A few details are enough to start. We'll reply to arrange a conversation with the person who runs the process today.
- We reply to set up a first call.
- We look at the process with your expert, and say whether it fits.
- If it does, we scope a pilot run alongside yours.