# Sumarity for accounting and tax firms

> One workflow for every client: VAT returns and bookkeeping checks that tie out, with only the doubts left for your team.

Source: https://www.sumarity.ai/for/accounting-firms/

For accounting, tax and Treuhand firms 

# Every client, every period. Just the doubts.

Describe how your firm prepares a return, or give Sumarity a period you've already filed. It designs one workflow for all your clients, ties each version out against what you signed off, and sends your preparers only what needs a professional's judgement.

Bring one process See the track record  
Code proves the numbers. People settle the doubts.

VAT return · quarter Q3 · client 14 of 62 Example run  

| Date | Line | Amount | Decided by  
| 02.08 | Invoice · Brakes, foodNet + VAT = gross · VAT number valid | 612.40 | Code · tied 
| 06.08 | Booking platform commissionReverse charge: supplier abroad · 0.94 | 184.00 | Engine · proven 
| 11.08 | Online order confirmationNot a VAT invoice: no VAT reclaimed · 0.92 | 86.97 | Engine · proven 
| 14.08 | Card refund to a guestReverses takings of 09.08, to the penny | −95.00 | Code · matched 
| 19.08 | KL-20817 and KL20817Look-alike invoices 3 days apart: the same one? | 1,140.00 | To a person 
| 22.08 | Invoice billed to another businessAddressed to someone else: not this client's | 430.00 | To a person  
|  

Settled 4 To a person 2 Settled wrong 0      Illustrative lines in the style of the pub VAT test quarters. Amber rows are the Judgement Engine's calls; each settles only above its cutoff and when its code check agrees.     

What we'd run for you 

## Start from a period you've already filed

Give Sumarity last period's files and the return you signed off. It asks about your files as they are, designs the workflow, and ties every version out against that return, saying where it differs and why.

Accounting firms, Treuhand firms, Steuerkanzleien, outsourced bookkeepers

VAT return preparationTestedWhat was this purchase used for, and is any VAT reclaimable?103 of 105 judgements right on a client held out of tuning; 15 of 15 problems reached the preparer. 0 settled wrong. 
Bank reconciliationTestedIs this receipt that invoice, paid days apart?Items for the team 236 → 73 over nine held-out months; 108 of 108 exceptions surfaced. 
Bookkeeping and account codingWhich account does this belong to? 
Client document chasingHas the client's answer settled the question? 
Swiss MWST and German VAT returnsWhich rate, and is the input VAT reclaimable?  
Sumarity for other work: Asset managers Restaurants and hotels Finance teams · everything

The number that matters most 
0 

## wrong answers settled without a person, in every test we have run.

Models are sometimes wrong, ours included. The workflow is built so that a wrong answer has to pass a measured cutoff and a check in code before it can settle. So far, none has.

| Test | Settled wrong  
| Bank reconciliation, nine held-out monthsWording tuning never saw · 108 real exceptions, all surfaced | 0 
| The same months, three different judgement modelsJudgement Engine and two frontier models in the same seat | 0 
| Against a 1,154-line program an AI wrote for the jobThree unseen months · 36 of 36 exceptions | 0 
| Custody statements from a live fund42 real lines first seen after tuning | 0 
| UK pub VAT returns, a pub held out of tuning105 judgements · 15 of 15 problems reached the preparer | 0 
| Deliberately wrong picks at 100% confidenceFed to a changed check before it went live · 36 of 36 rejected | 0  
| Total | 0      

How it works 

## Operations work is mostly rules. The rest is judgement.

Sumarity gives each part to whatever does it best: code for arithmetic, the Judgement Engine for lightning fast decisions, and people for what stays uncertain.

01

### Describe

The expert writes a paragraph. Sumarity asks the questions that change the design, then writes the workflow.
Expert revisions 

02

### Validate

Numbers only from code, every doubt to a person, every path handled. What fails is repaired before anyone sees it.
43 errors → 0 in one repair 

03

### Review

Every step on a canvas in plain words, editable and versioned. What runs is what was reviewed.
draft → validated → published 

04

### Run

Code matches and ties out. The Judgement Engine decides, and learns when it cannot. People get an inbox.
2,112 lines in 4.5 s 

05

### Tune

Rulings become an answer key. Better questions are proposed, tested on data they haven't seen, and lets you decide.
77% → 97% on unseen lines  
Every decision your team makes in the inbox goes back into step 5. It improves itself, with permission.    

What the platform enforces 

## In this work, mistakes are expensive. So the rules aren't optional.

These aren't settings a busy team can switch off. The validator refuses a workflow that breaks them.

Left alone, AI invents a plausible number.The AI never writes an amount or a date. Code works out every figure. 
It is sure, but not right.Nothing settles below the measured cutoff, or against a code check. 
It fails quietly.Every unsure or unproven call must reach a person. Every path is handled. 
It changes when its vendor updates it.The process changes only when an admin publishes a new version. 
It can't show its work.Every call is on record: what it saw, how sure, which version, who reviewed it. 
It answers questions it wasn't asked.Each judgement is one narrow question with a fixed set of answers. 
It doesn't learn from your team.Your rulings tune its questions and cutoffs, tested first on data they haven't seen. 
One person can approve their own work.Separation of duties and sign-off are enforced, not written in a policy.     

Industries and use cases 

## Wherever the work is mostly arithmetic, with judgement calls in between.

Every one of these has the same shape. Code does the matching and the sums, the Judgement Engine answers one narrow question per item, and whatever stays uncertain goes to a person.

Codematches, ties out, works out every figureJudgementone narrow question, measuredA personwhatever stays uncertain 

### Accounting and tax services
Outsourced bookkeepers, tax preparers, practices 
Many clients, the same process each period, and a filing at the end that has to be right.

VAT return preparationWhat was this purchase used for, and is any VAT reclaimable? 
Bookkeeping and account codingWhich account does this belong to? 
Expense and policy reviewWithin policy? 
Client queries and chasingHas the client's answer settled the question?  

### Restaurants and hotels
Restaurant groups, pubs, bars, hotels 
Every day's till, card, cash and bank to tie out, every supplier invoice to read, and a return at the end of the quarter.

Takings: till, card, cash and bankIs this deposit those days' cash, or something else? 
Supplier invoices and VATIs this a valid VAT invoice, and for what? 
Paid-outs and petty cashWhat was this paid-out for? 
Tips and troncWho is owed what from the card tips?  

### Finance and accounting
Controllers, shared services, treasury 
Matching and tying out, where most lines are easy and the rest need someone who knows the business.

Bank and cash reconciliationIs this receipt that invoice, paid days apart? 
Customer cash applicationWhich invoices does this remittance pay? 
Invoice and payment matchingA price change, or overbilling? 
Financial reporting tie-outDoes it tie out, and is the variance explained?  

### Risk and compliance
Compliance, financial crime, onboarding teams 
High-volume alerts where most are noise, every decision needs a record, and the real ones must never be missed.

Sanctions and screening alertsThe same party, or a namesake? 
Transaction monitoringWorth escalating? 
Vendor and client onboardingComplete and consistent? 
Periodic KYC reviewHas anything material changed?      

The library 

## A workflow for every process. In every industry.

Standard workflows, each tested against an answer key before it's published. Start from the closest one and describe how yours differs.

### Finance and accounting

18 workflows

### Accounting, tax and Treuhand firms

8 workflows

### Banking and payments

11 workflows

### Wealth, asset management and fund administration

9 workflows

### Insurance

8 workflows

### Healthcare

7 workflows

### Manufacturing and supply chain

7 workflows

### Retail, hospitality and consumer

6 workflows

### Logistics and transport

4 workflows

### Energy and utilities

4 workflows

### Telecoms

3 workflows

### Real estate

4 workflows

### Public sector

4 workflows

### Legal and compliance

3 workflows

### HR and payroll

4 workflows 

By processReconciliation and matching Payables, receivables and payments Filings, returns and reporting Documents, onboarding and KYC Claims, disputes and exceptions Screening, audit and compliance Close and accounting Billing and revenue assurance 
Browse the library    

Track record 

## Three processes, three domains. Every figure from a run.

Each workflow was authored through the product from an expert's description, then scored against a known answer.

UK pub VAT returnsaccounting service · a held-out pub Bank reconciliationfinance · nine held-out months Custody reconciliationfund operations · real statements  

Tested against the alternatives 

## Why not just hand it to an AI? We tried that too.

Same brief, same data, same scoring. These are the three things a capable team would try instead.

Give a frontier model the whole job 
Each model got the brief and both exports and returned the finished reconciliation, five times on each of two data sets.

|  | Sumarity | Opus 5.5 | GPT-6 Sol  
| Exceptions (of 108) | 108 | 108 | 95 
| Same result, 10 runs | 10/10 | 10/10 | 2/10 
| 2,112-line account | 4.5 s | 262 s | – 
| Record per line | Yes | No | No   
Past about 6,000 lines the job no longer fits in one reply.  

Ask an AI to write the code 
Opus wrote a careful 1,154-line program from the same brief and sample files. Both then ran on three months worded in ways neither had seen.

|  | Sumarity | AI-written code  
| Exceptions (of 36) | 36 | 36 
| Settled wrongly | 0 | 0 
| Judgement items to people | 3 | 14 
| False alarms | 3 | 14   
The program's rules were written from the sample's wording. Nothing in it learns from a review.  

Put a frontier model in the judgement seat 
The same workflow and checks, with only the judgement model swapped. 180 decisions with known answers.

|  | Engine | Opus 5.5 | GPT-6 Sol  
| Answered right | 100% | 99% | 100% 
| Median per call | 0.21 s | 1.76 s | 1.47 s 
| Per 1,000 calls | Included | $4.67 | $1.36 
| Settled wrong | 0 | 0 | 0   
The checks keep any model safe. What changes is the bill and the wait.   
Reasoning, audit trails and human approval are now common. Checks in code, measured confidence and tested changes are not.

### Checked, not just explained

Every call declares how it is checked. What code can't check, people audit by sample, and every doubt goes to a person.

### Confidence measured

Cutoffs come from backtests against your team's rulings, not from what a model says about itself.

### Changes tested first

Each improvement runs on data it hasn't seen, and against deliberately wrong answers, before an admin can publish it.

Start with a pilot 

## Bring one process. Run it alongside yours.

A few weeks in parallel on your own data, at our cost, then a line-by-line comparison: what it settled, what it sent to people, what it caught, and the time it took.

VAT and tax preparationBank and cash reconciliationInvoice and payment matching  
Contact us sumarity.ai    
- 01
Start from a tuned template, or describe itA paragraph from the person who runs the process today. 
- 02
Review the workflow with usOn the canvas, step by step, in plain words. 
- 03
Run it in parallelYour process stays the process of record. 
- 04
Your rulings tune its callsEvery inbox decision joins the answer key. 
- 05
Compare, line by lineThen decide, on your own evidence.     

Contact us 

## Tell us about the process. We'll come back to you.

A few details are enough to start. We'll reply to arrange a conversation with the person who runs the process today.

- We reply to set up a first call. 
- We look at the process with your expert, and say whether it fits. 
- If it does, we scope a pilot run alongside yours.
