Ask a partner in a small accounting, audit or tax practice where the year goes and the answer is rarely the accounting. It is the pile in front of the accounting. One client emails a PDF bank statement. Another photographs a stack of receipts on a kitchen table. A third sends a spreadsheet laid out the way a bookkeeper liked it years ago, with a merged header row and the tax column sometimes blank. A fourth posts an envelope. None of them are being difficult. They are sending what they have, in the format they have it in, and all of it lands on the same desk, where somebody opens it, reads it, and types it into the practice's system.
That typing is the real cost, and it appears on no invoice. It shows up instead as the reason a short job takes all afternoon, as the reason the deadline weeks are a wall, and as the reason the person hired to exercise judgement spends the first stretch of every morning doing transcription. It has a second cost that arrives later. Every re-keying is a fresh chance to transpose a digit, and a transposed digit in a filing is not an inconvenience. It is a correction, an apology, and sometimes a cost the practice quietly absorbs to keep a client. The intake pile is where both of those costs are made.
Which part of this does a machine actually do well?
Reading a document and putting its fields somewhere is the thing current models are genuinely good at, and it is worth being precise about why, because the reason is unglamorous. The task is narrow. The model is not being asked to know tax law. It is being asked to look at a page, find the date, the amount, the counterparty and the treatment the client has already indicated, and put each one into a field with a pointer back to the place on the page it came from. That is a bounded task with a checkable answer, which is the shape of work that automates cleanly. Anything broader than that is a different claim, and a much weaker one.
Routing is the other half, and in a practice it matters as much as the reading. A document that has been read still has to arrive somewhere: attached to the right client, in the right period, in the right ledger, flagged if it duplicates the one that client already sent last week. Most of the friction is not any single document. It is the pile's disorder. A system that reads each item, files it against a client and a period, and states plainly what is still missing before a deadline turns a pile into a queue. A queue can be worked through. A pile can only be dreaded.
What both halves have in common is that being wrong is cheap and caught early. If the model reads a nine as a four, the person reviewing sees it beside the source image and fixes it in a second. If it files a receipt into the wrong period, the review catches it long before anything is submitted. None of these failures compound, and that is the real test of what is safe to let a machine do routinely: not whether it is clever, but whether a mistake is visible and reversible at the moment it is made.
What has to stay with a qualified human?
Now the line, which is the more important half of this. Extraction is not an opinion. A filing is. The moment the work stops being what does this document say and becomes how should this be treated, it has left the machine's territory. Whether a cost is capital or revenue, whether a transaction sits inside a scheme, what a set of accounts asserts, what an auditor is willing to put a name to: these are professional judgements with a qualification behind them. A model can assemble the evidence and state the question. It cannot carry the responsibility, because responsibility is not a capability. It is an accountability, and it belongs to a person with a licence and a duty.
So the boundary is built in rather than mentioned in a footer. The system reads and drafts; a qualified person reviews and signs. Every number that will end up in a return or a set of accounts passes a human who can see the source beside the entry. Every approval is recorded, so months later it is possible to say who signed what, and on what evidence. This is not a hedge against a weak model. It is the design. In work a regulator, a client or a court may eventually read, the sign-off is the product, and the automation exists to make that sign-off fast rather than to remove it.
The sign-off is the product, and the automation exists to make that sign-off fast rather than to remove it.
Does this hold when the pile gets large?
We should be straight about where our evidence comes from. We have not run this inside an accounting practice. What we have run is the same shape of work at volumes a practice will never see, which is the more useful thing to know, because the question worth asking is whether the mechanism holds when the pile gets large. One of our principals spent five years building document intelligence at a national newsroom: a search index of thirty million documents, fifteen production applications sitting on top of it, and, measured against the prior workflow, recovery on the order of five hundred hours in every thirty-day window (the exact figure was around four hundred and ninety-seven).
Two other pieces of that record speak more directly to the intake problem. At a private-markets data platform the input was more than a hundred and fifty fund statements from more than six venture and private-equity firms, every one laid out differently, and the job was to read hierarchical tables reliably enough that the numbers could be trusted downstream instead of re-keyed by an analyst. That is the mixed-format problem exactly: the same kinds of fact, arriving in a hundred and fifty different shapes. At a clinical-intelligence company the work was standardising more than a hundred thousand clinical codes so a document's contents mapped onto a controlled vocabulary the profession recognises. A practice has its own controlled vocabulary too. The chart of accounts, the tax codes, a client's own naming for their own costs: mapping a messy document onto a fixed scheme is the same engineering in a different scheme.
None of that makes us your accountants, and we will not pretend otherwise. It means the part of the problem that is engineering has been solved well above a practice's scale, and the part that is professional judgement was never ours to take. What will be new is your last mile: your clients' habits, your filing calendar, the exceptions your seniors carry in their heads and have never written down. That part does not generalise from anyone else's build, and we would rather walk it with you than assume it.
What does a practice actually get?
What a practice actually gets is quieter than the pitch usually sounds. The inbox stops being a pile of attachments somebody opens one at a time and becomes a queue of drafted entries with the source on screen beside each one. Chasing becomes a list the system maintains: which clients still owe what, with the reminders drafted and a person sending them. The junior stops transcribing and starts reviewing, which is faster and much closer to the work they came to learn. And the partner still signs, on every number that leaves the building, with the evidence in front of them and the record of that signature kept.
That is the honest version. Nobody is getting an autonomous practice, and anyone offering one has not thought carefully about what happens the first time it is confidently wrong about a filing. What we would take off the desk is the reading, the filing and the chasing. What we would leave exactly where it is, is every judgement that carries a qualification, now reached faster because the paperwork arrives already assembled and already checked.