The software I build carried two accuracy numbers. Both measured, both true.
One came off a first import, before the thing had seen anything of a firm's books. The other came off books it had been running on for a while, once it had learned how that client's transactions behave. The second was higher. Of course it was. It had practice behind it.
I took the second one down. There's one number on the page now, and it's the day-one figure.
While both were up, anybody reading the higher one was reading a number taken under the best conditions I had. I chose which conditions they saw. Nobody lied anywhere in that.
Every AI number aimed at this profession is a measurement somebody chose.
Everybody with a product in this market publishes a number that way, me included. There is no version of publishing one that skips the choosing. So the useful move is a short set of questions that makes a chosen number answerable.
What does AI for bookkeepers actually do today?
Inside a firm right now the three I'd pay for are these. It sorts routine transactions into categories. It drafts text somebody was going to write anyway. And it reads a pile of something and tells you what's in it. Everything past that is a claim, and every claim carries a number taken under conditions the seller set. This is for the owner deciding what to keep paying for, not for somebody shopping. Four questions turn a claim back into something checkable: what it was measured on and against what, who catches it when it's wrong, what one wrong output costs, and what you keep when you stop paying. Two filters sit on top: whether it keeps working on a week you're busy, and what your firm has agreed can be sent outside it.
Key Takeaways
Every accuracy number was measured on something - and the something is usually not books that look like your clients' books.
Two of the four questions have no vendor answer - who catches an error and what it costs you are both settled inside your own firm.
A tool you operate is not a system that runs - most of what gets sold as automation is a faster manual step, which is worth money and adds no capacity.
What leaves your firm is your firm's call - one named person decides it in writing, and a vendor's terms page is not that decision.
Copy from a coded list instead of the client file - the names come off without anybody having to remember, which leaves a shorter list of things you still have to decide about.
This issue names no tools on purpose - I sell software into this profession, and a list from me would be shaped by that in ways you couldn't see from the outside.
The four questions to ask about AI for bookkeepers
Run these on anything sold to your firm, mine included. Two go to the vendor. Two you answer about your own firm, and those decide whether the vendor's answers matter at all.
What it was measured on, and against what. An accuracy claim is a fraction and the bottom of it is a pile of data somebody picked. Ask what was in that pile: whose books, how many, over how long, how messy at the start. Then ask what the number beat. Faster than what, run by whom, held to what standard, measured by which side. A percentage with no denominator and no baseline can't be checked by anybody, including the person who published it.
Who catches it when it's wrong. No vendor can answer this and you shouldn't ask them to. It's a question about your own workflow. When a categorization comes back wrong, where does it get caught, by whom, and how many days does it sit in the books first.
What one wrong output costs you. A tool that's right most of the time is cheap when the miss surfaces at review and expensive when it surfaces on a filed return. So price the miss. A wrong category on a small expense costs a correcting entry and the few minutes it takes to make it. A wrong number inside something a client forwards to their bank costs you the client. Same accuracy rate, two completely different products. Nobody was ever taught to ask this. Ask it anyway.
What you keep when you stop paying. Somebody will ask you for this work years from now, and the vendor is not the person who'll be asked. So find out what you hold: is the finished output a file in your own systems, or a page you can see for as long as you're a customer. Ask in writing, and ask before you renew.
A tool that can't answer the first and the fourth in writing inside two weeks isn't a tool you've evaluated. That doesn't make it bad. It makes it something you renew on faith, and you should at least know which ones those are.
The busy test
A tool you operate is not a system that runs.
Most of what gets sold to firms as automation is a faster manual step. That's genuinely worth money. It just doesn't add capacity, because it stops when the person stops.
The test takes a week. Pick one where you don't open the thing. If work was waiting when you came back, and you didn't start it, it runs. If nothing happened, it's a tool, competing for your hands with everything else you do with them.
The issue on month-end drew that line down the middle of a close. The capture address keeps collecting documents with nobody minding it. The weekly pass only happens because a person sits down.
What leaves the building
The prompts in this letter have carried the same warning from the start: take the names out first. That warning only ever covered the names.
Amounts attached to a person. Account numbers, EINs, socials. The client's own documents, which carry all of it at once plus a few things you forgot were in there. And the gap between pasting a paragraph and uploading a file, which are two decisions with two different sizes of mistake behind them.
Then the thing nobody has a habit for yet: what a tool your firm already uses, that already holds client data, starts doing with it after an update. Taking names out doesn't reach that. Somebody has to decide it, and when nobody does, it gets decided by whoever is closest to a deadline.
SDO CPA, where I'm a partner, keeps client files that look like yours. I had to answer this for us before I could write any of it down for you.
The move that survives a bad week sits upstream of the discipline. Build one coded list of your clients, once. A text file, one line each, A is a client and B is a client. Then the rule that follows from it: a prompt's input gets copied out of the coded list and never out of the client file. After that the clipboard has no names on it, whether or not anybody remembered to be careful.
The issue on picking a lane put a version of this in front of a client list. Same move, with the remembering taken out.
You don't need the annex for the first step. Write down the thing that never gets pasted anywhere, today, in whatever file your firm already reads.
Why this issue names no tools
Because I sell two pieces of software into this profession. Whatever I put on a list would be shaped by that, and you'd have no way to see where.
There's a harder version of the same problem, and you'd find it eventually, so here it is. The four questions above aren't neutral either. I picked them, and my own products happen to answer them reasonably well. Somebody with different products would hand you a different four, picked the same way, and it would read just as sensibly to you.
So add your own. Take the tool you'd least like to lose and find the question it would fail. That's the question missing from this list, and adding it is what makes the list yours.
The first issue, on getting clients without cold outreach made the case that an owner can build most of this stuff now without hiring anybody to do it. Still true. Building it yourself just doesn't spare you the buying decisions, and those have a salesman on the other side of them.
Do this before Friday: run the first and fourth questions at the AI tools your firm pays for, starting from the last invoice you can find. Then add the ones with no invoice, meaning anything anybody in the firm opened this month. The second and third you answer yourself, and they take longer.
Twelve of these now, which is also a chosen number, published by somebody who sells software into this profession. Run the four questions on the letter while you're at it. What it was measured on is one owner's firm. What you keep when you stop paying is all of it, because you were never paying.
What did you buy that stopped getting used, and what actually killed it? Reply with that. The product name doesn't help me. What made you stop opening it does, and it's worth more to me than any vendor page I could go read.
Operator annex
Send the email first. The rest of this is built around the reply.
The email you send before the renewal
Two of the four questions have a vendor answer, and the first one splits in half when you ask it, because a vendor answers the denominator and the baseline separately. Add the data question from What leaves the building, since you may as well get it all in one reply, and that's four things to send.
Subject: A few questions before we renew
Hi [name],
We're going through the tools we use at the firm and I have a few questions about [product]. Short answers are fine. I'd like them in writing.
1. The accuracy figure on your site: what was it measured on? How many businesses, what kind of books, over what period, and how clean to start with?
2. What was it compared to? If it beats something, tell me what, and who was running it.
3. If we cancel, what do we keep? Are the finished outputs files we hold in our own systems, or access that ends with the subscription? If the second, for how long after?
4. Where does our clients' data go when we use the product, and who else can read it? A pointer to the terms is fine.
Thanks,
[you]Send it and leave it alone. When the reply lands it goes wherever that tool's contract lives. The sent mail and the reply are the record, so there's nothing else to keep.
A tool still carrying two blanks two weeks after you asked is an unevaluated tool. Unevaluated tools go on the renewal list, not on the roster.
If you want them where you'll see them without opening your sent folder, it's a few lines per tool in a text file. Filled in, so an answer sits next to a non-answer:
[Tool]
Measured on: their own sample, size not stated, condition not stated
Compared to: nothing stated
On cancellation we keep: exports of finished output, 30 days of access after
Client data goes to: their processor plus one subprocessor, both named in their terms
Asked: 3 Mar Answered: 11 MarThe blanks are the finding. A sample with no size and no baseline next to it isn't an answer.
The busy test, written down
Take the week first. Then read this against what you found:
It runs if:
- Something it produced was waiting for you when you came back, and you didn't start it
- It fires on something happening outside your firm: a document arrives, a feed updates, a date passes
- When it finishes, the next thing happens without anybody carrying the output over
- It kept producing during the week when nobody in the firm touched it, not just when you personally didn'tFour yeses and it runs. Staff around it like capacity. Two or three and it's a partial system: find where a person is still carrying something across, because that handoff is capping it. One or zero and it's a tool. Price it like one. The hour it saves is an hour you'd have spent at the desk anyway, and on a week you're out it saves nothing.
What never leaves, written down
One page. It lives wherever your firm actually reads things, not in a policy binder.
[Firm] - what we don't paste
Client names, business names, and anything else that identifies a person go into the coded list first. Prompt inputs get copied out of the coded list, never out of the client file.
Never pasted or uploaded outside the firm, coded or not:
- Social security numbers, EINs, bank and card numbers
- A client's own documents, in any format
- Anything out of a client's payroll records
- Logins for anything
Pasting text and uploading a file are different decisions. A file brings everything in it, including whatever you forgot was in it. An upload needs a second person to agree first. Solo, the second person is tomorrow: an upload waits overnight, and if it still looks necessary in the morning it goes.
If a tool we already use turns on a feature that sends our data somewhere new, we stop using that feature until [name] has looked at it. Noticing is part of everybody's job, not [name]'s.
Decided by: [name]. Anybody can ask. Nobody else decides.
Last read: [date]And the file the first line of that page depends on:
Client codes - stays in [location], never leaves the firm
A = [client]
B = [client]
C = [client]
D = [client]
A new client gets the next code the day the engagement letter is signed.
A departed client keeps their code forever. Codes are never reused.
Past Z, go to AA, AB, AC.
This file never gets pasted into anything. It is the one file that stays home.What goes at the top of every prompt file
Prompt files go out of date and keep running anyway. One written against last spring's tool still returns something that looks right. A few lines at the top of each file, so the folder listing is the index:
Produces:
Reads:
Its input is copied from:
Never touches:
Last checked:Here it is with all five filled:
Produces: a first-pass categorization review note for one client month
Reads: a de-identified transaction export, one month
Its input is copied from: the coded list
Never touches: anything with an account number in it, and any payroll file
Last checked: 14 Feb"Its input is copied from" is the line doing the work. If the answer is anything other than the coded list, then something other than that page is choosing your inputs and nobody signed off on it. "Never touches" is yours to write. Don't let anything else write it.
Two prompts
These two only read. Nothing either of them touches changes. What both can do is sound certain about something they half-know, and the prompt that reads your prompt files is looking at a map of everything your firm sends outside. Read what comes back the way you'd read a junior's first draft: assume the confident parts are the ones to check.
Instead of reading the sales page for the third time. The email above goes to every vendor. This finds what one specific page left out, and hands you the extra sentences to add to it.
I run a bookkeeping firm and I'm evaluating a piece of software. Below is
the text of the vendor's own page, pasted by me.
Everything you need is in the block below. Don't go anywhere else for it,
and don't fill a gap with what you already know about this company. Where
the text doesn't say, write "not stated on this page."
[paste the page text]
Answer these one at a time:
1. What accuracy, speed, or time-saving figures does the page state, and
what does it say each one was measured on: whose data, how much, over
what period, in what condition?
2. What is each figure compared against? Name the baseline the page states.
Where there is none, say so plainly.
3. What does the page say about who reviews the output before it's used?
4. What does it say happens to finished work after a customer cancels?
Then, for every "not stated on this page," write the exact sentence I should
email the vendor to get it. One each, plain, no preamble.
Do not tell me whether to buy it.Instead of opening every prompt file to remember what's in it.
Below are prompts I use at my bookkeeping firm. Before pasting them I
replaced every client and business name with a letter, including the names
inside the example rows, because a saved prompt usually has a real one baked
into its example. Don't ask me for the real ones.
[paste the prompts]
For each prompt, tell me:
- What it produces, in one line.
- Every input it asks for, and whether each could carry a client's identity,
amounts tied to a person, or a document.
- Whether its output goes anywhere other than back to me.
- Anything in the prompt text itself that would be a problem if read by
somebody outside my firm.
Report what each one touches and stop. Do not write a "never touches" line
for any of them, and do not tell me which inputs are acceptable. That call
is mine, and it's the whole point of the exercise.
Give it back to me here as plain text. Don't create or edit any file.House ad. Both of these are mine, and the four questions get run on them below.
Growthy is bookkeeping software. It connects to QuickBooks Online, or you import bank statements, or upload a CSV. It categorizes the routine transactions and you review and approve the rest. About 85 percent accurate on a first import, measured before it has learned anything about a particular set of books. Running my own firm's books on it took our month-end from about four hours to about 30 minutes, on the categorizing half of the job. See what Growthy does.
TracePrep is the second product, built by a separate company, TracePrep Inc. It's a Reasonable Compensation Study platform for CPA and EA firms. Every figure in a Study traces back to the record it came from, your firm's own Reviewer signs off before a Study goes out, and the calculation underneath is deterministic and inspectable, which is a different kind of thing from what this issue has been about. It's in design-partner validation, not general release. See how TracePrep builds a Study.
Now the four questions, aimed at my own two, since these are questions about buying software and not about one category of it. Growthy's number comes with the conditions it was measured under, which is part of the first question and not all of it, because I haven't put a baseline next to it here. On the second it can only tell you where its own boundary sits, not who catches the miss: a person reviews and approves, so the catching is yours, and that's why it doesn't clear the back half of the busy test. On the fourth, ask me in writing, the same way you'd ask anybody else.
TracePrep's answer to the fourth is most of why it exists: your firm owns the finalized evidence and workpapers permanently, and finalized Studies remain downloadable in TracePrep for exactly three years after cancellation. Neither of them answers the second or the third for you, because who catches a miss and what it costs are facts about your firm and not about a product.
Nothing here runs on its own. The email gets sent, the week gets taken, the header lines get written once, the page gets read again in the spring. What the coded list changes is where the judgment sits. Stripping names is something you remember in a tired moment. Copying from one file is somewhere you go. In six months you'll be able to say what your firm sent outside instead of estimating it.


