Thinking

№ 14 · 9 JULY 2026 · 5 MIN

Fraud screening at 16 million decisions a month: what holds

Sixteen million decisions a month for more than five years. What keeps a screening system reliable is never the model.

A fraud screening system that makes sixteen million decisions a month, and has made them for more than five years, is a useful thing to study, because at that scale the illusions burn off. One of our principals ran exactly this: card-transaction fraud decisioning across three national retail portfolios, in production, for over five years. The question people expect to be interesting is which model it used. The question that actually determines whether such a system survives is quieter: what holds it up, month after month, when the model is only one component among many. The answer is never the model. It is the discipline around it. That is not modesty about models, which matter. It is an observation about where reliability comes from at scale. Across five years the models were revised more than once; the thing that stayed constant, and kept the system trustworthy through those revisions, was everything wrapped around them.

Start with the thing the model does, so it can be set aside. It produces a score, a number that says how likely this transaction is to be fraud. That is genuinely useful and completely insufficient. A score is not a decision. Sixteen million times a month, something has to turn that number into approve, decline, or hold, and that translation is where a screening system lives or dies. It is done by thresholds, set per portfolio, tuned to the trade-off each business will accept between blocking good customers and eating losses. The model is the sensor. The policy on top of the score is the system, and it is the part that has to be designed, documented, and owned.

The model is the sensor. The policy on top of the score is the system.

· № 14 · ¶ 03

The second thing that holds is a review lane for the cases the system should not decide alone. Most transactions are easy: clearly fine and cleanly cleared, or clearly bad and cleanly blocked. The danger lives in the uncertain middle, the transactions the score cannot confidently place. Sending those straight to a decision, in either direction, is how a system either bleeds fraud or infuriates good customers at scale. So the uncertain ones route to analysts, with the context assembled, and a human makes the call. This is not the system failing. It is the system working as designed: the human valve sits exactly on the band where being wrong is expensive, and the analysts' decisions feed back as labels that sharpen the next round. Sizing that lane is its own discipline. Route too little to people and the uncertain cases get decided badly at machine speed; route too much and you have rebuilt the manual process you were trying to escape. The threshold that decides what counts as uncertain is not a detail. It is one of the most consequential numbers in the system, set deliberately, reviewed, and moved only when the evidence says to.

FIG.01
LOWUNCERTAINHIGHLABELEDTransactionARRIVESModel scores itFRAUD LIKELIHOODAuto-clearCLEARLY FINEAnalyst reviewUNCERTAINBlockCLEARLY FRAUDAudit trailEVERY DECISION
Most transactions are easy and clear themselves in either direction. The uncertain middle, where being wrong is expensive, routes to an analyst. Every decision, human or automatic, lands in the audit trail. The human valve sits exactly on the band that needs it.

The third thing that holds is the audit trail, and at this scale it is not optional bookkeeping. Every decision has to be reconstructable months later: what the score was, which threshold applied, what the outcome was, and if a human touched it, who and why. Part of that is for regulators, who will eventually ask. The deeper reason is that a decision system you cannot inspect is a decision system you cannot trust, debug, or improve. The trail is what turns sixteen million opaque calls a month into something a team can actually reason about when one of them goes wrong, as some inevitably will. When one does, the trail is the difference between a cause found in an afternoon and a week of guessing. You can replay the exact inputs, see which threshold applied, and tell whether the failure was the model, the policy, or a human override. That is knowledge you do not have if the system only remembers its final answers.

The fourth thing, and the one most teams underweight, is the evaluation that catches drift before customers do. A model that was accurate when it launched is not permanently accurate. The population it screens shifts: new customers, new fraud patterns, new products, a holiday season that looks nothing like the training data. If nobody is watching for that drift, the first signal is a rise in complaints or losses, which means customers absorbed the failure first. The discipline is to monitor the population and the performance on a schedule, and to treat a meaningful shift as an incident to investigate, not a curiosity to note. Catching drift in the evaluation is the difference between a quiet fix and a public one. Customers should never be the drift detector. If the first place a shift becomes visible is the complaint queue, the monitoring was decorative.

We have written elsewhere that bank-grade is a discipline, not a badge, and this is the operational core of that idea. None of the four things that hold a screening system up is a model choice. Thresholds, a human review lane, an audit trail, and drift-catching evaluation are all system properties: designed, staffed, and maintained, and constant while models come and go underneath them. A better model dropped into a system without these buys you a slightly better score and none of the reliability. The same model inside a system that has these holds for five years and sixteen million decisions a month. This is the uncomfortable part for anyone selling a model as the answer. The durable work is not the part that demos well. It is the thresholds argued over in a room, the review lane someone has to staff, the trail nobody looks at until they need it, and the drift checks that run whether or not anything is wrong.

This is why, when the stakes involve money or eligibility, we build the system before we get precious about the model, and why the same shape shows up in adjacent work. The risk policy at a trade-finance lender, governing a billion dollars a year in transactions against a hundred-and-fifty-million-dollar portfolio, holds for exactly the same reasons: score, threshold, a human valve on the uncertain, an audit trail, and a watch for drift. The model is the part everyone wants to talk about and the part that matters least to whether the thing survives contact with real volume. What holds is never the single call the model makes. It is the system that decides what to do with it.