Pockets Full Systems

Case study — fig. 02

From failing every night to 2,245 of 2,245 records — in nine seconds.

The gap

A nightly automation the business depended on was failing at volume. Every morning started with the same question: what didn't run? Two previous attempts to fix it assumed the architecture was wrong and proposed rebuilding it. The failures kept coming — and each one meant records silently not processed, and a human downstream doing cleanup work the software was supposed to own.

The catch

Before rebuilding anything, I traced individual failing records through the job. The architecture was fine. The root cause was two data defects — bad records poisoning the batch and taking healthy records down with them. Fix the two defects, harden the flow against the pattern, and the "broken" automation turned out to have been sound all along.

The result

2,245 of 2,245 records processed. Nine seconds. Zero failures since — on the client's existing flow, still maintained by their own admin. No rebuild, no new platform, no ongoing dependency on me.

When an automation fails at volume, the expensive assumption is that the system is wrong. Sometimes the system is fine and the data is lying to it. Diagnosis before demolition.

Got a job that fails and nobody knows why? Tell me where work is falling through — I'll tell you the first three places I'd look. Free, no pitch.

Find the leak or just message me →