About Cleo
Cleo builds an AI financial assistant for people in the US and UK. Its main product is earned wage access: Cleo advances part of a customer’s wages and collects the advance on payday.
Subscriptions, advance payouts and collections each run as their own payment flow, across several vendors and across card, bank, real-time and ACH rails. One product alone can run through 25 routes, and each can fail on its own.
Every false alarm took hours to rule out
Cleo’s automated payment alerts often fired on drops that turned out to be normal. The team only found that out after hours of manual investigation.
Normal is hard to define for Cleo’s payments. Most US customers are paid every two weeks, some monthly and some daily, so a Friday looks nothing like a Monday. ACH does not run at weekends, and a public holiday moves paydays by a day.
Every alert sent an investigator chasing the cause, and many of those chases ended with no payments problem at all.
Cleo’s payments analytics team rebuilt the check around money that moved
The old alerts watched success rate. When fewer payments were attempted, the rate could hold steady while less money moved. The team switched to the count of payments where money actually moved, the number the business cares about.
They also put the payments knowledge into Rig: which tables hold each step, how success is defined on each route, and the usage rules and exceptions for each product. One Rig query now pulls every route at once.
A threshold for each route, on each day
Each morning, every route is scored against its own last 90 days, adjusted for the day of the week and for holidays. The thresholds are recalculated for each segment on every run, so a quiet Saturday is not mistaken for an outage.
Known weekend and holiday patterns sit on a list the check consults before it raises anything. Each time an alert turns out to be expected, the pattern is added, so the same false alarm does not come back.
Tracing a drop to the step that caused it
A payment passes through six steps, from falling due to the money moving. When a route drops, the check walks back up those steps one at a time and stops at the first one that explains it, whether that is fewer attempts, a vendor rejecting payments or a single rail failing.
The heavy investigation queries run only after an alert fires. On a normal day, the check reads one summary query and stops.
One Slack thread says which team looks at what
Between 9am and 10am, the check posts to Slack. Each red or amber alert comes with its cause and the team that should look at it, so payments, product or a vendor owner can act on it straight away.
Investigators no longer go around chasing false positives. When the thread says a drop is expected, nobody opens an investigation.
Investigation time goes to real problems
The check has run on 116 days since May. Each morning the payments analytics team starts from a thread that has already ruled out the expected drops and named an owner for the rest.