I built the platform,the dashboards, the modelsand the forecasts thatfraud decisions run on.
legitimateflagged ring
Six years across energy, automotive, fashion and cloud. Now at
AWS, on fraud prevention, where a wrong call costs
money in both directions. Miss the abuse and it is expensive. Shut
down a real customer and it is worse. So I measure both sides before
anyone acts.
Anyone can catch more abuse. The craft is the cost of catching it.
Move the detection threshold. Catching more abuse is the easy part. Drag
it up and watch the first bar climb. The hard part is the second bar,
because every step too far flags real customers. The job is living in the
narrow band where one is high and the other is near zero.
Abuse caught
86%
Legitimate customers flagged
2.4%
LenientAggressive
Who is doing this
I am Hani. Based in Berlin, six years in data.
The short version is four habits and the projects that came out of
them.
01
I start with the false positives.
A missed fraudster costs money. A wrongly closed customer costs money and trust, and turns up weeks later as a support ticket nobody connects to the rule that caused it. So the carve-outs get built before the bulk action does.
Usually that is infrastructure nobody asked for. Two company databases sat behind two different logins with a slow sign-in on every query, so I put them behind one interface and made queries about ten times faster. Every other project I have shipped runs on it.
03
I do not trust a number until something has tried to break it.
Headline figures get attacked before they reach anyone. On its founding case it caught a silently truncated query and moved the headline by 69 percent. Catching that in draft beats being wrong in a room full of stakeholders.
04
And the same instinct works outside fraud.
Before AWS I scored which customers were worth calling at an automotive marketplace, cut a Zalando pipeline from 34 minutes to a few, and put a euro figure on refund abuse across six markets. Same job underneath, different domain on top.
Hover any dot to isolate the group it belongs to. Toggle to hide the ordinary customers and leave only the suspicious groups. Representative shape. The live version runs on confidential data.
A line means two accounts share something meaningful, weighted by how rare that thing is and how strongly it actually predicts fraud. The larger dots are the accounts that bridge two groups, which are often the ones worth acting on first.
Same analysis, same cluster. The change was in what the job read, not what it ran on. Only the ~34-minute baseline is measured. The after bar is sized illustratively against it to show the shape.
What a value rank does to a week of callingrelative value carried
Top-ranked accounts
100
Next band
58
Middle of the base
31
Long tail
12
Illustrative magnitudes, since the live score runs on AUTO1's customer data. What the work established is the ordering, and that the base was not evenly valuable. Flat outreach treats these four bands as though they were the same bar.Show the data table
What a value rank does to a week of calling: ranked.
Category
Value (relative value carried)
Note
Top-ranked accounts
100
Worked first, because the rank put them first rather than because they called in.
Eight ordered gates before any shutdown 8 in order
01
Population scopingDefines the candidate pool a run may consider at all.
02
Unlabelled filterRestricts the pool to unlabelled accounts.
03
Confirmed-fraud linkageWeighted, and only over non-placeholder hard identifiers.
04
Not already enforcedDrops accounts an earlier action already covered.
05
Account-age floorA minimum account age is required to proceed.
06
Carve-outs → reviewLegitimate-customer patterns route to human review, never to closure.
07
Enforce-time re-checkState is re-verified at the moment of action, not just at selection.
08
Ranked per-run capRanked and capped, so one run has a bounded blast radius.
Order carries the meaning: the carve-out gate only ever sees accounts that cleared the five before it. Runs daily in production under a hard kill switch.
Candidate accounts enter at gate one. Per-gate survivor counts are not published.
Show the data table
Eight ordered gates before any shutdown: every stage in order.
Stage
#
What it checks
Population scoping
1
Defines the candidate pool a run may consider at all.
Unlabelled filter
2
Restricts the pool to unlabelled accounts.
Confirmed-fraud linkage
3
Weighted, and only over non-placeholder hard identifiers.
Not already enforced
4
Drops accounts an earlier action already covered.
Account-age floor
5
A minimum account age is required to proceed.
Carve-outs → review
6
Legitimate-customer patterns route to human review, never to closure.
Enforce-time re-check
7
State is re-verified at the moment of action, not just at selection.
Ranked per-run cap
8
Ranked and capped, so one run has a bounded blast radius.
Share of all tasks ever queued, from one ringof every task in the queue's history
71%29%
One coordinated ring71%
Everything else29%
Concentration this extreme is what disqualified fraud inflow as the planning variable. The distribution, not intuition, is what settled it.
Kept inside the baseline on purpose, after a ring-excluded companion scenario showed how different the forecast looks without it. A capacity plan should be conservative, and rings recur.
Show the data table
Share of all tasks ever queued, from one ring, of every task in the queue's history.
Refund / damage rate by customer value segmentrefund/damage rate (illustrative)
Each segment's excess over the trusted A/VIP benchmark, applied to its GMV, is the leakage estimate. Shape is illustrative. Live values come from the model.
Remaining Fraud Damage = Return Damage + Delivery Damage, as a share of GMV, tracked across DE · NL · BE · FR · IT · CH weekly and monthly.
What a handoff carries, in the order it is assembled 7 in order
01
Repo contextCaptured from the repo, not re-typed by hand.
02
Git stateThe branch and the diff travel with the task.
03
Decisions madeWhat was already settled travels with the task.
04
What is brokenThe failing test goes across too.
05
Next stepsWhat comes next is recorded, not reconstructed.
06
Continuation promptPaste-ready, and it works across 11 coding agents.
07
HandbackWork done elsewhere re-injects natively into Claude Code.
Order is the point: the continuation prompt is only useful because the five fields before it are already filled, and the handback closes the round trip so the original session knows what changed.Show the data table
What a handoff carries, in the order it is assembled: every stage in order.
Stage
#
What it checks
Repo context
1
Captured from the repo, not re-typed by hand.
Git state
2
The branch and the diff travel with the task.
Decisions made
3
What was already settled travels with the task.
What is broken
4
The failing test goes across too.
Next steps
5
What comes next is recorded, not reconstructed.
Continuation prompt
6
Paste-ready, and it works across 11 coding agents.
Handback
7
Work done elsewhere re-injects natively into Claude Code.
Where the base drained across the customer relationshipshare still active (illustrative)
Illustrative shape, since the live analysis ran on AUTO1's customer data. What the real work established is that drop-off concentrated at identifiable points in the relationship, which is what let a revenue figure be attached to each one.
What a moved number clears before it is read as fraud 4 in order
01
Comparable populationSame markets and same customer segments as the period it is measured against. A different population is a different measurement.
02
Soft exclusions accounted forCases kept out of the count change the denominator without ever appearing on the chart.
03
Base rate stableHow much fraud there was to find at all. More fraud around raises the flagged rate with no change in detection.
04
Now read the move as fraudOnly what clears the first three gets reported as a change in fraud, and only that is worth acting on.
No counts are published here. The order is the argument, because a base-rate check on a population that was never comparable settles nothing.
Applied to suspicious rate, detected rate and steer rate, weekly and monthly, across the top six markets: Germany, the Netherlands, Belgium, France, Italy and Switzerland.
Show the data table
What a moved number clears before it is read as fraud: every stage in order.
Stage
#
What it checks
Comparable population
1
Same markets and same customer segments as the period it is measured against. A different population is a different measurement.
Soft exclusions accounted for
2
Cases kept out of the count change the denominator without ever appearing on the chart.
Base rate stable
3
How much fraud there was to find at all. More fraud around raises the flagged rate with no change in detection.
Now read the move as fraud
4
Only what clears the first three gets reported as a change in fraud, and only that is worth acting on.
The chain the joined dataset runs along, in order 6 in order
01
LoginAuthentication events. Who signed in, and when.
02
Order placedOrder-placed events with order-position rows, so the basket is visible and not only the order.
03
Risk assessmentWhat the abuse-protection platform made of it. Read from the legacy source and the current one.
04
Risk decisionWhat was actually decided about that assessment.
05
Steering decisionThe customer-risk steering tables. How that customer was handled from then on.
06
Steer feedbackThe feedback recorded against that steer. Return steering carries its own feedback, joined alongside.
Each stage only sees what the one before it produced, which is why the order is the design and not a presentation choice. No counts are published here. The purpose of the chain is evaluation: how abuse-protection decisions performed, and where fraud damage remained.
Customer dimension data, sales channel, route-accessed events, and a table mapping fraud rules to fraud domains join alongside the chain as context.
Show the data table
The chain the joined dataset runs along, in order: every stage in order.
Stage
#
What it checks
Login
1
Authentication events. Who signed in, and when.
Order placed
2
Order-placed events with order-position rows, so the basket is visible and not only the order.
Risk assessment
3
What the abuse-protection platform made of it. Read from the legacy source and the current one.
Risk decision
4
What was actually decided about that assessment.
Steering decision
5
The customer-risk steering tables. How that customer was handled from then on.
Steer feedback
6
The feedback recorded against that steer. Return steering carries its own feedback, joined alongside.