Fraud measurement · counterfactual 2024 to 2026

A euro figure on the fraud that survived, and controls that can't flatter themselves

Remaining Fraud Damage sizes the logistics refund leakage that slips past every control, in euros, per segment, per market. Holdout and counterfactual framing then stops the controls taking credit for reducing it.

Senior Product Analyst, Risk & Abuse, Zalando

Impact

Produced the first euro-denominated estimate of logistics refund leakage surviving existing controls, by customer segment and across six markets, refreshed weekly and monthly for business reviews, and replaced flattering before/after reads with holdout / counterfactual evaluation so prevented damage was separated from remaining damage.

Refund / damage rate by customer value segment refund/damage rate (illustrative)
A/VIP benchmark segment rate trusted A/VIP floor → higher-risk value segments
Each segment's excess over the trusted A/VIP benchmark, applied to its GMV, is the leakage estimate. Shape is illustrative. Live values come from the model.

Remaining Fraud Damage = Return Damage + Delivery Damage, as a share of GMV, tracked across DE · NL · BE · FR · IT · CH weekly and monthly.

Estimated effect of the control -45%
Naive treated-vs-untreated
overstated
Holdout / counterfactual
true effect
Illustrative shape. Selection bias inflates the naive read. The holdout strips out what would have happened anyway, leaving prevented damage separated from remaining damage.
Context

Everyone tracks the fraud they catch. The damage that matters for the business is the part that slips past every control: manual refund leakage on missing-delivery, item-not-received, and parcel-missing claims, invisible precisely because nothing flagged it. The mirror-image error sits on the other side of the same ledger. Risk controls like Secure Delivery and refund-denial steering are applied exactly where abuse is most expected, so a naive treated-vs-untreated read lets selection bias do the talking and credits the control for a difference that was already there.

What I did
  • I built the Remaining Fraud Damage measure.
  • Trusted high-value customers (A/VIP) set the honest baseline refund and damage rate, and every other segment's excess over that benchmark, applied to its GMV, is the leakage estimate.
  • Remaining Fraud Damage = return damage + delivery damage as a share of GMV, computed across the top six markets on manual-refund and Salesforce case data in PySpark on Databricks.
  • A leakage figure is only worth the evaluation of the controls aimed at it, so I treated that evaluation as causal work rather than reporting.
  • Holdout and counterfactual framing separated prevented damage from remaining damage.
  • I stated plainly where a biased intervention group would inflate the result.
  • I also owned the fraud KPIs for weekly and monthly business reviews.
  • I read suspicious, detected, and steer rates against base-rate effects and soft exclusions, so a shift in population mix never got mistaken for a shift in fraud.
Outcome

'Fraud we can't see' became a number Risk could argue about, own, and set targets against, and Secure Delivery and refund-denial thresholds got debated against prevented damage the control had actually caused.

Without this work

The leakage stays invisible: no euro figure to budget against, no segment or market to target, and no way to tell whether refund damage is getting better or worse. Meanwhile the controls get credited for differences that existed before they ran, and threshold decisions get made on a number that was never real.

The full story

Caught fraud is the easy half. The number leadership actually needed was the other half: how much logistics refund damage (missing delivery, item-not-received, parcel-missing) was leaking past every control? By definition nothing had flagged it, so no figure existed at all.

I built one, and the whole method rests on choosing an honest baseline. Trusted high-value customers refund and claim at some natural rate that isn’t abuse, so treat that as the floor. Every other value segment’s refund/damage rate above the floor, applied to its GMV, estimates the excess that shouldn’t be there. Sum return damage and delivery damage and you have Remaining Fraud Damage as a share of GMV. The plumbing is ordinary e-commerce data work: GMV denominators before and after returns, refund reasons, and risk signals joined to value segments.

A leakage number invites the obvious follow-up: are the controls shrinking it? That is exactly where this kind of measurement goes wrong. Naive before/after flatters every control you will ever ship.

So I evaluated them causally instead. The honest effect is smaller than the naive one. That is not a disappointment. It is the difference between a number that survives scrutiny and one that doesn’t, and it is the only kind of number I would want to hand to someone who is about to move a threshold.

The same instinct ran through the fraud metrics I owned for business reviews, where the job was as much explaining why a number moved as reporting that it did.

  • PySpark / Databricks
  • Refund leakage
  • Benchmarking
  • GMV
  • Holdout
  • Counterfactual
  • Base-rate effects
  • WBR/MBR