Lailara Lift Math

Accuracy · estimate vs truth

← Back to the Scorecard

How wrong are these numbers?

Every incrementality tool asserts accuracy. This one measures it. The two baselines are scored against known ground truth — the error is shown, by regime, including where it is large. That is the whole claim, and it is a narrow one: this is the error a standard method makes under a realistic, fully-known world. It is not a prediction of the error on your data.

The estimators are provably blind — enforced in code, not promised. An AST gate runs over every estimation file on every push; the generator's own coefficients are banned from the estimation path; and the git history shows both methods frozen and tagged before this page's code first read truth. The blindness claim is scoped exactly there — to the code — and nowhere wider.

Median error, full population

Method 0 · pre-period

26.5%

median absolute error on incremental units

Bias +11.6% · 120 events scored

Method 1 · comparable-store

26.3%

median absolute error on incremental units

Bias +22.0% · 117 events scored

Scored covers the estimable events with a measurable true lift — 120 of 129 for Method 0, 117 of 129 for Method 1. The rest have true incremental below one unit — phantom and negligible-effect promotions, where a percentage of almost nothing is undefined — so they are set aside from the median, not hidden. The four seeded stories are included in these figures and also broken out below.

Both methods over-credit promotions — the bias is positive for each. And the more defensible comparable-store method over-credits by more, not less. Two conclusions follow, both honest: the true promo book is worse than either method shows, and a better baseline is not automatically a better-calibrated one. A demonstration engineered to make the naive method lose would show neither.

The four seeded stories, scored separately

These four events are planted outliers the tool is supposed to surface. They are reported here, apart from the headline median above — the honest denominator is the full population, not the outliers.

StoryMethod 0 errorMethod 1 error
Pure subsidy PRE-0002-35.7%+8.6%
Hero cannibal PRE-0043+4.6%+8.8%
Pantry trap PRE-0056+1.0%+2.2%
Clean winner PRE-0087-15.3%+25.3%

Error by regime

Median absolute error, cut by observed features only — promotion type, depth, season and the like. No cut uses a truth-derived label, and every bucket holds at least five events, so no number reads back to an individual promotion.

Retailer

RetailerMethod 0Method 1
RET-COSTCO13.7%30.7%
RET-KROGER27.2%25.9%
RET-REGIONAL39.0%39.4%
RET-SPROUTS36.6%30.2%
RET-WALMART16.1%20.5%
RET-WHOLEFOODS28.7%15.3%

Promotion type

Promotion typeMethod 0Method 1
BOGO12.9%8.0%
TPR_deep17.5%20.6%
TPR_shallow49.5%38.3%
ad_circular16.5%21.9%
digital_coupon132.2%69.5%
endcap16.1%22.8%

Coupon events have small true lifts, so percent error is measured against small denominators; the metric is stressed here, not just the method.

Product line

Product lineMethod 0Method 1
AS33.8%25.6%
DG26.0%23.8%
PS17.8%26.1%
SB16.1%25.3%
SC30.4%29.4%

Discount depth

Discount depthMethod 0Method 1
deep (20–30%)19.9%24.0%
moderate (10–20%)33.0%26.2%
shallow (<10%)27.0%37.0%

Duration

DurationMethod 0Method 1
1 week19.8%25.9%
2 weeks33.0%29.3%
3+ weeks26.9%26.5%

Season

SeasonMethod 0Method 1
Fall35.7%26.5%
Spring19.5%21.2%
Summer16.0%31.3%
Winter49.1%24.2%

Match relaxation — Method 1 only

Method 1 matches comparable stores by region, format class and volume; where the in-format pool is too thin it relaxes to region and volume alone. This cut asks whether the relaxation costs accuracy — and it does, though not where you'd look first: median error holds, but fully-relaxed events run +22.6% hot against +15.7% for mixed.

Match stratumMedian errorBiasEvents
fully relaxed26.3%+22.6%63
mixed26.5%+15.7%53

Synthetic data is the only honest testbed for this. It is at once the only world where truth is knowable and the only world that can be published: real client promotion data can never be shown, by any vendor, at any client count. Anyone claiming to demonstrate accuracy on real client data is either breaching confidentiality or making it up. The methodology and the deliverable are real; the data is synthetic, and that is the point.