A promotion is judged by comparing two comparable groups of guests over the same period: one that received it and one deliberately left out. The difference between their visit rates, multiplied by the size of the treated group, gives the incremental visits. A difference smaller than the ordinary swing of that same rate has not been measured at all.
Every figure in the worked example below is a teaching number, picked so that the arithmetic can be checked line by line. None of them is a benchmark, none describes a real venue, and none should be carried into your own calculation. There is no official statistic for how much a restaurant promotion lifts visits, and this page names none: the catalogue of Eurostat datasets carries nothing on campaign response or promotion effect for the accommodation and food service section — its only entries mentioning marketing are the old innovation-survey tables about types of marketing innovation — and the trade section of the Polish statistical office publishes turnover and a methodological handbook for trade and catering activity, not campaign results. So the page gives you the arithmetic instead of a number to copy.
Why "sales went up after the campaign" says nothing about the campaign
A before-and-after comparison compares two weeks, not two groups. Between those two weeks everything else in the city also moved: the weather turned, a public holiday landed, salaries were paid, a competitor two streets away closed for renovation, a local match filled the district for one evening. All of it lands inside the same number, and the campaign is given credit for the lot.
Confounder — anything that changed at the same time as the campaign and moved the whole restaurant together: weather, a holiday, a competitor closing, a review that happened to appear. A holdout absorbs confounders, because both groups live through the same week. A before-and-after comparison absorbs nothing.
Three ways to judge a promotion, and what each one mistakes for the effect
| Method | What it compares | What it mistakes for the effect | When it is still usable |
|---|---|---|---|
| Before and after | The campaign week against the previous week | Weather, calendar, payday, a competitor's closure | Nothing else moved and you can name what did not move |
| Year on year | The campaign week against the same week last year | A year of price changes, a changed menu, a changed neighbourhood | The venue has a long, stable, written history |
| Treated against holdout | Two groups of guests inside the same week | Only the difference between the two groups themselves | Your guests are identifiable and a share of them can be left out |
The third row is the only one whose error term is small enough to name. That is the whole argument of this page.
What a holdout group is, and why it costs almost nothing
Holdout group — guests deliberately excluded from a campaign so that their behaviour shows what would have happened anyway.
Treated group — guests who received the campaign. They are comparable to the holdout only if the split was made before the campaign started and by a rule unrelated to how much they spend.
A holdout is not an experiment you pay for. It is the part of the send you do not make. If the campaign goes to nine guests out of ten instead of ten out of ten, the entire cost is the margin you would have earned from the tenth — and only from the share of that tenth who would have answered at all. Against a campaign whose effect is unknown, that is a small price for the only number that survives an argument.
Whatever channel carries the offer, the holdout is defined the same way: it is exactly the people you did not send to. If the sending is automated, the exclusion has to be automated with it, or a well-meaning colleague will send the tenth guest the offer "so nobody is left out" and dissolve the measurement. Set the exclusion where the sending lives — in the message channels — and not in a note beside it.
How to split the base so the two halves are comparable
Two rules decide whether the split is honest, and both are about timing rather than statistics.
The split happens before the campaign, not after. A group assembled after the send — "let us compare the ones who redeemed with the ones who did not" — compares people who wanted the offer with people who did not want it. That difference existed before the campaign and will still be there after it.
The rule has nothing to do with spending. Sorting by last visit, by average bill or by loyalty tier puts the better guests on one side, and the better guests would have come back anyway.
Splitting rules, and what each one does to the comparison
| Splitting rule | Gives comparable groups | How it spoils the comparison |
|---|---|---|
| Last digit of the guest record number | Yes | Nothing, provided the number was not issued by spend or by tier |
| Alphabetical halves of the surname | Yes, in practice | Almost nothing; watch for a district where one surname group is concentrated |
| Everyone above the average bill gets the offer | No | The treated group is made of better guests before a single message is sent |
| Everyone who has not visited for three months | No | The two groups differ by exactly the thing being measured |
| Whoever opened the last newsletter | No | The treated group is made of people who already read you |
| Whoever the manager remembers fondly | No | Unrepeatable, and it changes with who is on shift |
The split lives in the guest records, so it is made where those records are kept and not in a spreadsheet copy that ages the moment it is exported — see guest records and, if the question of which system layer holds them is still open, the article on CRM or ERP.
Response rate: one formula, two sets, one window
Response rate — the share of a group that visited inside the observation window.
Response rate = Guests of the group who visited in the window ÷ Guests in the group
Guests of the group who visited in the window— the count of distinct guests from that group with at least one visit inside the window, guests;Guests in the group— the size of the group as it stood at the moment of the split, guests;- the window is one and the same for both groups and is fixed before the campaign starts.
Fixing the window in advance matters more than it sounds. A window chosen afterwards will always be the window that flatters the campaign, and the person choosing it will not notice they are doing so.
The worked example runs on a base of 1 200 guests with contact details, split in half by the last digit of the record number: 600 treated, 600 held out, a window of thirty days.
- treated: 96 of 600 visited, so
96 ÷ 600 = 0.16, that is 16.00 %; - holdout: 75 of 600 visited, so
75 ÷ 600 = 0.125, that is 12.50 %.
A lift in percentage points and a lift in visits are two different numbers
Incremental lift — the difference between the response rates of the treated and the holdout group. It is measured in percentage points, not in per cent.
Incremental lift (percentage points) = Response rate (treated) − Response rate (holdout)
- both terms are shares of their own group and are dimensionless;
- their difference is dimensionless too, and the name of that difference is percentage points.
16.00 % − 12.50 % = 3.50 percentage points.3.50 ÷ 12.50 = 0.28, a relative lift of 28 %.Both numbers are true and they are not interchangeable: one is a difference, the other a ratio.
It is the most common way a promotion number lies — the same class of trap described in the numbers an owner actually looks at.
Turning the lift into something a kitchen can feel takes one more multiplication:
Incremental visits = Incremental lift × Guests in the treated group
Incremental lift— the difference of the two shares, dimensionless;Guests in the treated group— the size of the treated group, guests;- the product is therefore in visits.
In the example, 0.035 × 600 = 21 incremental visits. That is the number to put on the screen next to the base, rather than the raw count of 96 — see dashboards.
How to label both numbers so the report still reads a year from now
A difference and a ratio get confused in reports not because anyone cannot divide, but because both numbers are handed the same label: "uplift". The label decides whether this report will line up with the next one, so it is worth settling once and keeping.
| Number | What it is | Unit | How to label it |
|---|---|---|---|
| 16.00 % and 12.50 % | The response rates of the two groups | Per cent of their own group | "response rate, treated / holdout" |
| 3.50 | The difference between those two rates | Percentage points | "lift, percentage points" — never with a per-cent sign |
| 28 % | The same difference divided by the holdout rate | Per cent of the comparison base | "lift relative to the holdout group" |
| 21 | The difference carried onto the size of the treated group | Visits | "incremental visits, treated group" |
The lift in money: contribution first, campaign cost second
Visits are not money until they are multiplied by what a visit leaves behind.
Incremental margin = Incremental visits × Contribution margin per visit − Campaign cost
Incremental visits— the visits the campaign added, visits;Contribution margin per visit— what one visit leaves after variable costs, PLN per visit;Campaign cost— everything the campaign consumed: the cost of what was given away, the channel fee, the design work, PLN;- visits multiplied by PLN per visit gives PLN, and PLN minus PLN is PLN.
The contribution margin is not defined on this page, because it already has an owner. Quoted from the break-even pillar word for word: "Contribution margin — the money left from net sales after variable costs, available to cover fixed costs and, once they are covered, to become profit. Expressed as a share of net sales it is the contribution margin ratio; expressed per guest it is the contribution margin per cover."
In the example, with a contribution of 38 PLN per visit and a campaign that cost 500 PLN in total: 21 × 38 = 798 PLN, and 798 − 500 = 298 PLN. Positive, and far smaller than almost anyone expects at the moment they approve the campaign. If the giveaway is a free dish or a loyalty reward, its cost belongs inside Campaign cost at its own variable cost, which is the subject of the loyalty programme page; what to do with the result afterwards belongs to return on advertising measured on margin.
How many guests you need before a difference is visible at all
A measured share wanders. Repeat the same campaign next month on a group of the same size, change nothing else, and both rates will land somewhere slightly different. The size of that wandering is the standard error, and it decides the smallest difference you are allowed to read.
Standard error of a proportion — how far a measured share would wander if the same experiment were repeated on a group of the same size. It sets the smallest difference that can be read at all.
Standard error of the difference = √( p1 × (1 − p1) ÷ n1 + p2 × (1 − p2) ÷ n2 )
p1,p2— the response rates of the two groups, dimensionless shares between 0 and 1;n1,n2— the number of guests in each group, guests;- a dimensionless quantity divided by a count and then square-rooted stays dimensionless, of the same nature as the share itself, so the result is read in percentage points.
In the example: 0.16 × 0.84 ÷ 600 = 0.000224, 0.125 × 0.875 ÷ 600 = 0.00018229, the sum is 0.00040629, and its square root is 0.02016 — about 2.02 percentage points.
Why a difference smaller than its own wobble cannot be read
The working rule of this page: a difference below twice the standard error is not a result. It is not a small result, not a weak result, not a result that needs one more week — it is a measurement that did not happen.
In the example, twice the standard error is 2 × 2.02 = 4.03 percentage points, and the measured difference is 3.50. So the honest sentence about that campaign reads: the promotion may well have worked, and this base is too small to say so. The 298 PLN calculated two sections above is arithmetic performed on a difference that has not been established.
This threshold is a rule of common sense stated out loud, not a statistical test. A formal test of a hypothesis carries conditions and caveats that do not fit on a page about running a restaurant, and dressing an approximation up as a formal criterion would be the same kind of lie as inventing a benchmark. What the threshold is good for is the decision it prevents: rolling a campaign out to the whole base on the strength of a difference nobody could read.
Turning the same inequality around answers the more useful planning question — how big the groups have to be before a difference of a given size becomes readable:
Guests needed in each group = 4 × ( p1 × (1 − p1) + p2 × (1 − p2) ) ÷ d²
p1,p2— the response rates you expect, dimensionless;d— the difference you want to be able to read, in the same dimensionless units;- a dimensionless numerator divided by a dimensionless square gives a pure count, guests.
For the rates of the example and a difference of 3.5 percentage points: 4 × (0.1344 + 0.109375) ÷ 0.001225 = 0.9751 ÷ 0.001225 = 796 guests in each group, about 1 592 in the base. The same pair of rates reads a difference of 7 percentage points on 199 guests per group: halve the difference you want to see and the group you need grows fourfold. That is the single most useful fact about experiment sizes, and it is the reason a small base should test a big change rather than a small one.
What to do when the difference does not clear the threshold
"Not measured" and "did not work" are two different sentences, and confusing them costs you either a campaign that worked and was dropped, or one that did not and was rolled out to the whole base. There are three ways out and all three are honest; the dishonest one is the fourth — writing in the report that the promotion succeeded.
Enlarge the groups and repeat. Suits you when the base is bigger than the one that went into the test and you intend to run the campaign again anyway. The group-size formula on this page tells you how many guests are needed at the rates you have already measured — and you did measure them, even though the difference could not be read.
Test a bigger change instead of the same one. Suits a small base, and it follows straight from the same formula: a difference twice as large needs a group four times smaller. A small discount on a small base is unreadable by arithmetic, not by execution.
Repeat the same campaign and pool the two rounds. Suits you when nothing in the restaurant changed between them and the split was made by the same rule each time. Two rounds on the same base are bigger groups, not two separate results to be compared with each other.
Whichever of the three you take, the report says what came out: the measured difference, the threshold, and the sentence that the difference did not clear it. That sentence sounds weak, and it is the only one that will still be true a quarter from now.
What to do when nobody can be held out
Sometimes a holdout is impossible: the offer is a poster in the window, the base is 200 people, the owner refuses to leave regulars out. Then the comparison gets weaker, and the honest move is to say by how much rather than to pretend otherwise.
Alternate the weeks. Run the offer in odd weeks and not in even ones, over a long enough stretch. It absorbs slow drifts but not a holiday that lands inside one of the halves.
Split the rooms instead of the base. Two venues in a chain, one running the offer and one not, then swapped over the following month. It absorbs the calendar and the weather, and it does not absorb a difference between the two neighbourhoods.
Compare against a forecast rather than against last week. A written forecast with calendar, weather and event factors gives a baseline that already contains the confounders — that is exactly what demand forecasting is for. It absorbs everything the model knows about and nothing it does not.
All three are worth doing and none of them is a holdout. Name in the report which of the three was used, and the report stays honest even when the answer is unclear. This is also the first of the situations listed in the five cases where we advise against automating: measuring nothing is cheaper than measuring the wrong thing, but it is not cheaper than measuring the right thing badly and knowing it.
What a holdout will not prove: the long tail and the reputation effect
A thirty-day window measures thirty days. A guest who kept the offer, forgot about it and came back in the eleventh week counts as a non-response, so the number you get is a floor rather than the whole effect. Widening the window does not fix this for free: the longer the window, the more of everything else falls inside it, and the difference between the groups gets harder to read, not easier.
Three things a holdout will not tell you, and they belong in the report rather than in an argument three months later:
- Whether the campaign moved people who never had a card. The measurement covers named guests. A poster works on the street, and the street is in neither group.
- What the offer did to full-price sales. Guests who would have come anyway and used the discount are not incremental, and the discount they used is a real cost — that arithmetic belongs to the discount break-even page, not to this one.
- Whether the campaign brought anyone new. New guests are not in the base at the moment of the split, so they can be in neither group; their cost is counted separately in the cost of acquiring a guest, and where the list of named guests comes from in the first place is the returning-guest page.
The report that comes out of all this is short and dull: two group sizes, two response rates, one difference in percentage points, one standard error, and a sentence saying whether the difference clears twice that error. Dull is the point — it is the same report next quarter, and next quarter it will compare. Automating the writing of it is worth more than automating the sending of the campaign: see automated reports and analytics, and, for where to start at all, the four thresholds and the overview of what a restaurant can automate.
Frequently asked questions
Why is a sales increase after a promotion not proof that it worked?
Because it compares two weeks rather than two groups, and between those two weeks the weather, the calendar, the neighbourhood and your competitors also changed. Every one of those differences lands inside the same number, and the campaign is given credit for all of them. A holdout group removes them, because both groups live through the same week under the same conditions.
What is a holdout group and how large should it be?
It is the share of your guests you deliberately do not send the offer to, so that their behaviour shows what would have happened without the campaign. Its size is not a matter of taste: it follows from the difference you want to be able to read. Put your expected rates and that difference into the group-size formula on this page, and the answer comes out in guests.
How do I split my guest base fairly?
By a rule fixed before the campaign and unconnected to how much a guest spends — the last digit of the record number does the job. Anything sorted by average bill, by loyalty tier or by "who has not been in for a while" puts the better guests on one side, and the better guests would have come back without your offer.
What is the difference between percentage points and per cent here?
A rate of 16.00 % against 12.50 % is a difference of 3.50 percentage points and a relative lift of 28 %. Both are correct, they answer different questions, and they must never share a label. Write "percentage points" every time the number is a difference of two shares, and the same report will still be readable a year from now.
How small a difference can I actually detect?
Roughly, anything above twice the standard error of the difference, which the formula on this page computes from your two rates and your two group sizes. Below that line the number is not a small effect but an unmeasured one. This is a stated rule of thumb rather than a formal statistical test, and it is used here to prevent one specific decision: a full rollout on the strength of a difference nobody could read.
What do I do when I cannot hold anyone out?
Alternate the campaign week by week, or run it in one venue of a chain and not in another and then swap them, or compare the result against a written forecast that already carries calendar, weather and event factors. All three are weaker than a holdout, all three are far better than a before-and-after chart, and the report should say which one was used.
Your first holdout group can be set up today and it costs nothing: leave every tenth guest out of the next send, and fix the observation window before the send goes out. One campaign later there will be a number on the table that survives being argued with — and the rest of the arithmetic that turns it into money lives in the restaurant section.