AURA

Closing Runs Late: Look for the Stage, Not the Culprit

Closing ran 24 minutes over the norm on 7 of the last 11 shifts with one staff combination, and evening load did not explain it. Here is how to turn that into a stage rather than a culprit, how to test the fix over five shifts without fooling yourself, and what to write down so the same problem is not solved again next year.

Published
23 min read4657 words
Aura editorialAuthor

Key takeaways

  • Seven late closings out of eleven prove nothing on their own: the same run turns up in 27.4 % of purely random sequences.
  • A log that records only the roster will always name a person; add one clock mark per closing stage and the answer becomes a stage.
  • No regulation sets a closing time, so the norm has to be measured from your own ordinary nights rather than declared in an office.
  • Exclude the easy explanation first: put covers or revenue after 21:00 in the column next to closing minutes for the same shifts.
  • Five shifts buy a direction, not a measurement: five of five the right way happens by chance 3.1 % of the time, three of five 50 %.
  • Minus 18 minutes against a 24-minute overrun removes 75 % of the deviation and leaves 6, so say which base the 18 was measured against.
  • Seven fields — problem, solution, reason, owner, expected result, actual result, reuse — are what stop a restaurant re-deciding this next year.

Closing ran 24 minutes over the norm on 7 of the last 11 shifts with one particular staff combination, and evening load did not explain it. That is a pattern, not negligence. To fix it you need the stage where the minutes are lost, not the name of a person — and then a test that tells a real improvement apart from luck.

Seven out of eleven is a comparison, not a count

The first thing to do with "7 of 11" is to stop admiring it. On its own the number says almost nothing, because it has no partner.

63.6%
Seven divided by eleven is 0.636, that is 63.6 % of those shifts.

The reverse check is trivial — 0.636 × 11 = 7 — and it is also the point: the arithmetic is not where the difficulty lives.

Here is why the bare count is weak. Suppose closing ran long as often as not, for no reason at all, the way a coin comes up heads. Out of the 2 048 possible sequences of eleven shifts (2 to the power of 11), 562 contain seven or more late closings. That is 27.4 %. In other words, if lateness were a coin, you would see "seven or more out of eleven" in better than one case out of four, and you would still be tempted to call a meeting about it.

So the strength of the observation does not come from the seven. It comes from two comparisons standing next to it:

  • the other shifts. How often does closing run long when this combination is not on? If the rest of the roster closes at the norm, then 63.6 % against roughly nothing is a real difference. If everybody runs late half the time, you have found nothing about these people at all — you have found that your norm is wrong.
  • the obvious explanation. In this restaurant the delays were not caused by high evening load. That is a separate check, and it is done before anybody talks about the crew.

There is a third thing worth saying out loud, because it changes the whole investigation. The pattern is attached to a combination of people, not to a person. A combination is not a character trait — it is a particular division of duties that happens when those two roles stand next to each other. Which means the honest object of the search is a stage of work, and the names are only the label on the file.

A log that records only names will always name a person

Take the log most restaurants actually keep at 22:30: the time the last person left, and the roster. Two columns. Now ask that log why closing takes too long. It has exactly one thing that changes from shift to shift, so it will hand you exactly one answer: the people who were on.

That is not a moral failure of the manager reading it. It is a property of the instrument. A measurement that records a single feature will name that feature as the cause every time, and it will sound convincing while doing it, because the correlation is genuinely there. The seven late shifts really do share those two names.

Add one column — a clock mark at the end of each stage of closing — and the thing that varies stops being a person and starts being a stage. Nothing else about the restaurant changed. The conclusion changed because the log got wider.

This is the whole difference between "staff are working badly" and "closing is regularly delayed at stage X". The first is unfixable: you can only replace people, and the next pair will inherit the same seam. The second has a repair.

The closing table: name the stages, mark the clock

Write the closing routine down as stages, in the order they actually happen — bar cash-up, kitchen shutdown and cleaning, cold-store checks, floor reset, waste and delivery bins, alarm and lock-up, the shift note. Eight to twelve lines is normal. Then, for every closing, three fields per line: stage — clock time it finished — who did it.

One A4 sheet per shift is enough. This is a notebook job before it is a software job, and it should stay a notebook job long enough for you to see whether the stages you wrote down are the stages that exist.

Two rules make this table worth keeping:

The norm is measured, not wished. A norm invented in an office ("closing should take forty minutes") will make every shift late and teach you nothing. Take your own last ten or fifteen closings, per stage, and use the middle value. A norm is a description of your restaurant on an ordinary night; it becomes a target only after you know what an ordinary night is. The same reasoning behind the numbers an owner actually looks at applies here: a number you did not measure cannot be missed.

The stages must add up. Sum the stage durations and compare with the gap between "last guest left" and "alarm set". Minutes plus minutes give minutes, so this is a legitimate check of dimension, and it catches the most common defect of a fresh table: time that belongs to no stage. If the stages sum to 52 minutes and the room was occupied for 71, the missing 19 minutes are not noise — they are a stage you have not named yet, and quite often they are exactly where the delay lives.

Rule out the easy explanation before you look at people

The delays here were not caused by high evening load, and that sentence is a claim, not a decoration. It is established by putting one more column next to the closing time: a measure of how heavy the evening actually was. Covers served after 21:00 will do; so will revenue after 21:00, or the clock time of the last table leaving. Any of them, consistently, for the same eleven shifts.

Then look at the two columns together. If the seven late closings are also the seven heaviest evenings, you do not have a checklist problem — you have a staffing problem, and it belongs with sales per labour hour and with demand forecasting, because the answer is likely to be one more person on the floor rather than a new order of tasks. If the late closings are scattered across ordinary evenings, load is off the list.

Two warnings, both cheap and both learned the hard way.

Ruling out one explanation does not promote the next one to "proved". It only means the next one is now worth testing. Load was excluded; the seam between two roles is a hypothesis with better standing than it had, and nothing more.

And check the other cheap confounders in the same pass, because they cost one column each: day of the week, delivery days, the night of the monthly stocktake, the time of the last accepted booking, whether a trainee was on. A pattern that survives all of them is worth acting on. A pattern that dissolves the moment you write down "Fridays" has saved you from changing a process that was never broken.

From a stage to a cause: look at the seam between two roles

Once the table points at a stage, the useful question is not who was slow on it. It is which two roles meet at that stage, because that is where a division of duties can quietly go wrong in two directions:

  • the orphan. Nobody owns the task. Each role has a defensible reason to think the other does it, so it gets done at the end, by whoever notices, after everything else is finished. It is never skipped — that is why it never shows up as a complaint — it just always lands last.
  • the twin. Both roles own it. It gets done twice, or it gets negotiated every night, and the negotiation is the delay. This one is easy to miss because both people will describe the shift as normal.

There is a five-minute diagnostic for both. Ask each role separately, in writing, who performs stage X and what has to be finished before it can start. Two different answers, or two people naming the same predecessor task, is your seam. You do not need agreement in the room; you need the disagreement on paper.

The probable cause in this case was exactly that: the distribution of duties between two roles. Note the word — probable. It stays probable until a change is made and measured. What follows from it is a rewritten closing checklist, with the disputed stage given one owner and one predecessor. What does not follow from it is a roster change. Rearranging people around a seam moves the seam; it does not close it.

What 24 minutes cost, in money and in the record

The money is small, and it is worth computing anyway, because a number you can say out loud survives a conversation better than an irritation you can only describe.

This restaurant already has a price for a minute of one person's time: extending one shift by 90 minutes adds 54 PLN of payroll. So one minute of one person costs 54 ÷ 90 = 0.60 PLN, and the reverse check holds — 0.60 × 90 = 54 PLN.

From there:

what is countedarithmeticresult
total overrun across the eleven shifts7 × 24168 minutes, that is 2 h 48 min
average overrun spread over all eleven shifts168 ÷ 1115.27 minutes per shift
cost of the overrun, per person who stays for closing168 × 0.60100.80 PLN

Multiply the last line by the number of people who actually stay to the end — that number is yours, and it is the only one in this table you have to supply. Two closers make it 201.60 PLN over eleven shifts. Set against a restaurant turning over about 500 000 PLN a month with 25–30 staff, that is not the reason to act, and pretending otherwise is how process changes get sold and then quietly abandoned.

The real reasons are two. First, it is the tail end of a shift, landing on people who have already worked a full evening; recurring unpaid-feeling minutes at 23:00 are one of the cheapest ways to lose a good closer, and what a departure costs is not measured in zloty per minute. Second, this is working time in the legal sense, not goodwill.

Two provisions of the Polish Labour Code are worth knowing here, and both were taken from the consolidated text of the act, not from the "as published" edition. Article 151 § 1 states that work performed beyond the working-time norms binding on the employee constitutes overtime. Article 149 § 1 obliges the employer to keep a record of the employee's working time for the purpose of correctly establishing pay and other work-related entitlements, and to make that record available to the employee on request. Those are translations of the Polish wording, not literal quotations; the full text of the act is at api.sejm.gov.pl.

What the act does not contain is equally useful, and it was checked in both editions with a control phrase in the same pass: there is no statutory duration for closing a restaurant, no statutory closing checklist, and no obligation to record work stage by stage. The record of working time is required; the stage marks are yours, and nobody will hand you the norm. That is precisely why it has to be measured rather than declared.

Write the prediction down before you change anything

Before the new checklist goes on the wall, one sentence goes into the record, and it has four parts: which stage should shrink, by roughly how much, over how many shifts, and what result would count as a failure.

Something like: "Stage X should lose most of its overrun within five shifts; if the average closing time has not moved by at least ten minutes after five shifts, the seam was not the cause."

This sentence costs nothing and does one very large thing: it makes the test capable of failing. Without it, every outcome confirms the change. Closing got faster — the checklist worked. Closing did not get faster — the crew needs more time to adapt. Closing got slower — the new order is still bedding in. A test that cannot come out badly has not tested anything, and the same trap appears everywhere from holdout groups for promotions to separating your own change from the weather in an evening's revenue.

Add one more line while you are there: what you will not change during the test window. Not the roster, not the closing hour, not the menu, not the number of people staying to the end. Every extra change you make in that week is a rival explanation you will not be able to eliminate afterwards.

Five shifts: what such a test can show and what it cannot

Five shifts is a small test. It is worth doing anyway, and it is worth knowing exactly what you are buying.

Start with the cleanest thing five shifts can give you, which is not the average at all — it is the direction. If the new checklist did nothing whatsoever, each shift would land on the faster side or the slower side like a coin.

3.1%
Five out of five landing on the faster side happens in 1 sequence out of 32, that is 3.1 % of the time.

That is a genuine signal, and you can read it off the sheet without any statistics.

50%
Now the number that stops most people: three out of five going the right way happens in 16 sequences out of 32 — exactly 50 %. Three of five is a coin toss.

If your five shifts come back three-good and two-bad, you have learned nothing, and the correct response is to keep measuring, not to declare victory and move on.

What five shifts cannot do, however the result comes out:

  • estimate the size of the effect with any precision. "About eighteen minutes" is the right way to say it; "17.6 minutes" is not.
  • separate the checklist from everything else that happened that week. A quieter week, a new supplier delivering earlier, a manager standing closer to the door — all of these are alive inside the same five nights.
  • survive a season. Five shifts in November say nothing about the same restaurant in July.
  • show whether the change holds after attention moves elsewhere. A crew that knows it is being measured closes faster; that part of the effect is real, it is not cheating, and it fades.

Two guards cost nothing and roughly double what you learn. First, keep measuring for five more shifts without announcing it — if the improvement is a process change it stays, and if it was attention it decays, and you will see which. Second, do not let the person who designed the new checklist be the only one reading the result. That is not distrust; it is that the author knows what the number is supposed to say, and knowing that is enough to bend a judgement about a stage that "basically finished on time".

And one rule with no exceptions: the norm may not be redefined during the test. Moving the baseline mid-test is the single most common way an honest measurement turns into a decoration.

Why minus 18 minutes against 24 still counts as a result

After the five shifts, average closing time was 18 minutes shorter. Set against the 24-minute overrun, that is 18 ÷ 24 = 0.75, three quarters of the deviation removed, with 24 − 18 = 6 minutes still on the table. Reverse check: 18 + 6 = 24.

Before anybody celebrates, the honest reader has to say what the two numbers are measured on, because they are not measured on the same thing:

  • 24 minutes is the overrun on the late shifts of one combination, against the norm.
  • 18 minutes is a shift of the average closing time across five shifts under the new checklist.

Comparing them is allowed, but only out loud. And there is a check that forces the issue. The average overrun spread across all eleven shifts is 7 × 24 ÷ 11 = 15.27 minutes. If your "before" figure for the five-shift comparison were that 15.27, then an improvement of 18 minutes would be larger than the entire overrun, meaning closing is now finishing ahead of the norm itself. That is possible — but if it is true, your norm needs re-measuring, and if it is not true, then the "before" you actually used was the late shifts, not all eleven. Either answer is fine. Not knowing which one you used is not.

With the base named, minus 18 counts as a result for three reasons, and none of them is the size of the number:

  1. The direction was predicted before the change, in writing, together with a condition that would have counted as failure.
  2. The effect is large next to the everyday scatter of the metric. Eighteen minutes is not a rounding of the difference between a Tuesday and a Friday; if it were, you would already know, because the scatter is in the same table.
  3. The remainder was named rather than absorbed. Six minutes are still over. That is the next question, not a smaller version of the same one, and it may well have a different cause.

The correct verdict, written in exactly these words, is: not refuted, worth keeping, re-check in a month. Not "proved". The difference matters, because "proved" closes the file and "worth keeping" leaves the next measurement scheduled.

The money side, for completeness: five shifts at 18 minutes saved is 5 × 18 = 90 minutes, and at 0.60 PLN per minute that is 54.00 PLN per closer — the exact price of one 90-minute shift extension, which is the sort of decision covered in an evening forecast with a price attached. Fifty-four zloty is not why you did this. The ninety minutes given back to people at 23:00 is.

The memory of decisions: seven fields that stop you solving this again next year

Everything above is worth one entry in a record that most restaurants do not keep. Not "I remember the manager's name" — that is address-book memory. The entry has seven fields:

  1. which problem was found — closing over the norm on 7 of 11 shifts with one combination;
  2. which solution was chosen — a rewritten closing checklist with one owner for the disputed stage;
  3. why that one — load was excluded, the stage table pointed at a seam between two roles;
  4. who carried it out — by role, and who checked;
  5. what result was expected — written before the change, with a failure condition;
  6. what actually happened — average closing 18 minutes shorter over five shifts, 6 minutes still over the norm;
  7. whether to use this approach again — yes, with the re-check scheduled.

Seven lines. Ten minutes. And it is the difference between a restaurant that accumulates management experience and one that re-derives it every year, because the reason businesses forget is not carelessness — it is turnover. The manager who ran this test leaves, the closer who knew the seam moves to another restaurant, and eighteen months later somebody looks at a fresh table, sees two names against seven late closings, and writes "staff are working badly" on a fresh sheet of paper.

Record the failures too, and record them in the same place. A solution that did not work is more expensive to repeat than one that did, because it comes back wearing the same reasoning that made it attractive the first time. Over a couple of years these entries stop being a diary and become the thing that makes the next recommendation start from "here is what already worked in this restaurant" rather than from a blank page. That value compounds hardest across several locations, where the same seam is being rediscovered independently in every one of them, and it is the same discipline that keeps a restaurant chain readable on one screen.

What a system does here, and what has to stay with you

Everything on this page can be done in a notebook. Slower, and only if somebody remembers to do it every night, which is the actual reason it usually is not done.

What a management system changes is narrow and worth stating precisely. It keeps the stage marks without anybody remembering to write them, because the marks come from the till, the door and the task list rather than from a pen. It notices "7 of 11" without being asked, because comparing a combination against the rest of the roster is exactly the sort of dull comparison that machines do not get bored of. It puts the load column beside the closing column automatically, so the obvious explanation is excluded on the first screen rather than in week three. It holds the written prediction and compares against it after five shifts by itself, which removes the single most common failure of these tests — that nobody ever comes back to look. And it keeps the seven-field record, so the next similar pattern arrives with the previous decision attached.

The pieces that do this work in practice are shift and team data and task tracking for the raw marks, analytics and dashboards for the comparison, and a decision engine for holding the prediction and reporting back on it. The wider picture of what such a system is for is in an AI restaurant management system and, at a more practical level, in restaurant automation.

What it must not do is decide. It should not name a culprit — it has the same data you do, and "combination of people" is a label, not a verdict. It should not change the roster. It should not adopt a process change because five shifts came out well. The boundary between what a system does alone, what it does with a person, and what stays a decision for the owner is a design choice made in advance, not an accident of configuration; it is set out in three levels of system autonomy, and it is the same boundary that keeps a single complaint from turning into a root-cause hunt before anyone has checked whether there is a pattern at all.

If you are wondering where to start, start with the stage marks and nothing else. One sheet per shift for two weeks. Most of what is written above becomes obvious once that sheet exists, which is roughly the point made in four thresholds for automating a small business — and the situations where a change is not worth making at all are worth reading before you make one, in five situations where we advise against it. The same discipline applies to the money side of a menu decision, where a dish that sells while earning less is found the same way: a pattern first, a cause second, a test third. If you want the metric tree this all hangs from, it is in the restaurant KPI guide and in labour cost as a percentage.

Frequently asked questions

Is 7 late closings out of 11 shifts enough to act on?

Enough to investigate, not enough to conclude. On its own, seven or more out of eleven appears in 27.4 % of purely random sequences, so the number carries weight only next to two other things: how often the remaining shifts run late, and whether the obvious explanation — evening load — has been excluded. Once both comparisons are in the table, acting on it is reasonable. Acting on the seven alone is how a process gets rewritten for no reason.

How do I set a norm for closing time if no regulation gives one?

You measure it. There is no statutory duration for closing a restaurant and no statutory closing checklist — this was checked in both editions of the Polish Labour Code, including the consolidated text. Take your own last ten to fifteen closings, break them into stages, and use the middle value per stage. A norm invented as a target makes every shift late and teaches you nothing; a norm measured from ordinary nights tells you immediately which shifts are unusual.

What if the delay really is caused by a heavy evening?

Then you have a different problem with a different repair, and that is good news, because staffing questions are easier to answer than process questions. Compare closing minutes against covers or revenue after 21:00 for the same shifts. If the late closings are the heavy ones, the answer is likely to be an extra pair of hands at the end of the evening rather than a new order of tasks — priced the way any shift extension is priced, in payroll minutes.

Why test on five shifts instead of thirty?

Because thirty shifts is six weeks, and a process change nobody can evaluate for six weeks usually stops being followed in week two. Five shifts buys you a direction, not a measurement: five out of five in the right direction happens by chance in only 3.1 % of cases, which is a real signal. Three out of five happens 50 % of the time and means nothing. Use the five shifts to decide whether to keep going, then keep measuring quietly.

The average dropped 18 minutes but closing is still 6 minutes over the norm. Did it work?

It worked in the sense that matters and did not finish the job. Three quarters of the deviation is gone — 18 ÷ 24 = 0.75 — and the remaining 6 minutes are a separate question that may have a different cause. Write the verdict as "not refuted, worth keeping, re-check in a month" rather than "solved". And say which base the 18 minutes was measured against: the late shifts, or all eleven. Those give different stories.

Do I have to name the two roles in the record, or is that blaming people?

Name the roles, never the people. The entry says "the seam between the bar close-down and the floor reset", not two surnames. That is not politeness — it is accuracy, because the pattern was attached to a combination, and combinations are divisions of duties. A record naming people ages badly: those two leave, the seam stays, and the next person reading the file learns nothing usable.

What do I write down if the five shifts show nothing at all?

The same seven fields, with "what actually happened" filled in honestly and "use again" set to no. A test that came out empty is worth more written down than remembered, because the reasoning that made the change attractive will come back and will look just as good the second time. Note what you would need in order to test it properly — more shifts, a comparison group, a cleaner week — and leave it. Then go back to the stage table and look at the second-largest gap.

Closing time is one of the few restaurant numbers that costs nothing to record and tells you something about the process rather than about the market. If you want to see how the rest of the operating numbers fit together — from staffing to margins to the evening forecast — the whole set lives in our restaurant section.

Related services

In this section

Let us look at your numbers

Tell us how enquiries are handled today — how many there are, who picks them up, where they get lost. Aura walks the process with you and shows what can be taken off a person, and what is better left alone.

Talk to Aura

The home page with Aura opens. Give a company name — she looks at it in public data and shows what a client sees. No promises of a result.

Prefer to write? marketing@auraglobal-merchants.com

Next step

Let us check whether Aura fits your place

We do not take everyone: first we look at your processes, sales and current systems and tell you honestly whether it makes sense for us to come in. A few questions, about five minutes.

Take the assessment →