AURA

Three Levels of Autonomy: What Runs Without Asking You

The evening report fits into nine lines, and two of them do not mean what the reader assumes. Here is the reverse check on both, the five signs that assign a case to one of three levels of autonomy, the price of a boundary drawn too low and too high, and a one-page register you can keep on paper.

Published
26 min read5129 words
Aura editorialAuthor

Key takeaways

  • The value is not the number of notifications but how many situations closed without disturbing management. An owner whose system writes more often has been given more work, not more control.
  • Reverse check: 21 480 ÷ 126 = 170.48 PLN, rounding to 170 — so the divisor was guests, not bills. The true average bill sits between 170.48 and 290.27 PLN (21 480 ÷ 74 bookings), and party size is 126 ÷ 74 = 1.70.
  • A percentage without a base is not a measurement: 21 480 ÷ 1.057 gives a forecast of about 20 320 PLN and a gap of about 1 158 PLN. The method is proven on a pair where both numbers are printed: (18 740 − 19 600) ÷ 19 600 = −4.4% and 860 PLN.
  • The four counts — 42 inquiries, 8 returns, 31 tasks, 4 deviations — must not be summed: four populations, four units, four bases. Each becomes a measurement only when divided by its own base.
  • A case's level is the highest level any single one of the five signs demands, and it belongs to the case, not to the channel. Responsibility stays human at all three levels — what changes is the number of decisions, not the responsibility.
  • A boundary drawn too high shows up in the refusal rate: an item at zero for a quarter is not a decision but a notification in a decision’s clothing.

A few words that show up in this text

Explained in plain language — you do not need to know the trade to read on.

CRM
One place holding clients and enquiries: who asked, about what, and what happened next.
follow-up
A planned return to the client after the first conversation or quote.

Sort every recurring situation into three levels: the system closes it alone, the system prepares it and a person acts, or the owner decides between prepared options. The level is assigned by written rules, reversibility and who carries the consequence — not guessed. Responsibility stays human at all three.

What the evening report says when the day was ordinary

At 22:30 the day is collected. The manager does not sit for another ninety minutes writing the owner a wall of text. The report that arrives when nothing went badly reads like this:

Day closed.

21 480 PLN in revenue, +5.7% against forecast.

126 guests.

Average check 170 PLN.

42 inbound inquiries handled without the team.

8 guests came back after CRM campaigns.

31 internal tasks closed.

4 operational deviations resolved automatically.

One situation worth a look tomorrow: cooking time is rising in one category.

The rest is within range.

Nine lines. What is remarkable about them is not what they contain but what they replace. The same day produced hundreds of small events — messages, confirmations, reminders, a supplier running late, a check that did not pass on the first ask. None of them are in the report, because none of them needed the owner.

That is the whole idea, and it is worth stating plainly because it is easy to get backwards. The value is not in the number of notifications. The value is in how many situations were closed without disturbing management. An owner whose system writes to him more often has not been given more control; he has been given more work.

Exactly one line here asks for a person: cooking time in one category. Everything else arrives as a count, and a count is not an event — it is a summary you either accept or open.

Which raises the first practical problem. A report this short is only readable if you know what each line was divided by, and two of these nine lines do not survive a reverse check in the form a reader assumes. The next two sections do that check on paper. Which figures deserve a place in front of an owner at all is worked through in the numbers an owner actually looks at, and the machinery that assembles a summary instead of a feed is automatic reports.

Reverse check: 21 480 divided by 126 against an average check of 170 PLN

Take the two largest numbers in the report and divide one by the other:

21 480 PLN ÷ 126 guests = 170.48 PLN per guest

Rounded to whole zloty that is 170, exactly what the fourth line says. Reverse the step: 126 × 170 = 21 420 PLN, sixty zloty short of the revenue line, and those sixty zloty are the rounding — 0.48 PLN across 126 guests is 60.5 PLN. The arithmetic closes.

The convergence is what should make you stop, because what converged is not what the label says. An average check, in the ordinary sense of the phrase, is revenue divided by the number of bills. A table of two leaves one bill; a family of four leaves one bill. Had the divisor been bills, the result would not have landed on the revenue-per-guest figure by coincidence. It landed there because the divisor was guests.

Nothing was falsified. A metric simply carries a name belonging to a different metric, and the two differ by a factor equal to the average party size.

How much does that matter here? The same evening had 74 bookings on the sheet, the number the shift forecast worked from. That sets the bounds:

divisorarithmeticresultwhat it would mean
every guest pays separately21 480 ÷ 126170.48 PLNrevenue per guest
every booked party pays one bill21 480 ÷ 74290.27 PLNrevenue per booked table

The true average bill sits between them. It cannot be under 170.48, because a bill covers at least one guest, and it is unlikely to reach 290.27, because walk-ins were part of that evening and each one adds a bill without appearing on the booking sheet. Party size follows from the same pair: 126 ÷ 74 = 1.70 guests per booking.

The consequence is not academic. One manager believes a guest spends 170 PLN and builds an upsell target on it; another believes a table spends 290 PLN and reads the evening as a strong one. They are reading the same line.

A second denominator hides in the same place. If delivery revenue sits inside the top line while delivery customers are not inside the 126, revenue per guest is overstated by the delivery share, and nobody notices, because both numbers are individually correct. The report does not say either way. Pin that down once, in the line itself:

Average check 170 PLN (revenue ÷ guests; delivery included / excluded)

Ten extra words remove an ambiguity that otherwise repeats every evening. Building a metric tree in which every figure names its own denominator is the subject of the restaurant KPI guide.

What forecast is hiding behind plus 5.7 per cent

The second line gives a percentage and withholds its base. A percentage without a base is not a measurement, it is a mood. Recover the base:

Forecast = Actual ÷ (1 + deviation) 21 480 ÷ 1.057 = 20 321.7 PLN

So the evening was measured against roughly 20 320 PLN, and the absolute gap is about 1 158 PLN.

Before trusting that, prove the method on a pair where both numbers are printed. That same morning the system reported on the previous evening: 18 740 PLN of actual revenue against a forecast of 19 600. Run it the other way:

4.4%
(18 740 − 19 600) ÷ 19 600 = −0.0439 = −4.4% absolute gap: 19 600 − 18 740 = 860 PLN

Minus 4.4 per cent and 860 PLN are exactly the two figures that morning message carried. So the percentage is computed against the forecast, not against the actual, and the recovery above is the same equation solved for the other unknown. That morning case is worked through in why evening revenue dropped.

One note on precision. Plus 5.7 per cent is rounded, so it pins the forecast to a band rather than a point:

21 480 ÷ 1.0575 = 20 312.1 PLN 21 480 ÷ 1.0565 = 20 331.3 PLN

Every forecast between about 20 312 and 20 331 PLN rounds to plus 5.7 per cent. The honest statement is "about 20 320 PLN, give or take ten". A recovered number inherits the precision of what it was recovered from, and printing it to one decimal is a small lie about how much you know.

Why bother? Because the same plus 5.7 per cent is a triumph against a cautious forecast and a disappointment against an ambitious one, and the line does not say which. A venue that quietly revises its forecast down every week will keep beating it every week while revenue falls. The forecast is a claim about the future and has to be scored as one — the method is restaurant demand forecasting.

Four lines that must not be added into one number

The middle of the report holds four counts: 42 inquiries handled without the team, 8 guests returned after campaigns, 31 internal tasks closed, 4 operational deviations resolved automatically. The temptation is immediate — add them and report 85 things handled. Do not.

linewhat it countsits natural base
42 inquiries handled without the teamexternal events with guestsinquiries received that day
8 guests returned after campaignsan outcome of a campaign run earlierguests contacted in that campaign
31 internal tasks closedinternal work items of unequal sizetasks opened
4 deviations resolved automaticallyexceptions caught by a ruledeviations detected

An inquiry is somebody outside the business asking for something. A task is a unit of work whose size the venue itself defines. A deviation is an exception a rule caught. A returned guest is not a closed situation at all — it is a result, and it belongs to a campaign that ran days earlier. Add them and you get a figure with no base, and a figure with no base cannot be compared with next week's.

Worse, it can be gamed without anybody intending to. If the total is what gets reported, the cheapest way to raise it is to cut tasks into smaller tasks. Nothing improves and the number climbs. That is what happens to every headline metric that sums unlike things.

Each of the four becomes a measurement when divided by its own base:

share handled alone = inquiries closed without a person ÷ inquiries received automatic resolution = deviations resolved automatically ÷ deviations detected same-day closure = tasks closed today ÷ tasks opened today

The report gives numerators; the denominators turn them into something you can watch move. And none of the three says a case was closed correctly — that is a separate check, done by sampling a handful of closed cases a week and reading them.

The money side needs the same discipline: additional revenue is not additional profit, contribution is what has to be counted, and revenue, profit and cost-equivalent lines never go into one total. The queue that turns scattered inquiries into countable items is described in handling inquiries as one queue, and the closing side of it is task management. What a returning guest is worth, and why eight of them needs its own base, is worked out in winning back lapsed guests and in CRM.

The three levels the system sorts its own work into

Every situation is classified before anything is done with it. There are three classes, and they are not about technology — they are about who carries the consequence.

LevelWhat happensWho decided
The system does it alonereplies, follow-up, standard calls, CRM, reminders, reports, checks, part of marketing, data collection, standard processesa person, in advance, in writing
The system plus a personthe system organises the process; the physical or accountable action is performed by an employeea person, in the moment
The owner decidesthe analysis is finished and several prepared options with their consequences are broughtthe owner, from prepared options

Read the right-hand column again: it is the same word in all three rows. The levels do not distribute responsibility, they distribute effort. What changes is how many times a day somebody has to stop and decide, not who answers for the result.

The system is not pretending that artificial intelligence should make every decision. The opposite — its job is to reduce the number of decisions a human has to make at all. A machine that asks about everything and a machine that asks about nothing are the same failure seen from two sides.

No external authority hands you this scale, either. The European regulation on artificial intelligence describes such systems as designed to operate with "varying levels of autonomy", but that phrase is a description, not a ladder: there is no numbered scale of levels in the text, and nothing about which of your situations belongs on which rung. The boundary is drawn inside the business, by the person who will carry the consequence of drawing it wrongly. What a system of this kind is and is not, at the scale of a whole venue, is set out in an AI restaurant management system; the narrowest and best-tested instance of level one is the phone, covered in what the phone system answers by itself.

Five signs that decide which level a case belongs to

The classification is not a matter of taste. Five questions settle it.

One. Is there a written rule, and does it cover this case completely? A guest asks for a bespoke banquet menu for 24 people at 180 PLN a head. Existing rules do not stretch that far. The correct behaviour is not to invent a price — it is to gather the requirements, draft the proposal, and hand a finished task to the kitchen manager for the one thing missing. Note what was handed over: not a thread of seventeen messages, but a prepared task with a single open question.

Two. Is the step reversible, and at what price? Restoring an advertising budget split to what it was three days ago costs only the time it was wrong. A price quoted to a guest cannot be unquoted. Reversibility is not a yes or no; it is a price, and the price decides.

Three. Does it need a hand or an eye in the room? When a supplier says one item will arrive two hours late, the stock figure in the system is a record; the physical shelf is a fact. So one line goes to the kitchen: confirm the physical stock of that item. Typical consumption, tonight's bookings, whether the remainder lasts — all of that is arithmetic that needs nobody.

Four. Does it commit money or a promise outwards? A delayed main course, a confirmed breach of the service standard, a dessert at the venue's expense. That such a compensation exists at all is written in advance by the owner; applying it to this table tonight is confirmed by the manager on the floor. Two decisions, two levels, one situation.

Five. Can the result be checked afterwards, and by what? A new closing checklist is not adopted, it is tested on the next five shifts and compared against the twenty-four minutes it was meant to remove. A change that cannot be scored afterwards is not a decision, it is a preference.

How to combine them: the level of a case is the highest level any single sign demands. One unwritten rule lifts a case out of level one however routine the other four look.

And the level belongs to the case, not to the channel. The same telephone line is level one when a returning guest wants a table for two at seven and the hall is open, and level three when the same call turns into a twenty-four-person banquet with a budget. Anyone assigning levels to modules — "the phone is automatic, the kitchen is manual" — has assigned them to the wrong noun. The logic that carries these rules per case is the decision engine; the difference between a system that answers and one that acts is unpacked in an AI agent versus a chatbot.

The middle level is the hard one: the system organises, a person acts

Level one and level three are easy to describe. The middle level is where most of a working day lives, and it is the one usually built badly.

Here it is working. Forty-five minutes before opening the team gets a short message: confirm the summer area is ready. An employee replies that it is. The confirmation is recorded. A second check does not come back, so it is asked again. If everything is in order, the owner learns none of this.

Three properties make that exchange work, and all three are easy to lose:

  • The request is narrow. One fact, answerable in a word. "Confirm the summer area is ready" gets an answer; "check the terrace" starts a conversation.
  • It goes to somebody who can see the thing. A request for a physical fact sent to a person who is not in the room turns into a forwarded message — the same information logistics the levels were supposed to remove.
  • The answer is recorded and closes the item. An answer that lands in a chat and nowhere else has to be asked for again tomorrow, and the third time nobody replies.

Here is the same level built badly: a stream of "check this, check that" going to the manager all day. It looks like organisation and is the opposite — it does not reduce the number of decisions, it multiplies them and moves them onto one person's phone. A venue can be turned into a permanent inspection regime by exactly this mistake, and the people inside it will describe the system as a nuisance, correctly.

The rule is easy to state and easy to violate: at the middle level the system must arrive with the work already done and one specific gap left open. Arriving with the gap and none of the work is not organisation, it is delegating its own job upward. The front-desk end of this level is the AI reception.

A boundary drawn wrong costs in both directions

Both errors are cheap to make and expensive in different currencies.

Too low: the system acts where the rule was incomplete

A case is marked as needing nobody and the rule covering it turns out to have a hole. The system answers something it should have escalated, promises a dish the kitchen cannot do tonight, applies a discount nobody authorised, or gives a confident answer about an allergen.

The damage is not the mistake itself — venues make mistakes with people too. The damage is when it surfaces. Nobody was watching, because the case was classified as needing nobody, so the error is found by the guest rather than by the business. The defence that works is built into behaviour, not intent: when the rules are not sufficient, the correct output is a prepared task, never an invented answer. A system that fabricates when it lacks a rule cannot hold level one for anything, however narrow.

There is a subtler form. A case reversible in general is not reversible for this guest: a standard reply is fine for the twentieth booking of the evening and wrong for the guest whose last visit ended in a complaint. History is part of the rule, and a rule written without it is incomplete in a way that looks complete.

Too high: a confirmation that is always approved is not a control

The opposite error produces no incidents at all, which is why it survives for years. Everything is escalated; the owner confirms things a written rule could have settled; messages arrive all day. It feels like control, and it is work.

It also degrades: a person asked to approve forty routine things a week stops reading them by the second week. The approval is still clicked, the audit trail still shows a human decision, and the human decided nothing. Level three has become level one with a delay and a signature on it — the worst of both, because the responsibility is now documented where the attention was not.

This one is measurable, and the measurement is the single most useful instrument on this page:

refusal rate = confirmations refused or changed ÷ confirmations requested

Run it per item over a month or a quarter. An item at zero across a long enough stretch is not a decision — it is a notification in a decision's clothing, and it belongs one level down with its rule written out. An item refused or amended often is the reverse: either it belongs a level up, or, more usually, the rule underneath it is wrong and the refusals are the venue correcting it by hand every time.

Both boundary errors show up in that one number from opposite ends, which is why it is worth keeping even in a paper register. Cases where automating at all is the wrong answer, rather than a matter of level, are listed in five situations where we advise against it.

Responsibility does not move between levels, the number of decisions does

On level one a person decided in advance and in writing, when the rule was set. On level two, in the moment. On level three, from prepared options. In all three cases a human answers for the outcome, and in all three the outcome traces back to a named human choice. Nothing was transferred except effort.

So "the system decided it" is never an explanation of a result. Somebody wrote that rule. Somebody assigned that level. Somebody chose not to review it after it went stale. If a register cannot name those three people for a given situation, the register is not finished.

What makes this checkable rather than pious is memory. A system worth the name records, for every situation of consequence: what problem was found, which option was chosen, why, who executed it, what result was expected, what actually happened, and whether the approach should be used again. A business without that record re-decides the same question every few months, from scratch, each time believing it is the first.

Testing whether a decision worked, rather than assuming it did because the numbers moved, needs a comparison group — the method is in did the promotion work. Sorting guests into groups that deserve different treatment, itself a level-one decision about who is contacted and who is better left alone, is guest segmentation.

What this changes for the manager, and what for the owner

The manager does not disappear. Before, roughly half the day goes into information logistics: did you do it, did they reply, what happened with the delivery, why is there no report, call that guest, look at that review, how much did we do yesterday, tell the kitchen, remind the waiter. None of it requires judgement; all of it requires a person, because there is no other way for information to move.

After the boundary is drawn, that half of the day is not eliminated — it moves to levels one and two, where it belongs. The manager gets it back and spends it on what a person is genuinely needed for: the team, service quality, atmosphere, conflicts, physical operations, non-standard situations, and rolling out improvements rather than announcing them. A strong manager is not removed by this. One strong manager becomes considerably more powerful, because the ceiling on how much a good manager can hold was never judgement — it was message volume.

For the owner the change is psychological before it is operational. The state before is a belief: if I stop watching, this starts falling apart. That belief is usually correct, which is why it is hard to argue with. The state after a good implementation is a different one: if something genuinely requires me, I will be told. What follows is ordinary and large — not checking the phone every ten minutes, not needing to be physically present, not sitting in every planning meeting, holding more than one venue, spending attention on growth rather than on dispatching a business that already exists.

Two honest limits. None of that is a promise of a result; it is what changes when the boundary is drawn, written down and honoured — draw no boundary and the same system produces more messages than the chat thread it replaced. And the second belief is earned only once the first has been tested against a stretch that includes something going wrong, because the belief is not about good days. Which routines can move and which cannot is set out in process automation: what a system can handle.

The autonomy register: one page you can keep on paper

None of this needs software to start. The boundary is a table, and its first version fits on one sheet:

SituationLevelThe rule that decidesWho confirmsHow we check afterwardsReviewed
Booking within standard rules1free table, party under 8, no special requestsnobodyweekly sample of 10 bookings01.08
Booking outside standard rules2above 8 people or a custom menukitchen managerevery case read within 24 h01.08
Late supplier, stock sufficient1remaining stock covers bookings plus buffernobodystop-list incidents per month01.08
Compensation to a delayed guest2breach confirmed, item from the approved listfloor managerrefusal rate per month01.08
Change to a standard process3any change affecting more than one shiftownertested over five shifts01.08

Four passes fill it in, and only the first takes real effort:

  1. Collect a week of your own inbox. Every question that reached you personally — messages, calls, someone catching you in the doorway. Write them down as they arrive, not from memory afterwards: memory keeps the dramatic ones and drops the repetitive ones, and the repetitive ones are what the register is for.
  2. Assign a level to each with the five signs. Highest sign wins. Expect the first pass to put too much on level three; the fourth pass fixes that.
  3. Write the level-one rules out in full. The test is that a competent stranger could apply the rule without asking you anything. If you cannot write it that way, the case is not level one yet, and the honest move is to leave it at level two until the rule is finished.
  4. Review monthly against the refusal rate. Zero refusals for a quarter drops the item a level. Frequent refusals mean the rule is wrong, not that the person is difficult.

When this gives nothing, and it is worth knowing in advance:

  • The rules live only in somebody's head. Then nothing can be level one, everything lands on levels two and three, and the register accurately describes the current load instead of reducing it.
  • Nobody reviews it. A rule right in March is wrong in August, and a system applies a stale rule with exactly the confidence of a fresh one. An unreviewed register is more dangerous than none, because it looks like a control.
  • Confirmations are clicked without being read. Then levels two and three are decorative, and the refusal rate will say so within a month.
  • The report's denominators are unknown. If the owner cannot trust what a line was divided by, he re-checks everything by hand and the number of decisions does not fall however well the levels are drawn. That is why the reverse checks at the top of this page come before the levels.
  • The venue's situations are mostly one-offs. A place whose days genuinely do not repeat has nothing stable to write a rule about. The register still clarifies who decides what, but no work is removed.

With one venue the register saves a manager's afternoon. With several, it is the only thing separating an owner who reads everything from an owner who reads what needs a decision — a different subject, taken up in several venues and three owner decisions.

Frequently asked questions

Does an average check of 170 PLN mean one guest spends 170 PLN?

In this report, yes — the divisor was guests, not bills. That is what the reverse check shows: 21 480 ÷ 126 = 170.48, which rounds to the figure printed. Had the line meant an average bill, it would sit between 170.48 and 290.27 PLN, the second being revenue divided by the 74 bookings that evening. Write the divisor into the line and the ambiguity is gone permanently.

How do I recover a forecast when the report gives only a percentage?

Divide the actual by one plus the deviation: 21 480 ÷ 1.057 is about 20 320 PLN. Verify the method on a pair where both numbers are printed — 18 740 against a forecast of 19 600 gives minus 4.4 per cent and a gap of 860 PLN, which is what that report said. A rounded percentage recovers a band, not a point: anything from about 20 312 to 20 331 PLN rounds to plus 5.7 per cent.

Can I add inquiries, tasks and deviations into one total?

No. Four populations, four units, four different bases, and their sum cannot be compared with anything, including itself next week. It also rewards splitting tasks into smaller tasks. Divide each count by its own base — inquiries received, deviations detected, tasks opened — and read three ratios instead of one meaningless total.

Which decisions should never sit on the first level?

Anything committing money or a promise outwards without a rule already covering it; anything irreversible at a price the venue would notice; anything whose necessary fact is physical rather than recorded. Those are signs two, three and four. The shortcut: if you could not write the rule so a competent stranger could apply it, the case is not level one yet.

How do I know the boundary has been drawn too high?

Count refusals. Take every item where a person confirms something and divide the confirmations refused or changed by the confirmations requested. An item at zero for a quarter is not a decision but a notification in a decision's clothing, and it should drop a level with its rule written down. The symptom you notice first is people approving without opening.

Does this replace the manager?

No, and a venue treating it as a headcount exercise gets the worst of it. What is removed is information logistics — chasing, relaying, reminding, asking how much we did yesterday. What is left is the team, service quality, atmosphere, conflicts, physical operations and non-standard situations, which is what a manager was hired for. One strong manager becomes able to hold considerably more.

Who is responsible when the system acted alone and got it wrong?

The same people as otherwise: whoever wrote the rule, whoever approved that level for that case, and whoever chose not to review it when the venue changed. Responsibility does not move between levels, only the number of decisions does. If your register cannot name those three people for a situation, that is the gap to close before any further automation.

Take one week of the questions that reached you personally, sort them with the five signs, and count how many sat on level three only because the rule for them was never written down. That count is the size of the problem, and it is usually the largest of the three. The rest of the restaurant management series is collected in the restaurants hub.

Related services

In this section

Let us look at your numbers

Tell us how enquiries are handled today — how many there are, who picks them up, where they get lost. Aura walks the process with you and shows what can be taken off a person, and what is better left alone.

Talk to Aura

The home page with Aura opens. Give a company name — she looks at it in public data and shows what a client sees. No promises of a result.

Prefer to write? marketing@auraglobal-merchants.com

Next step

Let us check whether Aura fits your place

We do not take everyone: first we look at your processes, sales and current systems and tell you honestly whether it makes sense for us to come in. A few questions, about five minutes.

Take the assessment →