Every year a new tool lands in the restaurant tech stack and someone at the top of the operation has to decide the same question. Do we pilot this at one unit for 90 days, or do we push it to every location on the same Monday.

The default answer at most operators I have worked with is "pilot everything." The default answer at most vendors is "you already know it works, go wide." Both are wrong most of the time. There is a real decision tree here, and getting it wrong costs quarters.

I have run this decision on Toast, on 7shifts, on QuickBooks Enterprise, on Salesforce, on Airtable, on a stack of custom GPTs, and on a homegrown labor forecaster. Here is the framework I would apply on the first day at a new operator, and the mistakes I made across 21 franchise units learning it.

The five-question test

Before you decide pilot or rollout, score the tool on five questions. Any single high-risk answer earns a pilot. All five green means you can consider a coordinated rollout.

The five-question pilot vs rollout tree New tool proposal score five questions 1. Guest impact? order, pay, wait, menu 2. Reversibility? out in a weekend? 3. Integration surface? POS, payroll, accounting 4. Cost per unit? install and license 5. Training burden? minutes or full shift? Any high-risk? PILOT one unit, 90 days, written exits All five green? ROLLOUT coordinated wave, staggered training Hit exit criteria 60 straight days then approve rollout wave 1 miss twice? another 30 or stop Wave of 5, then 8, then remainder two weeks between waves rollback plan named before wave 1

Fig. 1 · One high-risk answer earns the pilot. Skipping the tree earns a rollback.

1. Guest-experience impact

Anything a guest can feel earns a pilot. New POS, new payment terminal, new online ordering interface, new menu system, new reservation platform, new kitchen display. If the guest can tell it changed, you pilot. The Resy-to-Toast reservation switch we ran at one unit looked identical on the operator side and was a mess on the guest side because the confirmation emails looked different enough that a subset of regulars thought their reservation had been cancelled.

2. Reversibility

Can you rip this tool out in a weekend if it fails. If the answer is yes, you can rollout faster because the downside is bounded. If the answer is no because of data migration, hardware installation, contract terms, or dependency on other systems, you pilot. QuickBooks Enterprise migration is not reversible in a weekend. A new scheduling tool from 7shifts to a competitor mostly is. Behavior should not be the same across those two.

3. Integration surface

How many other systems does this tool touch. A tool that reads from Toast, writes to QuickBooks Enterprise, syncs with 7shifts, and pushes to Power BI has an integration surface that will surface bugs you cannot predict. Pilot. A standalone tool that runs in its own browser tab has no surface. Rollout.

4. Cost per unit

If the vendor charges $2,000 per unit for install, a 21-unit rollout is $42,000 that is only justified if you already know the tool works. If the vendor charges a flat platform fee and a per-user license, pilot is nearly free. The math changes the risk tolerance.

5. Training burden

How long does a competent line-level user need to become proficient. Under 30 minutes on a phone, rollout. A full shift per person of training, pilot. A three-day certification, do not adopt the tool. This is the question most operators underweight, because the training burden is the piece the vendor is worst at estimating and the operations team pays for.

The vendor's install estimate is a floor. The operator's training estimate is the number that decides whether the rollout survives.

What a real exit criterion looks like

The pilot is only useful if you wrote down what "pilot succeeded" means before it started. This sounds obvious. Almost nobody does it. Most pilots end with a conversation that goes "well the managers seem to like it" and a decision to roll out that is essentially political.

Written exit criteria have three properties. They are numbers. They have thresholds. They have a duration.

Some real examples from pilots I have shipped:

  • 7shifts scheduling pilot. Schedule build time under 90 minutes per unit per week. Labor variance to forecast within ±1.2 points for four consecutive weeks. No-show rate below 4 percent. All three, 60 days.
  • Toast Tasks rollout for shift-close. Completion rate above 95 percent nightly. Manager close time under 20 minutes. At least one flagged safety escalation in the first 30 days (proving the escalation path works). Sixty days.
  • Custom GPT for daily close reports. Report accuracy above 98 percent against the manual version. Time to first draft under 4 minutes. General manager reads the report five days out of seven. Thirty days on this one because the tool is fully reversible.

If the pilot hits every criterion for the duration, it earns rollout. If it misses one criterion twice in a row, it earns another 30 days or a stop. If it misses two criteria, it earns a stop and a rebuild. That decision is numerical, not political, which is the whole point.

The 7shifts story across 21 franchise units

At Hana Group we ran 21 franchise units across six states, all inside Walmart, Sam's Club, Whole Foods, or Target host stores, on a $36M P&L. When we moved from spreadsheet scheduling to 7shifts, I made every mistake in the book on the first pass.

The pilot was three units, 60 days. Labor variance improved. Manager scheduling time dropped by half. No-show rate held steady. All three criteria hit, on schedule. We approved rollout to the remaining 18 units in three waves of six each, two weeks apart. I was proud of the sequencing.

What I did not do was ask why I picked those three pilot units. The honest answer: they were the highest volume units, with the strongest general managers, with the cleanest existing schedules. The pilot passed because those three units did not have the problem the tool needed to solve.

The mid-volume units in wave two revealed the real issue. Their labor variance was worse to begin with, their general managers were less confident schedule builders, and the 7shifts recommendation engine kept over-suggesting labor on their slower days. What worked on autopilot at the top units required 90 minutes a week of manager judgment override at the mid units. We had to build a training tier, a manual-override protocol, and add a fourth exit criterion around forecast accuracy before the tool moved the needle.

Cost of that mistake: about six weeks and one franchisee who publicly complained to the franchise council. The fix was cheap once we saw it. The lesson was more expensive.

The second lesson from that rollout: the pilot exit criteria were the right shape, they were just measured at the wrong sites. If I had picked a median-volume unit and a lower-volume unit alongside one top unit, the pilot would have surfaced the recommendation-engine over-suggestion inside 30 days. Same criteria, different site selection, better outcome. Site selection is the piece the framework does not fully protect you from. A rigorous pilot at a comfortable unit still lies.

The third lesson was a management lesson, not a technology lesson. When the wave-two units struggled, the operations team reflexively blamed the general managers instead of the pilot design. That instinct is dangerous. If a well-designed rollout puts a large fraction of your operators into distress, the tool or the pilot design is usually the story, not the operators. We caught it in about two weeks and pivoted the training tier. Left uncorrected, that instinct burns real trust with the field.

Pilot-to-pilot mistakes

A few patterns I have seen (and made) repeatedly:

Piloting at the best unit

Covered above. The best unit is the wrong pilot site because the tool passes for the wrong reason. Pick a median unit. Not the worst either, because the worst unit has too many confounding failures. Median.

Piloting without a control

If you pilot at unit A for 90 days and unit A's labor cost drops 1.8 points, you do not know whether the tool did it or the season did it. Track the same numbers at two comparable non-pilot units for the same period. If those units also dropped 1.5 points, the tool did nearly nothing.

Piloting three tools at once

New POS, new scheduler, new inventory tool at the same unit at the same time. Attribution becomes impossible. If everything improves, you cannot say which tool did it. If everything regresses, you cannot say which one to rip out. One pilot at a time per unit.

Piloting for 30 days when the tool needs 90

Scheduling tools show their weaknesses over one payroll cycle, which is 30 days. But their real behavior shows over a holiday weekend, a promotional push, and one full month-end close, which adds up to 90 days. Setting the pilot too short is one of the most common ways to ship a bad rollout.

When rollout is actually cheaper than pilot

There are cases where a well-negotiated full rollout beats a pilot. Not many, but real.

  • Franchise agreement mandate. If the franchise agreement requires uniform systems across all units, you will roll out eventually. A pilot just delays the inevitable and adds parallel-cost drag.
  • Vendor discount tied to volume. If the vendor drops the per-unit price 30 percent for a same-week install at all locations, and the tool is proven at your scale, the math favors rollout.
  • Operations bandwidth limit. If your operations team can manage exactly one training push per year and this year's training was already scheduled, splitting into pilot then rollout costs a full year.
  • Proven at scale by peers. When Toast is already the POS at 200 comparable multi-unit groups, the pilot at your operation is not really discovering new information.

In all four cases, the rollback plan replaces the pilot as the risk control. Name the plan before wave one. Test the rollback on the first unit installed. If you cannot cleanly roll back within a week, you should have piloted.

The point

Pilot is a bet you place when you do not yet know what a tool will do in your environment. Rollout is a bet you place when you do. The framework is not "always pilot" and it is not "just ship it." It is a five-question read followed by written exit criteria followed by a wave plan.

Get the pilot site wrong and the pilot lies to you. Skip the exit criteria and the rollout decision goes political. Skip the pilot when you needed one and the rollback plan becomes the story. All three of those failures cost the same thing: quarters of drag on the operation and a franchisee or two who lose trust in the process.

Draw the tree. Score the questions. Write the numbers. Then decide.