I inherited a vendor scorecard once. It had 34 metrics, ran in a Google Sheet, and was updated quarterly by a rotating cast of coordinators. Every vendor received a score. Nobody read them. Behavior never changed.
I rebuilt it with five metrics, one page, monthly delivery, and a quarterly review with the account executive. Inside two quarters, fill rate across the top eight vendors moved from 89 percent to 96 percent. Same vendors. Same categories. Same volume. What changed was the design of the scorecard and the meeting behind it.
Here is how to build one that works the same way.
The scorecard exists for one reason
Not to document performance. Not to have data for the next RFP. Not to satisfy a compliance requirement. The scorecard exists to change vendor behavior. That is the whole reason to spend time on it.
Once you accept that, the design choices get simpler. Every feature has to earn its place by moving a vendor toward better performance. If a metric does not create a lever for behavior change, it does not belong on the page.
This is the filter that kills the 34-metric scorecard. Most of those metrics measure things nobody can act on. They document instead of moving.
The five categories that matter
After running this exercise across three operating groups, the same five categories keep earning their place. Weighted, so the operator's real priorities show up in the math.
Fig. 1 · One page, five lines, monthly.
Fill rate, weighted 30 percent
Lines shipped in full divided by lines ordered, calculated at the line-item level, not the invoice level. A vendor who fills 95 percent of cases but drops the tenth line item every Tuesday is not a 95 percent fill vendor. They are a 90 percent line-fill vendor and your prep cook feels it every morning.
Weight this the heaviest because a fill rate miss propagates fast. Missing product means substitutions, 86ing menu items, or emergency runs to Restaurant Depot. Every one of those has a cost that never gets tagged to the vendor invoice.
On-time delivery, weighted 25 percent
Delivered within the two-hour window you agreed on, or not. Late deliveries stack up in receiving, disrupt line setup, and push prep out of position for lunch. Chronic tardiness is more damaging than a single missed delivery, because it teaches the receiving team to expect the vendor to be late.
Measure this at the delivery level. Score at the vendor-month level. Any single delivery outside the window counts as a miss, regardless of the reason. Vendors will push back on this. Hold the line. The window is the window.
Invoice accuracy, weighted 20 percent
Every invoice error costs your AP team 15 to 30 minutes to research and correct. A vendor with 20 line items per week and a 5 percent error rate is costing you an hour of AP time every week. Multiply by 52 weeks and by however many vendors do the same, and this becomes a real number.
Score at the invoice-line level. Track types of errors: pricing off contract, missing credits, wrong units, duplicate lines. The category matters because the fix is different for each one, and the account rep needs to know which one to escalate internally.
Quality incidents, weighted 15 percent
Any product that arrives out of spec. Wrong grade, wrong cut, damaged case, temperature outside range. Count events, not dollars. A vendor who sends one bad case a week gets a lower score than a vendor who sends one bad pallet a quarter, even if the pallet was more valuable.
The reason: quality incidents are a signal about upstream discipline. Frequent small incidents mean the vendor has a systemic issue. Rare large incidents usually mean a one-off. Both matter, but the frequency signal is where behavior change lives.
Account responsiveness, weighted 10 percent
How fast does the account rep respond to a receiving issue, a credit request, a price question, a substitution ask. Track the median response time from your receiving inbox. Score against a target. Four hours is my usual target for the top eight vendors. Twenty-four hours for the rest.
The weight is smallest because response time itself does not move your P&L. But a slow-responding vendor turns every issue into a two-week resolution cycle, and the compounding cost of that becomes real fast.
One page or nobody reads it
The scorecard fits on one printed page in landscape. Vendor name, month, five lines, weighted score, status flag. Nothing else.
This constraint kills every "let's also track" impulse from procurement, from finance, from the safety team. It forces the discussion up to the categories that decide P&L impact.
The rep opens it in a meeting on their laptop, then prints it and takes it back to their office. It has to survive that trip. A two-tab spreadsheet with 30 hidden columns does not survive that trip. A one-page PDF does.
Monthly delivery, quarterly review
Both cadences are non-negotiable. Here is why each one exists.
Monthly delivery keeps the data fresh and prevents the surprise conversation. When the rep sees three consecutive months of declining fill rate, they know a serious conversation is coming, and they usually try to fix it before it gets there. The monthly send does most of the work of the quarterly meeting.
Quarterly review is the meeting where behavior actually changes. Sixty minutes. Both sides. Account executive from the vendor. Head of procurement or operations from you. Walk through the three months of scores, look at the trend, discuss the top two incidents, and agree on one action item each side will take by the next review.
The monthly send is the data. The quarterly meeting is the leverage. Both together are the mechanic.
Tie the score to a real consequence
Reports do not change behavior. Consequences do. The scorecard needs a defined ladder of outcomes before it goes live, and the vendor needs to see that ladder from day one.
The ladder I use:
- Score above 4.5 for four consecutive quarters: vendor is invited to bid on a new category or a new market.
- Score 3.5 to 4.5: continue at current volume, quarterly review as usual.
- Score 2.5 to 3.5 for two consecutive quarters: vendor loses one case pack or one product line to a secondary supplier. Notified in the quarterly review.
- Score below 2.5 for two consecutive quarters: vendor loses primary status. Their volume moves to a secondary vendor over the following two months.
Print this ladder on the back of the scorecard the first time you send it. The vendor knows the rules. The consequences are predictable. The behavior change follows.
Share the raw incident data
The score alone is not enough. The vendor needs to see what created the score, or they cannot fix it.
The monthly send includes a second page: the underlying incident log for that month. Every fill miss with the line item, every late delivery with the timestamp, every invoice error with the line reference. This transparency turns the scorecard from a verdict into a diagnostic tool.
Some vendors will push back on the level of detail. They will claim the incidents are wrong. They will ask for more evidence. Good. That means they are looking at the data. That is when behavior changes.
Who owns the scorecard
The head of procurement, if you have one. Otherwise the head of operations. Never the individual general managers, because they only see their store's experience with the vendor. A vendor who is a disaster in your Bay Area units might be great in Sacramento, and only the aggregated view will show that.
The score has to be centralized to be useful. The consequences have to be centralized to be credible. And the quarterly meeting has to be attended by someone senior enough that the vendor takes it seriously. Sending a coordinator to the quarterly review is a signal that the scorecard is a formality. Send someone with authority.
Where the scorecard fails
Three failure modes I have seen kill vendor scorecards, in the order I have seen them most:
Metric drift
Someone in finance wants a metric added. Someone in safety wants a metric added. The scorecard grows to twelve categories, then twenty, then thirty. Fill rate gets buried. The vendor stops responding to the top-line score. Refuse metric additions unless something else comes off.
Consequence avoidance
The vendor scores below threshold for two quarters. The operator does not want to have the hard conversation, so the case pack does not move. The vendor sees this. Every other vendor sees this. The scorecard loses all leverage inside a year.
Ownership diffusion
Nobody owns it. The scorecard is "the procurement scorecard" but three people have partial responsibility. The monthly send drifts to bi-monthly. The quarterly meeting gets rescheduled. Vendors stop responding. Assign one owner and defend the calendar.
What good looks like
Six months after launching a scorecard the way I have described, you should see:
- Top-eight vendor fill rate above 95 percent, consistently.
- Invoice accuracy above 96 percent across the top eight, with the primary error type shifting away from pricing disputes to unit corrections.
- Quality incident count down 40 percent versus the prior year.
- At least one vendor consequence exercised. Either a case pack pulled, or a primary status revoked. If nothing changed in six months, the scorecard has no teeth.
- At least one vendor promotion. Either into a new category, or into a new market. If nothing was rewarded, the vendors will read the scorecard as purely punitive.
The scorecard is not the whole procurement job. It is one instrument on the panel. But it is the one that turns supplier performance from a story you tell after the fact into a lever you can pull every month.