All articles

The Cheapest Supplier Finished Last: A Direct-Rating AHP Example

A purchasing manager leaning back at her desk with a pen at her lip, six supplier bid folders fanned out in front of her and a bar chart on her laptop
A purchasing manager leaning back at her desk with a pen at her lip, six supplier bid folders fanned out in front of her and a bar chart on her laptop

On Tuesday morning, six supplier bid folders were waiting on the desk of a purchasing manager I will call Dana. Dana is fictional. The arithmetic is not.

A yellow sticky note from the finance director sat on top: cheapest wins unless you can show me why not. Friday.

Dana did not need the quotes to tell her which folder was cheapest. Supplier E had made sure of that. What bothered her was the feeling that the cheapest folder was also the riskiest one, and she knew how much evidence a feeling carried in a budget meeting: none.

The lesson in three lines

Supplier E submitted the cheapest bid and finished last of six. Price carried 0.257 of the decision; quality and financial stability together carried 0.429.

Direct rating reached that answer in 35 inputs instead of 85 pairwise judgments, which is why it suits a long list. It carries no pairwise consistency check, so every 0 to 100 scale needs anchors before you trust it.

The ranking got Dana into the meeting. The breaking point let her leave it: A stays first until its quality rating falls to about 79.6, and below that D takes the lead, not E.

What Dana was really choosing

She opened a clean page and wrote Supplier A through F down the left. Across the top she wrote the five things that would decide whether cheapest really meant best:

  • Unit price, because someone has to sign the invoice.
  • Quality, measured the way her plant already measured it: incoming inspection pass rate.
  • Lead time, from purchase order to dock.
  • Financial stability, because a supplier that goes under mid-contract costs more than any unit price saving ever recovered.
  • Support responsiveness, the one everybody forgets until the line stops at two in the morning.

The problem was visible before she entered a number. Dollars, pass rates, days, financial risk, and response time did not share a unit, and they did not deserve equal influence. Dana could not simply total the raw columns. She first had to decide what mattered, then translate unlike evidence into comparable value scores.

Five criteria on a notepad. Not one of them is measured in the same units as another.
Five criteria on a notepad. Not one of them is measured in the same units as another.

Tuesday afternoon: eighty-five judgments

By Tuesday afternoon, Dana had rebuilt the decision as a hierarchy and reopened a method she remembered from an MBA course: compare two things at a time. Quality or price, and by how much? Supplier A or Supplier B on lead time, and by how much?

Then she counted. A set of n items requires n(n-1)/2 pairwise judgments. Six suppliers meant 15 comparisons under each criterion. Across five criteria, that became 75. The five criteria added 10 more. Total: 85.

At comparison thirty, the supplier names stopped feeling like companies and started looking like buttons. Dana caught herself choosing the answer that moved the wizard forward.

She took her hand off the mouse. A rushed matrix would not flash red. It would return polished weights, a tidy ranking, and the same air of authority as careful judgment. The danger was not that the model would crash. It was that it would succeed.

Comparison thirty of eighty-five. She has stopped reading the names.
Comparison thirty of eighty-five. She has stopped reading the names.

Wednesday morning: stop comparing, start scoring

Wednesday morning, Dana reopened the model and noticed a mode she had skipped: Direct Rating. Instead of asking A versus B, then A versus C, it asked for one 0-to-100 judgment at a time.

SpiceLogic calls this workflow Direct Rating AHP. It is not the standard pairwise-comparison form of AHP. The direct scoring idea is closer to the SMART family of methods; the software then normalizes the ratings and combines them with criterion weights. Dana was using it as a fast screening model, not pretending it offered the same consistency check as pairwise AHP. The first calculation turned each criterion rating into a weight:

wi = ratingi / Σ ratingj

Next, the software normalized each supplier rating within a criterion by the total of all six supplier ratings in that column:

rij = ratingij / Σk ratingkj

Finally, it multiplied each normalized supplier priority by the corresponding criterion weight and added the results: Σi wi × rij.

Five criterion ratings plus six suppliers scored on five criteria meant 35 numbers. By lunch, Dana had reached the results screen. The day before, she had not even finished the pairwise questions. The speed was real. So was the responsibility to make every rating defensible.

The mode she had skipped past: rate each item from 0 to 100 instead of comparing pairs.
The mode she had skipped past: rate each item from 0 to 100 instead of comparing pairs.

Step 1: turn the argument into weights

Dana did not assign those five importance numbers alone. She called the plant manager in, closed the door, and asked the question the spreadsheet could not answer: if price and quality conflict, which one gets to pull harder? Twenty minutes later, Quality had 100, Price 90, Lead Time 70, Financial Stability 50, and Support Responsiveness 40. Quality won by one argument: a bad part costs more to discover than a cheap part saves.

CriterionImportance (0 to 100)Weight
Quality1000.286
Unit price900.257
Lead time700.200
Financial stability500.143
Support responsiveness400.114
Total3501.000

The ratings totalled 350, so dividing each one by 350 produced the weights. Quality carried 0.286; price carried 0.257. Put differently, quality had about 11 percent more weight than price. The disagreement had not disappeared. It had become visible, reviewable, and changeable.

Twenty minutes with the plant manager over whether quality or price gets the 100. Quality won.
Twenty minutes with the plant manager over whether quality or price gets the 100. Quality won.

Step 2: the cheapest bid nearly got punished for being cheap

Before scoring a supplier, Dana wrote one rule above the grid: higher must always mean better. For each criterion, she also wrote anchors for what 0 and 100 meant. Otherwise a 70 would be a mood, not a measurement.

Price caught her anyway. On the first pass, she gave Supplier E a 40 because E's quoted price was low. Then she stopped. The quote was low, but the value of that price was high. If higher always means better, E's price score belonged near the top. She changed it to 90. Had she not caught the mistake, the model would have punished the cheapest supplier for being cheap and done so without complaint.

SupplierPriceQualityLead timeStabilitySupport
A6095709080
B8570605575
C4090858060
D7565906085
E9050554565
F5580657050
Column total405450425400415

The software divided each column by its own total. Supplier A's price priority became 60/405 = 0.148; its quality priority became 95/450 = 0.211. Each normalized column now summed to 1. That made the criteria combinable, but it also made these priorities relative to the six suppliers currently in the model. Change the shortlist, and Dana would need to rerun and review the result.

A low price is a good price, so its rating has to be high. She caught it. The model would not have.
A low price is a good price, so its rating has to be high. She caught it. The model would not have.

Step 3: let every criterion cast its vote

Now the model became almost disappointingly quiet. Each normalized supplier score was multiplied by the corresponding criterion weight, then the five contributions were added. Dana opened Supplier A's row because she wanted to see the winner assembled, not merely announced:

  • Price: 0.257 × 0.148 = 0.038
  • Quality: 0.286 × 0.211 = 0.060
  • Lead time: 0.200 × 0.165 = 0.033
  • Financial stability: 0.143 × 0.225 = 0.032
  • Support: 0.114 × 0.193 = 0.022

Supplier A totalled 0.186 after rounding. Repeating the same five-line calculation for the other suppliers produced the ranking:

RankSupplierScore
1A0.186
2D0.176
3C0.168
4B0.167
5F0.155
6E0.149

The unrounded six scores sum to exactly 1.000. The three-decimal values displayed in the table add to 1.001 because of rounding. That small discrepancy is not a model error; it is the price of showing shorter numbers.

Thursday: the cheapest supplier finished last

Supplier E sat at the bottom of the ranking. Dana checked the formula, then checked it again. E's price score was 90, the highest in the table; A's was only 60. Nothing was broken. Price had simply lost the argument when all five criteria spoke.

For Friday, she built one picture: rank by price on the left, rank by final score on the right, and a line showing where each supplier landed.

Supplier rank by price versus rank by final AHP score Supplier E ranks first on price but last on the final weighted score. Supplier A ranks fourth on price but first overall. Rank by price rating Rank by final AHP score Supplier E · 90 Supplier B · 85 Supplier D · 75 Supplier A · 60 Supplier F · 55 Supplier C · 40 Supplier A · 0.186 Supplier D · 0.176 Supplier C · 0.168 Supplier B · 0.167 Supplier F · 0.155 Supplier E · 0.149 Best price finishes last. Price carries 0.257 of the decision; quality and stability together carry 0.429.

The explanation fit on an index card. Price carried 0.257 of the decision. Quality and financial stability together carried 0.429. E scored only 50 and 45 on those two criteria, while A scored 95 and 90. The cheap bid's advantage was real, but it was too narrow to cover the gap.

Dana could have guessed on Tuesday that A might be the safer choice. What she could not have guessed was the exact sentence that would survive a sceptical meeting. The ranking was useful. The arithmetic behind it made the ranking defensible.

Friday, 9:00 a.m.: the question the ranking could not answer

The director studied the chart, nodded once, and asked the question the ranking could not answer: How far could A's quality score fall before A stopped being first?

Dana had a winner. She did not have its breaking point.

Back at her desk, she opened the sensitivity analysis and moved A's quality rating down from 95. The lead survived until the rating reached about 79.6. Below that point, Supplier D moved into first place. Supplier E did not suddenly become the winner; the sensitivity test revealed that D was the rival close enough to matter.

At 10:40, Dana sent a two-line reply: A remains first until its quality score falls by roughly 15 points. Below that, D becomes first. The ranking got her into the room. The breaking point let her leave it.

While she was there, she noticed another warning. Suppliers C and B scored 0.1676 and 0.1669 before rounding, a gap of less than seven ten-thousandths. Several ordinary five-point rating changes could reverse them. Dana marked them effectively tied and added a note: break the tie with evidence the model does not yet contain, such as near-term capacity, distance to the plant, or an existing working relationship. A model can calculate a difference more precisely than the evidence deserves.

The director points to the ranking and asks how far Supplier A's quality can fall before A stops ranking first.
The director points to the ranking and asks how far Supplier A's quality can fall before A stops ranking first.

What Dana bought with speed, and what she gave up

By Friday afternoon, Dana had a defensible result. She had not turned direct rating into standard pairwise AHP. The shortcut had bought time, and it had cost her something.

Pairwise AHP asks more questions because the redundancy is part of its safety net. If you say A is twice B, B is three times C, and A is roughly equal to C, the comparison matrix exposes the contradiction through its consistency ratio. That check can reveal preferences that sounded reasonable one question at a time but do not hold together as a whole.

Direct rating asks for each number once. There is no triangular cross-check, so a consistency ratio is not applicable. SpiceLogic reports it that way rather than displaying a reassuring but meaningless 0.000. The burden moves back to the scoring anchors and the reviewer. A backwards price score or an arbitrary 70 can still look perfectly tidy.

For many people, comparing two things is easier than defending an absolute score such as 70. Dana accepted that loss of elicitation support because 85 careful judgments were unrealistic by Friday. She had not discovered a method that was universally better than pairwise AHP. She had chosen a faster screen, with fewer internal checks, for a problem that was too wide for a rushed comparison matrix.

The rule Dana uses now

Price carries 0.257 of the decision. Quality and financial stability together carry 0.429.

Before she closed the file, Dana wrote a rule for the next sourcing decision:

  • Up to four or five options, and the decision is expensive or contested: standard pairwise AHP. The comparison count is still tolerable and the consistency ratio is worth having.
  • Six options or more, or a long list to screen: Direct Rating AHP first, then standard pairwise on the two or three survivors.
  • Experienced raters who are short on time: direct rating, because a rushed pairwise matrix is worse than a careful rating.

Screen wide, then compare deep. Use direct ratings to reduce a long field, then spend the pairwise effort where the final decision can still hurt.

A broad direct-rating screen narrows six suppliers to a shortlist for careful pairwise comparison.
A broad direct-rating screen narrows six suppliers to a shortlist for careful pairwise comparison.

When the spreadsheet stops being enough

For one model this small, a spreadsheet is enough. That is not a weakness in the method; it is a useful boundary. Software begins earning its place when the hierarchy grows, several people judge independently, the report must be repeatable, or someone asks for the breaking point while the meeting is still warm.

Our AHP Software keeps direct-rating and pairwise workflows in one model. You can screen a broad list with ratings, compare the finalists pairwise, check consistency where pairwise judgments exist, run sensitivity analysis, and combine stakeholder judgments. The point is not to replace thinking. It is to keep the arithmetic, assumptions, and changes visible.

For a small first experiment, the free online AHP calculator lets you try the pairwise mechanics in a browser before building a full supplier model.

Related product Analytic Hierarchy Process Software