
On Tuesday morning, six supplier bid folders were waiting on the desk of a purchasing manager I will call Dana. Dana is fictional. The arithmetic is not.
A yellow sticky note from the finance director sat on top: cheapest wins unless you can show me why not. Friday.
Dana did not need the quotes to tell her which folder was cheapest. Supplier E had made sure of that. What bothered her was the feeling that the cheapest folder was also the riskiest one, and she knew how much evidence a feeling carried in a budget meeting: none.
The lesson in three lines
Supplier E submitted the cheapest bid and finished last of six. Price carried 0.257 of the decision; quality and financial stability together carried 0.429.
Direct rating reached that answer in 35 inputs instead of 85 pairwise judgments, which is why it suits a long list. It carries no pairwise consistency check, so every 0 to 100 scale needs anchors before you trust it.
The ranking got Dana into the meeting. The breaking point let her leave it: A stays first until its quality rating falls to about 79.6, and below that D takes the lead, not E.
What Dana was really choosing
She opened a clean page and wrote Supplier A through F down the left. Across the top she wrote the five things that would decide whether cheapest really meant best:
- Unit price, because someone has to sign the invoice.
- Quality, measured the way her plant already measured it: incoming inspection pass rate.
- Lead time, from purchase order to dock.
- Financial stability, because a supplier that goes under mid-contract costs more than any unit price saving ever recovered.
- Support responsiveness, the one everybody forgets until the line stops at two in the morning.
The problem was visible before she entered a number. Dollars, pass rates, days, financial risk, and response time did not share a unit, and they did not deserve equal influence. Dana could not simply total the raw columns. She first had to decide what mattered, then translate unlike evidence into comparable value scores.
Tuesday afternoon: eighty-five judgments
By Tuesday afternoon, Dana had rebuilt the decision as a hierarchy and reopened a method she remembered from an MBA course: compare two things at a time. Quality or price, and by how much? Supplier A or Supplier B on lead time, and by how much?
Then she counted. A set of n items requires n(n-1)/2 pairwise judgments. Six suppliers meant 15 comparisons under each criterion. Across five criteria, that became 75. The five criteria added 10 more. Total: 85.
At comparison thirty, the supplier names stopped feeling like companies and started looking like buttons. Dana caught herself choosing the answer that moved the wizard forward.
She took her hand off the mouse. A rushed matrix would not flash red. It would return polished weights, a tidy ranking, and the same air of authority as careful judgment. The danger was not that the model would crash. It was that it would succeed.
Wednesday morning: stop comparing, start scoring
Wednesday morning, Dana reopened the model and noticed a mode she had skipped: Direct Rating. Instead of asking A versus B, then A versus C, it asked for one 0-to-100 judgment at a time.
SpiceLogic calls this workflow Direct Rating AHP. It is not the standard pairwise-comparison form of AHP. The direct scoring idea is closer to the SMART family of methods; the software then normalizes the ratings and combines them with criterion weights. Dana was using it as a fast screening model, not pretending it offered the same consistency check as pairwise AHP. The first calculation turned each criterion rating into a weight:
wi = ratingi / Σ ratingj
Next, the software normalized each supplier rating within a criterion by the total of all six supplier ratings in that column:
rij = ratingij / Σk ratingkj
Finally, it multiplied each normalized supplier priority by the corresponding criterion weight and added the results: Σi wi × rij.
Five criterion ratings plus six suppliers scored on five criteria meant 35 numbers. By lunch, Dana had reached the results screen. The day before, she had not even finished the pairwise questions. The speed was real. So was the responsibility to make every rating defensible.
Step 1: turn the argument into weights
Dana did not assign those five importance numbers alone. She called the plant manager in, closed the door, and asked the question the spreadsheet could not answer: if price and quality conflict, which one gets to pull harder? Twenty minutes later, Quality had 100, Price 90, Lead Time 70, Financial Stability 50, and Support Responsiveness 40. Quality won by one argument: a bad part costs more to discover than a cheap part saves.
| Criterion | Importance (0 to 100) | Weight |
|---|---|---|
| Quality | 100 | 0.286 |
| Unit price | 90 | 0.257 |
| Lead time | 70 | 0.200 |
| Financial stability | 50 | 0.143 |
| Support responsiveness | 40 | 0.114 |
| Total | 350 | 1.000 |
The ratings totalled 350, so dividing each one by 350 produced the weights. Quality carried 0.286; price carried 0.257. Put differently, quality had about 11 percent more weight than price. The disagreement had not disappeared. It had become visible, reviewable, and changeable.
Step 2: the cheapest bid nearly got punished for being cheap
Before scoring a supplier, Dana wrote one rule above the grid: higher must always mean better. For each criterion, she also wrote anchors for what 0 and 100 meant. Otherwise a 70 would be a mood, not a measurement.
Price caught her anyway. On the first pass, she gave Supplier E a 40 because E's quoted price was low. Then she stopped. The quote was low, but the value of that price was high. If higher always means better, E's price score belonged near the top. She changed it to 90. Had she not caught the mistake, the model would have punished the cheapest supplier for being cheap and done so without complaint.
| Supplier | Price | Quality | Lead time | Stability | Support |
|---|---|---|---|---|---|
| A | 60 | 95 | 70 | 90 | 80 |
| B | 85 | 70 | 60 | 55 | 75 |
| C | 40 | 90 | 85 | 80 | 60 |
| D | 75 | 65 | 90 | 60 | 85 |
| E | 90 | 50 | 55 | 45 | 65 |
| F | 55 | 80 | 65 | 70 | 50 |
| Column total | 405 | 450 | 425 | 400 | 415 |
The software divided each column by its own total. Supplier A's price priority became 60/405 = 0.148; its quality priority became 95/450 = 0.211. Each normalized column now summed to 1. That made the criteria combinable, but it also made these priorities relative to the six suppliers currently in the model. Change the shortlist, and Dana would need to rerun and review the result.
Step 3: let every criterion cast its vote
Now the model became almost disappointingly quiet. Each normalized supplier score was multiplied by the corresponding criterion weight, then the five contributions were added. Dana opened Supplier A's row because she wanted to see the winner assembled, not merely announced:
- Price: 0.257 × 0.148 = 0.038
- Quality: 0.286 × 0.211 = 0.060
- Lead time: 0.200 × 0.165 = 0.033
- Financial stability: 0.143 × 0.225 = 0.032
- Support: 0.114 × 0.193 = 0.022
Supplier A totalled 0.186 after rounding. Repeating the same five-line calculation for the other suppliers produced the ranking:
| Rank | Supplier | Score |
|---|---|---|
| 1 | A | 0.186 |
| 2 | D | 0.176 |
| 3 | C | 0.168 |
| 4 | B | 0.167 |
| 5 | F | 0.155 |
| 6 | E | 0.149 |
The unrounded six scores sum to exactly 1.000. The three-decimal values displayed in the table add to 1.001 because of rounding. That small discrepancy is not a model error; it is the price of showing shorter numbers.
Thursday: the cheapest supplier finished last
Supplier E sat at the bottom of the ranking. Dana checked the formula, then checked it again. E's price score was 90, the highest in the table; A's was only 60. Nothing was broken. Price had simply lost the argument when all five criteria spoke.
For Friday, she built one picture: rank by price on the left, rank by final score on the right, and a line showing where each supplier landed.
The explanation fit on an index card. Price carried 0.257 of the decision. Quality and financial stability together carried 0.429. E scored only 50 and 45 on those two criteria, while A scored 95 and 90. The cheap bid's advantage was real, but it was too narrow to cover the gap.
Dana could have guessed on Tuesday that A might be the safer choice. What she could not have guessed was the exact sentence that would survive a sceptical meeting. The ranking was useful. The arithmetic behind it made the ranking defensible.
Friday, 9:00 a.m.: the question the ranking could not answer
The director studied the chart, nodded once, and asked the question the ranking could not answer: How far could A's quality score fall before A stopped being first?
Dana had a winner. She did not have its breaking point.
Back at her desk, she opened the sensitivity analysis and moved A's quality rating down from 95. The lead survived until the rating reached about 79.6. Below that point, Supplier D moved into first place. Supplier E did not suddenly become the winner; the sensitivity test revealed that D was the rival close enough to matter.
At 10:40, Dana sent a two-line reply: A remains first until its quality score falls by roughly 15 points. Below that, D becomes first. The ranking got her into the room. The breaking point let her leave it.
While she was there, she noticed another warning. Suppliers C and B scored 0.1676 and 0.1669 before rounding, a gap of less than seven ten-thousandths. Several ordinary five-point rating changes could reverse them. Dana marked them effectively tied and added a note: break the tie with evidence the model does not yet contain, such as near-term capacity, distance to the plant, or an existing working relationship. A model can calculate a difference more precisely than the evidence deserves.

What Dana bought with speed, and what she gave up
By Friday afternoon, Dana had a defensible result. She had not turned direct rating into standard pairwise AHP. The shortcut had bought time, and it had cost her something.
Pairwise AHP asks more questions because the redundancy is part of its safety net. If you say A is twice B, B is three times C, and A is roughly equal to C, the comparison matrix exposes the contradiction through its consistency ratio. That check can reveal preferences that sounded reasonable one question at a time but do not hold together as a whole.
Direct rating asks for each number once. There is no triangular cross-check, so a consistency ratio is not applicable. SpiceLogic reports it that way rather than displaying a reassuring but meaningless 0.000. The burden moves back to the scoring anchors and the reviewer. A backwards price score or an arbitrary 70 can still look perfectly tidy.
For many people, comparing two things is easier than defending an absolute score such as 70. Dana accepted that loss of elicitation support because 85 careful judgments were unrealistic by Friday. She had not discovered a method that was universally better than pairwise AHP. She had chosen a faster screen, with fewer internal checks, for a problem that was too wide for a rushed comparison matrix.
The rule Dana uses now
Price carries 0.257 of the decision. Quality and financial stability together carry 0.429.
Before she closed the file, Dana wrote a rule for the next sourcing decision:
- Up to four or five options, and the decision is expensive or contested: standard pairwise AHP. The comparison count is still tolerable and the consistency ratio is worth having.
- Six options or more, or a long list to screen: Direct Rating AHP first, then standard pairwise on the two or three survivors.
- Experienced raters who are short on time: direct rating, because a rushed pairwise matrix is worse than a careful rating.
Screen wide, then compare deep. Use direct ratings to reduce a long field, then spend the pairwise effort where the final decision can still hurt.
When the spreadsheet stops being enough
For one model this small, a spreadsheet is enough. That is not a weakness in the method; it is a useful boundary. Software begins earning its place when the hierarchy grows, several people judge independently, the report must be repeatable, or someone asks for the breaking point while the meeting is still warm.
Our AHP Software keeps direct-rating and pairwise workflows in one model. You can screen a broad list with ratings, compare the finalists pairwise, check consistency where pairwise judgments exist, run sensitivity analysis, and combine stakeholder judgments. The point is not to replace thinking. It is to keep the arithmetic, assumptions, and changes visible.
For a small first experiment, the free online AHP calculator lets you try the pairwise mechanics in a browser before building a full supplier model.





