All articles

How to Use AHP in Your Thesis: From Research Question to Defense

A graduate student at a lectern answering an examiner's question during her thesis defence, an AHP hierarchy diagram projected behind her
A graduate student at a lectern answering an examiner's question during her thesis defence, an AHP hierarchy diagram projected behind her

Priya was fourteen slides into her thesis defence when the external examiner leaned forward, tapped his pen twice on the table, and asked the question she had been dreading since March.

Why these five criteria, and not seven?

She had an answer. That is the whole point of this article. But she only had it because of six separate moments over the previous year when she very nearly did not, and every one of those moments is a mistake I see in AHP theses again and again. So let me take you through them in the order she hit them, and then come back to the room.

October: choosing the method before the question

Priya's first draft proposal said, more or less, "I will use AHP to study renewable energy in my region." Her supervisor sent it back with one line: AHP to answer what?

AHP answers one shape of question. Given several alternatives and several criteria that are not equally important and not measured in the same units, which alternative is best, and by how much? Site selection, supplier selection, technology choice, policy prioritisation, risk ranking. If your question is "what causes X" or "is A significantly different from B", AHP is the wrong instrument, and an examiner will say so in the first five minutes.

Priya rewrote the question into that shape: which of four renewable options best fits the district, judged on five criteria by people who know the district. It took her a week. It was the most important week of the year.

One line circled in red: AHP to answer what? Rewriting the question took a week. It was the most important week of the year.
One line circled in red: AHP to answer what? Rewriting the question took a week. It was the most important week of the year.

November: criteria that seemed important

Her first hierarchy had seven criteria. When her supervisor asked where each one came from, the honest answer for three of them was that they had seemed important. That is exactly where a defence goes wrong.

Two rules survived the rebuild. First, every criterion needs a source: prior studies, an industry standard, a regulation, or a documented expert workshop. Second, criteria under the same parent must be genuinely distinct, because you are about to ask people to weigh them against each other. Two of her seven overlapped heavily, which means the underlying concept would have been counted twice and quietly gained weight in the result.

She ended up with five, each with a citation, and a paragraph explaining why two candidates had been merged. That paragraph is the answer she gave the examiner in the room.

A practical limit worth knowing: beyond about seven items under one parent, pairwise judgments stop being reliable, and your expert panel starts to lose patience. If you have more than that, use subcriteria to group them rather than flattening everything to one level.

Seven criteria became five, each with a source, and two merged because they would have counted the same thing twice.
Seven criteria became five, each with a source, and two merged because they would have counted the same thing twice.

January: how many experts is enough?

Priya's plan said she would survey one hundred people. Her supervisor asked why. She said it seemed like a respectable sample. That was the wrong frame, and it trips more AHP students than any other single point.

AHP is expert judgment elicitation. It is not a survey, so there is no power calculation and no minimum n that makes it valid. Published AHP studies commonly work with panels from a handful to a few dozen, chosen for expertise rather than for representativeness.

What she had to be able to defend was the selection rule: what made someone qualified to judge, how she verified it, and what perspectives she deliberately included. She ended up with eight people, an explicit inclusion rule written in the methodology chapter, and citations to the AHP studies in her field for the panel size norm. Eight defensible experts beat a hundred convenience responses every time.

Eight experts with a written inclusion rule beat a hundred convenience responses. This one ran a substation.
Eight experts with a written inclusion rule beat a hundred convenience responses. This one ran a substation.

February: the survey, and a consistency ratio of 0.23

Before anyone compared anything, each expert got the goal, the criteria and the alternatives explained, with which parent each subcriterion belonged to. Judgments collected from someone who did not understand the hierarchy are noise that looks like data.

Each comparison was one focused question: with respect to this criterion, which of these two is preferred, and how strongly? She used the software's questionnaire export so the panel could work on their own time, then imported the responses rather than transcribing by hand. Every transcription step is a place to introduce an error you will not notice.

Seven consistency ratios came back under 0.1. One came back at 0.23.

The mechanics, so you know what that number is. The software derives the priority vector, multiplies the comparison matrix by it, divides through, and averages to get the principal eigenvalue. The consistency index is that eigenvalue minus n, over n minus 1. The consistency ratio divides the index by a random index for the matrix size, which for three items is 0.58. Saaty's guideline is 0.1 or below.

Priya's instinct was to drop the eighth respondent. Her supervisor's instruction was better: go back, show the expert the specific comparisons that conflicted, and let them revise. In most high ratios, one or two judgments cause most of the inconsistency, and that was true here. The expert had inverted one comparison by misreading the scale. The revised ratio was 0.06. Priya wrote one sentence in her methodology chapter: one respondent exceeded 0.1 and was re-interviewed on the two conflicting comparisons. That sentence reads as rigour. Silently dropping a respondent reads as something else the moment anyone checks your response count.

The eighth respondent had inverted one comparison by misreading the scale. One call, and 0.23 became 0.06.
The eighth respondent had inverted one comparison by misreading the scale. One call, and 0.23 became 0.06.

A friend's thesis and the consistency ratio of zero

A friend of Priya's, defending a month earlier, had proudly reported a consistency ratio of 0.000 and been asked by her examiner how that was possible.

It was possible because she had used the transitivity rule. Some tools, ours included, can infer comparisons rather than ask for them, which cuts the count from n(n-1)/2 down to n-1. That is a real time saver. It also forces the consistency ratio to zero by construction, because inconsistent judgments become impossible to enter.

A ratio of zero is not a perfect result to report. It means consistency was structurally enforced rather than measured. If you use transitivity, say so, and say that. Priya's friend passed, but only after ten uncomfortable minutes explaining something she should have written down.

A consistency ratio of 0.000 is not a perfect result. It means consistency was enforced, not measured, and the examiner will ask.
A consistency ratio of 0.000 is not a perfect result. It means consistency was enforced, not measured, and the examiner will ask.

March: one sentence about aggregation

With eight experts, Priya had to combine their judgments, and there are two accepted ways that can give different answers.

Aggregation of Individual Judgments combines the panel's pairwise matrices cell by cell using a weighted geometric mean, then derives one set of priorities from the combined matrix. It treats the group as a single body speaking with one voice. Aggregation of Individual Priorities calculates each person's priorities separately and combines those. It treats members as separate stakeholders whose conclusions are pooled.

Priya's panel included a utility engineer, a district planner and a community representative. They were not one voice. She used individual priorities, reported both for comparison, noted that the ranking was stable across the two, and wrote one sentence saying which she used and why. That sentence took thirty seconds and pre-empted a question.

Individual judgments or individual priorities: two accepted methods, two slightly different rankings, one sentence saying which and why.
Individual judgments or individual priorities: two accepted methods, two slightly different rankings, one sentence saying which and why.

April: the number that ended the questions

If you do one thing beyond the basic calculation, do this one.

A ranking is a point estimate built on judgments that could reasonably have been slightly different. Sensitivity analysis asks how much a weight would have to change before the winner changes. Our software ranks every variable by exactly that, reporting a sensitivity rank of 100 minus the percentage change needed to flip the decision, so the fragile assumptions sort to the top. You can also switch a criterion off entirely and see whether the winner survives.

Priya's result held unless the weight on capital cost rose by more than 34 percent, and no respondent had come within half of that. She put that on slide 19.

The result holds unless the weight on capital cost rises by more than 34 percent. Nobody had come within half of that.
The result holds unless the weight on capital cost rises by more than 34 percent. Nobody had come within half of that.

Back in the room

The external examiner had asked why five criteria and not seven. Priya said that she had started with seven, that two had overlapped enough to double count, that she had merged them with a documented rationale, and that each of the remaining five had a source in the literature listed on slide 6. He nodded and wrote something down.

The internal examiner asked whether the result would change if the weights were a bit different. Priya went to slide 19. Thirty-four percent. That was the last methodology question of the afternoon.

Slide 19. Thirty-four percent. That was the last methodology question of the afternoon.
Slide 19. Thirty-four percent. That was the last methodology question of the afternoon.

What examiners actually ask

  • Why these criteria? Sources, not intuition.
  • Why this many experts, and who were they? Your inclusion rule.
  • What did you do about inconsistent responses? Your re-interview procedure.
  • Would the result change if the weights were slightly different? Sensitivity analysis.
  • Why AHP rather than another multi-criteria method? The shape of your question: several alternatives, unequal criteria, incommensurable units, expert judgment rather than measured data.

Every one of those is answerable before you walk in. None of them is about arithmetic.

The methodology chapter checklist

Work straight down this list: the hierarchy with a source for every criterion; the comparison scale used; how experts were selected and how many; how judgments were collected; each respondent's consistency ratio and what you did when one exceeded 0.1; which aggregation method you used and why; the priority weights with the calculation method named; and the sensitivity analysis.

Name the calculation method explicitly, because there is more than one and they can disagree. If your software offers approximate eigenvector, largest eigenvector, geometric mean and fuzzy variants, saying which you used is part of making your work reproducible.

Our AHP Software covers the whole chain Priya used: questionnaire export and import for panels, both aggregation methods, sensitivity analysis, and printable reports for the appendix. If you only need to check a small model or learn the mechanics first, the free online AHP calculator gives you weights and a consistency ratio in the browser.

Everything an examiner will ask is answerable before you walk in. None of it is about arithmetic.
Everything an examiner will ask is answerable before you walk in. None of it is about arithmetic.
Related product Analytic Hierarchy Process Software
Last modified Jun 13, 2026

© All content, photographs, graphs and images in this article are copyright protected by SpiceLogic Inc. and may not, without prior written authorization, in whole or in part, be copied, altered, reproduced or published without exclusive permission of its owner. Unauthorized use is prohibited.