All frameworks
consultingdiagnosisstructuring

5 Whys / Issue Tree

Ask why five times, each answer the cause of the last, until you hit a system you can actually fix.

On this page

The gist

  • Turn the prompt into one measurable root question, then split it MECE with an equation or process (3-4 branches) before analysing anything.
  • Go breadth first: test each branch with one killer number vs last year, plan and benchmark, and kill dead branches out loud.
  • On the surviving branch, keep asking why (3, 5 or 8 times) until the cause is a process, policy or incentive the client can fix — never a person.
  • Verify the chain backwards with therefore, prove fixing the root moves the original number, then size the fix in rupees.

The framework at a glance

5 Whys Root-Cause Chain
Symptom: EBITDA 9% to 3%
One metric, dated baseline
Issue tree locates branch
Why? Channel cost tripled
Rs 9 to Rs 31 per order
Milk inflation explains little
Why? Aggregator share doubled
20% to 45% of orders
Brand-funded discounts on top
Why? Managers chased volume
Bonus tied to order count
Why? Scheme never updated
Built for 2023 land-grab
No margin report existed
Root: nobody owns incentives
Systemic, inside client control
Fix, then re-check symptom

When to use it

Use it when the prompt is diagnostic - profits are down 20%, churn has doubled, the plant is missing delivery targets, costs are rising faster than sales - because those cases need you to find a cause before you can recommend anything. Use it as your default when the case fits no framework you know, which happens more often than students expect: unusual industries, non-profit or government problems, operational bottlenecks, or open-ended how-would-you-think-about-this prompts. Use it as scaffolding inside other frameworks too - a market entry case is really an issue tree whose Level 1 branches are market, competition, capability and economics. And use it in the middle of a case, not just at the start: when the interviewer hands you a data point that surprises you, an on-the-spot mini-tree plus a couple of whys is exactly how strong candidates convert a surprise into an insight instead of freezing.

What it is

The issue tree and the 5 Whys are two halves of the same skill: taking a vague, messy business problem and turning it into a small number of specific questions you can answer with data. The issue tree is the top-down half. You write the problem as one measurable question at the root, then split it into 3-4 branches that are MECE - Mutually Exclusive (no branch overlaps another) and Collectively Exhaustive (together they cover the whole problem, nothing left outside). Each branch splits again into sub-causes, and you stop at two or three levels, where each leaf is something you can test with a number. The 5 Whys is the bottom-up half. Once data points you at one branch, you keep asking why that thing is happening - three times, five times, eight times, however many it takes - until you reach a cause that is systemic and fixable rather than a symptom.

Keep reading ↓

The two came from very different places. The issue tree is the core of the McKinsey problem-solving method: structure first, hypothesis second, analysis third. The 5 Whys came out of Toyota, developed by Sakichi Toyoda and refined by Taiichi Ohno as part of the Toyota Production System. Ohno's famous example: a machine stopped, so the fuse blew, because the bearing was not lubricated, because the pump was not pumping enough, because the pump shaft was worn, because no strainer was fitted and metal scrap got in. The obvious fix was to replace the fuse. The real fix was to fit a strainer. Every additional why moved from a symptom to something structural.

Put together, they give you a discipline instead of a guess. The tree stops you from tunnelling into the first idea that pops into your head, because it forces you to see all the places the problem could live before you pick one. The 5 Whys stops you from stopping too early, because the first plausible explanation is almost never the one worth fixing. This is why almost every consulting case - profitability, market entry, operations, pricing - runs on issue trees underneath, whatever branded framework sits on top. Learn to build a clean tree and you no longer need to memorise frameworks; you can build the right one live, for a problem nobody has written a template for.

How to apply it, step by step

  1. 1

    Write the root as one measurable question

    Do not start with 'why is the company struggling'. Convert the prompt into a single question with a number, a direction and a timeframe: 'Why did EBITDA fall from 9% to 3% over the last four quarters?' If the client's real objective differs from the symptom they described, clarify it now. A sloppy root question guarantees a sloppy tree, because every branch below inherits its ambiguity.

  2. 2

    Split Level 1 using an equation, not adjectives

    The safest way to be MECE is arithmetic. Profit = Revenue - Cost. Revenue = Price x Volume. Volume = Customers x Frequency. Cost = Fixed + Variable. If no equation exists, use a process split (steps in the value chain, stages of the customer journey) or a segment split (product line, channel, geography, customer type). Keep it to 3-4 branches; a Level 1 with seven boxes is a list, not a structure.

  3. 3

    Go one level deeper on every branch before going deep on any one

    Breadth first, then depth. Under Revenue put price, volume and mix; under Cost put COGS, labour, rent, logistics, marketing. Stop at Level 2 or Level 3 - past three levels the interviewer loses the thread and you lose the clock. Every leaf should be phrased so a single number could confirm or kill it.

  4. 4

    State a hypothesis and pick a branch out loud

    A tree alone is completeness; a hypothesis is speed. Say which branch you think holds the answer and why: 'Revenue is flat but profit collapsed, so I would start on the cost side, specifically variable cost per order.' Interviewers score this heavily - it is the difference between a candidate who can list and a candidate who can lead.

  5. 5

    Test each branch with data and kill branches explicitly

    Ask for the one number that eliminates or confirms a whole branch, not ten numbers that nibble at it. Compare against three things: last year, plan, and the competitor or industry benchmark. When a branch is dead, say so aloud - 'revenue per store is flat year on year, so I am setting the revenue branch aside' - so the interviewer can follow your narrowing.

  6. 6

    Switch to 5 Whys on the surviving branch

    Once one leaf is clearly carrying the problem, stop drawing boxes and start asking why. Each answer must be a cause of the answer above it, not a restatement of it. Test the chain backwards with 'therefore': if reading it upwards does not make logical sense, one of your links is an assumption rather than a fact.

  7. 7

    Stop when the cause is systemic and inside the client's control

    You are done when the answer names a process, a policy, an incentive, a system or a capability gap - something the company can change. If your last why names a person ('the manager was careless') you have found a scapegoat, not a root cause; ask why the system allowed that behaviour. If it names something outside the client's control ('the rupee weakened'), ask why the business was exposed to it.

  8. 8

    Prove the fix removes the symptom, then quantify it

    Read the chain top-down one final time: if we fix the root cause, does the original number move? Then size it - 'this recovers roughly 3.5 points of margin, about 17 crore a year' - and add a check that would catch the problem recurring. A root cause without a quantified fix is an observation, not a recommendation.

Worked example

TapriCo is a chai-and-snacks cafe chain with 220 outlets across Maharashtra and Madhya Pradesh, mostly in Tier-2 cities like Nashik, Nagpur and Indore. Revenue over the last four quarters is flat at about 480 crore, but EBITDA margin has fallen from 9% to 3%. The promoters are convinced milk and packaging inflation is to blame and want to lock a three-year milk contract. They ask you to confirm before they sign.

Root question

Why did TapriCo's EBITDA margin fall from 9% to 3% over four quarters while revenue stayed flat at 480 crore? Because revenue is flat, the fall is arithmetically a cost or mix problem - but say that as a hypothesis to test, do not assume it.

Level 1 split (equation)

EBITDA = Revenue - Cost. Revenue splits into outlets x orders per outlet x average order value. Cost splits into COGS, store operating cost (labour, rent, utilities), and channel or fulfilment cost. Three branches, no overlap, nothing left out.

Level 2 split

Under COGS: milk, tea and sugar, snacks, packaging. Under store operating cost: staff, rent, power. Under channel cost: aggregator commission, brand-funded delivery discounts, and delivery-specific packaging. Every leaf is now a per-order number you can ask for.

Test with data

Revenue per outlet is flat, so the revenue branch is set aside out loud. COGS per order is up 6% - real, but worth only about 0.8 points of margin. Channel cost per order has gone from 9 rupees to 31 rupees, roughly 4.5 points of margin. The promoters' milk theory explains less than a sixth of the gap. Drill into the channel branch.

5 Whys on the surviving branch

Why has channel cost per order tripled? Because orders through Swiggy and Zomato went from 20% to 45% of volume, and each carries commission plus a brand-funded discount. Why did aggregator share rise so fast? Because city managers pushed deep app discounts all year. Why did they push discounts? Because their quarterly bonus is tied to total order volume. Why is the bonus tied to volume? Because the scheme was designed in the 2023 expansion phase when land-grab, not margin, was the goal. Why was it never updated? Because nobody owned the incentive scheme after the CFO transition, and no city-level contribution margin report exists.

Root cause and fix

The root cause is an incentive and reporting gap, not milk prices. Fix: re-base city manager bonuses on contribution margin per order, publish a monthly city-level contribution margin dashboard, and cap brand-funded discounts by channel. Sizing: pulling aggregator share back to about 30% and removing the deepest discounts recovers roughly 3.5 points of margin, around 17 crore a year, versus about 1 point from the milk renegotiation.

Takeaway: The tree showed the promoters were chasing a 1-point problem while a 4.5-point problem sat one branch over; the 5 Whys showed the real lever was a bonus formula, not a supplier contract. Structure finds where the problem lives, asking why finds what to actually change.

More worked examples

Worked example: Boeing 737 MAX — from "pilot error" to the real root cause+

Boeing's 737 MAX was the fastest-selling jet in the company's history, with a backlog reported at roughly 5,000 orders. Then two aircraft of the same brand-new type crashed within five months of each other — Lion Air Flight 610 in October 2018 and Ethiopian Airlines Flight 302 in March 2019 — killing 346 people between them. Regulators grounded the worldwide fleet for about 20 months, and Boeing has publicly disclosed programme charges reported in excess of 20 billion dollars. Early public commentary in both cases pointed at the airlines and the crews. Your task is not to assign blame but to find the cause that, if fixed, makes a repeat impossible.

Fatalities

346 (two accidents)

Fleet grounding

~20 months (Mar 2019–Nov 2020)

Reported programme cost

>$20bn (approx., publicly reported)

AoA sensors MCAS read

1 of 2

Sim-training clause avoided

~$1m per aircraft (reported, approx.)

Write the root as one measurable question

Do not start with "why did the MAX fail" — that inherits every ambiguity below it. Frame it as: why did a derivative aircraft certified as safe suffer two fatal loss-of-control events with the same activation signature within five months of entry into heavy service? Notice what the framing forces: "same signature, two operators, two continents" is already a data point, not a symptom. Also clarify the real objective with the client — the question is not "who was at fault on the flight deck" but "what inside Boeing's control has to change so this cannot recur," which is the difference between an inquiry and a recommendation.

Split Level 1 with a process split, not adjectives

There is no revenue equation here, so use the process split the guide recommends: any accident is caused in the design layer, in the safeguard layer, or in the operating layer. Design covers aerodynamics, control laws and sensor architecture. Safeguards cover the certification route, the hazard assessment, the flight manual and the training requirement — everything that is supposed to catch a design hazard before it flies. Operations covers airline maintenance and crew response. Three branches, mutually exclusive, and collectively they cover every path from drawing board to impact — no fourth box needed, and resisting the temptation to add a "culture" branch is what keeps the tree testable rather than rhetorical.

Test with data and kill a branch out loud

Take the operating branch first because it is the one the public had already convicted. Two different airlines, on two continents, with different maintenance organisations and different crew training regimes, produced the same failure sequence five months apart — a repeated common-mode failure across independent operators cannot have an operator-specific root cause. Maintenance of the angle-of-attack vane and crew response were contributing factors and belong in the recommendation, but they cannot be the root. Say it aloud: "I am setting the operator branch aside as a contributor, not a cause; the common factor sits in design or safeguards." That single sentence is worth more in an interview than ten more boxes.

Switch to 5 Whys on the surviving branch

Why did the nose repeatedly pitch down? Because MCAS commanded nose-down stabiliser trim, repeatedly. Why did MCAS command trim when the aircraft was not stalling? Because it took its input from a single angle-of-attack vane, that vane was giving an erroneous reading, and there was no cross-check against the second vane on the other side of the aircraft. Why was a single-sensor input acceptable for a function with that much control authority? Because the safety assessment had classified MCAS as a limited-authority handling augmentation, and when its authority was later increased for the low-speed case (reported as roughly 0.6 to 2.5 degrees of stabiliser travel) and made able to re-trigger repeatedly, the single-point-of-failure question was not re-opened.

Keep going past the first plausible stop

Most candidates stop at "single sensor, bad design" — that is the fuse, not the strainer. Why did the aircraft need MCAS at all? Because the LEAP-1B engine has a much larger fan and had to be mounted further forward and higher on the wing, which produces a pitch-up tendency at high angle of attack, and a software control law was the cheapest way to make the MAX feel like the previous NG. Why did it have to feel like the NG? To preserve a common type rating so airlines could convert crews with computer-based differences training instead of full-flight simulator sessions — a launch customer contract reportedly carried a penalty of about a million dollars per aircraft if simulator training became mandatory. Why was that constraint so binding? Because after a major US carrier's 2011 A320neo order, Boeing chose to re-engine a 1960s airframe on a compressed schedule rather than build clean-sheet, and schedule plus commonality became the programme's governing metrics.

Verify the chain backwards, and apply the stopping test

Read it upward with "therefore": competitive schedule pressure, therefore a derivative not a clean sheet, therefore a no-simulator commonality constraint, therefore an aerodynamic problem solved in software, therefore MCAS with expanded authority, therefore a single-sensor trigger, therefore uncommanded repeated trim. Every link is a cause of the one above, not a restatement — the chain holds. Now apply the guide's stopping test. "The crews should have run the runaway-stabiliser memory items" names people, so ask why the system allowed it: MCAS was not described in the flight manual and the AoA Disagree alert was tied to an optional feature package, so crews were being asked to diagnose a system they had never been told existed. That is systemic and inside Boeing's control, so it belongs in the chain, not outside it.

Prove the fix removes the symptom, then size it

Three root causes survive, and each is Boeing-controllable: a hazard-assessment process that let a function's authority grow without re-classification, a single-point-of-failure sensor architecture in a repeatedly-triggering function, and a governance model where schedule and training-cost commonality outranked design margin — reinforced by self-certification under delegated authority. The fixes map one-to-one: MCAS now compares both AoA vanes and disables itself on disagreement, activates once per event, is overridable by pilot trim, AoA Disagree became standard, the system is documented and trained, and the delegation process was reformed. Test top-down: with a two-sensor cross-check and a single activation, neither accident sequence can start. Then quantify the trade: the constraint being protected was worth on the order of a few hundred million dollars of simulator training across launch fleets, while the consequence has been reported above 20 billion dollars in charges plus about 20 months of lost production — roughly two orders of magnitude, before counting 346 lives.

Takeaway: The tree killed the "pilot error" branch in one move by noticing that two independent operators produced the same failure, and the 5 Whys walked from a stuck fuse (a bad sensor) to the strainer (a hazard-classification process and a commercial constraint that made a software patch preferable to a design change). The general lesson for a case: when your last why names a person, you have found a scapegoat — keep asking until it names a process, an approval gate or a metric, because that is the only kind of answer you can write a recommendation about.

Worked example (India, case-interview style): Sunvira Labs' tablet plant is scrapping batches+

Sunvira Labs runs an oral solid dosage plant outside Hyderabad, making generic tablets for regulated export markets and the domestic trade. Over the last nine months the batch rejection-and-rework rate has climbed from about 2.1% to about 8.6%, and on-time-in-full delivery to their largest US customer has slipped from 96% to 82%. That customer accounts for roughly 120 crore of annual revenue and its supply agreement has a service-level clause. The plant head believes the problem is API quality from a new supplier they onboarded eight months ago and wants to switch back at a higher price. The CEO wants a second opinion before signing.

Rejection rate

2.1% → 8.6% (9 months)

Scrap cost

~Rs 8.6 cr/yr (approx.)

OTIF

96% → 82%

Customer revenue at risk

~Rs 120 cr/yr

Cost of the fix

~Rs 25 lakh (approx.)

Root question with a number, a direction and a clock

Why has batch rejection risen from 2.1% to 8.6% over nine months at the Hyderabad OSD plant, and is that the same cause as the OTIF fall from 96% to 82%? Size the problem before touching it: the plant runs roughly 1,200 batches a year with an average COGS of about 11 lakh per batch, so the extra 6.5 percentage points of rejection is about 78 batches, or roughly 8.6 crore of product scrapped annually — plus the capacity those batches consumed, which is what is actually breaking OTIF. State the objective back: the CEO's real question is whether to pay a supplier premium, so our test is whether the API branch can carry 8.6 crore of scrap. Anchoring on that number stops you from investigating something that could at best explain a fraction of it.

Level 1 split using the process, not a checklist

Rejections happen somewhere in the manufacturing sequence, so split by process stage: dispensing and inputs, granulation and drying, compression, coating, and packing. That is MECE by construction because a batch passes through each stage exactly once and can only fail in one of them. Layer the classic 4M cut underneath as the Level 2 lens — material, machine, method, man — so each stage can be interrogated the same way. Refuse the plant head's framing at this point: "API quality" is one leaf under material at the inputs stage, not the top of the tree, and treating it as the root before testing it is exactly the tunnelling the framework exists to prevent.

Test each branch with one number and kill branches aloud

Ask for the rejection log cut three ways — by product, by line, and by failure reason — rather than for a general quality report. It comes back extremely concentrated: 71 of the 78 excess rejections are one high-dose extended-release product, all on compression line C-3, and the failure reason on almost all of them is tablet weight variation and content uniformity out of specification. That single cut kills the API hypothesis outright: the same API lots from the same new supplier ran through line C-1 for the domestic pack with rejection at 2.3%, so material cannot be the discriminator. Say it out loud — "the supplier is not the variable, because the variable that changes is the line, not the lot" — and note the compression stage is where the defect appears, which is not necessarily where it is caused.

Switch to 5 Whys on the surviving branch

Why do tablets fail weight and content uniformity on C-3? Because granule flow into the die is inconsistent, and the loss-on-drying moisture of the granules coming to that press ranges from about 1.4% to 3.8% against a specification of 2.0–2.5%. Why is granule moisture that variable? Because the fluid bed dryer feeding C-3, FBD-2, is being run to a fixed 45-minute drying time instead of being stopped on an in-line moisture reading. Why is it being run on time? Because the near-infrared moisture probe on FBD-2 has been throwing an error since a monsoon-season power event about seven months ago, and the SOP permits a time-based fallback, which the operators dutifully switched to. Note the timing match — seven months of a broken probe against nine months of rising rejections, with the ramp starting once the affected product's volume grew.

Do not stop at the broken probe

A broken sensor is a symptom with a work order, not a root cause, so keep going. Why was it not repaired for seven months? Because the probe is classified as a non-critical instrument in the asset register, which puts it on annual-shutdown maintenance rather than the monthly preventive list, and its replacement sensor costs about 4.2 lakh with a roughly ten-week import lead time and needs three quotes plus a capex approval because it crosses the one-lakh threshold. Why did nobody escalate it as a quality risk? Because the rejection dashboard the plant reviews is monthly and reported at plant level, so a problem concentrated in one product on one line stayed buried inside an average until we cut the data. And why was the workaround allowed to run indefinitely? Because the fallback clause in the SOP has no expiry date and no deviation trigger — a temporary method silently became the standard method.

Stop at the systemic cause and check it is inside the client's control

Two root causes survive, and both name a policy rather than a person. First, asset criticality in the register is classified by equipment cost and downtime impact, not by product-quality impact, so an instrument whose failure directly moves a critical quality attribute sits outside preventive maintenance and outside the critical-spares list. Second, an SOP fallback with no time limit and no escalation, combined with a monthly plant-level rejection report, means the organisation had no mechanism to notice a seven-month degradation. Resist the tempting stop at "the maintenance engineer should have raised it" — that is the scapegoat answer; the correct question is why the system let a quality-critical instrument stay dead for seven months without anyone being required to act.

Prove the fix removes the symptom, then quantify it

Fixes map directly onto the two roots: re-classify assets by quality impact using a short FMEA so every instrument tied to a critical quality attribute moves onto monthly PM and a pre-approved critical-spares kanban with a standing capex waiver; put a 72-hour expiry and an automatic deviation trigger on every SOP fallback clause; and move the rejection dashboard to daily, cut by line, product and failure reason. Read it top-down to test: with moisture end-pointed on a working probe, granule LOD returns to the 2.0–2.5% band, tablet weight variation stops, and the C-3 rejections disappear — the original number moves. Sizing: this recovers roughly 7 crore of the 8.6 crore scrap, returns OTIF to around 95%, and protects a 120 crore contract, against a fix cost of roughly 25 lakh for the sensor plus spares. Compare that with the plant head's proposal, which would have paid an API premium of several crore a year to fix a branch the data had already killed.

Takeaway: The tree took nine months of rising rejections and, with one three-way cut of the rejection log, collapsed it to one product on one line with one failure mode — killing the supplier hypothesis the plant was about to spend crores on. The 5 Whys then walked past the obvious fix (repair the probe) to the policies that let a quality-critical instrument stay broken for seven months: an asset-criticality rule that ignores quality impact, and a fallback SOP with no expiry. Replace the probe and the problem returns in a year; change the classification rule and the reporting cut, and it cannot.

Common pitfalls

  • Building branches that overlap. 'Marketing', 'digital' and 'customer acquisition' are not three separate cost buckets - they double count, so any number you allocate to them is meaningless. Use an equation or a process to force clean edges.
  • Diving deep on the first branch before finishing Level 1. If you spend six minutes on pricing and the answer was in supply chain, no amount of good analysis on pricing saves the case.
  • Treating the 5 Whys as exactly five questions. Sometimes three is enough, sometimes you need eight. Counting to five and stopping produces a tidy-sounding root cause that fixes nothing.
  • Ending on a person. 'The store manager forgot to reorder' is a symptom. Ask why the system let one person's memory be the control, and you reach the real cause.
  • Running one causal chain when several exist. Big problems usually have two or three contributing roots; if your chain explains only 30% of the gap you started with, you have lost part of the problem.
  • Drawing a beautiful tree and then abandoning it. If you never come back to say which branches you killed and what is left, the interviewer sees decoration, not reasoning.

Interview tips

  • Ask for 60-90 seconds to structure, and actually use it. Write the root question at the left of the page and lay branches out horizontally so you have room to add Level 2 without cramping.
  • Turn your paper towards the interviewer and walk it top-down, one clean sentence per branch. On a video interview, describe the tree in the same order and number the branches so they can follow without seeing it.
  • Always follow the tree with a hypothesis: 'I would start with branch two, because revenue is flat, which points to cost.' Structure plus a point of view beats structure alone.
  • Prefer equation-based splits whenever the case involves a quantity - profit, revenue, market share, capacity, utilisation. They are MECE by construction, so you never have to defend them.
  • Narrate your eliminations. Saying 'that closes the revenue branch' out loud is worth almost as much as the analysis itself, because it shows you are converging rather than wandering.
  • When a data point surprises you, do not go silent - ask why twice on the spot. A live mini 5 Whys is the single most reliable way to look like a consultant rather than a student.

Test yourself

Best video explainers

Go deeper

Now use it on a real case

Reading a framework isn't the same as applying it under pressure. Practise with an AI interviewer that pushes back.

Practise a case free