Our thinking

We don't bring a playbook. We bring a way of seeing.

This page is for the people who want to look under the hood: how we diagnose each layer, the lenses we carry through an engagement, why “gut versus data” is the wrong fight, and where the method comes from. The homepage is the short version.

Four layers, one rule.

Change happens bottom-up; each layer only holds if the one beneath it does. We score every layer from observable behaviour, act on the profile rather than the total, and intervene at the lowest layer that is weak. Select a layer to see how we work it.

ToolsPlatforms, frameworks,ritualsWays of workingHow work moves, learnsand reaches outcomesOperating systemDecision rights, structure,cadence, metricsLeadership and incentivesHow leaders decide, what's safe,what gets rewardedEach layer only holds if the one beneath it does.

Leadership behaviour and what gets rewarded

Do leaders believe evidence should be able to change their decision? Is it safe to be wrong? What gets celebrated: confident calls that worked, or good decisions regardless of outcome?

Diagnose
Sit in three to five leadership meetings. Watch how a surprising result is received, how the last misses were discussed, and whether anyone senior ever says “I don't know” out loud. Review promotion criteria.
Failure mode
Leaders promoted on instinct, watching the boss decide on gut, rationally copy it. Evidence reads as an identity threat, not a method question. Programmes that skip incentives revert within two quarters.
Intervene
Lower the threat and change what gets rewarded. Prediction before data. One senior sponsor who shares their own miss first. A safe sandbox, then reinforcement built into every review and into promotion criteria.
AI angle
Framed as augmentation it lowers threat; framed as automation it raises it. Always leave the leader an adjustable final step.

When something doesn't take, go one layer deeper.

The common mistake is escalating at the layer that failed: more training, more tools, more tests. The cause almost always sits one layer down. Before adding anything, we ask what the layer below is doing.

What you observe
Where it shows
Where the cause usually sits
A new platform sits unused
Tools
Ways of working: no rhythm requires it. Or leadership: nobody believes the numbers.
Agile rituals run, nothing gets faster
Ways of working
Operating system: decision rights are unclear, so every hand-off waits for someone upstairs.
The decision template exists but is filled in after the fact
Operating system
Leadership: people don't feel safe committing to a prediction they might get wrong.
Leaders agree in the room and revert the next week
Leadership
Incentives and sponsorship: what gets rewarded hasn't changed, and the boss still decides the old way. This is where we look before pushing behaviour.

Four lenses we carry through an engagement.

Together they give leaders a face-saving way to be wrong, a way to size rigour to the decision, and a recipe for change that sticks.

One-way doors and two-way doors

Not every decision deserves the same rigour. Classify the door first and most arguments about “how much analysis is enough” disappear. It also gives instinct a legitimate home: fast, reversible calls are exactly where experienced judgement should run.

Two-way doorReversible. Decide fast with about 70% of the information. One owner. A quick bet, or a cheap test if the stakes justify it.
One-way doorHard to reverse. Slow down: written hypothesis, pre-mortem, named decider, all available evidence.

Decision quality is not outcome quality

Judging a decision by how it turned out is called “resulting”. Separating the two turns “gut versus data” into “process versus luck”. Nobody has to admit they were wrong; they only have to agree that good process is worth rewarding.

Good outcome
Bad outcome
Good process
Deserved winReinforce the process, not just the result.
Bad luckProtect this person publicly. Safety is built or lost here.
Bad process
Lucky winThe dangerous cell. It gets copied. Talk privately.
Deserved lossThe only cell for a hard post-mortem.

Gut writes the hypothesis, evidence sizes the bet

We never ask a leader to stop trusting their instinct. We ask for their prediction and confidence before any data is shown, so instinct gets credit when right, they learn cheaply when wrong, and calibration builds over time. Pre-registering a leader's prediction is the single highest-leverage move we know.

The four ingredients of change that sticks

Decades of research converge on the same recipe, and programmes that use all four are roughly eight times more likely to succeed than programmes that use one. A staff function or an outside partner can supply two of them. The other two belong to line leaders and whoever sets pay and promotion, so we secure both before we push behaviour.

Role modellingA senior leader visibly changes how they decide first.
Understanding and convictionThe growth mandate and their own past decisions, not abstract method.
ReinforcementIncentives, promotion criteria, rituals. The most skipped ingredient.
CapabilityHands-on, co-designed. Training alone changes almost nothing.

From gut to evidence: why the resistance is normal.

Most companies we walk into have dashboards everywhere and evidence nowhere near a decision. That isn't a data problem; it's a status problem, and it needs a different fix than more reporting.

What is really going on

  • Leaders were promoted for good judgement under uncertainty. A method that says “your judgement may be wrong” threatens the thing that made them successful.
  • Being publicly proven wrong by data is a loss; being proven right is a small gain. Losses weigh about twice as much, so a rational manager avoids the test.
  • Nobody was promoted for a well-reasoned bet that lost. Managers copy what got the boss promoted. Until the rewarded behaviour changes, copying instinct is the locally rational choice.
  • Even at companies with world-class experimentation, only about one in three well-designed ideas improves the metric it targets. Experienced people are wrong most of the time; the difference is whether they find out cheaply.

What we do about it

  • Frame the ask as “your instinct writes the hypothesis; the evidence sizes the bet”, never as “stop trusting your gut”.
  • Use the owner's growth target as neutral ground: “we cannot afford to find out in 18 months that a bet didn't work” lands better than “let's test whether you're right”.
  • Find the one or two managers most likely to adapt and over-invest there. A peer's visible win moves the rest more than any memo.
  • Sequence: the sponsor shares their own recent miss first, then the champions, only then everyone. If a junior person shares a miss first and gets punished, the programme loses a year.
  • Report learning as a chain: tests run, things learned, decisions changed, money moved. Lead with the last two links, reconcile the money with finance, and keep a holdout on every scaled winner.

Where AI helps, and where it hurts.

AI doesn't settle the status question; it moves it. The framing decides which way.

Helps

  • Talking to data in plain language removes “I have to wait for an analyst” and “I don't understand the regression” at once, but only on a governed metric layer. On raw tables it manufactures confident wrong answers.
  • Agents draft the pre-registration, keep the decision log and play devil's advocate on one-way-door decisions. Good process becomes the default rather than an effort.
  • Throughput rises, with the largest gains for less experienced people. That quietly removes skill gaps and win-rate anxiety as excuses.
  • The efficiency track is the least threatening place to install baselines, holdouts and counterfactuals: exactly the discipline we want everywhere.

Hurts

  • Framed as automation, AI raises threat and damages performance. Framed as augmentation, it lowers it. We never use headcount-reduction language inside a company.
  • Domain experts are the group most likely to abandon an algorithm after one visible error, and your managers see themselves as experts. Always leave them an adjustable final step.
  • Perceived speed-up is systematically larger than real speed-up. Every AI gain gets measured with a control, or it isn't claimed.
  • Process theatre at machine speed: a beautifully drafted hypothesis that nobody reads is still theatre. The human states their view before seeing the draft; irreversible decisions stay human-decided.

Culture change is measurable if you measure behaviour, not sentiment.

Share of significant decisions with a pre-registered hypothesisThe one to watch. When it becomes normal, everything else follows; when it stalls, drop a layer and look at incentives.
Decisions changed by evidenceThe metric that proves data enters the decision. Reported with the manager's name on it.
Calibration by managerPredicted versus actual. Whether instinct is getting sharper and whether people trust the exercise.
Cycle time from idea to decisionWhether testing is cheap enough to be routine.
Claimed versus realised impactHoldout on every scaled winner, reconciled with finance quarterly. One inflated number hands the sceptics their best argument.
Maturity score per companyTwice a year on all four layers. Read the shape, not the total: strong tools on weak leadership is a red flag, not a head start.

Where the method comes from.

Nothing here is invented. The four layers tie three bodies of work together, and every claim above traces to one of them.

Change and leadership

  • Kotter, Leading Change. Eight steps; don't declare victory early.
  • Hiatt, ADKAR. Progress stops at the first missing step, usually desire.
  • Edmondson, The Fearless Organization. Psychological safety.
  • Kegan and Lahey, Immunity to Change. Hidden competing commitments.
  • Basford and Schaninger, McKinsey. The four building blocks of change.
  • Lewin, Unfreeze, change, refreeze.

Decisions and behaviour

  • Duke, Thinking in Bets. Resulting; decision quality versus outcome.
  • Kahneman and Tversky, Prospect Theory. Losses weigh twice as much as gains.
  • Bezos, Shareholder letters. One-way and two-way doors.
  • Rock, SCARF. Status threat and why people defend rather than reason.
  • Dietvorst, Simmons and Massey, Algorithm Aversion. Give people an adjustable final step.
  • Klein, Performing a Project Premortem.

Product and experimentation

  • Kohavi and Thomke, HBR. Only a third of well-designed ideas work.
  • Thomke, Building a Culture of Experimentation. Central infrastructure enables decentralised decisions.
  • Cagan, EMPOWERED. Teams given problems, not roadmaps.
  • Perri, Product Operations. Data, insight loop and cadence that make the model repeatable.
  • Kniberg, Lean from the Trenches. Make the whole flow visible.
  • Brynjolfsson et al., Generative AI at Work, and METR (2025) on perceived versus measured AI productivity.
Talk to us about where your transformation is stuck