The foundations

The frameworks

This product draws on twenty established bodies of research. They aren't a reading list — they're arranged along a single developmental spine, from how a leader makes sense of their role to how they see the whole system.

The through-line

Every body of work. One developmental spine.

Everything here maps onto the five stages of The Leadership Development Scale — the structural backbone of the product, drawn from constructive-developmental theory. Each framework explains a different part of how leaders move along it.

Practitioner
Stage 1
Driver
Stage 2
Cultivator
Stage 3
Architect
Stage 4
Steward
Stage 5
The structural spine

The Leadership Development Scale

Five developmental stages describing how leaders make sense of their role, from Practitioner through Steward. The primary structural framework of this product, developed independently and grounded in the constructive-developmental literature below.

Constructive-Developmental Theory — Robert Kegan (Harvard)

The psychology of adult development underneath the stages: how meaning-making evolves through adulthood, and what each stage transition actually requires.

Competing-Commitments Research — Kegan & Lahey (Harvard)

A diagnostic for why intelligent, motivated people don't change even when they want to — mapping the hidden commitments that protect the status quo.

Structural Scoring — Cook-Greuter, Fischer, Dawson

Your stage is a centre of gravity, not a total. Each answer is read for the complexity of reasoning it implies, following criterion-referenced developmental measurement — so the result reflects how you make sense of a situation, not how many points you accumulated.

What the questions measure — Rotter & Levenson, Miron-Spektor, Edmondson, Krumrei-Mancuso & Rouse

Each dimension is built on an established construct rather than an invented one: locus of control (Rotter; Levenson), integrative complexity, paradox mindset and tolerance of ambiguity (Suedfeld & Tetlock; Miron-Spektor et al.; Budner; McLain), perspective-taking, relational orientation and psychological safety (Davis; Edmondson; Owens & Hekman), and openness to revising one’s viewpoint — the intellectual-humility construct (Krumrei-Mancuso & Rouse; Leary et al.).

How the questions are written — Pulakos & Schmitt, Heifetz, Vroom & Yetton

Every item asks what you actually did the last time something happened, not what you would typically do — past behaviour is harder to answer aspirationally than a hypothetical (Pulakos & Schmitt). And a quarter of the items deliberately score the sophisticated-sounding answer LOW, because the situation called for a decision rather than more exploration. Development means applying complexity where it fits, not performing its vocabulary (Heifetz on adaptive versus technical work; Vroom & Yetton on participation being contingent rather than always right).

Why every option sounds like a good answer — forced-choice measurement

The five options in each question are written so that any of them could be said with pride, and their order is shuffled for every person. Both are deliberate: a forced-choice question resists being gamed only while its options are genuinely equally attractive — otherwise the respondent picks the flattering one and the format buys nothing (Brown & Maydeu-Olivares). We hold this as a rule the items were written to. It has not yet been checked by independent raters scoring each option for desirability, which is the study that would turn it from an intention into a finding.

Checking that answers are answers — Meade & Craig, Johnson

Because option order is shuffled per person, selecting the same position on the screen fifteen times produces fifteen unrelated strategies rather than one repeated view. We detect that pattern with the long-string index, the standard check for non-differentiation. It never changes anyone’s score, and it is not treated as a verdict about the respondent — the likeliest reading is that none of the options fitted, which is a fact about the instrument rather than the person.

How the four dimensions relate to each other

The dimensions share items by design — a single question can speak to more than one of them — so they are correlated rather than independent. They describe four facets of one position, not four separate scores. We do not claim statistical independence, which would require a factor analysis on a normative sample we do not yet have.

The relational & narrative work

Narrative Identity — Dan McAdams (Northwestern)

The role of personal narrative in shaping leadership identity — the tools for examining the formative stories leaders lead from, and re-authoring them.

Emotional Intelligence — Goleman, Salovey & Mayer

Self-awareness, self-regulation, motivation, empathy and social skill as leadership competencies. Applied through the relational orientation dimension.

Psychological Safety — Amy Edmondson (Harvard)

The conditions under which people feel safe to speak up, take risks and contribute fully. Central to Cultivator's relational work.

Cultural Adaptation — GLOBE (House et al.)

Nine societal clusters, following the GLOBE study's groupings, used to calibrate assessment wording, profile narratives and coaching tone — so the product doesn't default to US/UK norms. Wording only: the scoring is identical in every cluster, so the same answer means the same thing wherever you lead.

The depth & systemic work

Systems Thinking — Senge, Meadows

Structural, not linear, approaches to how organisations produce the behaviours they produce. Applied in the later, systemic phase of the coaching arc.

The Disowned Self — narrative identity research

The disowned aspects of the self that shape behaviour from below awareness. Applied to the shadow section of the profile — naming what each level gives and what it costs.

The professional standard

Does coaching work at all — Theeboom et al., Jones, Woods & Guillaume

The premise underneath everything else, and worth stating plainly rather than assuming. Two meta-analyses find coaching has real positive effects: Theeboom, Beersma & van Vianen (2014) report effect sizes from g = 0.43 on coping to g = 0.74 on goal-directed self-regulation, across performance, wellbeing, coping, work attitudes and self-regulation. Jones, Woods & Guillaume (2016) find δ = 0.36 overall and δ = 0.51 on affective outcomes. This is the best-evidenced claim on this page.

Behaviour Change — Gollwitzer, Locke & Latham, Bordin

The mechanics under the journey: implementation intentions (the if-then experiment format), goal-setting theory (specific-and-difficult calibration), and the working alliance — the best-replicated predictor of coaching outcomes.

The gap between conversations — Cepeda et al., Bruijniks et al.

Each conversation ends by designing an experiment to run in your actual working life, and the next one debriefs what happened. That only works if the experiment has had somewhere to happen, so the default gap is a fortnight — the same two weeks the daily prompts are written across. Two bodies of evidence set that length. The spacing effect is among the most replicated findings in learning research: distributed practice beats massed practice even when the total amount of practice is identical, and Cepeda and colleagues (2006) put the optimal gap at roughly 10–20% of the interval over which you want the learning to hold. For change intended to last months, that is weeks rather than days. Separately, coaching-specific research finds meeting every one to two weeks outperforms monthly or less across nearly every measured outcome, and that frequency matters more than the total number of sessions. The evidence is not one-directional, and pretending otherwise would be dishonest. Bruijniks and colleagues (2020) randomised 200 adults to once- or twice-weekly psychotherapy and found the more frequent schedule produced faster improvement and, notably, less drop-out — though the advantage had disappeared by 24 months. Frequent contact keeps people engaged. A gap that is too long loses people who were ready to work. So the gap is a floor, not a waiting room, and it can be shortened by doing the work rather than by waiting. Capturing leadership moments — or completing the experiment — opens the next conversation from day seven. The first three conversations use a shorter ten-day default and a five-day earned unlock, because they excavate the story and name the pattern rather than debriefing an experiment, and because the early weeks are where a journey is most often abandoned. What none of this research tests is a developmental arc compressed into a single week. Nobody has studied it because nobody designs one that way. The honest position is that the fortnight is reasoned from adjacent evidence rather than measured here, and the interval is recorded for every conversation so that it can eventually be answered with our own data instead of by argument.

Team Effectiveness — Hackman, Woolley

The conditions under which distributing authority works: real bounded teams with compelling direction, and the collective-intelligence findings on turn-taking and social perceptiveness.

ICF Core Competencies & AI Coaching Framework 2025

The International Coaching Federation's competency framework, and its 2025 standards for AI coaching systems. The framework is explicit that AI is a tool alongside a human coach rather than a substitute for one, and it sets out what an AI coaching system owes the person using it: disclosure that it is AI, and informed consent for what happens to the conversation. We hold to both. The habit of speaking little and asking much comes from coach training generally, not from an ICF ratio — there isn't one.

For the trials, the sycophancy finding and the professional endorsement behind these, see the Research page →

One developmental spineFor the full evidence base, see the Research page