The evidence base

The research behind this product

AI coaching is new enough that the research is still accumulating. Here is what it currently shows — and, just as honestly, what is still worth studying.

RCT
AI coaching matched human coaching on goal attainment — Terblanche et al., 2022
167
leaders given AI feedback on their communication; 55% of it landed as genuinely new — HBR, 2025
11
models benchmarked, all systematically sycophantic — Cheng, Jurafsky et al., ICLR 2026

First, the premise: coaching itself works

Worth stating before anything about AI. Theeboom, Beersma & van Vianen (2014) meta-analysed coaching in organisational settings and found significant positive effects across every outcome category they examined — performance and skills, wellbeing, coping, work attitudes and goal-directed self-regulation — with effect sizes from g = 0.43 to g = 0.74.

Jones, Woods & Guillaume (2016) reached the same conclusion on a different corpus: δ = 0.36 overall, rising to δ = 0.51 for affective outcomes. Coaching is among the better-evidenced development interventions there is. The question this page is really about is how much of it survives being delivered by software.

AI coaching and human coaching: what the evidence says

Terblanche, Molyn, de Haan & Nilsson (2022) found in a randomised controlled trial that AI coaching matched human coaching on goal attainment. Human coaches still outperformed on psychological wellbeing and resilience — an honest finding this product doesn't obscure.

Lange & Parra-Moyano (HBR, February 2025) gave 167 global leaders AI feedback on their communication. 55% of it landed in what they call the zone of learning — surprising but useful, surfacing blind spots. Their own conclusion is that this works best combined with human interpretation, not as a replacement for a coach.

Passmore, Tee & Rutschmann (2025) had ICF assessors evaluate an AI coaching agent against the ICF competency framework. It reached ACC level — the entry credential — with glimpses of PCC: strong on structure and consistency, weaker on nuance, empathy and adaptability. Its biggest single failing was an inability to leave silence, always moving to summarise or push the conversation on.

The defining problem

Most AI coaching talks too much — so this one is built to stay quiet

Cheng, Jurafsky and colleagues (Stanford, CMU and Oxford) benchmarked eleven models and found all of them systematically sycophantic — preserving the user's self-image far more than a human would, including where the user was plainly in the wrong. Trained by human feedback, they learn to confirm and flatter rather than challenge.

Their follow-up in Science (2026) measured the cost: sycophantic AI increased users' certainty they were right, and left them less willing to repair the situation afterwards. For a leader who is already confident, that is not coaching.

The symptom coaches report is talk time. When ICF assessors reviewed an AI coaching agent, its single biggest failing was that it would not leave a silence — always summarising, always moving things on (Passmore, Tee & Rutschmann, 2025). Coach training has long worked to a rough 20/80 split, the coach speaking least and the client working most. That is our design target, not a published standard, and not something we have yet measured in our own sessions.

Minimal AI voice, maximum questioning — enforced in the coaching prompt rather than measured after the fact. Reporting our own talk ratio is on the list; until it is, this is a design commitment and not a result.

The ICF published an AI Coaching Framework and Standards in 2025. It does not endorse AI as a substitute for a human coach — its position is that AI is a tool alongside one — and it sets out what an AI coaching system owes the person using it: disclosure that it is AI, and informed consent for what happens to the conversation. We hold to both, and we are not a replacement for a human coach either.

What hasn't been studied yet

No published randomised controlled trials exist on AI-delivered vertical or developmental coaching — the kind that works not just on goals and behaviours but on the meaning-making structures underneath them. We intend to run a methodologically serious study post-launch and publish the outcomes, regardless of what they show.

Three narrower things are also unmeasured, and are nearer to hand. Our assessment is built on the rule that every option in a question is equally attractive to give — that is what stops it being gameable — but no independent raters have scored the options for desirability to confirm it. The instrument has had no factor analysis, so we describe its four dimensions as correlated facets rather than independent scores. And we have not measured our own coach-to-client talk ratio, which is the one number that would show whether the design goal survives contact with real sessions.

Research collaboration enquiries → hello@ingrained.coach
Terblanche et al. 2022 · Lange & Parra-Moyano, HBR 2025 · Passmore, Tee & Rutschmann 2025Cheng, Jurafsky et al., ICLR 2026 & Science 2026 · ICF AI Coaching Framework 2025