The research behind this product
AI coaching is new enough that the research is still accumulating. Here is what it currently shows — and, just as honestly, what is still worth studying.
First, the premise: coaching itself works
Worth stating before anything about AI. Theeboom, Beersma & van Vianen (2014) meta-analysed coaching in organisational settings and found significant positive effects across every outcome category they examined — performance and skills, wellbeing, coping, work attitudes and goal-directed self-regulation — with effect sizes from g = 0.43 to g = 0.74.
Jones, Woods & Guillaume (2016) reached the same conclusion on a different corpus: δ = 0.36 overall, rising to δ = 0.51 for affective outcomes. Coaching is among the better-evidenced development interventions there is. The question this page is really about is how much of it survives being delivered by software.
AI coaching and human coaching: what the evidence says
Terblanche, Molyn, de Haan & Nilsson (2022) found in a randomised controlled trial that AI coaching matched human coaching on goal attainment. Human coaches still outperformed on psychological wellbeing and resilience — an honest finding this product doesn't obscure.
Lange & Parra-Moyano (HBR, February 2025) gave 167 global leaders AI feedback on their communication. 55% of it landed in what they call the zone of learning — surprising but useful, surfacing blind spots. Their own conclusion is that this works best combined with human interpretation, not as a replacement for a coach.
Passmore, Tee & Rutschmann (2025) had ICF assessors evaluate an AI coaching agent against the ICF competency framework. It reached ACC level — the entry credential — with glimpses of PCC: strong on structure and consistency, weaker on nuance, empathy and adaptability. Its biggest single failing was an inability to leave silence, always moving to summarise or push the conversation on.
Most AI coaching talks too much — so this one is built to stay quiet
Cheng, Jurafsky and colleagues (Stanford, CMU and Oxford) benchmarked eleven models and found all of them systematically sycophantic — preserving the user's self-image far more than a human would, including where the user was plainly in the wrong. Trained by human feedback, they learn to confirm and flatter rather than challenge.
Their follow-up in Science (2026) measured the cost: sycophantic AI increased users' certainty they were right, and left them less willing to repair the situation afterwards. For a leader who is already confident, that is not coaching.
The symptom coaches report is talk time. When ICF assessors reviewed an AI coaching agent, its single biggest failing was that it would not leave a silence — always summarising, always moving things on (Passmore, Tee & Rutschmann, 2025). Coach training has long worked to a rough 20/80 split, the coach speaking least and the client working most. That is our design target, not a published standard, and not something we have yet measured in our own sessions.
Minimal AI voice, maximum questioning — enforced in the coaching prompt rather than measured after the fact. Reporting our own talk ratio is on the list; until it is, this is a design commitment and not a result.
The ICF published an AI Coaching Framework and Standards in 2025. It does not endorse AI as a substitute for a human coach — its position is that AI is a tool alongside one — and it sets out what an AI coaching system owes the person using it: disclosure that it is AI, and informed consent for what happens to the conversation. We hold to both, and we are not a replacement for a human coach either.
No published randomised controlled trials exist on AI-delivered vertical or developmental coaching — the kind that works not just on goals and behaviours but on the meaning-making structures underneath them. We intend to run a methodologically serious study post-launch and publish the outcomes, regardless of what they show.
Three narrower things are also unmeasured, and are nearer to hand. Our assessment is built on the rule that every option in a question is equally attractive to give — that is what stops it being gameable — but no independent raters have scored the options for desirability to confirm it. The instrument has had no factor analysis, so we describe its four dimensions as correlated facets rather than independent scores. And we have not measured our own coach-to-client talk ratio, which is the one number that would show whether the design goal survives contact with real sessions.
Research collaboration enquiries → hello@ingrained.coach