If you’re wondering how to get better results from AI, treat its first response as a draft, not the finished answer. Give AI a clear job and useful context, define what good looks like, then review the output, challenge weak assumptions and give specific feedback. Prompting starts the work; coaching the output helps you improve it.
That shift sounds small. In practice, I think it’s one of the biggest differences between using AI and working well with it.
Coaching AI means actively managing and improving AI-assisted work after the initial prompt. You define what good looks like, inspect the response, challenge errors or assumptions, give targeted feedback and keep useful learning for future work. It’s a way of managing the work, not treating AI as a person.
Why the First AI Response Should Be a Draft
There’s a scene in the 1971 film Willy Wonka & the Chocolate Factory that keeps popping into my head when I watch people use AI. An enormous computer is fed a problem, lots of buttons are pressed and everyone stands back waiting for The Answer.
Funny in 1971. Less useful as a 2026 AI strategy.
A lot of early AI advice encouraged that mental model: write the right prompt, press Enter and judge whatever comes back. If it’s good, use it. If it’s rubbish, try a bigger prompt or decide that AI is rubbish.
The tools have moved on. Our habits need to move on too.
Microsoft’s 2026 Work Trend Index surveyed 20,000 workers using AI across 10 countries. Eighty-six per cent said they treat AI output as a starting point rather than a final answer. So the problem isn’t simply that everyone accepts the first thing AI says.
The more useful question is: what happens after the starting point?
Do you regenerate vaguely and hope for the best? Keep adding instructions? Or do you stay involved, diagnose what’s missing, challenge assumptions and deliberately improve the work?
That’s the territory I mean by coaching AI.
Prompting vs Coaching AI: What’s the Difference?
I’m not joining the “prompting is dead” brigade. Our Copilot prompt-engineering guide explains why a good prompt is still a good brief, and that still matters.
Microsoft’s own Copilot guidance recommends giving the tool a goal, context, expectations and sources. It also describes prompting as a conversation. That’s a useful way to think about it.
The problem comes when we put so much energy into crafting the mythical perfect prompt that we forget we’re allowed to have a second conversation.
| Prompting AI | Coaching AI |
| Defines the initial task | Improves the work after the first response |
| Supplies instructions and context | Diagnoses weaknesses and misunderstandings |
| Shapes the first output | Shapes subsequent outputs |
| Happens mainly before generation | Continues throughout the interaction |
| Asks “What should AI do?” | Asks “What needs to improve?” |
Imagine delegating a piece of work to a capable colleague. You’d explain what you need, why it matters, who it’s for and what useful looks like. But you wouldn’t expect them to disappear with a nine-page instruction sheet and return with the single correct answer forever.
You’d look at the work, say what’s strong and point out what’s been misunderstood. You might add context you forgot to mention or change your own mind as you see the work taking shape.
AI isn’t a colleague. It doesn’t carry human accountability and its memory depends on the tool and setup. But many of those work-management disciplines transfer surprisingly well.
The COACH Method for Improving AI Output
We’ve turned that way of working into a simple model we teach in our How to Coach AI workshop, where participants practise all five stages.
C: Clarify the Job
Start with the outcome, not a pile of instructions. What are you trying to achieve? Who’s it for? What matters? What are the constraints?
If you’re not sure what context would help, ask. “Before you start, ask me up to three questions that would materially improve the result” is often more useful than trying to predict every detail yourself.
O: Offer Useful Context
Share the background that changes the quality of the answer: source documents, examples, previous attempts, audience knowledge, decisions already made and constraints the AI can’t infer.
Do this within your organisation’s AI and information-security rules, of course. More context isn’t a licence to paste confidential information into a tool that shouldn’t receive it.
A: Agree What Good Looks Like
This is the step I think we most often skip.
What would make the output genuinely useful? Accurate? Specific? On-brand? Easy for the reader to act on? Appropriate for a senior audience? Short enough to send in Teams? And what would make you reject it?
Ask AI to summarise the quality bar back to you before it drafts. You’ll often spot misunderstandings earlier.
C: Critique and Coach the Work
Don’t say “make it better” if you can’t say what better means.
Try:
- “Keep the opening. The middle becomes generic. Rewrite that section around these two specific points.”
- “What assumptions have you made that I should check?”
- “Which claims here need a source before I use them?”
- “Against the criteria we agreed, where is this weakest?”
- “Give me two alternatives for this paragraph, and explain the trade-off between them.”
Specific feedback gives the next turn somewhere useful to go.
H: Harvest the Learning
At the end, don’t throw the whole conversation away.
Ask what you’ve learnt that would make the next piece of work better. Which preferences, examples, corrections or quality checks are reusable?
This could become a reusable skill or agent. Or it could simply become a short note, a saved example, or a line in a team template. If your tools and policies allow more advanced reuse later, brilliant. But start by capturing the learning somewhere you control, so the next piece of work is easier, sharper and less dependent on remembering what worked last time.
Want a copy to keep beside you while you work? Download the COACH guide and use the five-step framework on your next AI-assisted task.
Why Confident AI Answers Still Need Checking
Modern generative AI can produce something structured, confident and plausible before we’ve finished making a cup of tea. That speed is part of the appeal. It’s also why I like the provocation:
The moment AI impresses you might be the moment you most need to slow down.
Not because impressive answers are necessarily wrong. Because being impressed isn’t a quality check.
I had a useful reminder of this when one of my colleagues asked me who should be the primary human reader for VTT’s case studies. AI already had our audience profiles, yet it confidently assigned the needs and motivations of one profile to the wrong one.
I knew enough to spot that something was off, so I asked it to check against the source material. It corrected itself, but that still wasn’t the end. We worked through who would read the case study first, what they’d need and how to avoid writing for everybody. Only when the thinking was right did I ask AI to consolidate the answer. Then came one final note from me: “Come on, that doesn’t sound like me.”
There was no clever agent and no magical prompt. Just an ordinary conversation in which I stayed responsible for the outcome.
Microsoft Research studied 319 knowledge workers across 936 real examples of using generative AI at work. Higher task-specific confidence in GenAI was associated with less self-reported critical thinking, while confidence in people’s own ability to do and evaluate the task was associated with more.
That’s an association, not proof that confidence in AI causes critical thinking to disappear. But it’s a useful warning. A separate preregistered CHI 2025 study found that explanations from an LLM increased people’s reliance on both correct and incorrect answers. Helping people inspect sources or inconsistencies reduced reliance on incorrect answers.
Polish, confidence and explanation can make an answer easier to trust. None of them automatically makes it true.
Knowing What Good Looks Like Still Matters
AI can now help us do work that used to sit outside our own experience. That’s a fantastic opportunity. But it creates a difficult question: if I’ve never been particularly good at the task myself, how will I know whether the AI has produced something brilliant?
Research on the “jagged technological frontier” makes this concrete. In a large field experiment with 758 Boston Consulting Group consultants, participants using AI performed faster and better on tasks inside the capability frontier being tested. But on one task deliberately designed to sit outside that frontier, the AI groups produced correct solutions about 19 percentage points less often on average than the group without AI.
That doesn’t mean “don’t use AI for difficult work”. It means capability is uneven, and the human still needs to judge the task and the answer.
Sometimes the most intelligent next step isn’t another AI turn. It’s involving someone with expertise, accountability or context you don’t have. If the stakes are high, you can’t evaluate the answer properly, or the decision needs human responsibility, stop coaching the output and bring in the right person.
A VTT Example: We Managed the Work Better
Our coaches value the walkthrough videos we create to help them run our workshops. They bring the notes to life: how to flex an exercise, what to prioritise if time gets tight and how to make an example land. They’re also a lot of work to produce, especially for one-off or lightly used bespoke workshops.
So our design studio tested an AI-supported route. Strong coach notes became a script, and a voice tool narrated it using an AI clone of my voice.
The first version worked technically and was nowhere near good enough. The voice felt generic. The script was formal. It didn’t sound like VTT or like me.
We didn’t conclude that AI couldn’t do it…
We coached the work.
We improved the source notes, used transcripts of my real walkthroughs to identify how I actually speak and captured those patterns in guidance the team could reuse.
We made the use of my AI voice transparent, and we kept the human quality bar: a real human walkthrough may still be worth doing for core work.
For some lower-volume work that previously might not have had a walkthrough at all, we can now create something genuinely useful.
The technology didn’t dramatically improve between attempt one and attempt four. The way we managed the work did.
How to Get Better Results From AI Without Using More AI
Good AI use isn’t about feeding everything to the machine. It’s about making better decisions about where AI helps, what you need to contribute and where human judgement has to lead.
Microsoft’s 2026 Work Trend Index found that its most advanced users were more likely than others to pause before a task and decide what should be done by AI versus the human: 53% compared with 33%. They were also more likely to deliberately do some work without AI to maintain their skills: 43% compared with 30%.
You don’t need an agent to copy that behaviour. On your next ordinary AI-assisted task:
- Clarify the job and why it matters.
- Offer the context that would change the answer.
- Agree what a good result needs to do.
- Critique and coach the first response with specific feedback.
- Harvest the useful learning for next time.
No agent. No enormous prompt. No prompt-engineering doctorate required.
If your organisation wants people to practise those habits together, build shared quality standards and get more from the tools they already have, explore VTT’s How to Coach AI workshop.
Frequently Asked Questions
They overlap, but they’re not the same. Prompt engineering focuses on how you frame instructions and inputs. Coaching AI includes the work after the prompt: reviewing, challenging, refining, checking and reusing what you learn. A strong prompt is a good start, not the whole process.
Start with the outcome and the most relevant context, then ask AI to clarify what it needs. Agree a few quality criteria and review the first response against them. Give specific feedback such as “keep this”, “change that because…” and “check this assumption”. Several short, purposeful turns can be more manageable than trying to predict everything in one enormous prompt.
Bring in a human expert when the stakes are high, you don’t have enough expertise to judge whether the answer is good, the decision carries human or professional accountability, or the AI is repeatedly failing on the same important point. Good AI use includes knowing when not to delegate the next step to AI.
Not necessarily. Feedback can shape the current conversation, but what’s remembered across conversations depends on the product, settings and your organisation’s configuration. If something is genuinely useful, save it somewhere you control, such as a template, checklist or guidance note. More advanced tools may let you store reusable instructions, but they’re an optional extension.


