PLAYBOOK

Every piece of prompting advice starts from the same assumption, that you can describe the outcome you are after, so be specific, give examples, say what good looks like. That is sound advice, and it is no use at all on the days when the outcome is the thing you do not have. The proposal that feels wrong in a way you cannot name, the pricing page you would only recognise as right once it existed, the plan whose shape you will know on sight but cannot write down in advance. Those are the hardest prompts to write, and they sit under most of the work worth handing to these tools in the first place.

📍 In brief

When you cannot state the outcome, the prompt's job is not to specify the answer. Its job is to open a search that your judgement can close. You do not need to know what you want. You need to recognise it when it appears, and a method that makes it appear sooner. Five stages follow, drawn from the published research and from the firm's own build, fumbling included.

The Wayfinder Team

None of this is invented here, and the sourcing matters. Ethan Mollick, who has spent years studying how people actually work with these systems, puts it plainly, that you can open a conversation with these tools before you know what you are asking for, describe what you might need, and watch what comes back. OpenAI's enterprise guidance is equally blunt, the first prompt is not expected to work, and trying something, failing and refining is the normal path rather than the recovery path. Anthropic's guidance for its most capable models now describes starting with as little as an idea as a first-class way of working. What follows turns that research into a sequence, tested on the firm's own production system, most of which was built in sessions that began with no statable outcome.

When to use this. You have a real task, you would recognise a good result on sight, and you cannot yet write down what it is. That middle condition matters, because the method runs on your judgement, which is why the research consistently advises starting with AI inside your own domain, where you already know what good looks like. Two honest exclusions apply, because a playbook applied in the wrong situation fails and blames the playbook. If you can already specify the outcome, specify it and skip the method, and if the missing piece is a decision only you can make, no prompt will make it for you.

Five-stage flow diagram for prompting without a stateable outcome. Stage one, describe the situation, not the request. Stage two, make it interview you first. Stage three, ask for three rough directions. Stage four, react with reasons and loop until a version is right or near enough to describe. Stage five, bank the prompt you wish you had started with and run it clean.

The five stages at a glance. The loop at Stage 4 is where the specification you could not write gets said, one reasoned reaction at a time.

Stage 1, describe the situation, not the request

Open with a plain account of where you are, covering what you know, what you do not know, and what you think you might need. This is not an instruction, it is the paragraph you would say to a capable colleague before they ask their first question, who the work is for, why it matters, and what has already been tried.

The mechanism behind this is worth understanding, because it explains most of what follows. These systems default to the most statistically typical response to whatever they are given, so a thin request gets the average answer to that request, averaged across everyone who might have asked it. Everything you add about your situation, your customers and your constraints moves the response out of that average and towards your corner of the problem. Mollick describes context as the one trick that reliably improves a conversation, and Anthropic's guidance draws the same distinction in different words, that a constraint only tells the system what not to do, while context explains what the work is for, letting it make sensible calls in situations you never anticipated. When you cannot write the specification, the description of the situation is doing the specification's work.

The gotcha is writing a short prompt because short feels efficient. A thin request produces a generic answer, the generic answer reads as proof that the tool is mediocre, and the person who most needed the method concludes it does not work. The research names this exact trap, where naive prompting produces poor results that convince people the system is weak, and that verdict stops them learning the habits that would show them otherwise.

Stage 2, make it interview you before it produces anything

Before it generates a single word of output, add one instruction to the end of your description, asking it to repeat the ask back and then put to you every clarifying question it has, and answer them, including the awkward ones. Anthropic's practitioners call this the single most useful habit in working with these tools, and OpenAI's guidance recommends the same move under a duller name, using the model to interrogate your own request before it attempts the work.

The value is not in the answers you give, it is in discovering which questions you could not have thought to ask yourself. Which period are we looking at, what does good mean here, who sees this first. The trap being corrected is the assumption that the system already knows what is obvious to you, and it knows none of it. Thirty seconds of questions up front is cheap, while discovering the same gaps after three wrong drafts is not.

The gotcha comes when it asks a question you cannot answer. That is not the method failing but the method working, because it has just located the real gap, and the gap was never a prompting problem. It is a decision, it is yours, and the honest move is to go and make it, or to say so plainly and ask for both versions so the fork carries into the next stage.

Stage 3, ask for three rough directions, not one answer

Now ask for output, but not for the answer. Ask for three versions, deliberately different from each other and deliberately rough, because you are not commissioning a deliverable, you are laying options on the table to point at.

Good ideas come from having many ideas to choose between, and volume is the one thing these systems produce without complaint, a point Mollick's practical guidance has made since the earliest guides. It lands hardest exactly here, because although you cannot specify what you want, shown three different somethings you can say instantly which is least wrong and why. Recognition is available before specification, and this whole method is a machine for converting the one into the other.

Ask for rough on purpose. A polished single answer invites acceptance, and acceptance is precisely what you cannot yet trust, whereas three sketches invite comparison, and comparison is where your judgement starts producing information.

The gotcha is asking for the best option, or letting it converge too early. A single answer collapses the search back to the average you escaped in Stage 1, and the parallel failure is treating the three directions as finished candidates and picking a winner to ship. They are compass bearings, not deliverables.

Stage 4, react with reasons, and loop

Pick the least wrong direction and say why the others lost, then react to what is in front of you, specifically. Too formal for a customer who has already complained twice. Right structure but wrong emphasis, because the risk section is the point. This paragraph is the one to keep, and everything around it is filler. Then let it try again, and react again.

Watch what is happening to the specification you could not write. Every reaction with a reason attached is a fragment of it, stated at the exact moment you became able to state it, and three or four cycles of this usually mean you have said, in pieces, the brief you could not have produced in advance. This is why the loop is the method and not the fallback. The research says iteration is the expected path, and that the productive move each cycle is to diagnose the miss, whether missing context, wrong audience or wrong format, rather than simply rolling again. On the firm's own build, working sessions open with a situation description and converge in a handful of reasoned reactions, and the sessions that stall are reliably the ones where the reaction was a mood rather than a reason.

The gotcha is reacting without the reason. Asking for better moves the output randomly and teaches the system nothing about your corner of the problem, while a named miss moves it in a direction. If you cannot name what is wrong, say which part is closest to right and work outward from there.

Stage 5, bank the prompt you wish you had started with

At some point a version is right, or near enough that you can finally describe the target. The session's real product is not the draft, it is the fact that you can now write the specification you could not write this morning. So write it down before you close the tab, the situation, the task, and what good looks like, with the hard-won specifics from the loop folded in. The shape is the same handoff you would give a capable newcomer, and OpenAI's guidance treats that handoff shape as the core of a reliable prompt.

Keep one more thing alongside it, the pair of an early wrong draft and the final version. Anthropic's guidance notes that current systems can work out your standards from what changed between the two, which makes the pair a specification in its own right, one you never had to articulate.

Then run the banked prompt in a fresh conversation. Long exploratory threads accumulate half-retracted instructions and abandoned forks, and stale context is the quiet cause of most unexplained quality drops. The messy thread was scaffolding, so take the prompt out of it and let it go.

The gotcha is hoarding the thread as if it were the asset, when the asset is the prompt plus the before-and-after pair. On the firm's build, every durable prompt in the production system was written last, after the sessions had shown what good looked like, and each one runs cleanly precisely because the fumbling happened somewhere else.

📋 The checklist

  1. Describe the situation: what you know, what you do not, what you might need.

  2. Have it repeat the ask back and interview you before it writes anything.

  3. Ask for three rough directions, deliberately different.

  4. React with reasons, in a loop. Every reason is a piece of the spec.

  5. Write the prompt you wish you had started with, keep the before-and-after pair, run it clean.

Where it goes wrong

Treating the first answer as the answer. These systems are fluent before they are right, and fluency reads as finished. In this method the early outputs are navigation rather than deliverables, and the fix is built into Stage 3, since three rough directions cannot be mistaken for a finished answer.

Hunting for magic wording. A whole cottage industry sells secret phrasings and incantation templates. The research position is unambiguous, there are no secret prompts, and the reliable gains live in context, options and reasoned reaction. If a session is failing, the fix is almost never in the phrasing of the request but in Stage 1 or Stage 4.

Running it outside your judgement. The method's engine is your ability to tell better from worse on sight. Point it at a domain where you cannot, and Stage 4 produces confident cycles of nothing, because the reactions have no information in them. This is why the research advises building AI habits inside your own expertise first, and it is the honest boundary of this playbook.

The bottom line

If you keep one move, keep this one. Open with the situation, close with reasoned reactions. Everything else is scaffolding around the fact that your judgement is the specification, delivered in instalments, and the prompt's only job is to give it something to judge.

A prompt is not a request for the answer. It is the opening move in a search that your judgement finishes.

P.S. The prompts you bank in Stage 5 compound. Kept somewhere findable, they become a written record of what good looks like in your business, decision by decision, and six months from now that record will be worth more than any of the drafts it produced.