Conversation design vs. prompt engineering
Conversation design decides what the agent should do; prompt engineering gets a model to do it. Run them out of order and you tune the wrong behaviour.

Conversation design is deciding what an AI agent should do and say: the flows, the example turns, the branches where things go wrong, and the rules for handing off to a human, agreed with stakeholders before anything is built. Prompt engineering is getting a model to actually do it: writing and tuning the instructions, examples, and constraints that steer a specific model toward that behaviour. One is a business decision. The other is an implementation technique. Most teams need both, and the order matters more than people expect.
The two get conflated because both produce words about how an agent should behave, and because on small projects the same person often does both in the same afternoon. On any agent that faces customers, carries a brand, or has more than one stakeholder, they are different jobs, with different owners and different failure modes. Confusing them is how teams end up with a beautifully tuned agent that does the wrong thing.
What is conversation design?
Conversation design is the work of specifying an agent’s behaviour, turn by turn, in a form that people who never open a code editor can read and approve. It covers the flows, example turns for each situation, the data the agent pulls at each step, the branches for timeouts and declines and angry customers, and the exact conditions under which the agent stops and hands the conversation to a human.
The defining property is who it is made with. A conversation design is negotiated with stakeholders: the product owner who carries the project, the legal reviewer who worries about what the agent promises, the brand owner who cares how it sounds, the client who is paying for it. The output is a design document, and its job is to be read, disagreed with, revised, and eventually signed off. We wrote a fuller treatment in what is conversation design for AI agents.
What is prompt engineering?
Prompt engineering is the craft of writing and iterating the text that steers a model: system prompts, task instructions, worked examples, output-format constraints, tool definitions, and the guardrail clauses that keep the model inside its lane. It is done against a specific model, because that is the only place it can be done. The same prompt behaves differently on different models, and differently again on the next version of the same model.
Good prompt engineers maintain eval sets, run regressions before every model upgrade, know the difference between an instruction a model follows reliably and one it follows on good days, and can squeeze latency and cost out of a pipeline without losing accuracy. None of that is trivial, and none of it goes away as models improve; the target keeps moving. The mistake worth naming is narrower: prompt engineering is the wrong instrument for deciding what the agent should do in the first place, because the people who own that decision cannot review a prompt.
What actually separates them?
They differ in who owns the work, what it produces, who checks it, and how it breaks.
| Conversation design | Prompt engineering | |
|---|---|---|
| Owner | A designer, analyst, or product owner, working with stakeholders | An engineer or prompt author, working against a model |
| Artifact | A readable design: turns, branches, tool contracts, escalation rules | A system prompt, examples, and an eval set, living with the code |
| Reviewed by | Business owners, legal, brand, the client: people who read it end to end | Engineers, through evals, regression runs, and transcript review |
| Changes when | The business changes its mind: policy, scope, tone, a new scenario | The model changes: a version upgrade, a regression, a new capability |
| Fails how | The agent confidently does something nobody agreed to | The agent knows the target and misses it: ignored instructions, drift, wrong format |
The last row matters most. A design failure surfaces in a meeting: someone senior reads a transcript and says “since when do we offer that?” A prompting failure surfaces in an eval: the agent was told the target and missed it. The first is expensive because it is discovered late and litigated across departments. The second is cheap to find if you have evals, and invisible until production if you don’t.
Why does the order matter?
Prompt engineering is optimisation, and optimisation needs a target. Every hour spent tuning a prompt makes the agent better at whatever behaviour the prompt author currently believes is wanted. If that belief has never been written down and agreed, the tuning still works; it just converges on a private guess.
Say a prompt engineer is handed “build a refund agent, make it helpful.” They do excellent work: the agent empathises, resolves fast, and approves borderline refunds because helpfulness was the brief and generosity evals well. Weeks later finance reads the transcripts and discovers a refund policy nobody approved, invented one instruction at a time. The prompt engineer did the job right. The target was wrong, and no amount of prompting skill could have caught that, because catching it requires the people who own refund policy to have read the intended behaviour before it was implemented.
There is also a plain cost asymmetry. Changing a line in a design costs minutes and a comment thread. Discovering a wrong target after tuning costs the tuning, the re-tuning, and the argument about whose fault it was. Deciding first is cheaper even when the decision is hard, and it is usually cheaper because the decision is hard: the arguments happen over a document instead of over a live agent.
When does each matter most?
Conversation design earns its keep before the build and at every boundary where agreement is needed. It matters most when the agent faces customers in a regulated domain, when a client is paying and will eventually say “that’s not what we agreed”, when legal or compliance must sign off on wording, and when the interesting behaviour is in the failure branches rather than the happy path. On that last point, the branches are usually the whole game; we made that argument in the happy path is not a design.
Prompt engineering earns its keep once the target is fixed. It matters most when reliability is the bottleneck: getting tool calls formatted correctly every time, holding behaviour steady under adversarial input, keeping answers inside the agreed scope, and carrying the agent through a model migration without regressions. It also matters whenever the design asks for something the current model finds hard, which is exactly the information the design side needs to hear.
How do they hand off to each other?
The design is the prompt engineer’s requirements document and their eval source in one. Every branch in the design is a test case: the timeout branch, the decline branch, the angry-customer branch each describe an input and the agreed response. The voice section gives worked examples of on-voice and off-voice replies, which is few-shot material. The tool contracts say what each call needs and what happens when it fails. The escalation rules become the hard clauses in the system prompt. A prompt engineer holding a real design starts with the target, the constraints, and the test set already written.
The handoff runs the other way too, and this is the part teams skip. When tuning reveals that the model cannot reliably do what a turn asks, that is a design fact, and it belongs back in the design where stakeholders can see it and agree the fallback. Patching it silently in the prompt reopens the original gap: the agreed document and the actual behaviour drift apart, and the next “since when does it do that?” is already scheduled. The loop only closes when the design stays the reference both sides update. That reference is also what makes sign-off possible at all, which consultants in particular learn the hard way; see getting a client to sign off on an agent design.
In Klate, the turns, branches, and tool contracts are written, reviewed, and approved as one document, before anyone opens a playground. If your next agent has stakeholders, the cheapest moment to find out what they actually want is before the first prompt is written.



