How to design an AI agent conversation, step by step
Eight steps from a vague “we need an agent” to a design engineering can build from, each with the mistake that sinks it and the test for done.

To design an AI agent conversation, write it out as example turns before anyone builds: define the agent’s job and its escalation line, write the happy path, branch into every way it can fail, specify the system call behind each answer, set the voice with paired examples, then get the result reviewed and approved in writing. That is the whole method. The rest of this post is each step in working detail: what to produce, the mistake that most often sinks it, and the condition that tells you the step is done.
Two prerequisites before step one. First, you need the raw requirements: which systems exist, what customers actually ask, who owns the human queue. Our requirements guide for business analysts covers the questions that surface them. Second, it helps to know what the finished design contains; the section list is in our conversation design template.
Step 1: Define the job and the escalation line
Write one paragraph stating what the agent handles, what it never handles, and the precise condition on which it hands the customer to a human. That last sentence is the escalation line, and it is the most consequential sentence in the design: everything else is bounded by it.
Produce: a scope statement written in outcomes. “Explains a charge, cancels a subscription, takes a card dispute up to the point of filing” is a job. “Handles billing questions” is a topic, and topics cannot be tested.
Common mistake: leaving escalation as an afterthought, something like “falls back to an agent if confused”. Escalation is a designed moment with its own example turn, its own data handover, and a concrete trigger: a named request for a human, a failed retry, a sentiment threshold, a regulated topic.
Done when: someone outside the project reads the paragraph and correctly sorts ten real customer requests into “agent” or “human” without asking you anything.
Step 2: Write the happy path as real turns
Write the main scenario as alternating example turns, from opening message to resolution. Every line should be a sentence the agent could actually send. Write both sides: what the customer plausibly says, including the half-formed version, and how the agent replies.
Produce: a dialogue anyone can read top to bottom in a few minutes, for one scenario at a time. Resist merging scenarios; a dispute and a cancellation are two designs sharing a greeting.
Common mistake: summarising. A box that says “agent asks for the order number” hides every decision that matters: how it asks, whether it explains why it needs the number, what it does with a partial answer. Nobody can veto a summary, so the veto arrives during UAT instead, when changing it is expensive.
Done when: you can read the conversation aloud with a colleague and it holds, and a stakeholder can point at one specific sentence and object to it. Objections at this stage are the process working.
Step 3: Add the failure branches
Now branch. Four families cover most of what production will throw at the agent: the system behind an answer is down or slow; the customer is angry; the request is out of scope; the user is abusive or probing for a jailbreak. Each branch is written the same way as the happy path, as real turns with an explicit ending.
Produce: for each scenario, the branches that apply. The timeout branch decides what the agent says at second three and at second thirty. The angry branch decides where apology stops and escalation starts. The out-of-scope branch names where the agent redirects. The abuse branch decides what the agent declines to engage with and when it ends the conversation.
Common mistake: one generic apology reused everywhere. “Something went wrong, please try again later” as the answer to a timeout, a declined refund, and a policy refusal. Customers notice, and so does the steering committee. The second version of this mistake is delegating failures to the platform’s default fallback, which means nobody designed them at all.
Done when: every branch ends on purpose: resolved, redirected, or escalated. If a scenario has no failure branches yet, it is not designed yet; we have written elsewhere about why the happy path is not a design.
Step 4: Specify the tools behind each answer
Every agent turn that states a fact or performs an action rests on a system call. Write it down next to the turn: the system, the endpoint or operation, the parameters, and for each parameter whether the agent asks the customer for it, infers it from context, or receives it from the session.
Produce: a specification per call: what it reads or writes, whether it is safe to retry, what the agent says while waiting, and what it says when the call fails, which is where your step 3 branches attach.
Common mistake: the unnamed system. “The agent checks the order status” survives every review because nobody owns it. Then, days before go-live, it turns out “check my order” touches three systems, and nobody has decided what happens when one of them times out.
Done when: an engineer who was never in the room can read the design and list every integration the build needs, without a meeting.
Step 5: Set the voice with paired examples
Adjectives do not transfer. “Friendly but professional” produces a different agent from every writer who touches it. Paired examples transfer: this sentence is on-voice; this nearly identical sentence is off-voice, and here is why.
Produce: a short set of voice traits, each with an on/off pair, plus tone rules for the situations that need them: declining a request, delivering bad news, apologising for an outage. Five traits is plenty. Twenty is a style guide nobody will read.
Common mistake: a blocklist of banned words. A blocklist tells writers what to avoid and nothing about what to write, so every writer fills the gap with their own instincts and the agent’s voice drifts turn by turn.
Done when: two people each draft a new turn in the established voice and a third cannot tell who wrote which.
Step 6: Review with the people who will veto it later
Every agent project has people who can stop it at the door: legal or compliance, the operations lead who owns the queue you escalate into, the engineer who owns the systems your tools touch, security if customer data moves. Find them now and put the design in front of them, failure branches first.
Produce: a reviewed design in which each objection is resolved as a changed turn. “Legal wants softer language on the decline” is resolved only when the decline turn actually reads differently.
Common mistake: showing a demo instead of the design. Demos collect compliments; designs collect corrections, and corrections are what you came for. The other version of this mistake is inviting reviewers after the build starts, when accommodating each objection costs a sprint instead of a minute.
Done when: every named veto-holder has read the failure branches, and their objections exist in the document as edits.
Step 7: Get written approval
Agreement in a meeting evaporates. Get a named person to approve a specific version of the design, in writing, before the build starts.
Produce: a recorded approval: who approved, which version, on what date. The version matters as much as the name. A design that keeps changing after sign-off was never signed off.
Common mistake: taking workshop nods as consent. Month four of the build, someone senior watches the agent decline a refund and says “that’s not what we agreed”. If nothing was written down, they are right by default, and someone absorbs the rework. For the client-facing version of this problem, see getting a client to sign off on an agent design.
Done when: you can answer, without searching your inbox, who approved the design, which version, and when.
Step 8: Hand off, and keep the design current
The design now goes to whoever builds: an internal team, a vendor, a platform configuration. It is the reference the build gets measured against, and that only works if it stays true while the build runs.
Produce: a handover plus a working agreement: when a build constraint forces a change to wording or flow, and it will, the change lands back in the design the same week. The approved design and the shipping agent must not quietly diverge.
Common mistake: treating handoff as the finish line. A design that drifts from the build stops being consulted within weeks, and then every disagreement reopens from zero. A stale design is worse than none, because it lies with authority.
Done when: at UAT, testers test against the design, and every deviation is classified as either a bug in the build or a recorded change to the design.
How long does all this take?
Less than the alternative. The happy path for one scenario is an afternoon workshop. The failure branches and tool specifications take a few working sessions more, mostly because they force decisions people were deferring. Review and approval run at the speed of your organisation. Set that against one rebuilt sprint, or one steering-committee meeting where nobody can answer “what does it say when the API is down”, and the design is the cheap part of the project.
None of this requires particular software. Teams run the sequence in documents and spreadsheets, and step 8 is usually where those quietly die. We built Klate because the design this method produces (the example turns, failure branches, tool contracts, voice pairs, and the approval record) deserves better than being scattered across a deck and a spreadsheet. If you want to see the method held in one place, the product page walks through it.



