Agent flow documentation: what to write down before you build
Escalation rules belong in the design; retry timers belong in the platform. Draw that line and the design stays the record of what was agreed.

Agent flow documentation is the written record of the decisions an AI agent embodies: the flows it follows, example turns for each situation, where the conversation branches, the contracts of the tools it calls, and the conditions under which it hands off to a human. It is a separate document from the implementation that executes those decisions, and the line between the two is where most documentation efforts quietly die.
Why does chatbot flow documentation go stale?
It goes stale because it is usually a copy of the implementation rather than a record of decisions, and a copy always loses to its original. The sequence is familiar enough to feel like a law of nature. During discovery someone produces a flow: a diagram, a deck, a spreadsheet of intents. Everyone nods. The build starts. Within two weeks the build and the flow disagree: a branch was merged, legal changed a sentence directly in the platform, a tool got renamed. Nobody goes back to update the diagram, because updating it buys nothing; nothing checks it and nothing depends on it. Six months later a new analyst joins, reads the documentation, and learns an agent that no longer exists.
The conclusion teams draw is that documentation is a tax that fails to pay for itself. That conclusion is half right. The documentation failed because it recorded the wrong things at the wrong altitude. A diagram of boxes mirrors the shape of the build while holding no decision the build does not already hold, and when a document merely mirrors the build, the build wins and the mirror goes dark.
The fix is a division of labour: decisions in the design, implementation in the platform, each held to its own standard of freshness.
What belongs in the design document?
The design holds every decision that a person outside the engineering team must be able to read and veto. In practice that is four things.
- Example turns. How the agent handles each situation, shown as examples, including the unglamorous ones: the decline, the apology, the message on the way to a human. Any moment a stakeholder could object to belongs in the design. Read in full, the example turns are also the only version of the agent a legal or brand reviewer can meaningfully approve.
- Branching. Which situations get their own path, where each path ends, and which scenarios are deliberately out of scope. The happy path is a small fraction of this; the decisions that matter live on the failure branches, a case we make in full in the happy path is not a design.
- Tool contracts. What data the agent reads or writes at each step: the system it touches, the parameters it needs, where each parameter comes from (asked in conversation, inferred from context, or fixed), and what the customer hears when the call fails. This is the contract the conversation depends on; the API’s internals stay on the platform side of the line.
- Escalation. The exact conditions under which the agent stops and a human takes over, and what it says as it steps aside. Escalation is a designed outcome, reviewed like every other turn.
What belongs in the platform?
Everything that exists to make the model actually produce the agreed behaviour: prompts, model choice and configuration, retrieval setup, retry and timeout values, training utterances, deployment wiring. These are implementation. They change weekly, they are tuned by the people building, and nobody outside engineering signs off on a temperature setting.
The test for which side of the line something sits on: would a stakeholder need to approve a change to it? If yes, it belongs in the design. If a change to it is invisible so long as the behaviour holds, it belongs in the platform.
Prompt engineering, in this framing, is the craft of making the implementation hit the design, which is why the two are so often confused and so differently owned. We unpack that split in conversation design vs. prompt engineering.
| Decision or setting | Lives in | Why |
|---|---|---|
| How the agent handles a declined refund | Design | A stakeholder can veto it |
| Which scenarios get their own branch | Design | Scope is an agreement someone approved |
| A tool call’s parameters and where each comes from | Design | The conversation depends on the contract |
| When the agent hands off to a human | Design | An outcome someone must own |
| The system prompt | Platform | Rewritten freely to hit the design |
| Model, temperature, retrieval setup | Platform | Invisible while behaviour holds |
| Retry counts and timeout values | Platform | Tuning; what the customer hears at timeout is design |
How do you keep the design and the build aligned?
By making the design the reviewed source of intent and routing every change through it. Alignment is a workflow property, and it rests on four mechanisms.
- The design is what gets reviewed. Sign-off happens on the example turns, the branches, the tool contracts, and the escalation conditions. Once approval attaches to the design, the design acquires the gravity that keeps it current; people maintain what carries consequences.
- Changes go through it. When someone wants the agent to behave differently, the request lands as a change to the design: the turn is edited, the branch added, the delta reviewed, and then the build follows. The reverse order, change the build now and backfill the document later, has a completion rate every practitioner already knows.
- It is versioned and diffed. Named checkpoints, and a turn-by-turn diff of each version against the one before it. “What changed since the version the client approved” becomes a question you answer from the version history in minutes.
- Engineering builds from it. Nobody has to reconstruct the workshop from memory. Where a coding agent does part of the building, it can read the design the same way; in Klate that surface is MCP.
Some decisions are discovered in the build: the model cannot do the thing reliably, or latency forces a different shape, and the design has to move. The design leads on intent and the build sometimes leads on discovery; either way, no change counts as done until the design records it.
When the design and the build disagree, which wins?
The design, and the reason is definitional: the design is the record of what was agreed; the build is an attempt to realise it. When the two diverge, exactly one of two things is true. Either the build is wrong, in which case the fix is an engineering ticket, or the agreement has changed, in which case the design is updated and the delta re-approved by whoever owns that decision. There is no third state in which the build quietly redefines the agreement while the record shrugs.
Suppose the design for a refunds agent says a declined refund ends with an offer to escalate; in production, the agent apologises and stops. Which behaviour is correct? Without a reviewed design, that question is settled by whoever is most senior in the room and whoever remembers the workshop most confidently. With one, it is settled by reading: the escalation-on-decline turn is in the approved version, and a timeline nobody can edit shows who approved it and when. The build has a bug, the fix is unambiguous, and the conversation takes four minutes instead of a meeting. Getting that approval in place before the build is its own craft, covered in getting a client to sign off on an agent design.
The limit no tool removes
No tool synchronises intent. A platform export can tell you what the build currently does; only people making decisions and writing them down can say what it should do, and a person still has to check the build against the design. Anyone selling automatic design-to-build sync is describing another mirror of the build, and mirrors go dark.
What a tool can change is the price of the discipline. Keeping decisions written down happens only when it is cheap: editing a turn must take a minute, the diff must be generated for you, approval must be one emailed link, and the whole record must be readable end to end by the least technical person who holds a veto. When updating the record costs less than the arguments a stale record produces, teams keep the record current.
In Klate, example turns, branches, tool contracts, and escalation live as versioned, approvable objects, and the sign-off travels with the design. If your current home for those decisions is a diagram nobody has opened since the build started, start smaller: write one flow down properly, using the sections in our conversation design template. The Free plan is enough for that first flow. Client sign-off is on the paid plans.



