Guide

How to Build a Production-Ready AI Agent Workflow

Design the workflow before the agent. Define the job, tools, permissions, approval points, fallback paths, evaluation scenarios and operating metrics.

Start with the workflow

A production agent needs a defined job and an operating model.

This guide follows an example catalogue workflow: retrieve approved product information, prepare a structured draft, validate it and route it to a person before publishing. The same questions apply to support, sales research, reporting and internal knowledge workflows. The design choices below are a practical framework for scoping a business system.

Anthropic’s Building effective agents also distinguishes predefined workflows from systems that choose their own next steps and recommends starting with simple approaches. Use that distinction when deciding how much autonomy the job needs.

Step 1

1. Define the job

Start with the business result the agent is responsible for, the person who owns that result and a clear stopping condition. “Help with catalogue operations” is too broad. “Prepare missing product attributes from approved supplier records for an operator to review” is a boundary a team can implement and evaluate. Identify what remains outside the workflow, especially publishing, customer communication or commercial commitments.

Choose the simplest approach that can do the job. A fixed sequence of rules and software steps may be sufficient; add model judgment where interpretation is needed. Establish a baseline for the current process: completion time, correction rate, review effort and useful output. Those measures give the pilot a concrete purpose and prevent a convincing demo from becoming the only acceptance criterion.

Step 2

2. Define inputs and context

List the information required to act correctly. For each source, record its owner, access restrictions, expected freshness and how conflicts are resolved. A product document may describe a specification while a commerce API reports current availability; they answer different questions. Define which source takes precedence for each field instead of asking the model to reconcile everything implicitly.

Keep task instructions separate from retrieved content. Documents, messages and webpages can contain misleading instructions as well as useful facts. Treat that material as evidence to inspect, not permission to change the job. Give the workflow only the relevant context, retain source references and define what should happen when essential information is missing. A request for clarification can be a successful outcome.

Step 3

3. Define tools

Describe the systems the agent may read or modify through narrow tool contracts. A catalogue workflow might retrieve a product, look up an approved document and save a draft. Those are distinct operations with different consequences. Define accepted inputs, returned fields, errors and timeouts so the workflow can recognize whether a tool action actually succeeded.

Separate preparation from execution. A tool that previews proposed changes should not silently publish them. Validate tool arguments in application code and return explicit status rather than an ambiguous text response. Give external writes an identifier that lets the system recognize a repeated request, where supported. This makes a retry less likely to duplicate work after a timeout or interrupted response.

Step 4

4. Define permissions

Create an action matrix covering reads, drafts, updates and external effects. Specify what is allowed automatically, what requires review and what is forbidden. The agent should operate with the access needed for the current task, scoped to the appropriate account, dataset or product set. A user asking a question should not automatically grant the agent access to every company record.

Enforce permissions in the application and integration layer. A prompt explaining a restriction is useful context, but it cannot replace authorization checks. Recheck access when a paused workflow resumes and before a production action executes. Consider whether a change in user role, product ownership or approval state should invalidate a pending action. Keep the permission model understandable to the people responsible for the process.

Step 5

5. Add human approval

Place review where a person can make a meaningful decision. Show the proposed action, affected records, supporting evidence, uncertainty and expected effect. A reviewer should be able to approve, reject, edit or escalate the proposal. A vague confirmation dialog provides little control if the person cannot see what will actually change.

Approval should apply to the exact version of the proposed action. If the content, target records or relevant source data changes, decide whether fresh approval is required. Store who approved it and when; then verify that approval immediately before execution. In a Shopify PDP workflow, the approved product draft can move to publishing while a later edit returns to review. Define an owner for proposals that remain unanswered.

Step 6

6. Design state and memory

State describes where a particular run is now: the job identifier, completed steps, retrieved source versions, draft output, pending approvals and execution results. Persist enough information to resume safely after a restart. Use explicit states such as draft, validation failed, awaiting review, approved and published rather than inferring progress from the last chat message.

Memory across separate runs is a different design choice. Add it only when a durable preference or learned correction is needed for the job. Define ownership, retention, access and correction rules, and keep one customer’s context separate from another’s. Old information should not silently override current source data. A reliable first version may need durable workflow state without any long-term conversational memory.

Step 7

7. Design failure handling

Enumerate ordinary failures before the pilot: missing data, contradictory sources, low-confidence extraction, an unavailable tool, a rejected approval or a partially completed bulk update. Define an appropriate response for each. Some failures can be retried; others need a clarification, a smaller scope or human escalation. Unsupported information should remain unresolved instead of becoming a plausible-looking fact.

Bound retries, processing time and spending. Preserve successful steps so a resumed workflow does not repeat external effects unnecessarily. A bulk catalogue update needs status for each product, not only a single success message for the batch. Give the operator a useful recovery view showing what completed, what failed and which action is safe next. Define a manual path for work that must continue during an outage.

Step 8

8. Add evaluation

Build an evaluation set from representative work before expanding the scope. Include normal requests, ambiguous instructions, missing sources, conflicting values and attempts to exceed permissions. For each example, define acceptable output and expected behaviour, including when the correct response is to stop or escalate. Keep some cases separate from the examples used to tune the workflow.

Measure task completion, factual support, forbidden actions, corrections, latency, cost and human review effort. A high completion rate is not enough if operators spend longer repairing the output. Combine deterministic checks with expert review where judgment is required. Re-run the relevant scenarios after changing a prompt, model, tool, business rule or data source. Agree on release criteria and who decides whether a failure blocks rollout.

Step 9

9. Add observability

Record the workflow identifier, step transitions, tool calls, result status, source references, model or prompt version and approval decisions. Operators need to understand what happened and where to intervene. Track failures, retry loops, stale data and stalled approvals so problems become visible before they affect a large queue of work.

Logs should contain the information needed for diagnosis while respecting access and retention requirements. Avoid recording credentials or unnecessary sensitive content. Restrict who can inspect traces and define a process for reviewing recurring errors. A trace should explain actions and evidence; it does not need a hidden chain of thought. Connect operational alerts to a responsible owner and make stopping or rolling back the workflow a practical action.

Step 10

10. Production checklist

Before release, walk through one complete job with the business owner and the operator who will use it. Verify the source data, permissions, draft preview, approval, execution result and recovery path. Repeat the exercise with a tool outage, conflicting input and an approval that becomes stale. Confirm that the user can tell the difference between a prepared action and a completed one.

Check that the evaluation set passes agreed criteria, monitoring has an owner, spending and retries are bounded, and a manual fallback is documented. Start with a limited category or team, review actual outcomes and expand only when the evidence supports it. Production readiness is an operating commitment: the workflow needs maintenance as business rules, tools and data change.

Common mistakes

  • Starting with a model before the workflow is clear.
  • Ignoring data quality and access until implementation.
  • Leaving human approval and failure handling undefined.

Author

Kosma Lenar, Founder of Lumethica

Written from a product and agentic systems perspective for teams building AI products beyond the demo.

FAQ

Common questions

Who is this guide for?

It is for product, data, operations and leadership teams evaluating practical business AI use cases.

What should we prepare before a sprint?

Bring example workflows, existing tools, data sources, pain points and the business outcome you want to improve.

Can Lumethica help implement the outcome?

Yes. Lumethica can help with audit, productization, agent workflow design, data integrations and trust controls.

Related

Related services and definitions

Apply the framework

Need this workflow designed for your business?

Bring the current process, example inputs and the decisions that need human control. We can define the workflow and scope the first release.

Explore the AI Agent Workflow Sprint