Building an AI agent for business processes

Drie collega's buigen zich samen over een ontwerp op tafel
Date
October 8, 2026
Author
Isatis Group
Category
AI & Innovation
Read time
10 min read

Building an AI agent for business processes: how an agent differs from a chatbot and a workflow, its six building blocks, the pitfalls, and the build.

An AI agent for business processes is built from six parts: a language model, the tools it may call, the context it receives, limits on what it may do, a point where a person approves, and logging with evaluation. Choosing the model is the smallest part of the work. Most of the effort goes into the connections to your systems and the rules around them.

An agent is software in which a language model, the AI that reads and writes text, chooses its own next step. Below you'll see how it differs from a chatbot and a fixed workflow, what each building block involves, which processes fit, where agents go wrong, and how a build runs.

What an AI agent is, next to a chatbot and a workflow

For the build, one distinction matters: who decides the next step. Anthropic, the company behind the Claude models, calls a system a workflow when the route is fixed in code, and an agent when the language model directs its own process and tool use. Tools are the functions the model may call, such as looking up an order.

OpenAI draws the same line in its practical guide to building agents. A simple chatbot, where the model doesn't control how the work is carried out, falls outside that definition.

Feature Chatbot Workflow with a model step Agent
Who chooses the next step The user, with a new question The code, in a fixed order The model, within fixed limits
What the system does Writes an answer One defined step, such as reading or sorting Looks up, combines, and prepares actions
Access to your systems None, or read access only Fixed calls at fixed moments Tools the model chooses itself
Fits Questions about policy and documentation Work that always runs the same way Work where the number of steps varies

Take a customer service team that receives emails about open orders. A chatbot answers a question about the standard delivery time. A workflow pulls the order number from each email and puts the message in the right queue.

An agent picks up the email from a customer who wants delivery a week later and orders 20 extra units. It looks up the order in the ERP, the system that holds orders, stock, and invoices. Then it checks stock and planning and prepares a change proposal. The next email needs different steps.

AI agent architecture: the six building blocks

OpenAI names three core components of an agent: the model, the tools, and the instructions. In a business process, the agent works in systems where a mistake has consequences. So we use six building blocks here, each with what you decide before the build starts.

Building block What it is What you decide before the build
Model The language model that reads, reasons, and chooses the next step Which model you use and how you replace it later
Tools The functions the agent may call in your systems For each tool: the system, read or write, and the permissions
Context and memory The instructions, the data for the task, and what the agent remembers What the agent sees per task and what it keeps
Limits Rules in code about what the agent may do Which actions are forbidden and when the agent stops
Human approval The point where an employee approves a proposal Which actions a person confirms
Logging and evaluation The log of every step and a fixed set of test cases What you keep and what you test each version with

The model

Expect to replace the model with a newer version later. Build the agent so you can swap it without rebuilding the tools and the rules. OpenAI advises setting a baseline with the most capable model first, then testing whether a smaller model scores well enough.

Tools and APIs

A tool is a function the agent may call, such as "look up order" or "prepare draft reply". Usually an API sits behind it, the interface through which software requests or changes data in another system. OpenAI distinguishes, among others, tools that retrieve data and tools that take an action, such as sending a message. That second group needs the most care.

Keep tools narrow. A tool that changes the delivery date of one order line is safer than one that can edit any field in the ERP. Also check whether your systems have the APIs that AI needs.

Context and memory

Context is everything the model sees at a step: the instructions, the customer's email, the data a tool returned, and the steps so far. Memory is what the agent keeps for a later task, such as a customer who only accepts deliveries on Tuesdays.

More context doesn't automatically make an agent better. Anthropic describes how a model recalls information less accurately as the amount of text in its context grows. Give the agent only the data and work instructions it needs for that task.

Limits in code

A limit belongs in the software and in the permissions of the account the agent works with. A model can ignore or misread a sentence in its instructions. OWASP, known for its lists of software security risks, recommends enforcing authorisation in the downstream systems instead of letting the language model decide whether an action is allowed.

At a minimum, define:

  • which tools only read and which also write
  • a maximum per action, such as a number of units or order lines
  • a maximum number of steps per task, after which the agent stops and hands the task over

Human approval

For actions with large or lasting consequences, the agent prepares a proposal and an employee decides. OpenAI gives cancelling an order and making a payment as examples, and also has an agent hand over when it gets stuck after several attempts. For choosing who decides at each process step, see human approval in an AI workflow.

Logging and evaluation

Logging means keeping, for every task, what the agent saw, which tools it called, what they returned, and what it proposed. Without that log, you cannot explain a mistake afterwards.

You evaluate with a test set: a fixed set of real cases with the correct outcome, which you rerun after every change to the model, the instructions, or a tool. Assess both the final outcome and the route to it. A correct answer reached through the wrong tool is a mistake nobody has noticed yet.

Which business processes suit an AI agent

An agent is worth the effort where fixed rules fall short. OpenAI lists three characteristics of suitable work: decisions that need judgement, rules that have grown too extensive to maintain, and heavy reliance on unstructured data such as free text and documents. Anthropic recommends choosing the simplest solution and adding complexity only when it demonstrably improves outcomes.

An agent fits work where:

  • the number of steps varies per task
  • the information is spread across several systems
  • the input is free text, such as emails or documents
  • an employee can check the proposal quickly

A fixed workflow is the better choice when:

  • every task goes through the same steps
  • the rules are clear and rarely change
  • the outcome must be identical every time, as with an invoice posting

Matching an invoice to a purchase order is a fixed route. Research work suits an agent better, such as a complaint where you put the order, the delivery, and earlier contacts side by side.

Often a mix is the most sensible choice. A fixed workflow receives the task, checks it, and books the result, and the agent handles only the step that needs research. If you don't yet know which process qualifies, start with finding AI use cases.

Where agents go wrong

You can prevent the mistakes below in the design.

  • Too many tools and broad permissions. OWASP calls this excessive agency, with three root causes: excessive functionality, excessive permissions, and excessive autonomy. An agent that only looks up orders gets an account that can only read.
  • Outside text that gets read as an instruction. If an email contains a hidden instruction, a model may follow it. This is called prompt injection. OWASP describes an assistant that forwards sensitive information from the mailbox because of an incoming email. Narrow tools, limited permissions, and approval before sending limit the damage.
  • Errors that compound. Anthropic warns about compounding errors in autonomous agents and recommends extensive testing in a sandboxed environment. An order number misread in the first step carries through every step after it.
  • Tools that look alike. Anthropic calls a bloated set of tools one of the most common failure modes it sees. If an engineer cannot say for certain which tool fits a situation, neither can the agent.
  • A demo without measurement. Five successful examples say little about the hundreds of cases in a busy week. Without a test set, nobody notices when a new model version behaves differently.

Building an AI agent for business processes: from one narrow task to production

OpenAI writes in its guide that organisations typically have more success with an incremental approach than with a fully autonomous agent from the start, and recommends getting the most out of a single agent first. The build therefore runs in six steps.

  1. Pick one narrow task. For example, only questions about the delivery date of existing orders. Write down how an experienced colleague handles it today, exceptions included. That description becomes the instructions. Define scope and stop criteria as described in setting up an AI pilot.
  2. Collect real cases. Take dozens of completed tasks with the correct outcome, including the awkward ones. This becomes your test set.
  3. Build the tools, read access first. Each tool is a small piece of software with its own tests. Anthropic writes that for one of its own agents it spent more time on the tools than on the prompt.
  4. Let the agent run alongside without consequences. Staff work as usual, and you compare the agent's proposals with what they did.
  5. Release write actions one at a time. First with staff approval on every proposal, and without approval only once the measurements justify it over a longer period.
  6. Bring the agent under management. Name an owner, monitor errors and turnaround time, and rerun the test set after every change.

Building an agent is the same craft as other custom software: integrations, permissions, tests, and logging. We've been building software for more than 30 years, with 30+ engineers in Nijmegen and Sarajevo and more than 100 projects, and we're certified for ISO 9001 and ISO 27001. Read more about data and AI or book a conversation about the task you want to start with.

Frequently asked questions

What is an AI agent?

An AI agent is software in which a language model decides which step to take to complete a task. It calls tools to do so, such as looking up an order or preparing a draft reply, and works within limits that are fixed in the software.

When is a fixed workflow better than an AI agent?

When every task goes through the same steps and the rules are clear. A fixed workflow is then more predictable and easier to test. Choose an agent for work where the number of steps varies per task and the input is free text.

What do you need before you build an AI agent?

One clearly defined task with an owner, systems you can reach through an API, dozens of real cases with the correct outcome, and an employee who reviews proposals. If access to your systems is missing, build that first.

Do you need multiple AI agents for one process?

Start with one agent. OpenAI recommends getting the most out of a single agent first, because multiple agents add complexity. Split the work only when one agent no longer follows its instructions well or keeps choosing the wrong tool.

Share this article

Jack van Poll

Jack van Poll

Co-Founder, Isatis

Writes about nearshore engineering,
software partnerships and building teams that last.

Get in touch

Read next.

All articles