In custom software, AI works reliably today on well-defined text tasks: extracting data from documents, classifying messages, summarizing case files, and searching your own knowledge. Fully autonomous agents and decisions made without human review still fall short. That boundary is where AI hype and reality part ways.
The numbers tell the same story. Business use of AI is growing fast, and most of it involves working with text. Yet only a small share of organizations see it in their bottom line, and many projects built on autonomous agents are shut down before they deliver anything.
This article covers which applications in business software demonstrably work, where they still struggle, how to spot hype, and which questions to ask a vendor. We've been building custom software for more than 30 years. That practice is our starting point, with independent research to back it up.
What the numbers say about AI hype and adoption
Across Europe, business use of AI is rising quickly. According to Eurostat, 20.0% of EU enterprises with 10 or more employees used AI in 2025, up from 13.5% a year earlier.
The Netherlands, where we're headquartered, shows the same curve. Statistics Netherlands (CBS) reports that one in six Dutch companies, 17%, used at least one form of AI in 2025, double the share of two years earlier. Among companies with 50 to 250 employees, the share rose from 20% in 2023 to 45% in 2025.
What companies use it for is telling. Text mining, the analysis of written text, is the most widely used form of AI in both the EU and the Netherlands: 11.8% of EU enterprises and 12% of Dutch companies. In practice, the work AI does is mostly reading, sorting, and summarizing.
Returns lag behind adoption. In McKinsey's State of AI 2026, published in August 2026, 37% of respondents attribute any impact on earnings (EBIT) to AI, the same share as a year before. Only 6% get more than 5% of their EBIT from AI. Eight in ten respondents do say AI has improved their own productivity. Individual gains rarely add up to results for the organization yet.
What works reliably in custom software today
The applications that demonstrably work in business software have a few things in common. The task is well defined, the output can be checked, and a mistake is cheap to fix. These four fit that description:
- Document extraction. Invoices, order confirmations, inspection reports, and application forms go in, and structured fields come out. It works best on recurring document types, with checks built into the software: does the total add up, does the supplier exist, does the order number match.
- Classification and routing. Incoming email, tickets, and reports get a category and a priority and land with the right team. A wrong label is easy to spot and cheap to correct.
- Summarization. Long case files, call notes, and logs become a short summary, and then an employee decides. The summary speeds up reading and leaves the judgment with a person.
- Search across your own knowledge. Manuals, procedures, and past solutions become searchable in plain language, with a link to the source behind every answer. That source link is what makes the answer verifiable.
In each case, the language model is one part of a larger process. The value only appears when the output lands cleanly in your ERP, CRM, or case management system, with error handling and a queue for cases that raise doubt. That's software integration, and it's where most of the work in a good AI project goes.
What still falls short or needs close supervision
AI agents are systems that carry out a series of steps on their own: looking up information, operating an application, completing an action. They're improving fast. According to Stanford's AI Index 2026, the success rate on OSWorld, a benchmark of real computer tasks, rose from about 12% to about 66%. The same agents still fail roughly one in three attempts. In a business process with several steps, those failures compound.
Research firm Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. Gartner also estimates that only about 130 of the thousands of vendors selling agents offer real agentic capabilities. The rest rebrand existing chatbots and automation, a practice Gartner calls "agent washing."
"Most agentic AI projects right now are early-stage experiments or proof of concepts that are mostly driven by hype and are often misapplied." Anushree Verma, Senior Director Analyst at Gartner
Beyond agents, there are three areas where AI without supervision still disappoints:
- Decisions with consequences. Rejecting an application, placing an order, or making a commitment to a customer. A language model sounds just as convincing when it's wrong.
- Exact business rules. Discounts, rates, deadlines, and calculations belong in regular code. A language model doesn't give the same answer to the same question every time.
- Data of uneven quality. Missing fields, outdated documents, and conflicting sources come straight back in the answers.
AI hype and reality side by side
Many promises sound bigger than what runs reliably today. This table sets common promises against what works in custom software now and what people still do.
| The promise | What works reliably today | What people still do |
|---|---|---|
| AI processes all your invoices | Extracting fields and proposing a booking, with checks on totals | Reviewing exceptions and approving |
| The chatbot knows your whole business | Searching approved documents, with source references | Keeping sources current and spot-checking answers |
| An agent handles your purchasing | Preparing a draft order based on stock and history | Confirming the order |
| AI decides on applications | Sorting and summarizing applications and flagging missing documents | Making and explaining the decision |
The pattern is always the same. AI takes over the reading and preparation, and a person keeps the decision. That's where the time savings are, and it's what keeps mistakes manageable.
How to recognize AI hype
Hype usually shows in what's missing from the pitch. Watch for these signs:
- There's an impressive demo and no measurement on your own documents or data.
- Nobody can tell you how often the system gets it wrong.
- The word "agent" is attached to what is really a chatbot or a fixed workflow.
- The pitch is all about the model and says nothing about the process, the integrations, or the data.
- The promise is full automation, with no review step for uncertain cases.
- There's no answer to what a single transaction costs once volume grows.
One sign alone is no reason to walk away. Three or more at once usually means you're looking at an experiment rather than a product.
Questions to ask an AI vendor
These six questions quickly show whether a proposal rests on evidence:
- Which of our own documents or data has this been tested on, and with what result?
- How often does the system get it wrong, and how do we notice in the process?
- Where does human review happen, and what does the reviewer see on screen?
- How do you record what the model proposed and what the employee decided?
- What happens when the model is unavailable or gets a new version?
- How does the solution connect to our existing systems, and who maintains that integration?
A good vendor answers with figures from an earlier measurement or offers to run that measurement with you. For broader questions about choosing a partner, see our guide on how to choose a software company.
How we build AI in: with people close to it
Our view is simple: AI works best when people stay close to it. We use it where it delivers clear business value, in a process you already know and can measure. This is how we approach it within data and AI:
- Pick one process with a lot of reading. Incoming documents or reports, for example, with a clear owner.
- Measure first. How many items per week, how much time per item, how many errors. Without a baseline, there's nothing to prove afterwards.
- Build a test set of real examples. Every version of the model gets measured against it, just like regular software.
- Design the review step. The employee sees the proposal, the source, and the uncertain cases, and decides.
- Build it into your own software. Through integrations with the systems you already have, so the work stays in your existing process.
This is the same craft as any other custom software development: clear requirements, testing, and maintenance. Isatis brings 30+ years of experience, 30+ engineers in Nijmegen and Sarajevo, and 100+ completed projects, and is ISO 9001 and ISO 27001 certified. Our cases range from maintenance software (MRO) for aviation to a subsidy administration platform and RFID in the supply chain. Book a call and we'll look together at which of your processes suits a first step with AI.
Frequently asked questions
What is AI hype?
AI hype is the gap between what's promised about AI and what it reliably does in practice. The gap is widest for fully autonomous agents and decisions made without review. For well-defined text work, such as extracting data from documents and summarizing, the gap is small.
Which AI applications work reliably in business software today?
Document extraction, classifying and routing messages, summarizing case files, and searching your own knowledge with source references. They work because the task is well defined, the output can be checked, and an employee keeps the decision.
Are AI agents ready to run business processes on their own?
For most processes, not yet. Agents are improving fast, yet according to the Stanford AI Index 2026 they still fail about one in three computer tasks. Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027. An agent that prepares the work and a person who confirms it does work today.
How do you start with AI in custom software without falling for the hype?
Start with one process that involves a lot of reading and has a clear owner. Measure how much time and how many errors it costs today, build a test set of real examples, and design a review step for the employee. Only expand once the measurements show it works.





