An agent is a system that does not merely answer a question but takes steps on its own: retrieving information, calling an application, making a decision, triggering an action. The demonstrations are impressive, and the technology is real.
The question for a leadership team is still not whether it works, but where it works reliably enough today to take responsibility for it.
The difference between demonstration and operation
In a demonstration, a process runs once, under good conditions, with clean data and a well-disposed audience. In operation, the same process runs a thousand times, under all conditions, with incomplete information and people who want something other than what the process anticipates.
The decisive figure is not whether it works but how often. An agent with a ninety-five per cent success rate sounds excellent, and at a thousand transactions a month that is fifty errors, some of which can be expensive. So the question is always: what happens in those five per cent, and who notices?
It is not the success rate that decides deployment, but the cost of the error and the likelihood of spotting it.
Where agents work well today
- Narrow, frequent processes with a clear data basis: order status, opening hours, internal policy questions, initial qualification of enquiries.
- Preparatory work that a person then reviews: drafting a reply, assembling a case file, proposing a categorisation.
- Tasks with reversible consequences. A draft created in error is annoying. A payment triggered in error is something else.
Where it is still too early
Anywhere an error cannot be reversed, where the data basis is unclean, or where the process has so many exceptions that the exception is the rule. In those cases you build a system that helps in eighty per cent of instances and creates more work than it saves in the remaining twenty, and the organisation mainly remembers the twenty.
How I recommend starting
Start with an agent behind a person, not in front of the customer. The agent prepares, the person decides and approves. After a few weeks the data shows you how often approval happens unchanged. If that figure stays high, you can discuss more autonomy, with evidence instead of assumptions.
In the Microsoft environment this can be built in reasonable time with Copilot Studio, and with Azure AI Foundry for more demanding cases. The limiting factor is rarely the tool. It is the state of the data, the exceptions, and the question of who owns the failure case.
The question I ask of every initiative
What is the worst plausible error this agent can make, how quickly will we notice it, and who bears the consequences? If those three questions have answers, an agent is a very good idea. If not, it is a risk with good marketing.