Jargon
Agent is a word from one field being used to sell something from another
The term has a precise technical history, a loose current usage and a marketing usage, and telling which one is meant requires asking what the system is actually permitted to do.
By Daniel Okonkwo3 min read

The technical meaning came from somewhere specific
In the study of learning from consequences, an agent is simply the thing taking actions: it observes a situation, chooses an action, and receives the world back in a changed state. The word carried no implication of sophistication. A thermostat fits the definition, and textbooks have said so for decades.
A parallel tradition in software used it for a program acting on somebody’s behalf with some autonomy, persistence and ability to communicate with other such programs. That usage produced a substantial research literature and a certain amount of enthusiasm in the 1990s that did not translate into deployed systems on the scale predicted.
Both meanings are older than anything currently being sold under the name. Neither requires a language model, and both describe an architectural arrangement rather than a level of capability.
What the word covers now
Current usage generally describes a language model placed inside a loop, given the ability to call external functions, and permitted to continue until some condition is met. It receives a goal, produces an action, sees the result, and decides what to do next, repeating until it stops or is stopped.
The components are unremarkable individually. What is new is the arrangement: the model is choosing its own next step rather than producing a single response, and the number of steps is not fixed in advance. That is a real change in how the system is used and it is the substance behind the term.
Everything else attached to the word is variation. Some systems plan explicitly before acting; some maintain notes across steps; some spawn additional instances to handle subtasks; some are constrained to a fixed set of permitted sequences. These are different designs sharing a label.
Autonomy is a spectrum and the label hides where a system sits
The useful question about any such system is what it can do without asking. One that drafts an action and waits for approval is a different proposition from one that executes and reports afterwards, and both are commonly described the same way. The difference is the entire risk profile.
A second question is how far it can go before something checks it. A system permitted three steps within a narrow set of operations is a different object from one permitted to run indefinitely across whatever tools it can reach. Both are agents by any current definition.
A third is what happens when it is wrong. Reversible actions and irreversible ones differ absolutely, and a system permitted only the former can fail repeatedly at acceptable cost. Much of the sound engineering in this area consists of arranging that boundary carefully.
The word is doing commercial work
Describing a product as agentic signals capability without committing to any specific claim, which is convenient. A great many systems marketed this way are conventional workflows with a model in one step, which may be perfectly good products and are not what the word implies.
This is a familiar cycle. A technical term acquires favourable associations, spreads to cover things it did not originally describe, and becomes useless for distinguishing anything. The same happened to several earlier terms in this field, and there is no particular reason to expect a different outcome here.
The practical response is to ignore the label and ask about the loop, the permissions and the checks. Those questions have concrete answers, and a vendor unable to answer them is telling you something.
Why the arrangement is genuinely harder than a single response
A system taking many steps compounds whatever error rate it has per step, and errors made early alter what it observes later. This arithmetic is unforgiving in a way that is easy to underestimate from watching a successful demonstration, and it is the main reason these systems perform worse in extended use than short trials suggest.
There is also an exposure question. A system reading external content and choosing actions on that basis has an attack surface consisting of everything it might read, which is a structural property of the arrangement rather than a configuration mistake.
There is a third difficulty that gets less attention. When a sequence of steps goes wrong, working out which one caused the eventual failure is genuinely hard, because each step was reasonable given what the system believed at the time and the belief was formed by an earlier step. Debugging a chain of judgements is a different activity from debugging a program.
None of this makes the approach unsound. It makes the scope of what is permitted the most important design decision, which is precisely the information the word obscures.
Common questions
Is a system that calls one tool an agent?
Under most current definitions, only marginally. The distinguishing feature is usually taken to be an open-ended loop where the system decides how many steps to take. A single tool call within a fixed sequence is a conventional program with a model in it.
Do these systems set their own goals?
No. The goal is supplied, and what is delegated is the choice of steps towards it. Systems that generate their own subgoals are doing so within a supplied objective, which is a meaningful difference however open-ended the behaviour looks.
Why do demonstrations look better than deployments?
Demonstrations are usually short, chosen, and run in prepared conditions. Extended sequences in unprepared conditions accumulate errors and encounter situations the design did not anticipate, and neither of those is visible in a recorded example.
Contributing editor, AI Worth Knowing
Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.





