What agents can and cannot do today
Most AI material you have seen falls into one of two camps. Either everything is possible and you are already behind, or nothing works and it is all a bubble. Neither camp would survive a board meeting. This chapter takes the position a good advisor takes: here is what the technology does reliably today, here is where it fails, and here is a simple rule for telling the two apart in your own company.
A quick anchor before we sort anything. An agent is a digital worker, closer to a new hire you give an email account and instructions to than to software. It works through tasks on its own, deciding step by step what to do next. If that framing is new to you, the chapter What an agent actually is builds it from the ground up. Here we take it as given and ask the practical question: what can this new hire actually do?
What agents are genuinely good at
Agents today are extraordinarily good at one specific shape of work: reading, writing, sorting, comparing, summarizing, and following a procedure with several steps across documents and systems.
That sentence sounds modest. The volume behind it is not. An agent reads 200 supplier invoices in roughly 20 minutes and matches each against the order it belongs to. It screens 200 job applications and writes one short, reasoned note per candidate before your morning coffee has gone cold. It summarizes a 90-page procurement contract in two minutes and lists the five clauses that differ from your standard terms. It drafts the same reply in flawless English and flawless Swedish, and does not care which.
Just as important as what it does is how it works. An agent does not get bored on invoice number 147. It does not lose concentration on a Friday afternoon. It is available at 03:00 when the night shift in your warehouse finds a discrepancy, and it costs the same whether it works at 03:00 or 14:00. It reads every format your company produces: PDF invoices, spreadsheet exports, scanned delivery notes, the email thread where the actual agreement is buried in message fourteen.
For the 120-person logistics firm in Jönköping whose invoice flow opened What an agent actually is, that profile maps onto thousands of hours a year. Not exotic hours. Ordinary ones: the inbox triage, the data entry, the report nobody enjoys assembling, the contract nobody reads until something goes wrong.
- Read 500 pages overnight
- Reconcile two spreadsheets, line by line
- Draft in flawless English and Swedish
- Compare a contract against standard terms
- Know when a customer is truly angry
- Decide which rule to break, and when
- Sense that a deal feels wrong
- Take responsibility for an outcome
Where agents fail, said plainly
Now the part most AI material skips, and the reason this guide is worth your time. Agents have three failure modes, and you need all three on the table before you delegate anything.
They sometimes state falsehoods with full confidence. The technical word is hallucination: when an agent states something false with full confidence, like a new hire who guesses rather than says "I don't know." It will not look uncertain when it happens. The wrong figure arrives in the same fluent, confident prose as the right one. This is not a bug that next year's version quietly removes. It is a property you manage, with spot checks and sign-off rules, the same way you manage a talented junior who has not yet learned to flag their own uncertainty.
They can lose the thread on long, ambiguous tasks. Give an agent a crisp procedure and it follows it for hours without drift. Give it a vague, sprawling assignment, "look into our supplier situation and fix what needs fixing," and somewhere along the way it starts making assumptions you never approved. The cure is the same as with any delegation: a clear brief beats a long leash.
They have no skin in the game. This is the deepest limit and the easiest to forget. An agent cannot be accountable. It cannot apologize to a customer and mean it, cannot weigh its career against a shortcut, cannot stand in front of your board and own a mistake. When something goes wrong, responsibility lands on a person in your organization, exactly as it does when a junior employee makes a mistake. Whoever signs remains the signature.
Treat an agent, in other words, the way you would treat a brilliant new hire with no track record: real talent, zero history, and a probation period before anything important carries its name alone.
The three buckets
Everything above compresses into one management tool. Every task in your company belongs in one of three buckets.
Bucket one: the agent does it alone. Work where the volume is high, the rules are clear, and a wrong answer is cheap to catch. Example: the Jönköping firm's order-status questions. A customer asks where delivery 7042 is, the agent reads the actual system, answers in seconds, and logs what it said. If it gets something wrong, the customer asks again and a human looks. Cost of the error: minutes.
Bucket two: the agent drafts, a human signs. This is the workhorse bucket, and it has a name worth knowing: human-in-the-loop, a working mode where the agent prepares the work and a person approves it before it counts. The four-eyes principle, applied to a digital worker. Example: an accounting practice with 40 people lets an agent draft the monthly client reports. A senior accountant reviews each one in ten minutes instead of building it in three hours, signs, and sends. The client sees the accountant's name, because it is the accountant's judgment that went out the door.
Bucket three: human only. Work where judgment rests on thin information, where the relationship is the product, or where being confidently wrong is expensive. Example: a regional manufacturer renegotiating terms with its largest customer, worth 30 percent of revenue. An agent can prepare every number for that meeting. It does not attend.
Notice what the buckets really are: a mandate, the limits you set on what an agent may do alone, what it must ask about, and what it may never touch. The same idea as an attestation limit for a new employee. You already run this system for people. You are extending it, not inventing it.
The buckets move, and that is the point
Here is the part executives most often miss. The buckets are not a law of nature. They are a management decision, and the right answer changes.
Work that required a human signature in 2024, first-draft customer replies, routine reconciliations, standard contract review, sits comfortably in bucket one for many companies in 2026. The frontier moves every year, in one direction. Which means sorting tasks into buckets is not something you do once at an offsite and frame on the wall. It is a standing item you revisit quarterly, the way you revisit attestation limits when an employee proves themselves. An agent that has drafted 1,000 invoice matches with an error rate below your human baseline has earned a promotion conversation, exactly like a person on probation.
The skill you are building as a leadership team is not technical. It is the oldest skill in management, applied to a new kind of worker: knowing the difference between delegating execution and delegating judgment. Execution is moving fast into buckets one and two. Judgment stays in bucket three longer, and some of it stays there for good.
The rule of thumb you keep
Everything in this chapter compresses into one test:
If the task is high volume, the instructions are clear, and a wrong answer is cheap to catch, it is agent work today.
Three tests, ten seconds, any task in your company. The invoice flow passes all three. The supplier negotiation fails two. The monthly report draft passes with a human signature attached. You can run your entire org chart through this filter on one flight between Stockholm and Malmö, and the chapter Where agents fit does exactly that, function by function.
What you should not take from this chapter is either fantasy. Agents are not about to run your company, and they are not a toy. They are a new category of worker with a sharp, legible profile: superhuman at volume and procedure, unreliable at judgment and accountability. Companies that learn that profile early get to redraw their cost base task by task, calmly, while others are still arguing with the two camps from the first paragraph. What that profile is worth to a company your size, in hours and kronor, is the question What this means for a company like yours puts a number on.
There are three buckets: work an agent does alone, work an agent drafts and we sign, and work that stays human. The buckets move every year, and deciding what goes where is now a management task.