← All articles
Work & responsibility4 min read

The Work AI Leaves Behind

An agent can finish its part of a task while leaving the organisation with more work to do.

Imagine arriving at work to find that your AI agents have handled a hundred tasks overnight.

Eighty are marked complete. Fifteen need approval. Five are blocked.

That looks like a productive night. But before celebrating, someone needs to answer a few questions.

How many of the eighty were checked against the result the business actually needed? How long will the fifteen approvals take? Who owns the five blocked cases? And does the report include work that was scheduled but never started?

The dashboard tells you what the agents recorded. You still need to establish what happened to the work.

This is the question I keep returning to in conversations about digital labour: how much human work remains after the agent says it has finished?

There is plenty of attention on how quickly agents can produce code, research, documents and decisions. The work around those outputs deserves equal attention: reviewing, correcting, reconciling, chasing missing results and deciding what happens next.

Those activities belong in the productivity calculation.

Consider an agent that prepares a customer response. It produces a fluent draft in seconds. A person then checks the account history, discovers a missing exception, corrects a commitment and asks for another version. The draft may still be useful. Whether it saves time depends on the effort required to reach a response that can actually be sent.

The same applies to software. Generating a patch is one part of completing a change. Someone must establish that it meets the requirement, behaves correctly in its surroundings and can be maintained. If nobody understands it well enough to investigate a later failure, some of the work has been postponed.

This is why the distinction between output and accepted outcome matters.

Recent discussions have made that distinction more tangible. In Sebastian Mueller’s account of agent operating costs, human review has its own cost line, and measured figures are separated from estimates. That is a useful discipline: a business case should make its assumptions visible enough to challenge.

I would also want to know how the review work arrives. A manageable weekly average can conceal a difficult Monday morning. Ten straightforward checks and ten ambiguous cases are very different demands on the same person.

An organisation can create an impressive volume of drafts and still build a growing queue of unresolved decisions.

The practical limit may be the attention available to deal with that queue.

Our experimental work on Agent Control has brought me back to the same issue from the execution side. In one small, bounded comparison involving three local AI workers, both brokered routing and the cheapest routing strategy verified nine of twelve outcomes. Failed attempts and retries remained in the accounting. The conclusion on monetary savings was inconclusive.

That was a limited experiment, not evidence of production-wide savings. Its value was in forcing us to keep the unfinished work visible. A lower resource total would have meant little if it had been achieved by quietly completing fewer tasks.

For an organisation adopting agents, I think that translates into a fairly practical set of questions:

  • What was the agreed outcome, and what evidence shows it was achieved?
  • How much human time went into review, correction and recovery?
  • Which cases remain unresolved, how old are they, and who owns the next action?
  • What happened to the capacity that was genuinely released?

That last question matters. Time saved can become better service, more learning, shorter hours or additional work. The destination should be an explicit organisational choice. A claim of “hours saved” tells us very little about whether people’s working lives improved.

There is also a temptation to reduce the visible review burden by making approval easier. A faster click does not necessarily mean a better process. If the person lacks the context or time to judge the result, the queue may appear to clear while uncertainty moves further downstream.

The better response may be to narrow the task, improve the evidence, handle a predictable exception in the workflow, or return a responsibility to a person until the system can support it properly.

Sometimes a useful agent should stop and ask. That pause should come with enough information for someone to act: what was attempted, what remains uncertain, and what decision is needed. Otherwise the human has to reconstruct the whole task before they can help.

This changes what I would want from an agent dashboard. Alongside completed tasks, I would want accepted outcomes, unresolved cases, repeat reviews and human time spent putting things right. A task that safely stopped should be visible as such. A task that never started should not disappear.

Agents could give people substantial capacity back. Demonstrating that benefit requires following the work far enough to see where the remaining effort lands.

When your agents finish, is there less work left for people—or just a different queue?

Also published on Medium ↗.

This article reflects the evidence and development status at its original publication date.