When AI takes on real work, who remains responsible?
An AI agent can complete a task and still create an operational failure.
Imagine asking an agent to update a customer record. It finds a record, changes the requested field and reports success. Every tool call works. The change is recorded in a log.
But it has updated the wrong customer.
The important question is whether the surrounding system checked which record the agent was authorised to change, prevented it from crossing that boundary, and verified the intended outcome.
That is why I believe systems like Tsunora could become important as organisations delegate more work to AI.
Capability needs an operating framework
A useful way to think about an agent is as a capability to which we can delegate a bounded task. Calling it a “digital worker” does not give it human judgement, legal responsibility or an unrestricted mandate.
Before delegating, an organisation should be able to answer some ordinary operational questions:
- What exactly is the work, and what counts as successful completion?
- Which systems and records may be accessed or changed?
- Who owns the outcome and approves consequential actions?
- What is the spending limit, including retries and review?
- What happens if the worker fails, becomes unavailable or exceeds its authority?
These questions should have answers that the system can enforce. Putting them in a prompt is useful, but a prompt alone is not an access boundary.
My interest in this comes from HR, payroll and workforce technology. Assigning work has always involved more than finding someone capable of doing it. It also involves scope, permissions, supervision, handovers and accountability. Those disciplines remain relevant when software performs part of the work.
Completion needs evidence
An agent saying “done” is a claim to assess.
For the customer-record example, completion could require confirming the intended customer's identity, checking that only the authorised field changed, and retaining evidence of the final state. A successful tool response would be one piece of evidence, not the whole conclusion.
The checks also need protection. If the same agent can freely rewrite its acceptance criteria, passing them tells us less than it appears to. Independent checking can help, but a second model is not automatically an independent or reliable judge. The verifier itself needs appropriate tests and limits.
Oversight must be usable, too. A system that asks a person to approve every minor step could simply move the workload into an approval queue. I would want routine, bounded work to proceed within tested limits, with meaningful decisions escalated to someone who has the context to make them.
The economics must include the difficult parts
The cost of digital work is more than the price of a model call.
It includes unsuccessful attempts, verification, human review, integration, recovery and ongoing maintenance. A cheap attempt can produce an expensive outcome if somebody has to reconstruct what happened and repair it.
For me, a more useful question is: what does an accepted outcome cost, at the quality and level of risk the organisation requires?
Answering that requires records connecting the task, the worker, the resources consumed and the result. Even then, lower cost is something to demonstrate against a fair baseline. It should not be assumed from the presence of automation.
Where Tsunora fits
Tsunora is my experimental exploration of a governed operating environment for digital work, derived from Agent Control. Its public design connects scoped work with digital-worker capability, bounded authority, verification, evidence and internal accounting. Tsunora source and documentation
The status matters. The current workforce prototype is research software, and its physical demonstration uses deterministic workers. Production authentication, enterprise integrations and model-backed workforce execution remain separately unqualified. That evidence does not establish readiness for production HR or payroll operations. Current qualification boundaries
The larger proposition is that organisations will need these operating capabilities as delegation expands. They might be supplied by existing platforms, specialist systems or a combination. A separate product is not inevitable; the requirement for clear responsibility is the part I expect to endure.
Why this matters beyond the technology team
If digital work becomes a material part of an organisation's capacity, its effects will concern finance, HR, operations and leadership as well as IT.
What work was actually completed? What did it cost? How much human intervention remained? Did it improve service, reduce errors, or merely increase activity?
Credible answers would also improve the wider debate about productivity, employment and who benefits from AI. Operational records cannot settle questions of tax or distribution, but they can help distinguish measured outcomes from promotional claims.
That is the future I want to explore through Tsunora: organisations able to delegate useful work with clear boundaries, retain evidence, and intervene when needed.
Before an AI agent is allowed to act on your organisation's behalf, what evidence would you need to trust the arrangement?
Also published on LinkedIn ↗.
This article reflects the evidence and development status at its original publication date.