Execution Audit: Logs Prove Activity. Evidence Proves Control.

Why AI agent execution needs evidence, not just activity logs
“We keep all the logs.”
It is one of the most common answers in audit fieldwork. It is usually offered with complete confidence, and it rarely answers the auditor’s actual question.
Enterprises are now preparing to give the same answer about AI agents.
That answer will fail for the same reason it has always failed: logs can show that an event occurred without proving that the event was properly authorized, controlled, and attributable.
The difference is scale. A human operations team might perform dozens of consequential changes in a week. A fleet of agents could perform hundreds in a day.
Audit evidence assembled manually at human speed will not survive execution at machine speed.
The question logs do not answer
When an auditor tests a production change, the question is not merely whether something happened.
The question is whether it was permitted.
Who or what initiated the action? Was approval required? Who approved it? Which policy applied? What was the state before the change? Depending on the control being tested, was a recovery path available?
Operational logs rarely carry all of that context in one reliable record.
The action may appear in a cloud audit log. The approval may live in a workflow system. The change request may be stored in an ITSM platform. The relevant policy may have changed since the action occurred. The pre-change state may not have been captured at all.
Connecting those pieces becomes manual work, usually performed under deadline by the engineering, security, or compliance team responding to the audit request.
Every sampled action creates another reconstruction exercise.
That approach was expensive when people made the changes. It becomes unmanageable when agents execute continuously.
What auditors are required to establish
This distinction deserves precision.
ISA 500, the international standard addressing audit evidence, does not instruct auditors to reject information produced by an organization’s systems. It requires them to determine whether that information is sufficiently reliable for the intended audit purpose.
When the auditor uses information produced by the entity, that assessment includes obtaining evidence about its accuracy and completeness. Japan’s Audit Standard Report 500 follows the same underlying discipline.
In practice, this creates three recurring questions.
Is the record complete?
A log contains what the organization configured the system to collect, for as long as it retained that information.
Collectors can fail. Retention periods can expire. Logging levels can change. Systems can remain outside the logging boundary. An integration can stop forwarding events without anyone noticing.
Before a log-derived listing can represent the population of actions, the organization may first need to demonstrate that the logging configuration and collection path remained effective throughout the relevant period.
The existence of many log entries does not prove that no entries are missing.
Is the record reliable?
The auditor must also consider the circumstances under which the information was produced and maintained.
Who can alter or delete it? Are administrative actions recorded? Is retention enforced? Can timestamps, identities, or event values be changed? What controls protect the system that stores the records?
This is why a log’s evidentiary value does not come from its content alone. It also depends on the controls surrounding the system that produced and preserved it.
Does the record connect the action to authority?
This is the largest gap.
Most operational logging was designed to support debugging, security investigation, and system administration. It records activity. It does not necessarily preserve the complete relationship between an action and the organizational authority that permitted it.
A control test may need to connect:
the proposed action;
the identity that initiated it;
the policy in force at the time;
the approval decision, when one was required;
the state before execution;
the observed result;
the available recovery position.
If those elements live in separate systems, someone must join them after the fact and demonstrate that the linkage is correct.
The problem is not that logs are useless. Logs are essential.
They are simply not the whole answer.
Logs record activity. Evidence connects it to authority.
“Audit log” and “audit evidence” are not opposing terms defined by a regulator. This is our framing, based on a practical distinction that audit teams already encounter.
A log records an event.
Execution evidence connects that event to the control that allowed it to occur.
That evidence should not be assembled for the first time when an auditor requests it. It should be produced as part of the execution itself.
When an action passes through a governed execution layer, that layer already has much of the context an auditor may later need:
what the agent proposed;
which identity proposed it;
which policy was evaluated;
whether human approval was required;
who supplied that approval;
what recovery position was established;
what the target system reported afterward.
If the execution path preserves those elements as one connected record, the audit posture changes.
The organization no longer begins with an event and searches backward for its justification. The action and its authority are linked from the start.
Evidence is produced, not reconstructed
This is not simply better logging. It changes where the burden of proof lives.
With conventional operational logs, the organization must substantiate the record after the event. It must locate the related approval, identify the applicable policy, establish the relevant state, and demonstrate that the pieces belong together.
With governed execution, those substantiating elements are captured as part of the path through which the action occurs.
If every in-scope execution is required to pass through that path, its records can form the governed population rather than a listing reconstructed afterward.
That does not eliminate the auditor’s work. The organization must still demonstrate that the execution layer itself is appropriately controlled. It must establish coverage, retention, integrity, access control, and the absence or management of bypass paths.
It does, however, move the problem to one system deliberately designed to answer it.
Completeness becomes a question of execution-path coverage rather than a search across every operational tool.
Accuracy becomes a question of how the evidence record is produced and protected rather than a manual reconciliation between unrelated systems.
Authorization becomes an explicit part of the record instead of an inference drawn from an event log.
That is a meaningful reduction in audit effort.
Evidence does not guarantee acceptance
There is an important limit to this argument.
No software vendor can promise that an auditor will accept a particular record as sufficient evidence.
The sufficiency and appropriateness of audit evidence remain matters of professional judgment. They depend on the audit objective, the control being tested, the assessed risk, and the circumstances of the engagement.
A product that promises “automatic compliance” is promising a conclusion that belongs to someone else.
The defensible claim is narrower and still valuable:
A record produced with its identity, authorization, policy context, observed result, and relevant recovery information can reduce the cost and uncertainty of the procedures needed to test it.
The goal is not to replace the auditor’s judgment.
It is to stop wasting that judgment on archaeological work.
The deadline behind the deadline
Regulation is reinforcing the importance of traceability.
For systems classified as high-risk, Article 12 of the EU AI Act requires technical capabilities for automatically recording relevant events over the system’s lifetime. The purpose is to support traceability, operational monitoring, and post-market oversight.
Article 12 does not prescribe Aokumo’s model of execution evidence. It does demonstrate that “the system produced some logs” is no longer the end of the traceability discussion.
Japan’s internal-control regime has imposed a similar evidentiary discipline on IT-dependent controls for years. Organizations must demonstrate not only that controls were designed, but that they operated during the period under review.
These requirements were not written specifically for infrastructure agents. They do not need to be.
From an auditor’s perspective, an agent making consequential changes is another execution actor. It happens to operate faster, more frequently, and with less direct human involvement than the actors the control environment was designed around.
The real deadline is therefore not only a date in a regulation.
It is the first reporting period in which agent-driven changes become material to an audit.
When that period closes, “we keep all the logs” can turn a routine evidence request into weeks of reconstruction.
Build the evidence into the action
The organizations that will handle that audit calmly are not necessarily the ones with the largest logging platforms.
They are the ones that can answer, for every consequential agent action:
What happened?
Why was it permitted?
Which authority applied?
Who approved it, if approval was required?
What result was independently observed?
What recovery position existed when it ran?
Those answers should come from the execution record, not from a meeting assembled six months later.
As agents perform more production work, the distinction between observing activity and demonstrating control will become impossible to ignore.
Logs show what happened.
Execution evidence shows why the organization allowed it to happen.
Aokumo provides governed execution for AI agents operating cloud and Kubernetes environments. It helps enterprises apply enforceable policy, risk-based approval, recovery controls, independent verification, and audit evidence as part of the production execution path.
The agent acts. Aokumo preserves the authority and evidence behind the action.
Sources





