AI Agents Need an Undo Button

Imagine giving an AI agent access to production.
It can restart a Kubernetes workload. Change an IAM policy. Modify infrastructure. Patch a vulnerability. Update a database. Open and merge a pull request.
Now imagine it is wrong.
Not obviously wrong. Just wrong enough.
It restarts the wrong service.
It changes a security group that another system depends on.
It removes a resource that looked unused but was actually part of a critical recovery path.
This is the point where AI stops being a chatbot problem.
It becomes an operations problem.
And operations has always had one uncomfortable question:
How do we recover when something goes wrong?
AI is moving from answers to actions
Most AI systems were originally designed to generate things: text, code, analysis, recommendations.
That is relatively forgiving.
If an AI gives you a bad answer, you ignore it.
If it writes bad code, someone can review it.
But agents are increasingly being connected directly to tools, APIs, cloud environments, Kubernetes clusters, databases, and CI/CD systems.
They are no longer just recommending actions.
They are taking them.
That changes everything.
The important question is no longer:
“Was the model correct?”
It becomes:
“Was this action safe to execute?”
An approval button is not a safety system
The obvious answer is to put a human in the loop.
The AI proposes a change.
A human clicks Approve.
Problem solved.
Except it is not.
Before approving a production change, the person needs to understand:
What exactly will change?
What depends on it?
What could break?
How large is the blast radius?
Can we reverse it?
How do we know the change actually worked?
If the interface only shows:
Restart deployment?
Approve / Reject
then the human is not really approving the change.
They are approving a guess.
The hard part is not the button.
The hard part is giving both the agent and the human enough context to understand the consequences.
Production needs a safety layer
We think AI agents that touch production need a layer between intent and execution.
Something like:
Intent → Context → Plan → Policy → Approval → Execute → Verify → Recover
The exact implementation will vary.
But the principles should not.
Before execution, the system should understand what the agent is trying to change and what might be affected.
It should evaluate the action against policy.
Risky actions should require additional approval.
After execution, it should verify that the intended outcome actually happened.
And if something goes wrong, there should be a path back.
That last part matters.
A lot.
Because production systems are not deterministic enough to assume every AI-generated action will work exactly as planned.
AI agents need an undo button.
Undo is not Ctrl+Z
Of course, production rollback is much harder than undoing a document edit.
Imagine an AI agent changes a Kubernetes deployment from five replicas to two.
Rolling back might be simple.
Now imagine it changes an IAM policy, triggers a database migration, or deletes a cloud resource with downstream dependencies.
There may be no clean inverse operation.
The environment may have changed since the original action.
Other systems may already depend on the new state.
So “undo” cannot simply mean:
Run the opposite command.
A real recovery mechanism needs context.
It needs to understand:
what the state was before the action
what changed
what depends on that change
whether the action is reversible
what recovery options exist
how to verify the system after recovery
Sometimes the right answer is automatic rollback.
Sometimes it is a proposed recovery plan requiring human approval.
And sometimes the correct decision is:
Do not execute this action automatically at all.
That is also a successful outcome.
The missing part of AI governance
Most conversations about enterprise AI governance focus on questions such as:
Who can access which model?
What data can the model see?
Where is data stored?
What gets logged?
Those questions matter.
But they mostly govern what AI can know and say.
Agents introduce another problem:
What is AI allowed to do?
Once an agent can change real systems, governance has to extend to actions.
That means understanding:
what the agent can access
which actions it can execute
when approval is required
what the blast radius might be
whether the action is reversible
how the outcome will be verified
how the action will be audited
what happens when execution fails
This is not just model governance.
It is execution governance.
The agent does not need to be yours
There is another important implication.
Enterprises are unlikely to standardize on one AI agent.
Teams may use Claude, Codex, Gemini, internal models, open-source agents, or tools that do not exist yet.
The safety layer therefore should not depend on owning the agent.
The agent should be replaceable.
The policies, context, approvals, execution controls, audit trail, and recovery mechanisms should remain.
That is the approach we are taking at Aokumo.
We are building a governed execution layer between AI agents and production infrastructure.
The goal is not to build another agent.
It is to make agents safer when they start touching systems that actually matter.
Before an action reaches production, Aokumo can understand the environment around it.
During execution, it can apply policy and approval controls.
After execution, it can verify what happened.
And when something goes wrong, it can help determine the safest path back.
Because the future of AI in enterprise IT will not be determined only by how intelligent agents become.
It will also depend on whether enterprises can trust them to act.
And trust requires more than an Approve button.
It requires a way back.





