AI agent incident response and liability basics
What to do when an agent acts wrong: contain, revoke, notify, fix, and how contracts should assign responsibility.
Written by Dali
Dali is an AI agent systems studio. David leads engineering and product systems; Liana leads operations and workflow fit. We ship production agents inside tools teams already use.
David Hakobyan · Dali
Direct answer
When an agent fails: stop the agent, revoke credentials if needed, preserve traces, fix customer impact, then patch prompts/tools. Contracts should say who is responsible for business decisions vs software defects.
Contain
Kill switch, disable write tools, pause queues.
Investigate
Pull correlation ids, tool args, approvals - this is why observability matters.
Customer fix
Human outreach, refunds, corrections - with a real owner.
Prevent
Eval case from the incident, tighter gates, better idempotency.
Liability
Policies and human approval do not disappear because a model suggested text. Align vendor SOW with who owns outcomes.
SLA
Define response times for agent-impacted queues just like other production systems.
How Dali fits
Dali pilots define stop-switches and handoff owners: solutions.
FAQ
Follow law and brand policy; honesty often reduces damage when errors happen.