Forensics and evidence for AI agents. Two days on capturing, preserving and reconstructing what an AI system did, across local, hosted and frontier models, RAG, agents, and agents that move money. You leave able to stand up a forensic capability in-house.
Autonomous agents move money, delete data and call tools with no human in the loop. When one causes harm, you have to reconstruct exactly what happened and prove it. But an agent's most important evidence, the context and reasoning behind the action, is volatile and gone the moment the process ends. You cannot investigate what nothing recorded.
Every layer leaves different evidence, and is captured differently.
What you can capture when you run the model: prompts, outputs, the serving stack.
Working from the gateway view when the provider runs the model.
Retrieval evidence: what was retrieved, from where, and why.
Action evidence: every tool call, decision and side effect.
Payment and AML evidence: receipts, sanctions checks, non-repudiation.
Continuous, signed, tamper-evident logging so the evidence exists later.
Classes are custom built from the following learning modules. Instructors select ~16 hours for a 2-day delivery. Custom formats from 1 to 3 days available. (Times are approximate.)
Agents make decisions, call tools, delegate and act across systems. Traditional timelines assume a human operator. What changes when the actor is autonomous.
Where agent evidence lives: LLM call logs, tool execution records, memory stores, delegation chains, payment receipts, model state and infrastructure telemetry.
Agent context windows, ephemeral memory and container restarts destroy evidence fast. What to grab first and how.
Chain of custody for agent evidence. Admissibility standards (UK, EU, US). What regulators expect when an AI system is involved in a breach or harm event.
Hands on: extract and preserve full prompt-completion histories from model serving infrastructure. Timestamps, token counts, model version and request metadata.
Hands on: capture MCP tool call records, function invocations and external API calls. Correlating tool calls to the prompt that triggered them.
Hands on: reconstruct why an agent chose a particular action. Planning traces, reasoning steps and the chain from user intent to executed action.
Extracting vector store contents, conversation history, scratchpad state and retrieved documents. What the agent "knew" at the time of the incident.
Hands on: verify hash-chained receipt logs. Check signature validity, detect gaps or tampering, and reconstruct the sequence from the cryptographic evidence.
Container logs, DNS queries, network flows and API gateway records. The infrastructure layer that corroborates or contradicts the agent-level evidence.
Hands on: merge evidence from multiple sources into a single chronological timeline. Correlate LLM calls, tool executions, payments and external interactions.
When multiple agents are involved: tracing delegation chains, identifying which agent made each decision and where authority transferred.
What should the agent have done vs what it did. Identifying the point of deviation and whether it was adversarial input, misconfiguration or model behaviour.
Was this the agent acting autonomously, a prompt injection, a poisoned tool response or a human using the agent as a proxy? Techniques for distinguishing the source.
Hands on: package evidence with cryptographic hashes, timestamps and chain-of-custody metadata. Evidence bundles a third party can independently verify.
Report structure for agent incidents: executive summary, timeline, evidence inventory, analysis, conclusions and recommendations.
Translating agent investigation findings into language a board, regulator or judge can follow.
Turning forensic findings into control improvements. Feeding back into AISVS compliance.
Hands on: build a runbook your team can follow when an agent incident occurs. First responder actions, escalation criteria, evidence checklists and tool lists.
What tools to deploy now so evidence exists when you need it: logging requirements, receipt infrastructure, retention policies and storage.
What your IR team needs to know about agents that they do not know today. Bridging the gap between traditional IR and agentic IR.
A simulated agent incident. Collect evidence from LLM logs, tool records, receipts and infrastructure. Preserve it forensically. Present your findings.
Three agents, two incidents, one timeline. Merge evidence from multiple sources, identify the root cause and attribute the behaviour.
An agent audit trail has been tampered with. Detect the tampering, identify what was changed, and determine what the original evidence said.
Full mock investigation from alert to report. Collect, preserve, analyse, reconstruct and write up. Peer review with another team.
AgentBee puts a hardware root of trust on an agent's most critical actions. Every approval, and every forensic seal, is signed with FIDO-grade crypto (ECDSA P-256): a tamper-evident record of who authorised what, anchored in hardware.
In forensics, that hardware root is what makes the chain of custody hold. You learn to use it.
agentbee.co.uk →You leave with a blueprint to stand up an agentic forensics capability inside your own organisation: the flight recorder, the evidence pipeline, the chain of custody, and the playbook for an agent incident. Grounded in the IETF Agent Audit Trail Internet-Draft, authored by your instructor.
Add AI agents to your investigation practice.
Verify that agent evidence exists and holds up.
Build the recording and preservation into the stack.
Govern autonomous action with evidence that survives.
Stand up the capability in-house.
Extend digital forensics to autonomous agents.
Two days, hands-on, taught by the former OWASP-AISVS Co-Leader (v1.0) qualified in cyber-forensics and author of the IETF Agent Audit Trail draft. Places are limited.
Corporate & sovereign cohorts, and bespoke on-site delivery, on request.