Skip to content
AI Architect Academy

Production agent systems · stage 08 of 09

Incident simulation

Break it on purpose and find out whether your instrumentation notices.

The module

Lesson, exercise and rubric

Open the repository in a coding agent and run /module 08 for a Socratic session over this stage.

What you leave behind · markdown

Incident simulation report

Required sections:

  • injected failure mode
  • time to detection
  • what the telemetry showed
  • fix
  • guard added

Cohort-visible only; never published.

The eval · eval:incident-detected-by-telemetry

The incident was caught by instrumentation, not by reading the code

  1. injected-mode-namedThe report names which failure mode was injected. Checked by: The injected failure mode matches a FailureMode id with injectable = true.
  2. detection-source-is-telemetryDetection came from a signal that exists in the deployed system. Checked by: The telemetry section cites a log, metric, or eval that predates the injection.
  3. guard-addedA new guard or assertion was added so the same failure fails loudly next time. Checked by: The guard section links to a commit adding an assertion or check.

Every check must pass. There is no partial credit and no override.

Independent review · the author may never review their own work

Incident report honesty check

  • The detection signal existed before the injection, not added afterwards to look good.
  • Time to detection is measured, not estimated.
  • The guard would fail on a replay of the original injection.

Evidence that counts

Proof has to exist outside your own claim.

At least 2 of: a running deployment a third party can reach; a sign-off from a named reviewer who is not you. Evidence older than 90 days is stale and does not count.