The flagship path
Production agent systems.
Take one agent system from a written brief to a deployed system you have broken on purpose, and leave behind evidence a stranger can check.
9 stages. Each ends in an artifact a stranger can open and an eval that can be re-run. The order is the order production forces the decisions on you.
- 01
System brief
Whose job is this and what would make you stop building it.
- 02
Architecture decision
Which shape, and which credible alternative you are rejecting and why.
- 03
Tool and authority model
What the agent may do, as what principal, and how you take it back.
- 04
Eval harness
What a regression looks like, before you are allowed to ship a change.
- 05
Threat model
Which inputs you do not trust, and which risks you are consciously accepting.
- 06
Cost model
What one unit of work costs and what stops the bill when it runs away.
- 07
Deployment
Put it somewhere a stranger can reach, with a rollback you have tested.
- 08
Incident simulation
Break it on purpose and find out whether your instrumentation notices.
- 09
Public-safe portfolio proof
What you can show publicly without leaking an employer or a customer.
What the path grants, and what it does not
Ship a production agent system
Can take an agent system from brief to deployment with a bounded authority model, a harness that fails on regression, a stated cost ceiling, and an incident their own telemetry detected.
This is not a certification, an accreditation, or a statement about employability. It records that specific artifacts passed specific evals.