Runbook Sentinel · Software control trace
The model is not the control plane.
I designed and implemented a synthetic incident-response testbed in which retrieved text and model output can inform a structured proposal, while separate approval and fixed software checks retain every state-changing decision.
Inspect the pinned v0.0.20 source —f149ac24…- Problem
- Keep an evidence-retrieving synthetic incident agent bounded when evidence is unreliable—without letting retrieved text or model output authorize change.
- Intended reviewer
- Software and reliability reviewers assessing a bounded control architecture before any real-infrastructure connection.
- My role
- Repository author and release owner; designed the authority separation, implemented the runtime and surfaces, and authored the evaluation and release checks.
- Control system
- Dependency-free synthetic control system with bounded agent outcomes, structured proposals, separate approval and policy checks, synthetic state, and chained event logs.
- Result
- At pinned v0.0.20, 93 of 93 predefined synthetic attempts matched expected paths and final state; 9 of 84 tested-model outputs passed the required structure, so the candidate was excluded.
- Limit
- Synthetic fixtures and state only. No real-system connectors, arbitrary shell, production reliability, adoption, or operational-impact claim.
Case indexJump to a chapter
Movement II · Candidate rejected
The tested model failed the fixed contract.
In one source-gated local comparison, a tested three-billion-parameter local model produced only 9 valid structured outputs from 84 attempts. A failed component is not a trophy; it is a configuration decision.
Decision trip
What failed. The local 3B model produced only 9 valid outputs from 84 attempts; 75 failed schema validation and latency was 213.394 times the control.
84 fixed local outputs
- Valid structured output
- 9
- Invalid diagnosis identifier
- 67
- Invalid action argument
- 7
- Evidence outside context
- 1
- Total inspected
- 84
Exclude the candidate, retain deterministic control, and keep invalid model output outside proposal and execution authority.
The selection process visibly rejects a weaker candidate instead of treating model inclusion as progress.
Not evidence that the model was useful or safe, or that zero executed attacks proves universal resistance.
Movement III · Release progression
The failures rewrote the gate.
Two apparently passing records concealed different weaknesses. One changed what the evaluator requires; the other changed how the event log detects alteration and resumes.
Coverage trip
Headline coverage hid a held-out gap
What failed. Headline 3-of-3 action coverage hid that held-out tests never exercised deployment rollback; split-aware coverage was 5 of 6.
- 01
Headline view
3 / 3All three permitted actions appeared somewhere in the catalog. - 02
Split-aware review
5 / 6Held-out tests never exercised deployment rollback. - 03
Evaluator changed
+ 1 caseA predefined held-out rollback case was added; any missing development-or-held-out combination now fails the gate. - 04
Earned result
6 / 6Each action was covered in both development and held-out cases across 31 fixed cases and three trials.
Add one frozen held-out rollback case and fail the gate for any missing action-and-split pair.
All six action-and-split pairs are covered across 31 cases and three trials.
Still not established. Not production reliability; 31 cases remain below the separate 48-case target.
Trace trip
A passing event trace could be altered
What failed. A success value in a 150-event trace could be changed without breaking parsing, inspection, or release status.
- 150-event trace allowed an undetected value change
- 10 integrity cases
- preceding-event links + final anchor
- 165 contiguous events
Freeze ten integrity cases, chain each event, bind evaluation to the final anchor, and fail closed on incomplete resume.
The selected v0.0.20 trace has 165 contiguous chained events and an exact final anchor.
Still not established. No writer authentication, hostile-writer resistance, immutable storage, non-repudiation, or digital signature.
Movement IV · Proof room
What this release proves—and what it cannot.
The recruiter-scale story ends here. The exact release receipt, rendered checkpoint, source identities, and hard limits remain attached for deeper review.
- Fixed cases
- 31 × 3 trials
- Expected paths + final states
- 93 / 93 matched
- Action / no-action
- 36 / 57
- Action coverage
- 6 / 6 across development + held-out
- Selected trace
- 165 linked events
- Real systems
- 0 connected
Rendered checkpoint · supporting evidence
Rendered checkpoint, not a live console.
baseline-0020
The frozen dashboard corroborates the selected evaluation, exact test metrics, separate approval boundary, and disconnected real infrastructure. It is not an operations surface.
Limits that travel with the result
- Synthetic fixtures and state only; zero real systems, connectors, arbitrary shell, or operational adapters
- No production reliability, adoption, operational impact, or autonomous-remediation claim
- The loopback capability does not prove human presence, OAuth identity, or enterprise authorization
- The unkeyed trace proves bounded continuity, not writer identity or immutable storage
- One local-model comparison is not a general benchmark or safety result
- Docker, package-registry, hardware, and energy-cost claims remain excluded
- At v0.0.20, 31 fixed cases remained below the separately declared 48-case threshold.