Software flagshipVerified synthetic testbed · v0.0.20Real infrastructure · disconnected

Runbook Sentinel · Software control trace

The model is not the control plane.

I designed and implemented a synthetic incident-response testbed in which retrieved text and model output can inform a structured proposal, while separate approval and fixed software checks retain every state-changing decision.

Inspect the pinned v0.0.20 sourcef149ac24…
Problem
Keep an evidence-retrieving synthetic incident agent bounded when evidence is unreliable—without letting retrieved text or model output authorize change.
Intended reviewer
Software and reliability reviewers assessing a bounded control architecture before any real-infrastructure connection.
My role
Repository author and release owner; designed the authority separation, implemented the runtime and surfaces, and authored the evaluation and release checks.
Control system
Dependency-free synthetic control system with bounded agent outcomes, structured proposals, separate approval and policy checks, synthetic state, and chained event logs.
Result
At pinned v0.0.20, 93 of 93 predefined synthetic attempts matched expected paths and final state; 9 of 84 tested-model outputs passed the required structure, so the candidate was excluded.
Limit
Synthetic fixtures and state only. No real-system connectors, arbitrary shell, production reliability, adoption, or operational-impact claim.

Reading boundary. Software-control metaphor; no electrical or hardware implementation is claimed.

Case indexJump to a chapter

Movement I · Authority break

A proposal has no power by itself.

Retrieved text is treated as untrusted evidence and cannot directly approve or execute an action. Signal and authority travel on separate, inspectable software paths.

Signal railMay inform
  1. 01

    Untrusted evidence

    Fresh content stays distinguishable from stale identity and untrusted guidance.

  2. 02

    Bounded agent

    Diagnose, request evidence, propose one predefined test action, or abstain.

  3. 03

    Structured proposal

    A typed action request with no approval or execution authority.

Proposal alone: no state-change authority.
Authority railMay authorize synthetic change
  1. 01

    Separate approval

    A project-specific launch-scoped loopback credential—not proof of human identity.

  2. 02

    Fixed software checks

    Policy, arguments, replay, one-use approval, repeated-request, and state checks.

  3. 03

    Synthetic executor

    Only restart worker, roll back deployment, or warm cache in repository-local state.

The bounded diagnosis, proposal, and read interface has no approval or execution tool. Repeating an already consumed request cannot cause a second change; state is checked before and after the synthetic action.

Movement II · Candidate rejected

The tested model failed the fixed contract.

In one source-gated local comparison, a tested three-billion-parameter local model produced only 9 valid structured outputs from 84 attempts. A failed component is not a trophy; it is a configuration decision.

Decision trip

What failed. The local 3B model produced only 9 valid outputs from 84 attempts; 75 failed schema validation and latency was 213.394 times the control.

84 fixed local outputs

Valid structured output
9
Invalid diagnosis identifier
67
Invalid action argument
7
Evidence outside context
1
Total inspected
84
System changed

Exclude the candidate, retain deterministic control, and keep invalid model output outside proposal and execution authority.

What the selection supports

The selection process visibly rejects a weaker candidate instead of treating model inclusion as progress.

Still not established

Not evidence that the model was useful or safe, or that zero executed attacks proves universal resistance.

Movement III · Release progression

The failures rewrote the gate.

Two apparently passing records concealed different weaknesses. One changed what the evaluator requires; the other changed how the event log detects alteration and resumes.

Coverage trip

Headline coverage hid a held-out gap

What failed. Headline 3-of-3 action coverage hid that held-out tests never exercised deployment rollback; split-aware coverage was 5 of 6.

  1. 01

    Headline view

    3 / 3All three permitted actions appeared somewhere in the catalog.
  2. 02

    Split-aware review

    5 / 6Held-out tests never exercised deployment rollback.
  3. 03

    Evaluator changed

    + 1 caseA predefined held-out rollback case was added; any missing development-or-held-out combination now fails the gate.
  4. 04

    Earned result

    6 / 6Each action was covered in both development and held-out cases across 31 fixed cases and three trials.
Evaluator changed

Add one frozen held-out rollback case and fail the gate for any missing action-and-split pair.

What the gate now supports

All six action-and-split pairs are covered across 31 cases and three trials.

Still not established. Not production reliability; 31 cases remain below the separate 48-case target.

Trace trip

A passing event trace could be altered

What failed. A success value in a 150-event trace could be changed without breaking parsing, inspection, or release status.

  1. 150-event trace allowed an undetected value change
  2. 10 integrity cases
  3. preceding-event links + final anchor
  4. 165 contiguous events
Event interface changed

Freeze ten integrity cases, chain each event, bind evaluation to the final anchor, and fail closed on incomplete resume.

What continuity now supports

The selected v0.0.20 trace has 165 contiguous chained events and an exact final anchor.

Still not established. No writer authentication, hostile-writer resistance, immutable storage, non-repudiation, or digital signature.

Movement IV · Proof room

What this release proves—and what it cannot.

The recruiter-scale story ends here. The exact release receipt, rendered checkpoint, source identities, and hard limits remain attached for deeper review.

Fixed cases
31 × 3 trials
Expected paths + final states
93 / 93 matched
Action / no-action
36 / 57
Action coverage
6 / 6 across development + held-out
Selected trace
165 linked events
Real systems
0 connected

Rendered checkpoint · supporting evidence

Rendered checkpoint, not a live console.

baseline-0020
Runbook Sentinel dashboard screenshot showing the frozen evaluation pass, exact test metrics, a launch-scoped loopback approval boundary, and real infrastructure disconnected.

The frozen dashboard corroborates the selected evaluation, exact test metrics, separate approval boundary, and disconnected real infrastructure. It is not an operations surface.

Limits that travel with the result

  • Synthetic fixtures and state only; zero real systems, connectors, arbitrary shell, or operational adapters
  • No production reliability, adoption, operational impact, or autonomous-remediation claim
  • The loopback capability does not prove human presence, OAuth identity, or enterprise authorization
  • The unkeyed trace proves bounded continuity, not writer identity or immutable storage
  • One local-model comparison is not a general benchmark or safety result
  • Docker, package-registry, hardware, and energy-cost claims remain excluded
  • At v0.0.20, 31 fixed cases remained below the separately declared 48-case threshold.

Inspect the frozen record

Rejected-model comparison9 / 84 passed structure
Coverage-gap record5 / 6 → 6 / 6
Trace-integrity gapmutation probe