Justina Dong ← Controls

Controls testing · 2026-09-01 – 2026-09-30

Were the controls actually tested?Here is the method, the evidence, and the honest tally.

The controls page says what the 21 controls are. This page is the binder behind it: whether each was tested, how, against what population, with what sample, and with what result — including a worked example that passed and one that did not.

The approach

Six rules shaped every workpaper in the binder behind this page.

Every control is checked against five fixed design attributes — trigger, actor, action, record produced, and frequency — each pointing to the artifact that plays that role, never to a filename.

Each attribute is marked present, partly present, or absent, with the reason stated. A control can be well designed and still fail in operation, or vice versa — the two are scored separately.

Before anything is tested, the full population the control should have covered is defined and counted, from the system's own records.

An automated control that fires on every event gets one re-performed instance plus a configuration-unchanged check; a daily job gets five days sampled across the month; a manual control gets up to five events, or all of them if there are fewer.

Every control lands on exactly one of six conclusion strings, never a free-text verdict.

Effective · Effective with exception(s) · Deficient · Not operated in period · Not testable (policy only) · Design only (no TOE possible). The vocabulary is fixed in advance so a conclusion can't be worded to sound better than the evidence supports.

Every claim of what a control does, or what it found, is checked against the actual artifact — the running code, the log line, the committed history — not against a description of that artifact written somewhere else.

A document describing a safeguard is not the safeguard. Citing the document that makes a claim does not confirm the claim; the underlying mechanism has to be read or re-run directly.

Every automated safeguard this binder relies on was deliberately broken, in a throwaway copy, to prove its safety test actually turns red when something is really wrong — not just when a script runs.

The method: copy the safeguard outside the live system, break it on purpose, confirm the test fails, restore the original, and confirm nothing in the live system changed. A test nobody has ever watched fail is not evidence of anything.

This page is built from the private testing binder by a generator that runs a leak test on its own output, never by redacting the private binder in place.

A fixed list of private, licensed, or confidential paths was never read while building this page. The generator refuses to publish prose that leaks a name, a local path, a session identifier, or other sensitive material.

The lead sheet — all 21 controls

1 Effective · 14 Effective with exception(s) · 3 Not operated in period · 3 Not testable (policy only) · 1 Design only (no TOE possible). 35 exceptions · 5 observations · 168 evidence references in the private binder.

Underlined terms like hook are engineering shorthand; hover, focus, or tap one for a plain-language definition — or switch on Explain the tech above to read the whole page that way. Audit vocabulary (population, sample, exception, and the like) isn't translated; it's assumed familiar.

Area Conclusion

A detective, manual control: the builder re-scores every risk in the register each quarter and on trigger events, compares what is left against a written tolerance, and names a tracked treatment for each risk outside it.

Attributes tested
Trigger
RequiresQuarterly; on a risk realizing; on a new autonomous surface being added
Evidence rolethe risk register's header and its review-cadence section
PresentY
Actor
RequiresThe builder, stated as not independent
Evidence rolethe register's author line and the review entry
PresentY
Action
RequiresRe-score every risk before and after its control; compare to the written tolerance; name a treatment for each out-of-tolerance risk
Evidence rolethe 2026-09-29 review table (14 rows) and the stated-tolerance section
PresentY
Record produced
RequiresA dated review line in the register that sets the next due date; a public version with names and paths removed
Evidence rolethe register's dated review entry and the public governance page's register section
PresentY
Frequency / timing
RequiresQuarterly cadence stated, not assumed
Evidence rolethe register's review-cadence section — which states a different next-due date from the header
PresentY
Population

Every review instance required in the period: scheduled quarterly reviews due 2026-09-01 to 2026-09-30, plus trigger events (a risk realizing; a new autonomous surface added). Reviews performed: 1 (2026-09-29). Trigger events located: 1 (a partial realization of one risk, same day). Scheduled due dates in period: 0. Completeness cannot be fully established for the new-surface trigger, because no dated list of autonomous surfaces exists to enumerate against.

Sample

All: the single review instance of 2026-09-29.

Procedures performed
  1. Re-performed the re-scoring attribute: counted the review table's rows (14 of 14 risks) and confirmed each carries likelihood, impact, inherent and residual scores, a treatment and a tolerance verdict; spot-checked the arithmetic on three rows.
  2. Re-performed the register's own change summary against the table: 3 residual scores fell and 1 rose, matching the control description; re-added the tolerance verdicts: 7 outside tolerance, 3 in the band that needs a detective control and lacks one, 4 accepted — 14.
  3. Traced each of the 7 out-of-tolerance risks to a tracked work item in the owner's open-item registers by title.
  4. Checked the dated review line and its next-due date, and compared the public governance page (13 rows shown, 1 disclosed as withheld) against the private register (14 rows), including a live fetch of the page.
Results (10 pass / 2 exceptions)

10 attributes passed and 2 took exceptions: the review's next-due date, and the treatment tracking for one risk.

Exceptions (2)
  • The register states two next-due dates: 2026-12-29 in its header, review entry and on the public governance page, and 2026-11-07 in its review-cadence section (repeated in the change-management procedure). A quarterly control with two due dates has no single trigger. Severity: deficiency. Root cause: the 2026-09-29 review updated the header and the review entry but not the cadence section written on 2026-08-07.
  • One out-of-tolerance risk (untrusted content reaching an agent that runs with permissions skipped) has a treatment named in the register but no tracked work item located by title. The other 6 out-of-tolerance risks each trace to one. Severity: deficiency.
Observations (1)
  • No enumerable register of autonomous surfaces exists, so the 'new autonomous surface added' trigger cannot be audited for completeness.
Conclusion

Effective with exception(s)

Effective with exception(s) — the control operated once in the period as designed: 14 risks re-scored before and after control, compared to a written tolerance, published with the independence limit stated, and 6 of 7 out-of-tolerance risks traced to a tracked work item. The evidence does not support the 'sets the next due date' attribute unambiguously, or a tracked treatment for one risk; and nothing here tests the quality of the ordinal judgments, which the register itself disclaims.

Follow-ups filed
  • Reconcile the register's next-due date to one value across its header, its review-cadence section and the change-management procedure.
  • File a tracked work item for the one risk whose treatment has none (content-level detection on the messaging ingress and the permissions-skipped runtime); re-test at the next re-score.

9 evidence references cited in the private binder for this control (not shown here — see Provenance below).

A directive, manual control: plain-language procedures describe how a change is proposed, reviewed, built, verified and recorded; how each class of change is rolled back; which actions cannot be rolled back at all; and how an agent, job, hook or script is retired. They are reviewed when the watched-path table changes and quarterly.

Attributes tested
Trigger
RequiresReview on change to the watched-path table; quarterly
Evidence rolethe change-management procedure's review section
PresentY
Actor
RequiresThe builder/owner writing prose procedures
Evidence rolethe change-management procedure's preamble (prose on purpose; the hook wins on disagreement)
PresentY
Action
RequiresDocument propose→review→build→verify→record; rollback by class; irreversible actions; retiring an agent or job
Evidence rolethe change-management procedure (change flow, classes, irreversible list), the back-out procedure (restore by class, what has no back-out), the decommissioning procedure (retiring a job, an agent, a hook or script)
PresentY
Record produced
RequiresProcedure documents with dated review lines
Evidence rolethree procedure documents with a 'written 2026-09-13' date and no review-log slot
PresentN
Frequency / timing
RequiresQuarterly plus on watched-path change
Evidence rolethe change-management procedure's review section
PresentY
Population

Review events required in the period: (a) each change to the watched-path table 2026-09-01 to 2026-09-30 — 0 (the table's file had 2 commits in September; both diffs change the spec-matching logic, not the table); (b) each quarterly register review in period — 1 (2026-09-29). Total required reviews: 1. Complete: the table lives in one version-controlled file, and the register's review log is the only review record.

Sample

All: the single required review (the 2026-09-29 register review).

Procedures performed
  1. Listed the configuration folder and identified the three procedure documents; read each one's header and section list and mapped every element of the control activity to a section.
  2. Confirmed version-control provenance: all three were first committed on 2026-09-14 and have no later commit through 2026-09-30.
  3. Resolved the watched-path table to its real file in the agent-control repository (12 labels), and inspected both September commits to that file for changes to the table itself.
  4. Searched each procedure for a review line dated 2026-09-29 and checked whether the register review or the public governance page records reviewing the procedures.
Results (2 pass / 1 exception)

2 attributes passed (the procedures exist, dated and version-controlled; they cover every element), 1 attribute was not applicable (no watched-path change in period), and 1 took an exception (no review line at the 2026-09-29 register review).

Exceptions (1)
  • The register's quarterly review ran on 2026-09-29. None of the three procedures carries a review line for it, and the change-management procedure still points at the superseded 2026-11-07 date. By the procedure's own trigger a review was due. Mitigating: the documents were 16 days old and nothing in them is known to be stale. Severity: deficiency. Root cause: the 2026-09-29 re-score was run as a publishing step for the public governance page, not as the full quarterly sitting, and the procedures' review was never on that checklist.
Observations (0)

None.

Conclusion

Effective with exception(s)

Effective with exception(s) — the procedures the control requires exist, were written and committed inside the period, and cover every element the register names, including the irreversible-action list and job/agent/hook retirement. The evidence does not support the review attribute: the one review trigger that fired in period left no dated review line, and the documents have no place to record one. The watched-path-change trigger had no instance to test.

Follow-ups filed
  • Add a dated review-log block to the three procedures and record the 2026-09-29 review, or record that it was deliberately skipped.
  • Align the change-management procedure's next-due date with the register.
  • Update the register's own 'tested' field for this control from 'design walkthrough only' to this result.

7 evidence references cited in the private binder for this control (not shown here — see Provenance below).

A directive, manual, point-in-time control: a privacy impact assessment inventories categories of personal data (never the data itself), maps the surfaces data flows to and from, and grades perimeter, minimisation and flow awareness separately.

Attributes tested
Trigger
RequiresPoint-in-time; one assessment, no re-run trigger stated
Evidence rolethe assessment's date line and the governance crosswalk gap that originated it
PresentY
Actor
RequiresThe builder
Evidence rolethe assessment's authorship (authored in-session; no independent assessor named)
PresentY
Action
RequiresInventory categories, never the data; map flows in and out; grade the three dimensions separately
Evidence rolethe assessment's category inventory, its flow map grounded in live configuration, and its grade table (perimeter B, minimisation D, flow awareness C; overall C+)
PresentY
Record produced
RequiresThe assessment, graded by dimension
Evidence rolethe assessment document (326 lines) with 10 scored residual exposures and a mitigation per exposure
PresentY
Frequency / timing
RequiresPoint-in-time; the register's limit admits no repeat
Evidence rolethe assessment — no re-run date or trigger anywhere in it
PresentY
Population

Assessment instances (full or re-run) performed 2026-09-01 to 2026-09-30: 0. One maintenance edit in period (2026-09-07; 3 lines added, 3 removed, noting the retirement of a scheduled content watcher). Complete for the artifact: an assessment performed elsewhere and not written into this document would not be a privacy impact assessment under the control's own evidence definition.

Sample

Not applicable — no instance in period to sample. For record, the pre-period instance of 2026-08-07 was walked through in the test of design.

Procedures performed
  1. Located the assessment by searching the vault for the control's own terms and confirming it is the one document that grades perimeter and minimisation together; ruled out an earlier public-surface inventory that grades neither.
  2. Read the scope, the shareability statement, the grade table, the section headings and the residual-exposure table in full; tested 'never the data itself' with negative searches for a national-ID number pattern and for address and email fragments (0 hits each — a sample, not a sweep).
  3. Inspected the one in-period edit and compared the assessment's top-ranked residual exposure against the risk register's 2026-09-29 record of the same risk.
Results (1 pass / 0 exceptions)

No instance in period. For record, the 2026-08-07 instance passed the three design attributes tested (categories not data; flow map; three separate grades); the standing artifact drew 1 observation on currency.

Exceptions (0)

None.

Observations (1)
  • One residual the assessment still lists as open is recorded in the risk register's 2026-09-29 re-score as partly treated, with a remaining step pending. The assessment has not been re-run since it was written on 2026-08-07, so it is stale on at least one of its own items — the failure mode the control's stated limit predicts. No re-run trigger exists anywhere in the document.
Conclusion

Not operated in period

Not operated in period — the assessment exists (2026-08-07), is well-formed against the control's own activity text, and contains no raw personal data on the samples tested; nothing re-ran it in September, and its top-ranked residual exposure no longer matches the register's record of the same risk. The evidence supports the artifact; it does not support the artifact as a description of the system on 2026-09-30.

Follow-ups filed
  • Add a re-run trigger to the assessment — at minimum, at each quarterly register re-score or on any new ingress or egress surface; the control has no clock.
  • Update the assessment's top-ranked exposure to the register's current state, or mark the section superseded with a pointer.
  • Update the register's own 'tested' field for this control to reflect the documented negative-control sample on the 'never the data itself' attribute.

6 evidence references cited in the private binder for this control (not shown here — see Provenance below).

A directive, manual control: the mission and a short list of objectives are written down, each objective has a measure and a status taken from data, they are published on the governance page, and an objective the data cannot support is marked unmeasured or unmet.

Attributes tested
Trigger
RequiresWritten once; reviewed at each quarterly register re-score
Evidence rolethe governance data's oversight statement, rendered on the public governance page
PresentY
Actor
RequiresThe owner
Evidence rolethe governance data block — drafted by the system's agent and selected by the owner, per the showcase build's rulings
PresentY
Action
RequiresMission plus objectives, each with a measure and a data-derived status; unmeasured or unmet marked as such
Evidence rolethe governance data: one-sentence mission; objectives O1–O4 each with target, measure, status and note; O3 sourced from the register
PresentY
Record produced
RequiresPublished on the governance page
Evidence rolethe public governance page's objectives section (statuses: O1 not yet measured, O2 not yet measured, O3 not met — 9 of the 13 published risks outside tolerance after the 2026-09-29 re-score, O4 not yet measured)
PresentY
Frequency / timing
RequiresQuarterly, with the register re-score (next 2026-12-29)
Evidence rolethe public governance page's oversight line and the register's header
PresentY
Population

Quarterly reviews of mission and objectives due in the period: 0. The objectives were written on 2026-09-29; the register re-score that day is their baseline, not a review. Next review due 2026-12-29. Complete: the governance data and its rendered page are both version-controlled and there is no other place a review could be recorded.

Sample

Not applicable — no instance to sample. The 2026-09-29 creation was walked through in the test of design.

Procedures performed
  1. Loaded the governance data and checked each objective for its three required parts (measure, status, note), and the mission for presence.
  2. Opened the page generator and confirmed it renders status from data, rejects an unknown status, and refuses to render O3 as 'Met' while any published risk sits outside tolerance; the showcase build's record of that refusal being shown to fail on a planted 'Met' was cited, not re-run.
  3. Re-performed O3's data-derived status from the register: 7 outside tolerance plus 3 in the band lacking a detective control = 10 of 14; 1 row withheld from the page gives 9 of 13, matching the page.
  4. Searched the page's and the data file's change history after 2026-09-29 for any review or status change; fetched the live page to confirm the oversight sentence is served.
Results (2 pass / 0 exceptions)

No review instance in period. For record, the 2026-09-29 creation passed (mission and four objectives with measure and status, published) and O3's data-derived status re-performed to the register.

Exceptions (0)

None.

Observations (2)
  • The control text says 'the owner writes'; the record shows the mission and objectives were drafted by the system's agent and selected by the owner. Authorship is owner-approved, not owner-written. Not an exception against the objective, but the control wording overstates.
  • Three of four objectives are 'Not yet measured' — as the control's stated limit says. Only O3 has a data source; O1, O2 and O4 have none yet, so the first review on 2026-12-29 can only repeat 'Not yet measured' unless one is defined first.
Conclusion

Not operated in period

Not operated in period — the mission and four objectives exist, are published, carry a measure each, and report status honestly (three 'Not yet measured', one 'Not met' that re-performs to the register). No quarterly review has fallen due since they were written on 2026-09-29, so there is no operating instance to test; the first will be the 2026-12-29 re-score.

Follow-ups filed
  • Define a data source for the O1, O2 and O4 measures before 2026-12-29 (O1: share of builds that went through the build gate; O2: recurrence count of failure classes after their fix; O4: an audit for state held only by a vendor).
  • Reword the control's activity from 'the owner writes' to 'the owner approves', unless the owner authors the next revision.
  • Calendar the 2026-12-29 review as the first operating instance for both this control and the register re-score.

7 evidence references cited in the private binder for this control (not shown here — see Provenance below).

A preventive, automated gate that blocks any write to a privileged path unless a written spec names the path or a reviewer agent has recorded a clean audit within the last two hours.

Attributes tested
Trigger
RequiresFires before every file write, file edit, and shell command the agent issues
Evidence roleThree hook registrations in the agent host's settings file (one per tool)
PresentY
Actor
RequiresAll three entry points read one shared table of watched paths
Evidence roleThe shared path-resolver module's watched-path table, consumed by both gate hooks
PresentY
Action
RequiresBlock unless a spec file lists the path or a live reviewer marker (two-hour expiry) exists; the settings file accepts only a live marker
Evidence roleThe write/edit gate's marker check, size rules and spec lookup; the shell gate's marker check
PresentY
Record produced
RequiresA decision record the reviewer can re-perform from
Evidence roleOnly the block message in the session record and temporary marker files; no decision log
PresentN
Frequency / timing
RequiresEvery write; edits under 200 characters and rewrites under 30% are exempt for most labels
Evidence roleThe gate's size-exemption branch and the per-label exemption flag
PresentY
Population

Two populations from a scan of every session record in the period (1,688 records): 58 gate denials (36 by the shell gate, 14 on edits, 8 on new-file writes) and 1,378 passed writes to watched paths (1,253 edits, 125 new files), of which 327 edits were under the 200-character exemption and 1,051 needed a spec or a marker. A first, looser scan found 288 hits; 230 were reads of the gate's own source or probe output, not decisions. Session records are the only record of decisions, so the population is complete only to the extent the records are.

Sample

Fifteen items: 5 passes that needed a spec or marker and 5 size-exempt passes drawn haphazardly with a fixed seed, plus 5 denials chosen to cover both gate hooks and three watched labels.

Procedures performed
  1. Re-performed all 15 decisions through the live gate with payloads of the recorded tool, path and edit length, with no reviewer marker present, and compared the exit code to the recorded outcome.
  2. For the 5 spec-cleared passes, identified the clearing spec file and dated its first commit against the write; for the 5 denials, checked whether a clearing spec exists today and when its path list gained the file.
  3. Searched the shell-command audit log for the marker-override command in the period, and confirmed the gate hooks' configuration history for the month.
Results (14 pass / 1 exception)

All 10 passes still clear today on the rule that cleared them (a dated spec line or the size exemption); 4 of 5 denials now clear through a spec added after the block, in order; 1 denial, on a messaging-integration instructions file dated 2026-09-18, was followed under a minute later by a successful write with no recoverable reason; 0 marker overrides were used in the period.

Exceptions (3)
  • A denied write passed under a minute later with no recoverable reason: the clearing mechanism (a temporary marker) is not logged and expires, so the decision cannot be re-performed from retained evidence.
  • 17 of the 36 shell-gate denials in the period were read-only commands: the in-place-edit pattern in the path resolver spans the whole command, so a later flag-shaped token plus a watched path reads as an edit. Reproduced on 2026-10-01 when the gate blocked this testing's own read-only suite run. Fail-closed, but it trains operators to route around the gate.
  • The block message prints the two-hour override command to the actor it blocks, labelled owner-only, with nothing mechanical restricting it to the owner. No use was observed in the period.
Observations (0)

None.

Conclusion

Effective with exception(s)

Effective with exception(s) — across 58 denials and 1,378 passes every re-performed decision matched the gate's rule, two of three test suites went red when the gate was broken in a sandbox, and no override was used; but the control keeps no decision record, so the one anomaly cannot be explained and a future reviewer will have less population than this one.

Follow-ups filed
  • Add a one-line-per-decision gate log (timestamp, tool, path, label, verdict, clearing mechanism) so the population is no longer the session-record store.
  • Bound the in-place-edit match to the same command segment as its verb, and add the 2026-10-01 read-only command as a negative fixture in the shell-gate suite.
  • Remove the override recipe from the block text, or key the marker to something an agent cannot produce, and record every marker creation.

10 evidence references cited in the private binder for this control (not shown here — see Provenance below).

A preventive, automated hold: when the weekly reconciliation job auto-applies a change to a foundational rule, it commits with a hold tag, and the nightly commit-and-push job skips the offsite push while any held commit is under 48 hours old.

Attributes tested
Trigger
RequiresThe reconciliation job auto-applies a low-risk foundational change and tags the commit
Evidence roleThe reconciliation script's hold-tagged commit step; the job is loaded and runs weekly
PresentY
Actor
RequiresThe nightly commit-and-push job
Evidence roleThe nightly commit job, loaded with last exit 0, scheduled daily
PresentY
Action
RequiresSkip the push while a held commit is under 48 hours old
Evidence roleThe commit script's hold-detection block (48-hour cutoff, subject-tag test) and skip branch
PresentY
Record produced
RequiresA hold notice to the owner and a reminder before expiry
Evidence roleThe commit script's messaging-channel notice and 33–47-hour reminder
PresentY
Frequency / timing
RequiresEvery nightly run; 48-hour window
Evidence roleThe cutoff constant and expiry calculation in the commit script
PresentY
Population

Repository commits in the period whose subject carries both the hold tag and the reconciliation-job prefix: 0 in the period and 0 in the repository's whole history. The reconciliation job ran on 2026-09-06, 2026-09-13, 2026-09-20 and 2026-09-27 and auto-applied nothing on each run, so no held commit could exist.

Sample

No items; the population is empty.

Procedures performed
  1. Queried the repository history for commits carrying both tokens, and searched the nightly job's log for a hold line.
  2. Read the reconciliation job's log to confirm it ran in the period and auto-applied nothing on all four runs.
  3. Ran the nightly commit job's test suite (28 of 28 pass), whose hold test plants a tagged commit and asserts 0 pushes.
Results (0 pass / 0 exceptions)

No operating instance to test; the design is in place and the skip branch is covered by a unit test that was not mutated in this testing.

Exceptions (0)

None.

Observations (0)

None.

Conclusion

Not operated in period

Not operated in period — the design is in place and the skip branch is unit-tested, but no held commit has ever existed, so there is no evidence the control works end-to-end in production; reliance rests on the test and on reading the code.

Follow-ups filed
  • Add a joined test that runs the tagger's real subject format through the commit job's match, so a change on either side goes red.
  • Decide whether a control that has never fired should carry a 'unit only; never operated' status on the public register until the first production hold.

7 evidence references cited in the private binder for this control (not shown here — see Provenance below).

A detective, automated control: every night a scheduled job commits the working system to a private remote and commits the agent-control layer to its own local repository.

Attributes tested
Trigger
RequiresNightly schedule
Evidence roleThe nightly commit job's launch definition, loaded, scheduled nightly in the small hours
PresentY
Actor
RequiresOne job commits both repositories
Evidence roleThe nightly commit script — the only script carrying the agent-config commit subject
PresentY
Action
RequiresStage, commit and push the vault; commit the agent-control repository locally
Evidence roleThe commit script; the vault's private remote on the source-control service; the agent-control repository has no remote
PresentY
Record produced
RequiresCommit history and a run log
Evidence roleBoth repositories' histories; the run log is written to a temporary directory
PresentY
Frequency / timing
RequiresDaily at the scheduled hour
Evidence roleSampled commits stamped within three minutes of the scheduled hour
PresentY
Population

The 30 calendar nights of the period, each expected to produce one dated commit in each repository. Vault: 26 nights with a nightly commit (one night has two), 4 nights without — 2026-09-01, 2026-09-03, 2026-09-09, 2026-09-28. Agent-control repository: the same 26 nights and the same 4 missing. Of the 55 scheduled jobs loaded on the machine, 10 job definitions are under version control and 45 are not.

Sample

Every seventh run night plus the period end (5 of 26 run nights) for push and content, and all 4 missed nights examined.

Procedures performed
  1. For each sampled commit, read its timestamp, author and file count, and confirmed it is contained in the remote's main branch.
  2. For the missed nights, searched the run log and both repository histories for any record.
  3. Compared the nightly-commit day lists of the two repositories.
Results (5 pass / 2 exceptions)

All 5 sampled commits ran at the scheduled hour under the job's identity and are on the remote. 2026-09-03 failed on a repository lock held by a live session, was recovered by a manual sweep the same morning, and a retry was shipped that day. 2026-09-01, 2026-09-09 and 2026-09-28 have no commit in either repository and no retained run record.

Exceptions (2)
  • The 2026-09-03 nightly commit failed on repository-lock contention with a live session; recovered manually the same morning and a bounded retry was added that day. Remediated in period; verify next period.
  • 3 of 30 nights (10%) have no commit in either repository and no retained run record, because the run log is written to a temporary directory and had rolled; whether the job did not run or ran with nothing to commit cannot be distinguished.
Observations (0)

None.

Conclusion

Effective with exception(s)

Effective with exception(s) — on 26 of 30 nights the job committed both repositories at the scheduled hour and pushed the vault, and every sampled commit is on the remote; one failed night is explained and remediated, three are unexplained because the only run log is volatile. The record of changes is good; the record of the record is not.

Follow-ups filed
  • Move the job's run log to the durable log directory and write one line per run including the nothing-to-commit branch, so a silent night is distinguishable from a dead one.
  • Put the 45 untracked scheduled-job definitions under version control so this control covers them.

8 evidence references cited in the private binder for this control (not shown here — see Provenance below).

A detective control, half automated: new skills and commands default to test-first with a companion suite, an advisory hook nudges any new one written without a suite, and the builder's standing method proves each check can fail by breaking a sandbox copy and confirming the suite goes red.

Attributes tested
Trigger
RequiresFires before a new skill or command file is written
Evidence roleThe nudge hook's registration in the agent host's settings file; its new-file and no-companion checks
PresentY
Actor
RequiresThe hook (advisory) and the builder (runs suites and mutations)
Evidence roleThe nudge hook; the builder's written mutation method, which is policy
PresentY
Action
RequiresDeliver the nudge to the actor; builder runs suite and a red mutation before relying
Evidence roleThe hook writes three lines to its error stream and exits successfully — a channel the agent does not read
PresentN
Record produced
RequiresSuite output and mutation transcripts
Evidence roleProduced ad hoc by sessions, not by the hook
PresentN
Frequency / timing
RequiresEvery new skill or command; every change
Evidence roleThe hook per write; the every-change half is the builder's policy
PresentY
Population

Skills and commands added to the agent-control repository in the period, excluding 13 vendor-synced skill files: 4 authored items (2 commands, 2 skills). Day-granular, since the repository is committed nightly. For context, the command library holds 28 commands, 6 with a companion suite and 22 without.

Sample

All 4 (population of five or fewer).

Procedures performed
  1. For each item, confirmed a companion suite exists and compared the first-commit dates of the skill file and its suite for test-first ordering at day granularity.
  2. Counted mutation and red-test mentions in each suite as a proxy for a built-in negative control; the four suites were not run here.
  3. Searched every session record in the period for the nudge's own output string, and ran the nudge hook's suite (8 of 8) and a sandbox mutation that silenced the hook (6 of 8, red).
Results (3 pass / 1 exception)

3 of 4 items were born with a suite on the same day, each suite carrying its own mutants; 1 skill was created on 2026-09-07 and gained its suite on 2026-09-14, seven days later, with no nudge delivered; 0 rendered nudges appear in 1,688 session records.

Exceptions (2)
  • 1 of 4 new skills (25%) shipped without a suite for 7 days and was relied on in that window; the nudge did not reach the builder and the policy half was not followed. Closed by the later suite; root cause open.
  • The advisory hook has never been seen by an agent: it writes to its error stream and exits successfully, a channel the agent host does not feed back to the model (inference from the hook's form and the zero-delivery count; not re-verified against vendor documentation in this testing). Its suite asserts the text, not the delivery, so it is green while delivery is dead.
Observations (0)

None.

Conclusion

Effective with exception(s)

Effective with exception(s) — three of four new automations in the period were born with a suite that carries its own negative control and the fourth caught up within a week, but the automated nudge that makes this the default has an output channel the agent never reads, and the only evidence it works is a suite that asserts the text rather than the delivery.

Follow-ups filed
  • Re-emit the nudge on the agent host's advisory-context channel, fix the companion-path rule for skills, and add a delivery assertion to the suite.
  • Add a scheduled suite runner for the suites that back register controls, so green carries a date newer than the last hand run.

6 evidence references cited in the private binder for this control (not shown here — see Provenance below).

Two pre-execution hooks inspect every shell command the agent issues and block recursive force-deletes, bulk-delete idioms, and any deletion of a cloud-document pointer file before the command can run.

Attributes tested
Trigger
RequiresFires before every shell tool call, with no narrower filter
Evidence roleThe hook registry's shell-tool entry for both hooks
PresentY
Actor
RequiresThe hook decides; the model cannot waive the block
Evidence roleThe hooks' hard-block exit path
PresentY
Action
RequiresBlock recursive+force delete, bulk find-delete, interpreter tree-delete; block any delete of a cloud-document pointer or a folder holding one
Evidence roleThe destructive-command hook's command classifier; the pointer hook's two-layer pattern and path-aware check
PresentY
Record produced
RequiresA block notice the session keeps
Evidence roleThe session record of the refused tool call
PresentY
Frequency
RequiresEvery shell command, before execution
Evidence roleThe registry's pre-execution event binding
PresentY
Population

Every shell command in the period that one of the two hooks blocked: 172 blocked destructive commands plus 1 blocked cloud-pointer deletion (173), found by scanning every session-record file for the period (963 files, 367,904 records). A second, negative population: the 3 executed commands in the period whose text contains the recursive-force delete idiom, from the executed-command audit log. Completeness cannot be established: the session record is the only place a block is written, and sessions run on another model provider fire no hooks.

Sample

Every-34th item of the 172 (5), the single cloud-pointer block, all 3 audit-log items, and 2 live instances observed on the fieldwork date 2026-10-01: 11 items.

Procedures performed
  1. For each sampled block, re-extracted the full command from the session record and fed it to the live hook on standard input; compared the exit code and first message line to the period record.
  2. For each executed audit-log command, classified the idiom as prose (inside an echo or a search pattern) or executable, and fed the executable one to the live hook.
  3. Confirmed via source-control history that neither hook changed in the period and that both stayed registered at period start and end.
  4. Mutation test: copied each hook to a sandbox outside the vault, broke its detector, re-fed the payload from a file, confirmed the block was lost, and byte-compared production to the untouched copy.
Results (10 pass / 1 exception)

All 6 sampled period blocks and both fieldwork-date instances reproduce against the live hooks; 2 of the 3 executed audit-log commands carry the idiom only as prose; 1 executed command is a true bypass.

Exceptions (1)
  • A recursive force-delete executed on 2026-09-29 inside a case-statement arm. The classifier takes the first non-wrapper token of each command segment as the command; in a case arm that token is case, so the delete is never examined. The target was a throwaway scratch directory, so no vault or cloud data was at risk, but the control objective was not met for this shape. The shape is in neither the hook's documented gap list nor its test suite's known-gap block, and the suite's audit-log replay uses the same classifier, so it cannot see this class. Fix filed 2026-10-01; not yet landed.
Observations (1)
  • The suite's audit-log replay certifies zero fail-open receipts using the hook's own classifier; it is a plumbing check for bypass completeness, not evidence of it.
Conclusion

Effective with exception(s)

Effective with exception(s) — Both hooks were registered for the whole period with unchanged configuration, every sampled block reproduces, and every suite relied on was shown red under mutation; one recursive force-delete executed through a case-arm shape the classifier does not parse, so the objective is not fully met and the exception stands.

Follow-ups filed
  • Teach the classifier to strip case-arm patterns, add the vector to the suite's must-block list and known-gap list until then, and give the audit-log replay a second, independent detector so it can find classes the classifier misses. Fix filed 2026-10-01; not yet landed.

10 evidence references cited in the private binder for this control (not shown here — see Provenance below).

A pre-execution gate blocks any shell command that would change many files at once until a manifest keyed to that exact command lists each file with its own reason; sends, publishes, deploys, and spend require the owner's explicit yes as a written policy with no hook behind it.

Attributes tested
Trigger
RequiresFires before every shell tool call
Evidence roleThe hook registry's shell-tool entry for the bulk-change gate
PresentY
Actor
RequiresThe gate decides; the model cannot waive it
Evidence roleThe gate's verdict-and-exit path
PresentY
Action
RequiresBlock when the command names at least the threshold number of files or pairs a move verb with a glob, unless a manifest keyed to the command's digest exists with enough per-file lines
Evidence roleThe gate's classifier: threshold test, glob-shape trigger, digest-keyed manifest lookup
PresentY
Record produced
RequiresThe manifest (one line per file with its own reason) and a block notice in the session record
Evidence roleThe manifest store (12 files on 2026-10-01) and the gate's block text
PresentY
Frequency
RequiresEvery shell command, before execution
Evidence roleThe registry's pre-execution event binding
PresentY
Send and spend approval
RequiresAn explicit owner yes per send, publish, deploy, or spend
Evidence rolePolicy text in the standing rules only; no hook inspects send, deploy, or messaging tool calls
PresentN
Population

Every shell command in the period the bulk-change gate blocked: 19, from the same session-record scan (963 files, 367,904 records). An independent second trace: 9 manifest files carry September names in the manifest store. Completeness carries the same caveat as AC-01.

Sample

Every-4th item of the 19 (5) plus 1 live instance on the fieldwork date 2026-10-01: 6 items.

Procedures performed
  1. Re-extracted each sampled command in full and fed it to the live gate; recorded exit code and first message line.
  2. Where the live gate now passes a command it blocked in the period, computed the command's digest and looked for a matching manifest, the only designed unblock path; compared manifest write time to the block.
  3. For each matched manifest, counted per-file lines, duplicate file names, and reason strings reused across files.
  4. Mutation test: sandbox copy outside the vault with the threshold comparison made unreachable; payload read from a file; confirmed the block was lost; byte-compared production.
Results (6 pass / 0 exceptions)

3 of 5 sampled period blocks reproduce today; the other 2 are explained by digest-matched manifests of 21 and 69 per-file lines written 65 seconds and 22 seconds after their blocks; the fieldwork-date instance blocked this tester's own test harness and no manifest was written.

Exceptions (2)
  • The item-level human yes in the objective is not a pre-execution human step: in both manifest-backed samples the same agent wrote the manifest 22 to 65 seconds after the block and the command then passed, with no record of the owner seeing it first. The gate enforces per-file written reasons; the instruction to surface the manifest to the owner before re-running is text, not mechanism.
  • The send, publish, deploy, and spend half has no mechanism to test; the register already states this. No test of operating effectiveness was performed on it.
Observations (2)
  • One September manifest lists a file name twice; the gate counts lines, not distinct files.
  • The gate classifies command text, so a command that merely contains many file-like tokens is blocked: a conservative false positive, observed on the fieldwork date against this tester's own harness.
Conclusion

Effective with exception(s) / Not testable (policy only)

Effective with exception(s) for the bulk-change gate: registered all period, configuration unchanged, sampled blocks reproduce or are explained by the designed manifest path, suite shown red under mutation, but the objective's human-approval wording is not what the mechanism enforces. Not testable (policy only) for the send and spend half, as the register itself discloses.

Follow-ups filed
  • Either reword the control's objective and activity on the public register to what is built (a per-file written reason, reviewable), or add a review marker the gate requires before unblocking; schedule the reason-collision audit over the live manifest store rather than only inside the suite.

9 evidence references cited in the private binder for this control (not shown here — see Provenance below).

The messaging-channel listener acts only on messages from a single allowlisted sender, and a remote run-now trigger resolves only to a fixed map of scheduled jobs that must also be loaded, with a fixed argument list that never includes input text.

Attributes tested
Trigger
RequiresEvery message the listener receives and every run-now queue entry
Evidence roleThe listener's update handler and its queue poll on each cycle
PresentY
Actor
RequiresA resident listener process under the scheduler
Evidence roleThe scheduler's listener job, running with keep-alive; alive on 2026-10-01
PresentY
Sender allowlist
RequiresDrop any message whose sender id is not in the access list
Evidence roleThe listener's sender check, placed before every handler; access list of 1 sender
PresentY
Fixed job map, two-sided
RequiresA run-now label runs only if in the hard-coded map and currently loaded, and maps to a fixed scheduler argument list
Evidence roleThe listener's default-deny job map (5 jobs), live-job lookup, and fixed kickstart arguments
PresentY
Queue-side allowlist
RequiresThe producing endpoint carries the same default-deny set and a purpose-scoped consumer secret
Evidence roleThe dashboard's run-now endpoint and its consumer check
PresentY
Record produced
RequiresA listener log and rejected-trigger notices
Evidence roleRun-now rejections produce an acknowledgement and a notice; sender rejections produce nothing; the listener log carries no timestamps or per-message lines
PresentN
Timing
RequiresBefore any handler, on every message
Evidence roleThe sender check precedes every branch of the handler
PresentY
Population

Inbound messages in the period and those dropped as non-allowlisted; run-now entries and those rejected. Both unobtainable: sender drops write nothing, the listener log has 0 period lines and no per-message records, and the run-now queue lives on a remote endpoint outside this binder's read scope. Operation in the period is evidenced only indirectly: a chat-state file written on 2026-09-30 by code that runs only after the sender check, and a listener process alive since load.

Sample

No period items possible. Fieldwork-date re-performance on 2026-10-01 only: 5 synthetic attribute checks and 1 mutation.

Procedures performed
  1. Attempted to build the population from the listener log (0 lines dated in the period, 0 rejection lines) and from session records (the listener is not a hook, so none exist there).
  2. Ran a sandbox copy of the listener's logic with every outbound call stubbed (messaging API, model invocation, job kickstart, queue requests); nothing was sent or run. Checked: a stranger reaches no handler; the owner reaches a handler; an unknown label is rejected with no shell; an allowlisted loaded label yields exactly one fixed kickstart argument list; an allowlisted but unloaded label is rejected.
  3. Mutation test: removed the sender check in the sandbox copy and confirmed a stranger then reaches a handler; confirmed the production listener's digest unchanged before and after.
  4. Confirmed via source-control history that the 2 commits to the listener in the period touched neither the sender check nor the job map.
Results (6 pass / 0 exceptions)

All 5 synthetic attributes pass and the mutation shows red; no period instance could be tested.

Exceptions (2)
  • No record is produced when a non-allowlisted sender is dropped; the register's evidence claim of rejected-trigger notices covers only the run-now path. Operating effectiveness of the sender allowlist is untestable over any period until logging is added.
  • The listener's only log is untimestamped and lives in temporary storage that is wiped at boot; the last boot was 2026-09-28, so 27 of 30 period days are unrecoverable.
Observations (2)
  • A second inbound channel exists outside this control: the model host's remote-control feature accepts prompts authenticated by the host account rather than by this access list. It is not in the register.
  • The allowlisted sender can issue arbitrary prompts by design; the no-arbitrary-command property holds for the job path specifically. A phone-side root-cause reply is not a job-map entry but an ordinary owner message.
Conclusion

Design only (no TOE possible)

Design only (no TOE possible) — The sender allowlist (1 id) and the two-sided, fixed-argument job trigger are implemented where the register says, were unchanged in the period at the lines that matter, and were re-performed and shown red on 2026-10-01 in a stubbed sandbox; no operating population exists because sender rejections are not recorded and the listener log is untimestamped and boot-scoped.

Follow-ups filed
  • Log sender rejections (date, hashed sender id, chat reference) so the allowlist has a testable population, and correct the register's evidence text.
  • Move the listener log out of temporary storage and timestamp its lines.
  • Register decision: whether the model host's remote-control channel belongs inside AC-03 or under the vendor-change control.
  • Put the access list under a change record; the agent-control repository does not track it, so access-list changes are undated.

10 evidence references cited in the private binder for this control (not shown here — see Provenance below).

An hourly watchdog reads the fleet health table, restarts any scheduled job that has gone silent or stale, escalates once if the restart fails, and a status dashboard regenerates every 15 minutes.

Attributes tested
Trigger
RequiresRuns every hour and every 15 minutes without a person starting it
Evidence rolethe watchdog's and the dashboard generator's scheduler entries, both loaded
PresentY
Detect
RequiresReads the fleet health table and flags any job past 1.5× its expected cadence
Evidence rolethe watchdog's sweep stage and its single cadence table (18 jobs)
PresentY
Restart
RequiresReloads a stale job's scheduler entry once
Evidence rolethe watchdog's reload step
PresentY
Escalate, then mute
RequiresA second miss escalates to the owner; from the third on, the state is logged only
Evidence rolethe watchdog's escalation counter
PresentY
Fail closed
RequiresA sweep that crashes or returns nothing is reported as blind, never as all-clear
Evidence rolethe watchdog's blind branch, written after a measured fail-open defect
PresentY
Record
RequiresOne log line and one health-table row per run; a message to the owner only when something is actionable
Evidence rolethe watchdog's run log, its health-row writer, and the messaging channel sender
PresentY
Population

Watchdog runs with a run header dated in the period: 549 of 720 scheduled hours, with two days having no run at all. Dashboard regenerations: 269 recorded from 2026-09-28 to 2026-10-01; earlier runs are unobtainable because the generator logs to temporary storage that the 2026-09-28 reboot cleared.

Sample

Every one of the 549 run summaries was classified; the single stale-job event in the period was walked end-to-end; all 7 blind runs were enumerated.

Procedures performed
  1. Classified every period summary line: 459 all-clear, 55 'no actionable jobs' with muted escalations, 7 blind. Confirmed each blind run wrote an error health row, not an ok row.
  2. Walked the one stale event: a nightly loop-maintenance job was flagged on 2026-09-28 at 36.2 hours against a 36-hour threshold, reloaded in the same run, escalated the next hour when it stayed stale, and muted from the hour after. The job's own logs confirm it has not completed a run since 2026-09-26.
  3. Read the status dashboard rendered on 2026-10-01 for that job: badge 'healthy', row text 'log 3 days old (expected within 24 hours)'. The job writes its health row under a retired key the renderer does not look up, so the badge falls through to the default.
  4. Ran the two watchdog test suites (67 of 67 and 17 of 17 passing) and demonstrated the first suite red by making the blind branch unreachable in a sandbox copy (37 pass, 30 fail).
Results (6 pass / 3 exceptions)

Detection, restart, single escalation, fail-closed blind handling and the dashboard's 15-minute cadence all passed; hourly coverage, the mute after one escalation, and the dashboard's healthy badge on a dead job are exceptions.

Exceptions (3)
  • 171 of 720 hourly slots have no run, including two whole days; no host-uptime record exists to separate the machine sleeping from the watchdog failing.
  • A job dead since 2026-09-26 produced one escalation on 2026-09-28 and then 66 silent muted escalations through 2026-10-01; by design the owner stops hearing about a persistent failure after the second hour.
  • The status dashboard renders the dead nightly job as healthy while its own row says the log is three days old; eight nights of error rows between 2026-09-15 and 2026-09-22 rendered green for the same key mismatch. Already filed as a priority item on 2026-09-29; confirmed still live on 2026-10-01.
Observations (2)
  • A successful message to the owner leaves no record; only failures are logged. The escalation on 2026-09-28 is evidenced by the code path reached and the absence of a logged send failure.
  • Since 2026-09-29 every watchdog run emits an unhandled shell error from one line of the script; the run still completes. Outside the period; noted.
Conclusion

Effective with exception(s)

Effective with exception(s) — the watchdog detected the one stale job in period within the hour, reloaded it, escalated once when the reload failed, and its fail-closed branch is mutation-tested; it then went silent on a job that is still dead, the dashboard rendered that job healthy, and 24% of hourly slots have no run.

Follow-ups filed
  • Add a re-notify cadence or a dashboard-visible muted count so a persistently dead job is not silent after the second hour.
  • Record host uptime in each run header so a coverage gap can be attributed.
  • Re-key the nightly loop-maintenance job's health row to the name the dashboard renderer looks up (already a filed priority item).
  • Fix the unhandled shell error emitted on every run since 2026-09-29.

15 evidence references cited in the private binder for this control (not shown here — see Provenance below).

A scheduled detector classifies every automation run by its actual output, not its reported status, catching runs that claimed success but produced nothing useful.

Attributes tested
Trigger
RequiresScheduled, scanning all automation output every three hours.
Evidence rolea scheduled job that scans every automation's recent output
PresentY
Actor
RequiresA script, not a person.
Evidence rolea detector script with no retry capability built in
PresentY
Action
RequiresClassify each run by its actual output, not by its reported status: output that is too small, or output that contains a provider-level error at any size, is marked hollow.
Evidence rolea two-arm classifier — a size check and a content check, combined with OR
PresentY
Record produced
RequiresA scan line per run of the detector, a dedicated line per hollow hit, and a seen-set so each hit alerts only once.
Evidence rolea run log, a hit log, and a one-time-alert mechanism
PresentY
Scope
RequiresEvery automation's output directory.
Evidence rolea sweep over every subdirectory under the automations store
PresentY
Detect only, no retry
RequiresThe control stops at detection; whether to retry is a separate decision by design.
Evidence rolean absence of any retry or re-run call in the script
PresentY
Population

The population was every scheduled scan of the real automation output during the period that the control existed (it was installed partway through the period) — 167 scans across 22 days — covering 82 real automation output files.

Sample

All six hollow-run hits that occurred in the period were re-read in full, and the detector's run cadence was checked across every day it existed.

Procedures performed
  1. Counted scans per day from the detector's own run log and confirmed the schedule held (roughly eight scans a day, matching a three-hour cadence) for every day the control existed.
  2. Read each of the six hollow-run hits and checked which classification arm caught it — the size check or the content check.
  3. Confirmed the one-time-alert mechanism worked: each of the six hits produced exactly one alert line, with no repeats across roughly 160 later scans.
  4. Copied the detector to a sandbox outside the live system, inverted its size-comparison logic on purpose, and confirmed the test suite caught the break; restored and confirmed the live copy was untouched.
Results (4 pass / 0 exceptions)

From the day it was installed, the detector scanned on schedule every three hours and caught all six real failed automation runs in the period — runs that had reported themselves as successful despite producing almost nothing useful, or an outright provider error. Five of the six hits were larger than the simple size threshold and were only caught because a second check inspects the output's content for an error, not just its size. Each hit produced exactly one alert, with no repeat noise. A deliberate break-it test on a sandboxed copy of the detector's logic confirmed the safety check actually catches a real mistake, not just a cosmetic one.

Exceptions (0)

None.

Observations (0)

None.

Conclusion

Effective

Effective — from the day it was installed, the detector scanned on schedule, correctly classified six real failed runs by their actual output rather than their claimed status, alerted once for each, and its classification logic is proven to catch a real break when deliberately tested. The first part of the period predates the control's existence.

Follow-ups filed
  • Point any future test runs of the detector at a separate, non-production log path, so practice runs stop mixing into the real operating record.

7 evidence references cited in the private binder for this control (not shown here — see Provenance below).

A shared wiring check lets a hook's test suite assert that the hook is registered under a matcher that fires for today's tool name, and fails closed when the checker is missing.

Attributes tested
Trigger
RequiresRuns whenever a covered test suite runs
Evidence rolethe wiring block inside each adopting suite
PresentY
Single source
RequiresOne shared checker; suites do not copy its logic
Evidence rolethe shared wiring checker (98 lines)
PresentY
Assert registration and matcher
RequiresReports wired, matcher-miss, unwired, or error
Evidence rolethe checker's four-outcome output
PresentY
Invoke through the configured command
RequiresReturns the command string exactly as the configuration specifies
Evidence rolethe checker's wired output and the suite's use of it
PresentY
Fail closed
RequiresA missing or unrunnable checker turns the suite red, not green
Evidence rolethe adopting suite's wiring block
PresentY
Population

Hook test suites that call the shared checker: 6 of 30 hook suites (70 suites exist in total across hooks and scripts; 46 distinct hook scripts are registered).

Sample

Both suites the register names (the destructive-command blocks and the bulk-change gate) run in full; the checker re-performed against the live configuration and two mutated copies.

Procedures performed
  1. Ran the destructive-command suite (82 of 82) and the bulk-change-gate suite (49 of 49); each printed its wiring line confirming the hook is routed for the shell tool.
  2. Re-performed the checker: live configuration reports wired; a copy with the tool matcher renamed reports matcher-miss; a copy with the hook entry removed reports unwired.
  3. Counted callers across every suite and checked the checker's own change history for the period: unchanged, and it has no test suite of its own.
Results (5 pass / 2 exceptions)

The checker is correct on live and mutated configurations and both named suites pass; coverage is 6 of 30 hook suites and no record exists of when any suite last ran.

Exceptions (2)
  • 24 of 30 hook suites, including the suites behind the completion-claim block (OI-01) and the build gate (CM-01), do not assert their own wiring; the incident the checker was built for — a gate inert for about six weeks after a tool rename — can recur on any of them.
  • No ledger records which suites ran when or what the wiring line printed; the operating record can only be re-performed, not observed.
Observations (0)

None.

Conclusion

Effective with exception(s)

Effective with exception(s) — the shared checker correctly reports wired, matcher-miss and unwired on live and mutated configurations, and the two suites the register names use it and pass; it covers 6 of 30 hook suites, and nothing records when those suites last ran.

Follow-ups filed
  • Add the wiring block to the suites behind OI-01 and CM-01, then to the remaining hook suites.
  • Have each suite append one line (name, date, result, wiring outcome) to a run ledger; check first whether an existing suite-health file already does this.

6 evidence references cited in the private binder for this control (not shown here — see Provenance below).

Significant failures get a structured root cause — criteria, condition, cause, consequence, corrective action — and the corrective action becomes a tracked item.

Attributes tested
Trigger
RequiresA significant failure, or the automatic dispatch at every session close
Evidence rolethe root-cause command's invocation rule and the session-close dispatcher
PresentY
Actor
RequiresThe model, following a written doctrine
Evidence rolethe root-cause doctrine (128 lines)
PresentY
Five Cs
RequiresCriteria, condition, cause with evidence, consequence, corrective action
Evidence rolethe doctrine's Five-Cs table and leg rules
PresentY
Record and tracked item
RequiresA root-cause record plus a loop item; invocations logged
Evidence rolethe root-cause record files, the loop registers, and the skill-invocation ledger
PresentY
Two fixes
RequiresSymptom fix and root fix both executed and proven, not described
Evidence rolethe doctrine's scope note (corrected 2026-09-15 so the run itself executes the fix)
PresentY
Blameless postmortem
RequiresIncidents written up
Evidence rolethe public incidents page (not opened by this testing)
PresentN
Population

52 root-cause invocations in the period, from the skill-invocation ledger; 11 root-cause record files with period dates; 128 closed dispatched records all-time, of which 74 carry a date and at least one named artifact.

Sample

Full re-run of the produced-versus-wired measurement over all 74 adjudicable records, plus a haphazard sample of 5 of the 11 period records screened for the Five Cs and for the disposition of each guard they name.

Procedures performed
  1. Re-ran the measurement script on 2026-10-01: 74 adjudicable records, 52 produced a new script on the record's date, 22 cited only pre-existing files, 5 produced a script that is registered in the hook configuration.
  2. For each sampled record, counted the Five-Cs headings and checked every named guard: present as a prototype, registered in configuration, or absent.
  3. Cross-checked the two records with zero keyword hits by reading their heading lists.
Results (5 pass / 3 exceptions)

Runs happen, produce artifacts, and each sampled record has a tracked item; the corrective action reaches a registered guard in 5 of 74 cases, two sampled records cite guard files that no longer exist, and one lacks labelled consequence and corrective-action sections.

Exceptions (3)
  • 5 of 74 adjudicable root-cause runs ended in a guard registered in the hook configuration; 47 built a script that sits as a prototype. The measured rate is 7%, up from 1 of 54 on 2026-09-22. Already a filed priority item; this testing adds the re-measurement.
  • Two sampled records cite guard scripts that do not exist at the cited name.
  • One sampled record (2026-09-16) has no labelled consequence or corrective-action section; the Five-Cs structure the control requires is not followed.
Observations (0)

None.

Conclusion

Effective with exception(s)

Effective with exception(s) — root-cause runs happen (52 in September), produce structured records with the Five Cs in 4 of 5 sampled, and each has a tracked item; the corrective action reaches a registered guard in 5 of 74 cases, so the objective is met and the risk the control addresses — the same failure recurring — is largely not.

Follow-ups filed
  • Add the re-measured 5 of 74 to the existing priority item; the terminus step (register the guard or close the item) is the fix.
  • Locate or retire the two cited guard files that no longer exist.
  • Add a template check (the Five Cs as required headings) over root-cause records.
  • The blameless-postmortem attribute is unverified; the public incidents page was not opened.

7 evidence references cited in the private binder for this control (not shown here — see Provenance below).

Before a file-edit tool call on a shared register, a hook snapshots the register to temporary storage; the nightly commit provides the daily offsite copy.

Attributes tested
Trigger
RequiresBefore a file-edit tool call on a register
Evidence rolethe hook's registration under the write and edit tools (not the shell or multi-edit tools)
PresentY
Action
RequiresCopy the register to a timestamped snapshot
Evidence rolethe snapshot hook's copy step and register name pattern (7 registers)
PresentY
Scope inside the script
RequiresWrite, edit, multi-edit, notebook-edit and shell paths handled
Evidence rolethe hook's tool-name branches; the shell and multi-edit branches are unreachable because the registration does not route them
PresentY
Record
RequiresA snapshot file per edit
Evidence rolethe temporary snapshot folder
PresentY
Offsite daily
RequiresNightly commit and push
Evidence rolethe nightly commit job (population tested under CM-03)
PresentY
Never blocks
RequiresExit cleanly on every path
Evidence rolethe hook's return statements
PresentY
Population

Snapshots surviving with a period date: 62, all from 2026-09-28 to 2026-09-30; snapshots from 2026-09-01 to 2026-09-27 are unobtainable, cleared by the 2026-09-28 reboot. Offsite: 26 of 30 nights with a daily commit; missing 2026-09-01, 09-03, 09-09 and 09-28.

Sample

Re-performance of the hook with a synthetic edit on a register (expected one new snapshot) and on a non-register (expected none); full enumeration of surviving snapshots and of commit nights.

Procedures performed
  1. Synthetic edit on a register: snapshot count 79 to 80. Synthetic edit on a non-register: 80 to 80 — the hook can say no.
  2. Listed surviving snapshots by register and date and compared against the last boot time; none predate it.
  3. Read the hook's change history: until 2026-09-14 it was registered for the write tool only, and its own header records three silent register deletions under that design; the edit-tool registration landed on 2026-09-14 and the shell-tool line was deliberately not shipped.
Results (3 pass / 4 exceptions)

The hook snapshots a register on edit and ignores a non-register, and the 62 surviving snapshots match register edits; the 2026-09-01 to 09-27 snapshots are gone, the hook could not fire on edit before 2026-09-14, shell and multi-edit rewrites are never snapshotted, and 4 nights lack an offsite commit.

Exceptions (4)
  • No pre-write snapshot exists for any register edit between 2026-09-01 and 2026-09-27; restore-to-point for that window falls back to day granularity because the snapshot folder is temporary storage cleared at boot.
  • For 13 of 30 days the hook was registered for the write tool only and could not fire on the edit tool that actually changes registers; remediated on 2026-09-14.
  • Registers rewritten from the shell or with the multi-edit tool get no snapshot; the registration covers the write and edit tools only.
  • 4 of 30 nights have no offsite commit (owned by CM-03).
Observations (0)

None.

Conclusion

Effective with exception(s)

Effective with exception(s) — the hook takes a pre-image on write or edit of a register, re-performed with a negative control, and the nightly commit provides a daily offsite point for 26 of 30 nights; snapshots for 2026-09-01 to 09-27 are gone with the 2026-09-28 reboot, the hook could not fire on edit before 2026-09-14, and shell and multi-edit rewrites are never snapshotted.

Follow-ups filed
  • Move the snapshot folder out of temporary storage into persistent state with a size cap.
  • Widen the hook's registration to the shell and multi-edit tools (needs a fresh reviewer-agent clearance per the hook's own header).
  • The hook has no test suite; the negative control in this testing is the first recorded demonstration that it can say no.

7 evidence references cited in the private binder for this control (not shown here — see Provenance below).

An end-of-turn hook blocks any agent turn that changed something, claimed the work was done, and verified nothing in the same turn.

Attributes tested
Trigger
RequiresFires at the end of every agent turn
Evidence rolethe end-of-turn hook registration in the host configuration
PresentY
Actor
RequiresThe hook decides, not the model
Evidence rolethe hook's three-condition rule (change made, completion claimed, nothing verified)
PresentY
Action
RequiresBlock the turn when all three conditions hold and no hedge is present
Evidence rolesix block-emission sites in the hook, one per detector
PresentY
Record produced
RequiresEvery decision logged
Evidence rolethe hook's decision log
PresentY
Failure mode
RequiresOn a parse error, let the turn through but log it loudly; documented bypasses are logged
Evidence rolethe hook's fail-open-loud branch and its bypass variables
PresentY
Population

Every end-of-turn decision the hook logged between 2026-09-01 and 2026-09-30: 6,017 decisions, counted twice by independent commands with the same result.

Sample

Every 7th blocked turn in log order, first 5 taken; plus the full set of 192 blocks tabulated by day.

Procedures performed
  1. Counted outcomes by class over the period: 3,262 passes, 192 blocks, 372 re-stop escalations (a second stop after a block is not re-blocked, by design), 181 headless bypasses, 2 parse failures, 0 dry-run would-blocks.
  2. Read each sampled block line and confirmed it records a change, a completion claim, and either no verification or an unverified named entity, with the claim text and the change that triggered it.
  3. Ran the hook's test suite today: 52 of 52 green. Then copied the hook to a sandbox outside the system, neutered all six block sites, and re-ran the suite: 8 tests went red.
  4. Confirmed blocks were logged on both sides of the hook's three in-period logic extensions (123 blocks before 2026-09-14, 69 after).
Results (6 pass / 1 exception)

All five sampled blocks fired as designed; blocks are present on 25 of 30 days; 2 of 6,017 turns passed on a parse failure, which is the exception.

Exceptions (1)
  • 2 of 6,017 turns (0.03%) passed on a parse failure: the hook could not evaluate the claim and let the turn through, logging the failure as designed. Deficiency, low; recorded, no corrective item.
Observations (2)
  • 2 of 5 sampled blocks were false positives: completion vocabulary inside non-build prose, with an unrelated edit in the same turn. The control fired as designed; precision is a design limit the register already states.
  • The hook's logic was extended three times inside the period (2026-09-13, 2026-09-16, 2026-09-17), carried to the configuration repository only by the nightly auto-commit, with the job as the recorded author.
Conclusion

Effective with exception(s)

Effective with exception(s) — the hook is registered at the end of every turn, logged 6,017 decisions and 192 blocks across 25 of 30 days, the sampled blocks each show the three-condition rule applied, and the suite goes red when the block is removed. Two fail-open turns are the exception; the false-positive rate in the sample and the unrecorded mid-period changes are observations.

Follow-ups filed

None.

7 evidence references cited in the private binder for this control (not shown here — see Provenance below).

Public pages are built from curated material, never by redacting the private system in place; where a build has a leak test, the test blocks the build on real names, local paths, or account identifiers.

Attributes tested
Trigger (policy)
RequiresA standing rule: separate public version, never strip in place
Evidence rolethe standing-rules memory entry
PresentY
Trigger (learnings page)
RequiresThe generator refuses to write the page when its scrub gate finds a span
Evidence rolethe scrub gate inside the public learnings generator
PresentY
Trigger (controls and governance pages)
RequiresA leak test on the rendered page for agent names, hook names, local paths, account and employer identifiers, and private source names in raw HTML
Evidence rolethe leak check in the controls-page test runner
PresentY
Record produced
RequiresRunner output per check, red per planted defect; generator abort message
Evidence rolethe runner's pass/fail output and the generator's refusal line
PresentY
Negative control
RequiresEach check must go red on its own planted defect in a sandbox outside the system
Evidence rolethe runner's mutation stage and the scrub gate's refusal test
PresentY
Population

Builds of public pages committed to the site repository in the period, for the surfaces that carry a test: the learnings page 4 commits, the controls page 8, the governance page 3, the home page 9, the play page 1. Denominator for coverage: 21 distinct public pages in the site's public allowlist.

Sample

All tested surfaces re-performed today on a fresh render, plus the two in-period ship records (2026-09-29 and 2026-09-30) cited for their own leak-test runs.

Procedures performed
  1. Ran the controls-page test runner on fresh renders of the controls and governance pages: the leak check passed on both, and five planted leaks (a bare agent name, a local path, a private-source sentence, an altered pinned sentence, a private source in an HTML comment) each went red.
  2. Ran the learnings scrub gate's test suite: 24 of 24 passed, including the case where the generator exits without writing the page and names the gate as the reason, and the case where it writes on a clean source.
  3. Enumerated the 21 public pages and classified each by leak-test coverage: 3 have a standing, build-blocking test; 5 more had a one-off hand-run test at their 2026-09-30 ship; 13 have none.
  4. Checked the runner's overall state: 60 of 80 assertions pass today. The 20 failures are outside the leak check: two cross-page checks broken by the 2026-09-30 home-page launch, and 17 that report 'cannot measure' because two runtime dependencies are missing in this shell.
Results (6 pass / 2 exceptions)

The leak test is green on every in-period build of the three covered surfaces and goes red on planted leaks; the runner as a standing control is red since 2026-09-30 and partly unmeasurable today.

Exceptions (2)
  • The standing runner is red (60 of 80) as of 2026-10-01: the 2026-09-30 home-page launch removed the sentence the control-count check looks for and links to a section of the controls page that no longer exists. A launch was not followed by a green run of the cross-page suite. Deficiency; follow-up filed.
  • The runner's shipped-layout chain and its font-floor check cannot execute in this environment (two runtime dependencies absent), so 17 of 80 assertions report 'cannot measure' instead of a verdict. Deficiency; follow-up filed.
Observations (2)
  • Measured coverage: 3 of 21 public pages (14%) have a standing leak test, 13 of 21 (62%) have none. The register already states this limit; this is its number.
  • The leak check knows identifiers by construction (agent names, hook names, six literal patterns); a person's name that is not also an agent name is outside its vocabulary.
Conclusion

Effective with exception(s)

Effective with exception(s) — on the three surfaces that carry a standing leak test, the test is build-blocking, was green on every in-period build, and goes red on planted leaks today. The exceptions are that the cross-page runner has been red since the 2026-09-30 ships and that two of its check families cannot execute in the current environment; the measured coverage gap (13 of 21 pages untested) is a design observation the register already states, now with a number.

Follow-ups filed
  • Restore the control-count sentence and the section anchor on the home page, or retarget the checks, then obtain a green run of the runner.
  • Pin the runner's two runtime dependencies so all 80 assertions can execute in a plain shell.

8 evidence references cited in the private binder for this control (not shown here — see Provenance below).

A standing rule that every item the system surfaces carries an origin label (requested by the owner, raised by an agent, fixed by self-healing, produced by a hook or scheduled job, or origin unknown); nothing mechanical enforces it.

Attributes tested
Trigger
RequiresEvery surfaced item
Evidence roleone line in the standing-rules file naming the six labels
PresentY
Actor
RequiresThe model, at generation time
Evidence rolethe same rule, addressed to the model
PresentY
Action
RequiresPrefix one of six labels
Evidence rolethe six labels enumerated in the rule
PresentY
Record produced
RequiresNone beyond the session output itself
Evidence rolenone designed
PresentN
Enforcement
RequiresA mechanical check
Evidence rolenone in the hook registry; the register states none exists
PresentN
Population

Not definable: 'surfaced item' has no mechanical boundary, so no population can be drawn.

Sample

None; test of operating effectiveness not performed. An advisory scan was run instead and is not asserted on.

Procedures performed
  1. Located the rule as a single line in the standing-rules file and confirmed it is the only occurrence of the six-label set there.
  2. Reviewed the hook registry: no hook is named for origin labels, consistent with the register's own statement that nothing enforces them.
  3. Advisory scan, not a test: over 777 session records from the period, 3,908 substantial assistant replies were found and 16 of them (0.4%) carried one of the six labels. The denominator is loose, so the figure is directional only.
Results (0 pass / 0 exceptions)

Not tested; the advisory scan found 16 labelled replies among 3,908.

Exceptions (0)

None.

Observations (1)
  • The labels appear in well under 1% of substantial replies in the period, consistent with the register's own statement that nothing enforces them.
Conclusion

Not testable (policy only)

Not testable (policy only) — the rule is written and nothing records, counts, or enforces it, so there is no population to test and no check that can fail. The advisory scan (16 of 3,908 substantial replies labelled) suggests the policy is seldom applied; that is an observation for the register, not a tested result.

Follow-ups filed

None.

3 evidence references cited in the private binder for this control (not shown here — see Provenance below).

Workflows billed per API call must state an estimated cost per unit and get a yes before they run; a scheduled audit checks daily for exposed API keys and weekly for token burn, but does not measure spend.

Attributes tested
Trigger (preventive)
RequiresAny metered workflow states an estimate and asks before running
Evidence rolethe cost-flagging line in the standing-rules file
PresentY
Actor
RequiresThe model estimates and asks; the owner says yes
Evidence rolethe same rule; no hook enforces it
PresentY
Detective audit, on demand
RequiresA spend audit that queries real spend and tallies token burn
Evidence rolethe on-demand cost-audit command and its two tested scripts
PresentY
Detective audit, scheduled
RequiresDaily key-exposure, new-scheduler and dead-target checks; weekly full audit
Evidence rolethe daily security-mode scheduled job and the weekly ecosystem audit it is folded into
PresentY
Record produced
RequiresOne line per run prepended to the cost log
Evidence rolethe cost log in the records folder
PresentY
Population

Preventive half: metered API runs in the period — cannot be established, because no spend measurement exists to enumerate them from. Detective half: scheduled audit runs in the period — 29 daily security-mode lines (expected 30; 2026-09-28 missing, corroborated by the job's own stdout log) and 4 weekly full-audit lines.

Sample

Preventive: whole-population search of September session records for a real cost estimate. Detective: all 33 audit lines read.

Procedures performed
  1. Searched all 777 September session records for the cost-flag marker: 569 records contain it, but the hits are the rule's own text echoed into context. Filtered to a flag followed by an actual dollar figure: 0 instances in the period.
  2. Read every September audit line: all 29 daily lines report key exposure clean, no keys in the shell profile, 7 fast schedulers, 0 dead targets; all 4 weekly lines report real spend 'not checked (no admin key)' and token burn of about 4.1 million tokens per month across 10 jobs (3.8 million across 9 on 2026-09-06).
  3. Confirmed the administrative billing credential is absent from the shell profile (0 matches), so the spend query cannot run. Ran the scheduled job's test suite: 13 of 13 passed.
Results (1 pass / 2 exceptions)

The weekly audit ran 4 of 4 times but never measured spend; the daily audit ran 29 of 30 days; the preventive estimate-and-approve step could not be tested because no metered run can be identified and no estimate was found.

Exceptions (2)
  • Metered spend was not measured on any day of the period: the administrative billing credential has never been issued, so the spend script cannot query the vendor's cost report. The control's objective has no detective backstop. Design deficiency; the command's own text already names this as the pending human step.
  • The daily security-mode audit missed 2026-09-28 (29 of 30); not root-caused, consistent with a fleet-wide missed-fire pattern noted elsewhere. Deficiency, low; recorded.
Observations (1)
  • The preventive half has no record: an estimate and approval exist only in the transcript of the session that asked, and zero such estimates were found in the period.
Conclusion

Not testable (policy only)

Not testable (policy only) — the preventive half (estimate + approval) is a rule with no record and no identifiable population of metered runs; the detective audit ran 29 of 30 days and 4 of 4 weeks but, with no administrative credential, never measured spend. The control as written cannot be shown to operate against its own objective.

Follow-ups filed
  • Owner action: issue the administrative billing credential so the spend script can measure real spend; confirm whether an open work item already exists before filing a new one.

9 evidence references cited in the private binder for this control (not shown here — see Provenance below).

A daily scheduled job diffs the remote host's tool interface against the previous snapshot and flags tools renamed, added, or removed; state lives in plain files so vendors can be swapped, and a written procedure covers model deprecation.

Attributes tested
Trigger
RequiresDaily, morning
Evidence rolea loaded scheduled job with a fixed morning fire time
PresentY
Actor
RequiresAn unattended job pulls today's inventory; a tested script computes the diff
Evidence rolethe job's wrapper and the name-level diff script it calls
PresentY
Action
RequiresDiff against the most recent prior snapshot; append a work item on delta; stop loudly on any unusable verdict
Evidence rolethe wrapper's prior-snapshot lookup, its append step, and its fatal-exit branches
PresentY
Record produced
RequiresA daily snapshot file and a log line
Evidence rolethe snapshot directory and the job's log
PresentY
Flap exclusion
RequiresNames whose presence depends on connector connection state are excluded from the diff
Evidence rolethe exclusion predicate in the diff script, moved into testable code on 2026-09-13
PresentY
Portability design
RequiresState in plain files, not vendor memory
Evidence rolethe portability rule in the standing-rules file
PresentY
Deprecation procedure
RequiresA written, dated procedure
Evidence rolethe model-migration procedure document (updated 2026-09-13)
PresentY
Population

Daily runs from 2026-09-01 to 2026-09-30: 31 fires on 30 distinct days (one re-fire hit the idempotence skip), with 28 snapshots written; no day missing.

Sample

All 30 days classified by outcome; every delta line and every fatal line read.

Procedures performed
  1. Classified the September log: 28 clean completions, 20 'no delta' exits, 9 delta lines on 8 days, and 3 days (2026-09-01, 2026-09-08, 2026-09-29) where the inventory pull returned non-JSON and the job stopped loudly without a snapshot or diff.
  2. Compared the delta lines to the 2026-09-13 predicate fix: before it, four days (2026-09-03, 2026-09-05, 2026-09-09, 2026-09-10) flagged the known connector-state flap; on 2026-09-10 and 2026-09-11 it detected two real host changes (a new third-party connector tool, then a host release adding eight tools and removing ten); after the fix, one core tool name flapped on 2026-09-12, 2026-09-17 and 2026-09-18, producing three work items for no vendor change. The three items exist in the work-item registers.
  3. Ran the diff script's test suite: 45 of 45 passed, including the two fail-loud cases (malformed snapshot yields empty output and a non-zero exit). The suite carries its own two mutation arms and three deliberately-wrong predicates, with a mutation run recorded 2026-09-13; not re-mutated here.
Results (4 pass / 3 exceptions)

Ran 30 of 30 days, completed 27 with a diff, caught both real host changes; 3 days produced no diff, 4 pre-fix flap lines, 3 post-fix flap lines.

Exceptions (3)
  • 3 of 30 days (10%) produced no diff because the inventory pull returned non-JSON; the job failed loud and the next day diffed against the last good snapshot, so a change on a failed day is caught one day late, not missed. Deficiency, low; recorded.
  • After the 2026-09-13 fix, a core tool name flapped on 2026-09-12, 2026-09-17 and 2026-09-18, appending three work items for no vendor change; the exclusion predicate covers connector-state names only. Deficiency, low (false positives, not misses); follow-up filed.
  • 4 pre-fix flap lines on 2026-09-03, 2026-09-05, 2026-09-09 and 2026-09-10: the known defect the 2026-09-13 change fixed. Deficiency, closed in period.
Observations (1)
  • The diff is name-level only; a vendor changing a tool's parameters without renaming it is not detected. A delta is flagged by appending a work item; there is no notification path in the wrapper's logic as read.
Conclusion

Effective with exception(s)

Effective with exception(s) — the diff ran on all 30 days, detected the two real host changes in the period (2026-09-10 and 2026-09-11) and flagged them, and its predicate is covered by a suite with its own mutation arms. Three days produced no diff (loud failures), and the post-fix flap shows the exclusion predicate still admits one class of false positive.

Follow-ups filed
  • Add a stability rule to the diff (for example, require a name to be absent on two consecutive runs before flagging it) and a suite case for it.

10 evidence references cited in the private binder for this control (not shown here — see Provenance below).

Two worked examples

One control where the test passed cleanly, and one where it found a real gap. Both are walked through the same five steps: define the population, select the sample, re-perform, result, conclude.

Worked pass

OP-02 · Hollow-run detection

1 / 5

Define population

The control's scheduled scan ran every three hours against every automation's recent output; the period in scope was every scan from the day the control was installed through the end of the test period, covering 82 real output files.

2 / 5

Select sample

Every one of the six real hits the detector caught in the period was read in full, rather than a smaller slice, because six was a small enough number to check completely.

3 / 5

Re-perform

Each hit was read against the detector's two-arm rule — output too small, or output containing a provider-level error at any size — to confirm the rule, not just the detector's label, actually applied.

4 / 5

Result

All six hits were confirmed correctly classified: five were caught by the content check even though they were larger than the simple size threshold, and each produced exactly one alert, with no repeat noise across roughly 160 later scans.

5 / 5

Conclude

Effective — from the day it was installed, the detector scanned on schedule, correctly classified six real failed runs by their actual output rather than their claimed status, alerted once each, and a deliberate break-it test on a sandboxed copy of its logic confirmed the safety check catches a real mistake, not a cosmetic one.

Worked failure

AC-01 · Destructive-command blocks

1 / 5

Define population

Every shell command in the test period that this safeguard actually stopped: 172 blocks of a recursive force-delete, plus one block of a cloud-document deletion.

2 / 5

Select sample

Eleven items were sampled: five blocks spread across the month, the one cloud-document block, three lines found in a separate command log that happened to mention a force-delete, and two fresh re-performances on the day of testing.

3 / 5

Re-perform

Each sampled command's exact text was re-run against the live safeguard. Of three suspicious log lines, two turned out to be harmless descriptive text; one was a real destructive delete that had executed without being stopped. The real delete was wrapped inside a case-statement arm — a conditional branch in the shell script — and the safeguard's classifier only inspects the first word of each command segment, which in this shape is the branch's own keyword, not the delete itself.

4 / 5

Result

Six of six sampled blocks reproduced correctly when re-fed to the live safeguard today. The one real miss executed because of the case-statement-arm shape. The safeguard's own standing proof that it has 'zero fail-open results' uses the same first-word logic that missed this case, so it cannot, by construction, see this class of gap.

5 / 5

Conclude

Effective with exception(s) — the safeguard was in place all period, unchanged, and every sampled block reproduces today; one recursive force-delete executed through the case-statement-arm shape, so the objective was not fully met for this one pattern. A fix has been filed (2026-10-01) and has not yet landed as of this page's most recent render.

What the testing changed

Testing is not neutral — doing it changed some claims this system made about itself.

  • The register's own claim of how three controls were tested moved from a read-through to a real re-performance with sandboxed, deliberate-break proof (AC-01, AC-03, and the shared wiring checker behind OP-03).
  • One control's register entry carried two different claims about its own testing history in two different places; this binder's lead sheet now carries the reconciled, single figure.
  • A leak test now covers only 3 of 21 public pages on a standing, build-blocking basis; 5 more were checked once by hand at publish time; 13 have never had an automated leak test at all. That gap was previously stated qualitatively; it is now a number.
  • A shared wiring check — confirming a safeguard is actually connected under today's naming, not just present in a config file somewhere — is used by only 6 of 30 safeguard test suites.
  • Of 52 structured root-cause investigations run across the period, roughly 70% produced a real fix of some kind, but only about 7% of those fixes across the system's whole history are actually switched on and running; the rest sit built but unused.
  • 26 of 30 nightly offsite backups succeeded; one failure was explained and fixed the same day, three are unexplained because the only run log does not survive a restart.

Provenance

This page is built from the private testing binder by a generator that runs a leak test on its own rendered output. No fenced or licensed source was read while building it. Numbers on this page are cited to the binder by section, in a separate claims manifest, not retyped by hand.

Tester the system's own agent, inside the system it governs.

Reviewer vacant. No independent review has been performed on this binder.

Build structured fields (ids, titles, conclusion strings, counts) are extracted mechanically from the private workpapers and hashed; prose is authored by hand at role level and checked for required plain-language pairs before this page is generated.

Glossary

Every technical term used on this page, defined in plain language. Hover, focus, or tap any underlined term above for the same definition inline.

Agent
An instance of the AI working on a task, able to use tools on its own within limits it's been given.
Allowlist
A short list of the only things allowed through; anything not on the list is rejected.
Classifier
The part of a program that decides which category something falls into — here, usually "safe" or "dangerous."
Commit
A saved, dated record of a change, kept permanently in a change-history system.
Dashboard
A page that shows the current status of many automated jobs at a glance.
Escalation
Raising a problem's visibility — telling someone, or trying harder to fix it — because a first attempt didn't work.
Fail-closed
When something goes wrong inside a safeguard, it defaults to blocking, not letting things through.
Fail-open
When something goes wrong inside a safeguard, it defaults to letting things through rather than blocking — done on purpose only where blocking everything would be worse.
Fence
A boundary marking certain material as off-limits, so it's never read or used when building something else.
Generator
A program that builds a finished page or document automatically from source material, rather than someone writing it by hand.
Hollow run
A scheduled task that reported itself as successful but actually produced nothing useful.
Hook
A small program that automatically runs at a specific moment to check or stop an action.
Leak test
An automatic check that scans a finished public page for anything private that shouldn't be there.
Manifest
A written list, generated for one specific action, naming each item it covers and why.
Marker
A small file or flag that records a temporary state — for example, that something was recently approved.
Matcher
The part of a safeguard's setup that decides which actions it should even look at.
Messaging channel
The app or service used to send and receive messages with the system remotely.
Model
The underlying AI system that generates responses and decisions.
Mutation test
Deliberately breaking a safety check on purpose, in a safe copy, to prove the check would actually notice if something really went wrong.
Noindex
A setting telling search engines not to list a page, so it doesn't show up in search results even though it's publicly reachable.
Offsite copy
A backup kept somewhere separate from the main system, so a problem with the main system doesn't destroy the backup too.
Pattern match
Checking a piece of text against a known shape or wording, rather than understanding its meaning.
Pre-execution hook
A small program that runs automatically right before a command executes, with the power to stop it.
Remote host
A computer or service the system talks to over the internet, rather than running locally.
Runtime
The software environment a program actually runs inside, as opposed to how it reads on paper.
Sandbox
A safe, throwaway copy of something, used for testing without any risk to the real version.
Scheduled job
A task set up to run automatically on a timer, with no person needing to start it.
Self-healing
A system noticing its own problem and automatically trying to fix it, without a person stepping in.
Session
One continuous stretch of the system working on something, start to finish.
Snapshot
A saved copy of something at one point in time, so it can be restored later if something goes wrong.
Source control
A system that keeps a dated history of every change to a file, so old versions can always be found again.
Stdin payload
The data fed into a program as its input, the way a safeguard actually receives the command it's checking.
Suite
A collected set of automated tests for one piece of the system.
Tool call
One single action the system takes — running a command, editing a file, sending a message.
Transcript
The full saved record of one working session, including everything done and every result produced.
Watchdog
A program whose only job is to check that other automated jobs are still running and working.