Justina Dong · AI Governance · Control register

The control register.21 controls, and where each one falls short.

My working attempt at governing my own lab of 37 AI agents: 21 controls, mapped to ISO/IEC 42001, NIST AI RMF and four other frameworks, with the gaps left in.

21 controls
5 areas
6 frameworks
5 known gaps

How far each control is proven

  • 7 enforced and tested
  • 11 partly proven
  • 3 rule only

Nobody has settled how to implement or test AI governance yet, me included. This is one attempt. AI drafted the controls, mapped them to the frameworks and ran the tests. I directed it, decided what risk to accept, and I own the result. Nobody else has reviewed it. It shows how I designed the controls, not that anything is compliant.

What they're for: Governance

02 · Risk register

What could go wrong, and how much I'll live with.

I run this as a one-person lab. Most of what I build can only break things for me, so I take on a lot of risk on purpose. The line sits where harm could reach someone else, or can't be undone.

Where I drew the line

  • ✓Accepted

    Anything whose worst realistic outcome is rework, noise, or a degraded page that someone notices within a week.

  • −Accepted if caught within seven days

    Anything that can fail while still reporting healthy. A detective control has to catch it, because silence, not failure, is what this system can't tolerate.

  • !Never accepted, at any score
    • Irreversible loss of data with no offsite copy
    • Disclosure of sensitive personal material to any surface the owner didn't approve
    • An autonomous action against a third party (a send, a publish, a purchase, a delete) with no human checkpoint
    • Any breach of an employer's client confidentiality

Where the risks sit

13
risks shown
9
outside tolerance

1 more is withheld from this version, because describing it would disclose what it protects. It is outside tolerance.

  • R1Outside tolerance
  • R4Inside tolerance
  • Low
  • Moderate
  • High
  • Critical

Likelihood and impact are each my own 1 to 5 rating, on scales I defined, and the score is the two multiplied. The scales are ordinal, so a 10 isn't twice a 5, and where a risk sits tells you more than its number. Before means without the current controls; after means with them. I first scored these on August 7, 2026 and re-scored them on September 29, 2026. Nobody else has checked them yet. The next review is due December 29, 2026.

The risks

Tap a risk for its scores, how I'm treating it, and the controls that cover it.

  • ✓Inside tolerance
  • !Outside tolerance
!R1A guard against an irreversible deletion goes inert after a tool is renamed, and nothing notices10 High
Tolerance
Outside tolerance.
Score
Before: likelihood 3 × impact 5 = 15, CriticalAfter: likelihood 2 × impact 5 = 10, High
Treatment
Mitigate
Controls
✓AC-01, −OP-03. Both deletion guards now have tests that check they are still wired in. The tests run when something changes, not on a timer.
Change
Down from 15. The last open deletion path was closed in August.
!R2Instructions hidden in content the agents read reach a runtime that runs with permission prompts switched off10 High
Tolerance
Outside tolerance.
Score
Before: likelihood 3 × impact 5 = 15, CriticalAfter: likelihood 2 × impact 5 = 10, High
Treatment
Mitigate
Controls
−AC-03, −AC-02. Only one sender can reach the runtime. Nothing inspects the content itself.
Change
Unchanged.
!R3Private personal data reaches a public page10 High
Tolerance
Outside tolerance.
Score
Before: likelihood 2 × impact 5 = 10, HighAfter: likelihood 2 × impact 5 = 10, High
Treatment
Mitigate
Controls
−OI-02. The build-don't-strip rule, plus automated leak tests on two public surfaces. Most hand-written pages have no leak test.
Change
Unchanged score; two surfaces now checked by a test.
✓R4A model the system names by version is retired in the middle of a pipeline6 Moderate
Tolerance
Inside tolerance.
Score
Before: likelihood 4 × impact 2 = 8, ModerateAfter: likelihood 3 × impact 2 = 6, Moderate
Treatment
Accept
Controls
−OI-05, ✓OP-02. A written migration procedure, and a measured exposure of two scripts. Scheduled runs that hit an API error are flagged.
Change
Down from 12. Measuring the exposure lowered its impact.
!R5The machine is lost while part of the system has no current offsite copy8 Moderate
Tolerance
Outside tolerance.
Score
Before: likelihood 3 × impact 4 = 12, HighAfter: likelihood 2 × impact 4 = 8, Moderate
Treatment
Mitigate
Controls
✓CM-03, −OP-05. The main store is pushed offsite nightly and was current on the review date. The agent-control layer has no offsite copy.
Change
Down from 12. The backlog of unpushed changes seen in August is cleared.
✓R6A self-healing job moves or rewrites many files on a bad heuristic6 Moderate
Tolerance
Inside tolerance.
Score
Before: likelihood 2 × impact 4 = 8, ModerateAfter: likelihood 2 × impact 3 = 6, Moderate
Treatment
Accept
Controls
−AC-02, ✓CM-03. Nightly versioning makes a mass change reversible.
Change
Unchanged.
!R7Several credentials expire at once while the owner is away8 Moderate
Tolerance
Outside tolerance: no reliable detective control.
Score
Before: likelihood 4 × impact 3 = 12, HighAfter: likelihood 4 × impact 2 = 8, Moderate
Treatment
Mitigate
Controls
✓OP-01, ✓OP-02. Re-authentication procedures and a job watchdog. A single-credential version happened on the review date: jobs failed for about two hours, recovered on their own, and were found by a manual check.
Change
Unchanged score. Partly realized.
!R8An operating-system upgrade breaks many scheduled jobs at once9 Moderate
Tolerance
Outside tolerance: no reliable detective control.
Score
Before: likelihood 3 × impact 3 = 9, ModerateAfter: likelihood 3 × impact 3 = 9, Moderate
Treatment
Mitigate
Controls
✓OP-01. No post-upgrade smoke test exists.
Change
Unchanged.
!R9Work items are marked closed when the work was never done9 Moderate
Tolerance
Outside tolerance.
Score
Before: likelihood 3 × impact 3 = 9, ModerateAfter: likelihood 3 × impact 3 = 9, Moderate
Treatment
Mitigate
Controls
None. No closed item is ever sampled and re-performed.
Change
Unchanged.
!R10An unattended loop runs up metered spend9 Moderate
Tolerance
Outside tolerance.
Score
Before: likelihood 3 × impact 3 = 9, ModerateAfter: likelihood 3 × impact 3 = 9, Moderate
Treatment
Mitigate
Controls
!OI-04. Spend needs a yes before a metered workflow runs. The spend audit runs on demand only; the daily scheduled audit checks for exposed keys, not spend.
Change
Moved out of tolerance. The August scoring assumed a daily spend check that doesn't exist.
✓R11Agents act on standing rules that have since changed6 Moderate
Tolerance
Inside tolerance.
Score
Before: likelihood 4 × impact 2 = 8, ModerateAfter: likelihood 3 × impact 2 = 6, Moderate
Treatment
Accept
Controls
None. A weekly reconciliation job and two agent audit jobs.
Change
Unchanged.
✓R12Two writers update the same register file and one set of changes is lost4 Low
Tolerance
Inside tolerance.
Score
Before: likelihood 3 × impact 3 = 9, ModerateAfter: likelihood 2 × impact 2 = 4, Low
Treatment
Accept
Controls
−OP-05, ✓CM-03. Snapshots before each edit and nightly versioning.
Change
Unchanged.
!R14The owner is unavailable while scheduled jobs keep acting8 Moderate
Tolerance
Outside tolerance: no reliable detective control.
Score
Before: likelihood 2 × impact 4 = 8, ModerateAfter: likelihood 2 × impact 4 = 8, Moderate
Treatment
Accept
Controls
None. No pause switch, stand-down procedure, or successor.
Change
Unchanged. Accepted explicitly: one person can't be engineered away.

Where agents touch my real data

What an agent can do in each place, and what stops a mistake there.

−My working filesDocuments synced to cloud storageRecoverable, not reviewed
Access
Read and write. The whole system runs on these files.
Limits
Hooks block recursive deletes and deleting cloud documents (AC-01). Changing many files at once needs a manifest with a reason for each file (AC-02). Every night the files are committed and pushed offsite (CM-03), so a bad change can be rolled back.
Review
Not reviewed first. Recoverable after.
✓PhotosHome network storageCan't change it
Access
Read only.
Limits
The storage server gives the agents' account read access only, so the limit is enforced by the server, not by the agent.
Review
Not needed. It can't write.
−Financial accountsAn account aggregatorRead only, by the tool
Access
Read only, through the tool the agents are given.
Limits
The tool has read calls only, with no way to move money or change anything. The login behind it isn't read-only, so this rests on the tool, not the credential.
Review
Not needed for reading. Anything involving money needs my yes (AC-02), and that part is a rule, not code.
!Email and messagesMail and chat accountsRule only
Access
Can send.
Limits
Every send needs my yes, one approval per send (AC-02). That's a rule the model follows. No hook checks it yet.
Review
Yes, by rule only.

03 · The controls

Five areas, twenty-one controls.

Tap a control to see the risk, the control, and the gaps.

How far each control is proven

  • ✓Enforced and testedRuns as code, and a test has exercised it.
  • −Partly provenRuns, but only a walkthrough or partial test shows it works, or part of it is a rule the model follows.
  • !Rule onlyA rule the model is asked to follow with nothing checking it, or not running yet.

I work the grade out mechanically from each control's own "how it runs" and "tested" lines, so anyone can check my working. It's one way to grade a control, not a standard. Tested means a test ran when the code changed; nothing is re-tested on a schedule yet (known gap 02).

Governance & risk

4 controls
−GV-01Stated risk tolerance and a forward-looking risk registerForeseeable risks are identified, scored, and compared against a stated tolerance before they occur.
The risk

Risks get noticed only after they happen. The system gets good at closing incidents and never states which risks it has decided to accept.

The control

Each quarter, and whenever a risk realizes or a new autonomous surface is added, the builder re-scores every risk in the register: likelihood and impact before and after its current control, and whether what is left sits inside the written tolerance. Each risk outside tolerance gets a named treatment, tracked as a work item until it closes. The review is recorded in the register as a dated line that sets the next due date.

The gaps

Written by the builder it governs. The scores are ordinal judgment anchors, not measured rates, and the register says so.

suite coverage, surveyed 2026-09-29Testing: Operated once; the published copy is tested →

  • Detective
  • Manual
  • Quarterly, and on trigger events
Grade
Partly proven. Runs, but only a walkthrough or partial test shows it works, or part of it is a rule the model follows.
Evidence
The risk register, published on the governance page with names and paths removed: each risk's scores before and after its control, its treatment, and the review date. Read it in chapter 02 · Read it on /governance
How it was mapped to the six frameworks
NIST AI RMF
  • PartialMANAGE 1.4Negative residual risks (defined as the sum of all unmitigated risks) to both downstream acquirers of AI systems and end users are documented.Residual scores are ordinal judgments, not a sum of unmitigated risk.
  • FullGOVERN 1.3Processes, procedures, and practices are in place to determine the needed level of risk management activities based on the organization’s risk tolerance.
  • FullMAP 1.5Organizational risk tolerances are determined and documented.
GenAI Profile
No mapping
ISO/IEC 42001
  • PartialCl. 6 PlanningRisk assessment and tolerance; the objectives are in GV-04.
COSO 2013
  • FullP7 identifies and analyzes risk
SOX ITGC
Out of scopeThis is entity-level risk assessment, a COSO component that a SOX program scopes from. It isn't an IT general control.
OWASP LLM Top 10
No mapping
Every framework item, in the appendix →
−GV-02Written change, back-out, and retirement proceduresChange, rollback, and decommissioning are documented, including the actions that can't be reversed.
The risk

The gates exist only as code. Nobody can say how a change is supposed to go, how to undo one, or how to retire an agent or a job without breaking whatever depended on it.

The control

Plain-language procedures cover how a change is proposed, reviewed, built, verified, and recorded; how each class of change is rolled back; which actions can't be rolled back at all (sends, publishes, deletions of cloud pointers); and how an agent or scheduled job is retired. They're reviewed whenever the watched-path table changes.

The gaps

This documents the gates; it doesn't add any. The procedure says so itself: where it and a hook disagree, the hook is right and the document is stale.

suite coverage, surveyed 2026-09-29Testing: Read through, no record kept →

  • Directive
  • Manual
  • On change to the gates, plus quarterly
Grade
Partly proven. Runs, but only a walkthrough or partial test shows it works, or part of it is a rule the model follows.
Evidence
Procedure documents with dated review lines
How it was mapped to the six frameworks
NIST AI RMF
  • PartialMANAGE 4.1Post-deployment AI system monitoring plans are implemented, including mechanisms for capturing and evaluating input from users and other relevant AI actors, appeal and override, decommissioning, incident response, recovery, and change management.Covers change, recovery and retirement; no monitoring plan or user input.
  • FullGOVERN 1.7Processes and procedures are in place for decommissioning and phasing out AI systems safely and in a manner that does not increase risks or decrease the organization’s trustworthiness.
GenAI Profile
No mapping
ISO/IEC 42001
  • PartialCl. 8 OperationDocuments how changes are planned; CM-01 enforces it.
  • PartialA.6 AI system life cycleChange, rollback and retirement stages only.
COSO 2013
  • FullP12 deployed through policy and procedure
SOX ITGC
  • PartialIn scope: Program changeThe written policy; the operating evidence comes from CM-01 and CM-03.
The written change-management policy an ITGC walkthrough starts from.
OWASP LLM Top 10
No mapping
Every framework item, in the appendix →
−GV-03Privacy impact assessmentThe categories of personal data, where they're stored, and every flow in and out are inventoried and assessed.
The risk

Personal data spreads to places no one mapped, because the system grew faster than anyone's picture of where the data sits and where it flows.

The control

A privacy impact assessment inventories categories of personal data (never the data itself), maps the surfaces data flows to and from, and grades perimeter, minimisation, and flow awareness separately.

The gaps

Done once and not yet repeated. An assessment that isn't re-run describes the system as it was.

suite coverage, surveyed 2026-09-29Testing: Read through, no record kept →

  • Directive
  • Manual
  • Point-in-time
Grade
Partly proven. Runs, but only a walkthrough or partial test shows it works, or part of it is a rule the model follows.
Evidence
The assessment, graded by dimension
How it was mapped to the six frameworks
NIST AI RMF
  • FullMEASURE 2.10Privacy risk of the AI system – as identified in the MAP function – is examined and documented.
GenAI Profile
  • PartialData PrivacyAssesses the exposure; prevention sits in OI-02.
ISO/IEC 42001
  • PartialA.5 Assessing impacts of AI systemsPrivacy impacts only, not wider effects on people or society.
  • PartialA.7 Data for AI systemsInventories the data agents read; not its quality or provenance.
COSO 2013
  • PartialP7 identifies and analyzes riskOne risk domain, assessed once.
SOX ITGC
Out of scopePrivacy is a separate objective from financial reporting. It belongs in a privacy program or SOC 2 scope.
OWASP LLM Top 10
No mapping
Every framework item, in the appendix →
!GV-04Stated mission and measurable objectivesThe mission and objectives are written down, each objective has a measure, and each reports its current status, including when it isn't met.
The risk

The system grows toward whatever is interesting that week. With no written purpose or targets, nobody can say whether a control serves anything, or what a risk would cost if it happened.

The control

The owner writes the mission and a short list of objectives, each with a measure and a status taken from data. They're published on the governance page and reviewed at each quarterly register re-score. An objective whose status the data can't support is marked unmeasured or unmet.

The gaps

Written by the person it governs. Three of the four objectives have no measurement yet, and the register's risks were scored before the objectives existed.

suite coverage, surveyed 2026-09-29Testing: Not operated yet; the published copy is tested →

  • Directive
  • Manual
  • Quarterly, with the register re-score
Grade
Rule only. A rule the model is asked to follow with nothing checking it, or not running yet.
Evidence
The mission and objectives on the governance page, each objective with its measure and status. Read it on /governance
How it was mapped to the six frameworks
NIST AI RMF
  • FullMAP 1.3The organization’s mission and relevant goals for AI technology are understood and documented.
GenAI Profile
No mapping
ISO/IEC 42001
  • PartialCl. 6 PlanningObjectives only; risk planning stays with GV-01.
COSO 2013
  • PartialP6 objectives clear enough to assess riskRegister risks were scored before these objectives existed and aren't tied to them yet.
SOX ITGC
Out of scopeMission and objectives are entity-level, the part of COSO a SOX program scopes from. They aren't an IT general control.
OWASP LLM Top 10
No mapping
Every framework item, in the appendix →

Change management

4 controls
✓CM-01Spec-and-review build gateChanges to privileged paths are specified and independently reviewed before they're made.
The risk

An agent changes load-bearing configuration, an automation, or a public page without a written spec and without a second look. Or the builder decides for itself that the change is too small to need either.

The control

Pre-execution hooks intercept file writes and edits, plus shell commands that write through redirects or in-place edits, whenever the target is on a shared table of watched paths. The action is blocked unless a spec file lists the path, or a separate reviewer agent recorded a clean audit within the last two hours. A narrow size exemption covers trivial edits, and the builder can't waive the gate for anything more.

The gaps

The reviewer is a second agent, not a second person, so the segregation of duties is agent from agent. Some native system calls that rewrite scheduled jobs aren't matched, and sessions on another model provider fire none of these hooks.

suite coverage, surveyed 2026-09-29Testing: Shell path tested; the file-edit path isn't →

  • Preventive
  • Automated
  • Every write to a watched path
Grade
Enforced and tested. Runs as code, and a test has exercised it.
Evidence
Gate decision log, timestamped review marker, and the spec file listing the paths it authorizes
How it was mapped to the six frameworks
NIST AI RMF
  • FullGOVERN 1.4The risk management process and its outcomes are established through transparent policies, procedures, and other controls based on organizational risk priorities.
  • PartialMANAGE 4.1Post-deployment AI system monitoring plans are implemented, including mechanisms for capturing and evaluating input from users and other relevant AI actors, appeal and override, decommissioning, incident response, recovery, and change management.Change management only.
GenAI Profile
  • PartialHuman-AI ConfigurationStops the agent certifying its own change; the reviewer is also an agent.
ISO/IEC 42001
  • FullCl. 8 Operation
  • FullA.6 AI system life cycle
COSO 2013
  • FullP11 general controls over technology
SOX ITGC
  • PartialIn scope: Program changeThe approver is a second agent, not a second person.
Authorization and approval of a change before it reaches production.
OWASP LLM Top 10
  • PartialLLM06 Excessive AgencyLimits unreviewed changes; doesn't narrow the agent's tool permissions.
Every framework item, in the appendix →
✓CM-02Hold window on rule changesAutomated changes to foundational rules are held back from the offsite record until a human has had time to look.
The risk

A nightly job pushes a change to the system's own governing rules before anyone awake has seen it.

The control

When the reconciliation job changes a foundational rule, it commits the change with a hold tag. The nightly commit-and-push job looks for any held commit less than 48 hours old and skips the push until that commit ages out.

The gaps

The change takes effect locally the moment it's written. The hold only delays the offsite copy, and it requires no sign-off: if nobody looks within 48 hours, the push goes through.

suite coverage, surveyed 2026-09-29Testing: Tested, but never shown it can fail →

  • Preventive
  • Automated
  • Daily
Grade
Enforced and tested. Runs as code, and a test has exercised it.
Evidence
Hold-tagged commits, and push-skipped entries in the job log
How it was mapped to the six frameworks
NIST AI RMF
  • PartialMANAGE 4.1Post-deployment AI system monitoring plans are implemented, including mechanisms for capturing and evaluating input from users and other relevant AI actors, appeal and override, decommissioning, incident response, recovery, and change management.A review window over rule changes, with no sign-off.
GenAI Profile
No mapping
ISO/IEC 42001
  • PartialCl. 8 OperationDelays the offsite record, not the change.
COSO 2013
  • PartialP11 general controls over technologyHolds the backup, not the change.
SOX ITGC
Out of scopeThe rule is live as soon as it's written, and the push is a backup, so the hold isn't a promotion control. It's a review window over the record.
OWASP LLM Top 10
No mapping
Every framework item, in the appendix →
✓CM-03Versioned change recordConfiguration changes are recorded, dated, and recoverable.
The risk

Nobody can reconstruct what the system's rules were on a given date, or when they changed.

The control

Every night a scheduled job commits the working system (standing rules, procedures, registers, build scripts) and pushes it to a private remote. The same job commits the agent-control layer (agent specs, hooks, settings) to its own repository, locally.

The gaps

The agent-control repository has no offsite remote, and most scheduled-job definitions aren't under version control. Granularity is one day, and the recorded author is the job, not whoever made the change.

suite coverage, surveyed 2026-09-29Testing: Automated test that can fail →

  • Detective
  • Automated
  • Daily
Grade
Enforced and tested. Runs as code, and a test has exercised it.
Risks
Treats R5, R6 and R12 in the risk register.
Evidence
Commit history
How it was mapped to the six frameworks
NIST AI RMF
  • PartialMANAGE 4.1Post-deployment AI system monitoring plans are implemented, including mechanisms for capturing and evaluating input from users and other relevant AI actors, appeal and override, decommissioning, incident response, recovery, and change management.Keeps the change record; it doesn't manage change.
GenAI Profile
No mapping
ISO/IEC 42001
  • FullCl. 7 Support
  • PartialA.6 AI system life cycleRecords configuration changes, not operating event logs.
COSO 2013
  • FullP11 general controls over technology
SOX ITGC
  • PartialIn scope: Program changeThe job is the recorded author, so the record doesn't show who made each change.
The population of changes a tester would sample from to test CM-01.
OWASP LLM Top 10
No mapping
Every framework item, in the appendix →
−CM-04Test before promote, including a test that proves the check can failNew and changed automations are tested before they're relied on, including against an input that should fail.
The risk

A change ships because it ran once, and nothing shows that the check it adds would ever say no.

The control

New skills and commands default to test-first, with a companion test suite, and a hook nudges any new one written without it. The standing method for proving a check can fail is to copy the code to a sandbox, break it on purpose, confirm the test goes red, and restore the original byte for byte.

The gaps

Nothing re-runs the suites on a schedule. A suite shows green on the day someone runs it, and that's all it shows.

suite coverage, surveyed 2026-09-29Testing: Tested, but never shown it can fail →

  • Detective
  • Tests run by the builder; automated nudge
  • Every change
Grade
Partly proven. Runs, but only a walkthrough or partial test shows it works, or part of it is a rule the model follows.
Evidence
Suite output and mutation transcripts
How it was mapped to the six frameworks
NIST AI RMF
  • PartialMEASURE 2.3AI system performance or assurance criteria are measured qualitatively or quantitatively and demonstrated for conditions similar to deployment setting(s). Measures are documented.Shows a check works and can fail; no performance criteria are set.
GenAI Profile
No mapping
ISO/IEC 42001
  • PartialA.6 AI system life cycleVerification before use; no validation against stated criteria.
COSO 2013
  • FullP11 general controls over technology
  • PartialP16 ongoing and separate evaluationsEvaluates a control before it's relied on, on change only.
SOX ITGC
  • PartialIn scope: Program developmentThe builder runs the tests; nothing re-runs them on a schedule.
Testing of new and changed programs before they go into use.
OWASP LLM Top 10
No mapping
Every framework item, in the appendix →

Access & privileged actions

3 controls
✓AC-01Destructive-command blocksAn agent can't execute irreversible destructive operations.
The risk

An agent runs a recursive delete, or deletes cloud-document pointers from the synced drive, and the loss spreads to the cloud copy.

The control

Pre-execution hooks inspect every shell command, and block recursive force-deletes and deletion of cloud-document placeholder files.

The gaps

The hooks match patterns. A deletion done another way, through a script or a different tool, isn't caught.

suite coverage, surveyed 2026-09-29Testing: Automated test that can fail →

  • Preventive
  • Automated
  • Every shell command
Grade
Enforced and tested. Runs as code, and a test has exercised it.
Risks
Treats R1 in the risk register.
Evidence
Block notices in the session transcript
How it was mapped to the six frameworks
NIST AI RMF
  • PartialMANAGE 2.4Mechanisms are in place and applied, and responsibilities are assigned and understood, to supersede, disengage, or deactivate AI systems that demonstrate performance or outcomes inconsistent with intended use.Blocks specific destructive actions; it can't disengage an agent.
GenAI Profile
No mapping
ISO/IEC 42001
  • PartialA.9 Use of AI systemsKeeps use within intended actions, by pattern.
COSO 2013
  • FullP11 general controls over technology
SOX ITGC
  • PartialIn scope: Access to programs and dataRestricts one class of privileged function by pattern, not by permission.
Restricting privileged functions that can destroy data.
OWASP LLM Top 10
  • PartialLLM06 Excessive AgencyPattern-based; a deletion by another route isn't caught.
Every framework item, in the appendix →
−AC-02Human approval for bulk and outward-facing actionsBulk changes and actions that leave the machine require an explicit, item-level human yes.
The risk

The agent acts on a whole class of items at once, or sends, publishes, or spends, on its own reading of an instruction.

The control

A hook blocks a shell command that looks like it will change many files at once until a manifest exists listing each file with its own reason. Sends, publishes, deploys, and anything involving money need the owner's explicit approval, and one approval covers one action.

The gaps

The send-and-spend half is a rule the model applies, and no hook checks it. By this system's own standard, that half is a policy, not a control.

suite coverage, surveyed 2026-09-29Testing: Bulk part tested; sends and spend aren't →

  • Preventive
  • Automated for bulk changes; model-applied policy for sends and spend
  • Every such action
Grade
Partly proven. Runs, but only a walkthrough or partial test shows it works, or part of it is a rule the model follows.
Risks
Treats R2 and R6 in the risk register.
Evidence
Manifest files keyed to the exact command; approvals in the transcript
How it was mapped to the six frameworks
NIST AI RMF
  • FullGOVERN 3.2Policies and procedures are in place to define and differentiate roles and responsibilities for human-AI configurations and oversight of AI systems.
  • PartialMAP 3.5Processes for human oversight are defined, assessed, and documented in accordance with organizational policies from the GOVERN function.Defined and documented; nothing assesses the send-and-spend half.
GenAI Profile
  • PartialHuman-AI ConfigurationEnforced for bulk changes; model-applied for sends and spend.
ISO/IEC 42001
  • PartialA.9 Use of AI systemsHuman oversight of use; half of it is model-applied.
COSO 2013
  • FullP10 control activities that mitigate risk
  • PartialP12 deployed through policy and procedureThe send-and-spend half is a policy nothing checks.
SOX ITGC
Out of scopeApproving individual actions is an application or business-process control in SOX terms. It would be tested inside the process, not as an IT general control.
OWASP LLM Top 10
  • PartialLLM06 Excessive AgencyEnforced for bulk changes; model-applied for sends.
  • PartialLLM01 Prompt InjectionLimits what an injected instruction can do; doesn't detect it.
Every framework item, in the appendix →
−AC-03Inbound channel restrictionOnly the owner can issue remote instructions, and a remote trigger can't carry an arbitrary command.
The risk

Anyone who can message the remote channel can steer the agents or trigger production jobs.

The control

The remote messaging listener acts only on messages from an allowlisted sender ID. A remote job trigger resolves to a fixed list of scheduled jobs that must also be loaded, and it maps to a fixed argument list, never a command string taken from the input.

The gaps

It controls who can send instructions. It doesn't stop instructions hidden in content the agents read, such as web pages, email, and documents. That exposure is handled by AC-02's approval rules and is a scored risk in GV-01.

suite coverage, surveyed 2026-09-29Testing: Read through, no record kept →

  • Preventive
  • Automated
  • Every inbound message
Grade
Partly proven. Runs, but only a walkthrough or partial test shows it works, or part of it is a rule the model follows.
Risks
Treats R2 in the risk register.
Evidence
Listener log and rejected-trigger notices
How it was mapped to the six frameworks
NIST AI RMF
No mapping
GenAI Profile
  • PartialInformation SecurityCloses the remote channel to strangers; injection through content stays open.
ISO/IEC 42001
No mapping
COSO 2013
  • FullP11 general controls over technology
SOX ITGC
  • FullIn scope: Access to programs and data
Authentication and authorization over who can run production jobs.
OWASP LLM Top 10
  • PartialLLM01 Prompt InjectionBlocks instructions from unknown senders only.
  • PartialLLM06 Excessive AgencyLimits what a remote trigger can run; other paths aren't covered.
Every framework item, in the appendix →

Operations & monitoring

5 controls
✓OP-01Job monitoring and self-healing by exceptionScheduled jobs that go silent or stale are caught within the hour, and the owner hears only about problems.
The risk

A scheduled job stops running, or runs and fails, and nobody notices for days.

The control

Every hour a watchdog reads the fleet's health table. For any job that has gone silent or stale it tries a restart, then reports the problem, the action taken, and what it means. A status dashboard regenerates every 15 minutes.

The gaps

It sees whether a job ran and reported in. A job that exits cleanly after doing nothing looks healthy here; OP-02 covers that case.

suite coverage, surveyed 2026-09-29Testing: Alerts tested; the restart isn't checked →

  • Detective and corrective
  • Automated
  • Hourly
Grade
Enforced and tested. Runs as code, and a test has exercised it.
Risks
Treats R7 and R8 in the risk register.
Evidence
Dashboard snapshots and alert history
How it was mapped to the six frameworks
NIST AI RMF
  • PartialMEASURE 2.4The functionality and behavior of the AI system and its components – as identified in the MAP function – are monitored when in production.Watches that jobs ran, not what they produced.
  • PartialMANAGE 4.1Post-deployment AI system monitoring plans are implemented, including mechanisms for capturing and evaluating input from users and other relevant AI actors, appeal and override, decommissioning, incident response, recovery, and change management.Monitoring and recovery for scheduled jobs only.
GenAI Profile
No mapping
ISO/IEC 42001
  • PartialA.6 AI system life cycleOperation and monitoring of scheduled jobs only.
COSO 2013
  • FullP11 general controls over technology
SOX ITGC
  • FullIn scope: Computer operations
Job scheduling and monitoring, with failures followed up.
OWASP LLM Top 10
No mapping
Every framework item, in the appendix →
✓OP-02Hollow-run detectionA run counts as successful only if it actually produced output.
The risk

The status field reports success for runs that produced nothing.

The control

Model-executed scheduled automations were reporting success even when the model died before doing any work. A detector now classifies each run by its output rather than its status: output that's too small, or that contains an API error at any length, marks the run hollow.

The gaps

It only detects. Whether to retry is deliberately left as a separate decision.

suite coverage, surveyed 2026-09-29Testing: Automated test that can fail →

  • Detective
  • Automated
  • Scheduled, over each run
Grade
Enforced and tested. Runs as code, and a test has exercised it.
Risks
Treats R4 and R7 in the risk register.
Evidence
Detector log
How it was mapped to the six frameworks
NIST AI RMF
  • PartialMEASURE 2.4The functionality and behavior of the AI system and its components – as identified in the MAP function – are monitored when in production.Scheduled model runs only.
GenAI Profile
No mapping
ISO/IEC 42001
  • PartialA.6 AI system life cycleOperation and monitoring of scheduled model runs only.
COSO 2013
  • FullP11 general controls over technology
SOX ITGC
  • FullIn scope: Computer operations
A job that reports success with no output is the textbook operations exception.
OWASP LLM Top 10
No mapping
Every framework item, in the appendix →
−OP-03Control wiring checksEnforcement hooks are shown to be registered against the tool names the runtime actually uses.
The risk

A control is counted as operating when it can't fire at all.

The control

A shared wiring check lets a hook's test suite assert that the hook is registered, under a matcher that fires for today's tool name, and invoke it through the exact command the configuration specifies. It was built after an enforcement gate sat unwired for about six weeks because a tool it watched had been renamed. The suites for the destructive-command blocks (AC-01) and the bulk-change gate (AC-02) use it; most hook suites don't yet.

The gaps

Coverage is partial: most hook suites, including the one behind OI-01, don't yet assert their own wiring. And suites run only when something changes (CM-04), so a rename between runs goes unnoticed until the next one.

suite coverage, surveyed 2026-09-29Testing: No test of its own →

  • Detective
  • Automated, inside the test suites that use it
  • When a covered suite runs
Grade
Partly proven. Runs, but only a walkthrough or partial test shows it works, or part of it is a rule the model follows.
Risks
Treats R1 in the risk register.
Evidence
Suite output
How it was mapped to the six frameworks
NIST AI RMF
  • PartialMEASURE 1.2Appropriateness of AI metrics and effectiveness of existing controls are regularly assessed and updated, including reports of errors and potential impacts on affected communities.Runs on change, and most hook suites don't use it yet.
GenAI Profile
No mapping
ISO/IEC 42001
  • PartialCl. 9 Performance evaluationEvaluates whether controls can fire; coverage is partial.
COSO 2013
  • PartialP16 ongoing and separate evaluationsRuns on change, with partial coverage.
SOX ITGC
Out of scopeThis is management testing whether controls operate, which COSO calls monitoring. It isn't an ITGC itself; in a SOX program it's the evidence that supports relying on the ITGCs.
OWASP LLM Top 10
No mapping
Every framework item, in the appendix →
−OP-04Root cause and corrective actionEvery significant failure gets a root cause, and a tracked corrective action.
The risk

The same failure recurs because the fix addressed the symptom.

The control

Significant failures get a structured root cause: criteria, condition, cause, consequence, corrective action. The corrective action becomes a tracked item, and incidents are written up as blameless postmortems.

The gaps

A tracked item shows the fix was written down, not that it shipped. Some corrective actions produce a guard that isn't wired in.

suite coverage, surveyed 2026-09-29Testing: Only the alert version is tested →

  • Corrective
  • Manual, model-assisted
  • Per incident
Grade
Partly proven. Runs, but only a walkthrough or partial test shows it works, or part of it is a rule the model follows.
Evidence
Root-cause records, tracked corrective items, published postmortems
How it was mapped to the six frameworks
NIST AI RMF
  • FullMANAGE 4.3Incidents and errors are communicated to relevant AI actors, including affected communities. Processes for tracking, responding to, and recovering from incidents and errors are followed and documented.
  • PartialGOVERN 4.3Organizational practices are in place to enable AI testing, identification of incidents, and information sharing.Incident identification and sharing; testing is CM-04's.
GenAI Profile
No mapping
ISO/IEC 42001
  • FullCl. 10 Improvement
COSO 2013
  • FullP17 evaluates and communicates deficiencies
SOX ITGC
  • FullIn scope: Computer operations
Incident and problem management.
OWASP LLM Top 10
No mapping
Every framework item, in the appendix →
−OP-05Register snapshots and offsite copyThe shared registers can be restored to a point before a damaging write.
The risk

A bad write corrupts a register that several sessions share, and no earlier copy exists.

The control

Before a file-edit tool call, a hook snapshots the shared registers that agents read and write. The nightly commit puts the vault offsite (CM-03).

The gaps

Snapshots sit in temporary storage and don't survive a restart. They fire on the file-edit tools only, so a register rewritten from the shell relies on the nightly copy.

suite coverage, surveyed 2026-09-29Testing: Read through, no record kept →

  • Corrective
  • Automated
  • Every edit tool call; daily offsite
Grade
Partly proven. Runs, but only a walkthrough or partial test shows it works, or part of it is a rule the model follows.
Risks
Treats R5 and R12 in the risk register.
Evidence
Snapshot copies and commit history
How it was mapped to the six frameworks
NIST AI RMF
  • PartialMANAGE 4.1Post-deployment AI system monitoring plans are implemented, including mechanisms for capturing and evaluating input from users and other relevant AI actors, appeal and override, decommissioning, incident response, recovery, and change management.Recovery only.
GenAI Profile
No mapping
ISO/IEC 42001
No mapping
COSO 2013
  • FullP11 general controls over technology
SOX ITGC
  • PartialIn scope: Computer operationsSnapshots don't survive a restart, and the offsite copy is daily.
Backup and recovery.
OWASP LLM Top 10
No mapping
Every framework item, in the appendix →

Output integrity & data

5 controls
✓OI-01Unverified completion-claim blockA claim that work is done stands only if the same turn verified it against real state.
The risk

The agent reports work as done when it isn't, and a success log line stands in for the artifact.

The control

At the end of every agent turn, a hook checks three things: whether the agent changed something, whether it then claimed success, and whether anything in that turn verified the claim. If there's a change and a claim with no verification, it blocks the turn and sends the agent back to check.

The gaps

It recognizes claims by their wording. The same false claim, phrased in words the hook doesn't know, gets through.

suite coverage, surveyed 2026-09-29Testing: Automated test that can fail →

  • Preventive
  • Automated
  • Every agent turn
Grade
Enforced and tested. Runs as code, and a test has exercised it.
Evidence
Block records in the transcript
How it was mapped to the six frameworks
NIST AI RMF
  • PartialMEASURE 2.4The functionality and behavior of the AI system and its components – as identified in the MAP function – are monitored when in production.Checks completion claims by their wording.
GenAI Profile
  • PartialConfabulationCatches unverified completion claims, not other false statements.
ISO/IEC 42001
  • PartialA.6 AI system life cycleA verification step inside operation.
COSO 2013
  • PartialP13 relevant, quality informationProtects one kind of information: claims that work is done.
SOX ITGC
Out of scopeThe nearest SOX idea is completeness and accuracy of information the entity produces (IPE). That's tested report by report inside business processes, not as an ITGC.
OWASP LLM Top 10
  • PartialLLM09 MisinformationCompletion claims only.
Every framework item, in the appendix →
−OI-02Public surfaces are built, never strippedNothing public is made by redacting the private system in place.
The risk

Personal data reaches a public page through a copied artifact or a delegated task.

The control

The rule is that public artifacts are built from curated or fabricated material, never by stripping the private system in place. The public learnings reader, which is generated from private notes, runs a scrub gate. This page's build runs a leak test that fails on real names, local file paths, or account identifiers.

The gaps

Coverage is per surface, not universal: several hand-written public pages have no automated leak test. And a test only knows the identifiers it was given.

suite coverage, surveyed 2026-09-29Testing: Automated test that can fail →

  • Preventive
  • Policy, plus automated tests on some surfaces
  • Per public build, where a test exists
Grade
Partly proven. Runs, but only a walkthrough or partial test shows it works, or part of it is a rule the model follows.
Risks
Treats R3 in the risk register.
Evidence
Scrub-gate and leak-test output
How it was mapped to the six frameworks
NIST AI RMF
  • PartialMEASURE 2.10Privacy risk of the AI system – as identified in the MAP function – is examined and documented.Tested per public surface, not everywhere.
GenAI Profile
  • PartialData PrivacyCovers what reaches public pages.
ISO/IEC 42001
No mapping
COSO 2013
  • PartialP12 deployed through policy and procedureA policy everywhere, tested on some surfaces.
SOX ITGC
Out of scopeConfidentiality of personal data is a privacy objective, not a financial-reporting one.
OWASP LLM Top 10
  • PartialLLM02 Sensitive Information DisclosureTests published pages for identifiers; not model output generally.
Every framework item, in the appendix →
!OI-03Decision-origin labelsEvery item the system surfaces says where it came from.
The risk

Delegation turns into abdication. The owner can no longer tell what she asked for from what the system decided on its own.

The control

The rule is that every surfaced item carries an origin label: requested by the owner, raised by an agent, fixed by self-healing, produced by a hook or a scheduled job, or origin unknown.

The gaps

Nothing mechanical checks the labels. By this system's own rule, something the model consults and nothing enforces is a policy, not a control. It's listed so the gap stays visible.

suite coverage, surveyed 2026-09-29Testing: Not tested →

  • Detective
  • Model-applied policy
  • Every surfaced item
Grade
Rule only. A rule the model is asked to follow with nothing checking it, or not running yet.
Evidence
Labels in session output
How it was mapped to the six frameworks
NIST AI RMF
  • PartialGOVERN 3.2Policies and procedures are in place to define and differentiate roles and responsibilities for human-AI configurations and oversight of AI systems.Labels who started what; nothing enforces them.
GenAI Profile
  • PartialHuman-AI ConfigurationUnenforced labels.
ISO/IEC 42001
  • PartialA.3 Internal organizationAccountability for AI-initiated work, unenforced.
COSO 2013
  • PartialP14 internal communicationInternal communication, unenforced.
SOX ITGC
Out of scopeAccountability for what the AI initiated has no financial-reporting assertion. The part that's SOX-adjacent, the audit trail, is covered by CM-03.
OWASP LLM Top 10
No mapping
Every framework item, in the appendix →
!OI-04Spend controlMetered spend is estimated and approved before it's incurred.
The risk

An unattended loop runs up metered API spend.

The control

Workflows billed per API call have to state an estimated cost per unit and get a yes before they run. A spend audit runs on demand. The daily scheduled run of the same audit checks for exposed API keys, not for spend.

The gaps

The estimate is the model's own arithmetic, applied as policy.

suite coverage, surveyed 2026-09-29Testing: Only the new spend flag is tested →

  • Preventive
  • Model-applied policy
  • Per run
Grade
Rule only. A rule the model is asked to follow with nothing checking it, or not running yet.
Risks
Treats R10 in the risk register.
Evidence
Cost estimate and approval in the transcript; cost log
How it was mapped to the six frameworks
NIST AI RMF
  • PartialMANAGE 2.1Resources required to manage AI risks are taken into account – along with viable non-AI alternative systems, approaches, or methods – to reduce the magnitude or likelihood of potential impacts.Weighs cost per run; non-AI alternatives aren't considered.
GenAI Profile
No mapping
ISO/IEC 42001
No mapping
COSO 2013
  • PartialP10 control activities that mitigate riskA policy the model applies, with no check.
SOX ITGC
Out of scopeThis is personal operating spend, with no financial statement behind it. In an enterprise it belongs in procurement and FinOps controls.
OWASP LLM Top 10
  • PartialLLM10 Unbounded ConsumptionApproval per run; no hard spending cap.
Every framework item, in the appendix →
−OI-05Vendor change watch and portabilityVendor-side changes to the agent host are detected, and no vendor holds load-bearing state.
The risk

A vendor changes a model, tool, or interface underneath the system, and the controls that depended on it stop working without a sound.

The control

Every morning a scheduled job diffs the agent host's tool interface and flags tools that were renamed, added, or removed. The system is designed so models and tools can be swapped out, with its state kept in plain files rather than a vendor's memory. A written procedure covers model deprecation.

The gaps

The diff watches one vendor's tool interface. Other connectors are covered by the portability design, not by monitoring.

suite coverage, surveyed 2026-09-29Testing: Automated test that can fail →

  • Detective
  • Automated
  • Daily
Grade
Partly proven. Runs, but only a walkthrough or partial test shows it works, or part of it is a rule the model follows.
Risks
Treats R4 in the risk register.
Evidence
Diff reports
How it was mapped to the six frameworks
NIST AI RMF
  • FullGOVERN 6.2Contingency processes are in place to handle failures or incidents in third-party data or AI systems deemed to be high-risk.
  • PartialMANAGE 3.1AI risks and benefits from third-party resources are regularly monitored, and risk controls are applied and documented.Monitors one vendor's tool interface.
GenAI Profile
  • PartialValue Chain and Component IntegrationOne vendor's interface is watched; the rest rely on portability.
ISO/IEC 42001
  • PartialA.10 Third-party and customer relationshipsWatches a supplier; no supplier assessment.
COSO 2013
  • PartialP9 assesses significant changeDetects one kind of external change.
SOX ITGC
Out of scopeSOX reaches vendors through SOC 1 reports and complementary user-entity controls, and no service organization here supports financial reporting.
OWASP LLM Top 10
  • PartialLLM03 Supply ChainDetects interface changes; doesn't vet components.
Every framework item, in the appendix →

04 · Testing

How each control was tested.

AI ran every test here, on controls AI also built, at my direction. So none of it is independent yet. Tap a control for what I did, what it found, and one way it could be tested.

One way to test a control

I borrowed the two tests financial auditors use. Whether they fit AI systems well is part of what I'm trying out.

1 · Test of design

Walk one real instance from start to finish, and confirm the control as built deals with the risk.

2 · Test of operating effectiveness

Show it worked across a period. For a manual control, inspect a sample sized by how often it runs. For an automated one, test each case once, try to make it fail, and rely on change control to show the logic didn't change.

In an audit, the person testing is not the person who runs the control. Here, for now, it is.

One published example of minimum sample sizes, for a manual control
  • 1Yearly
  • 2 to 3Quarterly
  • 2 to 4Monthly
  • 5 to 10Weekly
  • 15 to 30Daily
  • 30 to 60Many times a day

From KPMG, Sarbanes-Oxley Section 404: Management's Assessment Process (2005), p. 10. The guide calls these examples and says each company should size its samples on its own facts.

How far my testing goes

  • Not testedNo test and no recorded review.2
  • Read throughThe design was read against the risk. No record was kept.4
  • Tested in partAn automated test covers part of the control, or never shows it can fail.9
  • Test can failAn automated test covers it and goes red when the control is broken.6
  • Independent, over a periodSomeone other than the builder tested it across a period. None yet.0

The grade in chapter 03 says how a control runs. This says how hard it has been tested.

Governance & risk

GV-01Stated risk tolerance and a forward-looking risk registerOperated once; the published copy is tested · suite coverage, surveyed 2026-09-29
What I did
  1. Re-scored the risks in the register on 2026-09-29, before and after each control, against the written tolerance.
  2. Wrote the review into the register as a dated line that sets the next due date.
  3. The page's test suite checks the published copy: a risk marked inside tolerance when it isn't, a wrong score, or a link to a control that doesn't exist must each fail.
What it found

register, pre-binder: Operated once: re-scored on 2026-09-29 by the builder, not independently. One residual score rose and three fell. Next due 2026-12-29.

binder conclusion, 2026-10-01: Effective with exception(s) · how →

One way it could be tested

Walk through one quarterly re-score. To test that it operates, inspect two quarterly reviews and re-score a few risks yourself. Ideally the person checking isn't the person who scored.

← Back to GV-01 in the controls
GV-02Written change, back-out, and retirement proceduresRead through, no record kept · suite coverage, surveyed 2026-09-29
What I did
  1. Read the change, back-out, retirement and model-migration procedures against the risk. No record of the read-through was kept.
What it found

register, pre-binder: Design walkthrough only

Checked 2026-09-29: none of the four documents has a dated review line yet, and one line in the back-out procedure still describes the snapshot step as it used to be.

binder conclusion, 2026-10-01: Effective with exception(s) · how →

One way it could be tested

Inspect the procedures for a dated review. Then pick changes made to the gates during the period and check that each one followed the written change and back-out steps.

← Back to GV-02 in the controls
GV-03Privacy impact assessmentRead through, no record kept · suite coverage, surveyed 2026-09-29
What I did
  1. Read the assessment, dated 2026-08-07. No record of the read-through was kept.
What it found

register, pre-binder: Design walkthrough only

binder conclusion, 2026-10-01: Not operated in period · how →

One way it could be tested

Inspect the assessment against the data the system actually holds, and check it was redone after any new data source was added.

← Back to GV-03 in the controls
GV-04Stated mission and measurable objectivesNot operated yet; the published copy is tested · suite coverage, surveyed 2026-09-29
What I did
  1. Not operated yet. The first review is due with the register re-score on 2026-12-29.
  2. The page's test suite checks how it is published: a mission changed by one character, or an objective marked met while a risk sits outside tolerance, must fail.
What it found

register, pre-binder: Not yet operated; written 2026-09-29.

binder conclusion, 2026-10-01: Not operated in period · how →

One way it could be tested

Inspect that each objective has a measure, then inspect two quarterly reviews that compare each measure with its target.

← Back to GV-04 in the controls

Change management

CM-01Spec-and-review build gateShell path tested; the file-edit path isn't · suite coverage, surveyed 2026-09-29
What I did
  1. The gate's suites try writes to protected paths through a shell redirect and expect each one blocked, and try plain reads and expect them allowed.
  2. They break the gate on purpose, by making its path table crash or switching the guard off, and expect the suite to go red.
  3. A parity test adds a fake protected path and checks that all three entry points pick it up.
What it found

register, pre-binder: Test suites, run on change: the shell-redirect path, and scope parity across the three gate entry points

All green in the full run on 2026-09-13.

Not covered: the file-edit entry point has no suite of its own, and the gate keeps no decision log, so that piece of evidence doesn't exist yet.

binder conclusion, 2026-10-01: Effective with exception(s) · how →

One way it could be tested

Automated, so test one of each case: a write to a watched path with no approved spec should be blocked, and one with approval should go through. Then rely on change control over the gate itself to show the logic did not change during the period.

← Back to CM-01 in the controls
CM-02Hold window on rule changesTested, but never shown it can fail · suite coverage, surveyed 2026-09-29
What I did
  1. The commit job's suite builds a throwaway repository, commits a rule change tagged as held, runs the job and checks nothing was pushed.
What it found

register, pre-binder: Test suite for the commit job, run on change

28 of 28 passed on 2026-09-14.

Not covered: nothing tests that a hold ends after 48 hours. The hold has never fired for real, so there are no held commits to inspect.

binder conclusion, 2026-10-01: Not operated in period · how →

One way it could be tested

Force the condition once: commit a rule change inside the hold window and confirm the push is held. Then read the job log for the period and confirm no held commit was pushed early.

← Back to CM-02 in the controls
CM-03Versioned change recordAutomated test that can fail · suite coverage, surveyed 2026-09-29
What I did
  1. The same suite checks that the nightly commit comes out the same with the network up or down, that another session's staged work stays out of it, and that a failing pre-commit check is not retried.
What it found

register, pre-binder: Test suite for the commit job, run on change

28 of 28 passed on 2026-09-14. The 2026-09-29 nightly run committed both repositories.

binder conclusion, 2026-10-01: Effective with exception(s) · how →

One way it could be tested

Pick files that changed on the live system and trace each one to a commit. Then check whether anyone reviews the record, since a detective control usually needs a reviewer.

← Back to CM-03 in the controls
CM-04Test before promote, including a test that proves the check can failTested, but never shown it can fail · suite coverage, surveyed 2026-09-29
What I did
  1. The nudge's suite writes a new command file with no test beside it and expects the reminder, then writes files that already exist or already have a test and expects silence.
What it found

register, pre-binder: The nudge hook has its own test suite

Green on 2026-09-13.

Not covered: the nudge watches new commands but not new skills, and its suite never breaks the nudge to prove it can fail.

binder conclusion, 2026-10-01: Effective with exception(s) · how →

One way it could be tested

Sample changes from the period, about 25 for a control that runs on every change, and inspect each for a passing suite and a mutation run that failed as it should, before promotion.

← Back to CM-04 in the controls

Access & privileged actions

AC-01Destructive-command blocksAutomated test that can fail · suite coverage, surveyed 2026-09-29
What I did
  1. The suites feed in dozens of destructive commands, such as recursive force-deletes, find with delete, and deletes of folders that hold Drive files, and expect each one blocked. Safe commands must pass.
  2. They break the command classifier on purpose and expect the dangerous command to stay blocked, and they check the hooks are wired as registered.
What it found

register, pre-binder: Test suites, run on change

All green on 2026-09-13.

binder conclusion, 2026-10-01: Effective with exception(s) · how →

One way it could be tested

In a sandbox, try each class of blocked command and confirm the block. Compare the block list with the destructive commands the system can actually run, and rely on change control for the rest of the period.

← Back to AC-01 in the controls
AC-02Human approval for bulk and outward-facing actionsBulk part tested; sends and spend aren't · suite coverage, surveyed 2026-09-29
What I did
  1. The bulk-change suite replays four real past incidents and expects each one blocked, then checks that nine named files pass, ten need a manifest, and a manifest written for a different command doesn't count.
What it found

register, pre-binder: Test suite for the bulk-change gate, run on change

Green on 2026-09-13.

Not covered: the gate counts the manifest's lines but doesn't check that each file has its own reason, and the approval rule for sends and spend has no test.

binder conclusion, 2026-10-01: Effective with exception(s) / Not testable (policy only) · how →

One way it could be tested

Automated part: try a bulk change with no manifest (blocked) and with one (allowed). Policy part: sample about 25 sends and paid runs from the logs and inspect each for an approval recorded before the action.

← Back to AC-02 in the controls
AC-03Inbound channel restrictionRead through, no record kept · suite coverage, surveyed 2026-09-29
What I did
  1. Read the listener code: the sender allowlist and the fixed list of jobs a message can start. No record of the read-through was kept, and no test sends a message from an unlisted sender.
What it found

register, pre-binder: Design walkthrough only (read in code 2026-09-29)

binder conclusion, 2026-10-01: Design only (no TOE possible) · how →

One way it could be tested

Send a message from an unlisted sender and from a listed one, and confirm only the listed one triggers anything. Inspect the allowlist and who can change it.

← Back to AC-03 in the controls

Operations & monitoring

OP-01Job monitoring and self-healing by exceptionAlerts tested; the restart isn't checked · suite coverage, surveyed 2026-09-29
What I did
  1. The watchdog's suites run it in a sandbox with system calls stubbed: a sweep that fails silently must be reported as blind, and a third blind run in a row must send an alert.
  2. They plant a known bug in a copy and expect the check to catch it.
What it found

register, pre-binder: Test suite for the watchdog, run on change

67 of 67 and 17 of 17 passed on 2026-09-29.

Not covered: the suite records the restart attempt but never checks that it happened.

binder conclusion, 2026-10-01: Effective with exception(s) · how →

One way it could be tested

Break one job on purpose and confirm the alert fires and the fix runs. Then take the alert history for the period and trace each alert to its resolution.

← Back to OP-01 in the controls
OP-02Hollow-run detectionAutomated test that can fail · suite coverage, surveyed 2026-09-29
What I did
  1. The detector's suite feeds in runs that finished empty or hit a rate limit and expects an alert, and healthy runs just over the size line and expects none. Each alert rule is also tested on its own.
What it found

register, pre-binder: Test suite, run on change

Green on 2026-09-13.

binder conclusion, 2026-10-01: Effective · how →

One way it could be tested

Feed in a run that finishes with empty output and confirm it is flagged. Then sample the detector log and trace each flag to its follow-up.

← Back to OP-02 in the controls
OP-03Control wiring checksNo test of its own · suite coverage, surveyed 2026-09-29
What I did
  1. The wiring check runs inside six other suites and confirms each hook there is registered the way the suite expects.
What it found

register, pre-binder: This control is itself a test

Not covered: the check has no test of its own, so nothing shows it would catch a hook that isn't wired.

binder conclusion, 2026-10-01: Effective with exception(s) · how →

One way it could be tested

List which controls have a wiring check and which do not. Break one wire and confirm its check fails.

← Back to OP-03 in the controls
OP-04Root cause and corrective actionOnly the alert version is tested · suite coverage, surveyed 2026-09-29
What I did
  1. No test of the root-cause process itself. The alert version has a suite, and the completion-claim block (OI-01) checks that a fix reported as done was actually made.
What it found

register, pre-binder: Design walkthrough only

binder conclusion, 2026-10-01: Effective with exception(s) · how →

One way it could be tested

Take every incident in the period, or a sample of about 25 if there are more, and inspect each for a root-cause record and a corrective item that was closed and checked.

← Back to OP-04 in the controls
OP-05Register snapshots and offsite copyRead through, no record kept · suite coverage, surveyed 2026-09-29
What I did
  1. Read the snapshot hook. No record of the read-through was kept, and no restore test is on record.
What it found

register, pre-binder: Design walkthrough only

Checked 2026-09-29: 55 snapshots were taken between 2026-09-28 and 2026-09-29.

binder conclusion, 2026-10-01: Effective with exception(s) · how →

One way it could be tested

Restore a register file from the snapshot and from the offsite copy, and compare both with the original.

← Back to OP-05 in the controls

Output integrity & data

OI-01Unverified completion-claim blockAutomated test that can fail · suite coverage, surveyed 2026-09-29
What I did
  1. The suite replays made-up sessions: a change followed by a claim of success with no check must be blocked, while the same words inside a plan document must not be.
  2. A fix reported as done in a session that changed nothing must still be blocked.
What it found

register, pre-binder: Test suite, run on change

Green on 2026-09-13, and run again on 2026-09-17 after the last change.

binder conclusion, 2026-10-01: Effective with exception(s) · how →

One way it could be tested

Send a turn that claims the work is done with no proof and confirm it is blocked; send one with proof and confirm it passes. Rely on change control for the rest of the period.

← Back to OI-01 in the controls
OI-02Public surfaces are built, never strippedAutomated test that can fail · suite coverage, surveyed 2026-09-29
What I did
  1. This page's test plants an internal agent name and a home-folder path into a copy of the built page and expects the check to fail, then checks that the real page passes.
  2. The learnings page has its own scrub gate with its own test.
What it found

register, pre-binder: This page: leak test with a planted negative control

71 of 71 passed on 2026-09-29.

Not covered: the leak check runs when someone runs the test, not on every build or push, and it looks for internal names and paths, not a list of real people.

binder conclusion, 2026-10-01: Effective with exception(s) · how →

One way it could be tested

List every public surface and whether it has a leak test. Where one exists, plant a known private string and confirm the build fails. Where none exists, inspect the surface by hand.

← Back to OI-02 in the controls
OI-03Decision-origin labelsNot tested · suite coverage, surveyed 2026-09-29
What I did
  1. None. This is a rule the model is asked to follow, with nothing checking it.
What it found

register, pre-binder: Not tested

binder conclusion, 2026-10-01: Not testable (policy only) · how →

One way it could be tested

Sample about 25 surfaced items from session output and check that each carries an origin label and that the label is right.

← Back to OI-03 in the controls
OI-04Spend controlOnly the new spend flag is tested · suite coverage, surveyed 2026-09-29
What I did
  1. The approve-before-spend rule has no test. Since 2026-09-29 the daily cost job has a suite for one new part: when spend can't be measured it must raise a flag, and a security-only run must leave that alone.
What it found

register, pre-binder: Design walkthrough only

The line above is out of date: a suite now exists for the flag, though not for the approval rule.

binder conclusion, 2026-10-01: Not testable (policy only) · how →

One way it could be tested

Sample paid runs from the cost log and inspect each for a cost estimate and an approval recorded before the run. Reconcile the cost log to the provider's bill.

← Back to OI-04 in the controls
OI-05Vendor change watch and portabilityAutomated test that can fail · suite coverage, surveyed 2026-09-29
What I did
  1. The diff tool's suite compares two real snapshots and expects no change, injects fake tool names and expects them reported, and treats an empty or crashed result as a failure.
What it found

register, pre-binder: Design walkthrough only

The line above is out of date: this suite has been green since 2026-09-13.

The 2026-09-29 daily run failed to save its snapshot.

binder conclusion, 2026-10-01: Effective with exception(s) · how →

One way it could be tested

Compare the watch list with the vendors actually in use. Confirm the diff report ran every day of the period and that each flagged change was looked at.

← Back to OI-05 in the controls

05 · Known gaps

What I haven't solved yet.

The gaps I know about, measured against the same frameworks. There are probably others I haven't found.

01

Segregation of duties between people

One person owns, builds, operates, reviews, and approves. There's no provisioning or access review, because there's only one user. Agent-from-agent review (CM-01) narrows this gap but doesn't close it.

NIST GOVERN 2.1 · NIST MEASURE 1.3 · SOX: segregation of duties · COSO 10

02

Scheduled operating-effectiveness testing

The test suites and wiring checks run when something changes, never on a timer. No control is re-performed on a sample over time, so every 'tested' line on this page means tested the last time someone ran it.

NIST MEASURE 1.2 · COSO 16

03

Risk metrics with thresholds

A few detectors have measured thresholds (OP-02). There's no set of risk metrics with tolerances tracked over time; the risk register scores on ordinal judgment.

NIST MEASURE 1.1 · COSO 16

04

Gate blind spots

Native calls that edit or unload scheduled jobs aren't matched by the change gate, and sessions on another model provider fire none of the hooks. Both appear in the change procedure itself.

NIST GOVERN 1.4 · SOX ITGC: program change

05

Independent review of this mapping

The person who built the system mapped it. Nobody who knows these frameworks independently has reviewed it yet, so it shows control design, not assurance.

NIST MEASURE 1.3

Sources
  • NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), 2023. Subcategory text is quoted verbatim from the official PDF. Category headings are shortened.
  • NIST AI 600-1, Generative AI Profile, 2024. Risk names from §2.
  • ISO/IEC 42001:2023. Clause and Annex A objective numbers only, with short labels of my own. No text from the standard is reproduced, and no control-level numbers are cited.
  • COSO, Internal Control: Integrated Framework (2013). Principle numbers, paraphrased.
  • OWASP Top 10 for LLM Applications 2025.
  • SOX IT general control domains as commonly scoped: program change, program development, access to programs and data, computer operations.

Appendix · Frameworks

The controls, mapped to six frameworks.

Pick a framework to see which of its items I think a control covers, and how well. Every fit here is my own judgment, open to challenge.

  • ✓MetIn my mapping, a control fits it fully, and that control is enforced and tested.
  • −Partly metIn my mapping, a control fits it fully but isn't enforced and tested yet, or controls cover only part of it.
  • !Not metI haven't mapped any control to it.

I left out an overall letter grade. Many items in these frameworks assume an organization with staff, leadership and affected communities, so I think a letter would grade the scope more than the work.

NIST AI RMF

72 requirements
  • ✓1 met
  • −18 partly met
  • !53 not met

19 of 72 subcategories have a control · 20 of 21 controls map here. 9 of the 19 have a full fit.

GOVERN

GOVERN 1 · Policies, processes, procedures and practices

!GOVERN 1.1No control
“Legal and regulatory requirements involving AI are understood, managed, and documented.”
!GOVERN 1.2No control
“The characteristics of trustworthy AI are integrated into organizational policies, processes, procedures, and practices.”
−GOVERN 1.3Full fit · 1 control
“Processes, procedures, and practices are in place to determine the needed level of risk management activities based on the organization’s risk tolerance.”

Partly met. GV-01 fits it fully, but GV-01 is only partly proven.

  • Full fit−GV-01 Stated risk tolerance and a forward-looking risk register
✓GOVERN 1.4Full fit · 1 control
“The risk management process and its outcomes are established through transparent policies, procedures, and other controls based on organizational risk priorities.”

Met. CM-01 fits it fully and is enforced and tested.

  • Full fit✓CM-01 Spec-and-review build gate
Known gap: Gate blind spots
!GOVERN 1.5No control
“Ongoing monitoring and periodic review of the risk management process and its outcomes are planned and organizational roles and responsibilities clearly defined, including determining the frequency of periodic review.”
!GOVERN 1.6No control
“Mechanisms are in place to inventory AI systems and are resourced according to organizational risk priorities.”
−GOVERN 1.7Full fit · 1 control
“Processes and procedures are in place for decommissioning and phasing out AI systems safely and in a manner that does not increase risks or decrease the organization’s trustworthiness.”

Partly met. GV-02 fits it fully, but GV-02 is only partly proven.

  • Full fit−GV-02 Written change, back-out, and retirement procedures

GOVERN 2 · Accountability structures

!GOVERN 2.1No control
“Roles and responsibilities and lines of communication related to mapping, measuring, and managing AI risks are documented and are clear to individuals and teams throughout the organization.”Known gap: Segregation of duties between people
!GOVERN 2.2No control
“The organization’s personnel and partners receive AI risk management training to enable them to perform their duties and responsibilities consistent with related policies, procedures, and agreements.”
!GOVERN 2.3No control
“Executive leadership of the organization takes responsibility for decisions about risks associated with AI system development and deployment.”

GOVERN 3 · Workforce diversity, equity, inclusion and accessibility

!GOVERN 3.1No control
“Decision-making related to mapping, measuring, and managing AI risks throughout the lifecycle is informed by a diverse team (e.g., diversity of demographics, disciplines, experience, expertise, and backgrounds).”
−GOVERN 3.2Full fit · 2 controls
“Policies and procedures are in place to define and differentiate roles and responsibilities for human-AI configurations and oversight of AI systems.”

Partly met. AC-02 fits it fully, but AC-02 is only partly proven.

  • Full fit−AC-02 Human approval for bulk and outward-facing actions
  • Partial fit!OI-03 Decision-origin labelsLabels who started what; nothing enforces them.

GOVERN 4 · A culture that considers and communicates AI risk

!GOVERN 4.1No control
“Organizational policies and practices are in place to foster a critical thinking and safety-first mindset in the design, development, deployment, and uses of AI systems to minimize potential negative impacts.”
!GOVERN 4.2No control
“Organizational teams document the risks and potential impacts of the AI technology they design, develop, deploy, evaluate, and use, and they communicate about the impacts more broadly.”
−GOVERN 4.3Partial fit · 1 control
“Organizational practices are in place to enable AI testing, identification of incidents, and information sharing.”

Partly met. OP-04 covers part of it. No control fits it fully.

  • Partial fit−OP-04 Root cause and corrective actionIncident identification and sharing; testing is CM-04's.

GOVERN 5 · Engagement with relevant AI actors

!GOVERN 5.1No control
“Organizational policies and practices are in place to collect, consider, prioritize, and integrate feedback from those external to the team that developed or deployed the AI system regarding the potential individual and societal impacts related to AI risks.”
!GOVERN 5.2No control
“Mechanisms are established to enable the team that developed or deployed AI systems to regularly incorporate adjudicated feedback from relevant AI actors into system design and implementation.”

GOVERN 6 · Third-party software, data and supply chain

!GOVERN 6.1No control
“Policies and procedures are in place that address AI risks associated with third-party entities, including risks of infringement of a third-party’s intellectual property or other rights.”
−GOVERN 6.2Full fit · 1 control
“Contingency processes are in place to handle failures or incidents in third-party data or AI systems deemed to be high-risk.”

Partly met. OI-05 fits it fully, but OI-05 is only partly proven.

  • Full fit−OI-05 Vendor change watch and portability

MAP

MAP 1 · Context is established and understood

!MAP 1.1No control
“Intended purposes, potentially beneficial uses, context-specific laws, norms and expectations, and prospective settings in which the AI system will be deployed are understood and documented. Considerations include: the specific set or types of users along with their expectations; potential positive and negative impacts of system uses to individuals, communities, organizations, society, and the planet; assumptions and related limitations about AI system purposes, uses, and risks across the development or product AI lifecycle; and related TEVV and system metrics.”
!MAP 1.2No control
“Interdisciplinary AI actors, competencies, skills, and capacities for establishing context reflect demographic diversity and broad domain and user experience expertise, and their participation is documented. Opportunities for interdisciplinary collaboration are prioritized.”
−MAP 1.3Full fit · 1 control
“The organization’s mission and relevant goals for AI technology are understood and documented.”

Partly met. GV-04 fits it fully, but GV-04 is rule only.

  • Full fit!GV-04 Stated mission and measurable objectives
!MAP 1.4No control
“The business value or context of business use has been clearly defined or – in the case of assessing existing AI systems – re-evaluated.”
−MAP 1.5Full fit · 1 control
“Organizational risk tolerances are determined and documented.”

Partly met. GV-01 fits it fully, but GV-01 is only partly proven.

  • Full fit−GV-01 Stated risk tolerance and a forward-looking risk register
!MAP 1.6No control
“System requirements (e.g., “the system shall respect the privacy of its users”) are elicited from and understood by relevant AI actors. Design decisions take socio-technical implications into account to address AI risks.”

MAP 2 · Categorization of the AI system

!MAP 2.1No control
“The specific tasks and methods used to implement the tasks that the AI system will support are defined (e.g., classifiers, generative models, recommenders).”
!MAP 2.2No control
“Information about the AI system’s knowledge limits and how system output may be utilized and overseen by humans is documented. Documentation provides sufficient information to assist relevant AI actors when making decisions and taking subsequent actions.”
!MAP 2.3No control
“Scientific integrity and TEVV considerations are identified and documented, including those related to experimental design, data collection and selection (e.g., availability, representativeness, suitability), system trustworthiness, and construct validation.”

MAP 3 · Capabilities, usage, goals, benefits and costs

!MAP 3.1No control
“Potential benefits of intended AI system functionality and performance are examined and documented.”
!MAP 3.2No control
“Potential costs, including non-monetary costs, which result from expected or realized AI errors or system functionality and trustworthiness – as connected to organizational risk tolerance – are examined and documented.”
!MAP 3.3No control
“Targeted application scope is specified and documented based on the system’s capability, established context, and AI system categorization.”
!MAP 3.4No control
“Processes for operator and practitioner proficiency with AI system performance and trustworthiness – and relevant technical standards and certifications – are defined, assessed, and documented.”
−MAP 3.5Partial fit · 1 control
“Processes for human oversight are defined, assessed, and documented in accordance with organizational policies from the GOVERN function.”

Partly met. AC-02 covers part of it. No control fits it fully.

  • Partial fit−AC-02 Human approval for bulk and outward-facing actionsDefined and documented; nothing assesses the send-and-spend half.

MAP 4 · Risks and benefits mapped for all components

!MAP 4.1No control
“Approaches for mapping AI technology and legal risks of its components – including the use of third-party data or software – are in place, followed, and documented, as are risks of infringement of a third party’s intellectual property or other rights.”
!MAP 4.2No control
“Internal risk controls for components of the AI system, including third-party AI technologies, are identified and documented.”

MAP 5 · Impacts on people, organizations and society

!MAP 5.1No control
“Likelihood and magnitude of each identified impact (both potentially beneficial and harmful) based on expected use, past uses of AI systems in similar contexts, public incident reports, feedback from those external to the team that developed or deployed the AI system, or other data are identified and documented.”
!MAP 5.2No control
“Practices and personnel for supporting regular engagement with relevant AI actors and integrating feedback about positive, negative, and unanticipated impacts are in place and documented.”

MEASURE

MEASURE 1 · Methods and metrics identified and applied

!MEASURE 1.1No control
“Approaches and metrics for measurement of AI risks enumerated during the MAP function are selected for implementation starting with the most significant AI risks. The risks or trustworthiness characteristics that will not – or cannot – be measured are properly documented.”Known gap: Risk metrics with thresholds
−MEASURE 1.2Partial fit · 1 control
“Appropriateness of AI metrics and effectiveness of existing controls are regularly assessed and updated, including reports of errors and potential impacts on affected communities.”

Partly met. OP-03 covers part of it. No control fits it fully.

  • Partial fit−OP-03 Control wiring checksRuns on change, and most hook suites don't use it yet.
Known gap: Scheduled operating-effectiveness testing
!MEASURE 1.3No control
“Internal experts who did not serve as front-line developers for the system and/or independent assessors are involved in regular assessments and updates. Domain experts, users, AI actors external to the team that developed or deployed the AI system, and affected communities are consulted in support of assessments as necessary per organizational risk tolerance.”Known gap: Segregation of duties between peopleKnown gap: Independent review of this mapping

MEASURE 2 · Evaluation for trustworthy characteristics

!MEASURE 2.1No control
“Test sets, metrics, and details about the tools used during TEVV are documented.”
!MEASURE 2.2No control
“Evaluations involving human subjects meet applicable requirements (including human subject protection) and are representative of the relevant population.”
−MEASURE 2.3Partial fit · 1 control
“AI system performance or assurance criteria are measured qualitatively or quantitatively and demonstrated for conditions similar to deployment setting(s). Measures are documented.”

Partly met. CM-04 covers part of it. No control fits it fully.

  • Partial fit−CM-04 Test before promote, including a test that proves the check can failShows a check works and can fail; no performance criteria are set.
−MEASURE 2.4Partial fit · 3 controls
“The functionality and behavior of the AI system and its components – as identified in the MAP function – are monitored when in production.”

Partly met. OP-01, OP-02 and OI-01 cover part of it. No control fits it fully.

  • Partial fit✓OP-01 Job monitoring and self-healing by exceptionWatches that jobs ran, not what they produced.
  • Partial fit✓OP-02 Hollow-run detectionScheduled model runs only.
  • Partial fit✓OI-01 Unverified completion-claim blockChecks completion claims by their wording.
!MEASURE 2.5No control
“The AI system to be deployed is demonstrated to be valid and reliable. Limitations of the generalizability beyond the conditions under which the technology was developed are documented.”
!MEASURE 2.6No control
“The AI system is evaluated regularly for safety risks – as identified in the MAP function. The AI system to be deployed is demonstrated to be safe, its residual negative risk does not exceed the risk tolerance, and it can fail safely, particularly if made to operate beyond its knowledge limits. Safety metrics reflect system reliability and robustness, real-time monitoring, and response times for AI system failures.”
!MEASURE 2.7No control
“AI system security and resilience – as identified in the MAP function – are evaluated and documented.”
!MEASURE 2.8No control
“Risks associated with transparency and accountability – as identified in the MAP function – are examined and documented.”
!MEASURE 2.9No control
“The AI model is explained, validated, and documented, and AI system output is interpreted within its context – as identified in the MAP function – to inform responsible use and governance.”
−MEASURE 2.10Full fit · 2 controls
“Privacy risk of the AI system – as identified in the MAP function – is examined and documented.”

Partly met. GV-03 fits it fully, but GV-03 is only partly proven.

  • Full fit−GV-03 Privacy impact assessment
  • Partial fit−OI-02 Public surfaces are built, never strippedTested per public surface, not everywhere.
!MEASURE 2.11No control
“Fairness and bias – as identified in the MAP function – are evaluated and results are documented.”
!MEASURE 2.12No control
“Environmental impact and sustainability of AI model training and management activities – as identified in the MAP function – are assessed and documented.”
!MEASURE 2.13No control
“Effectiveness of the employed TEVV metrics and processes in the MEASURE function are evaluated and documented.”

MEASURE 3 · Tracking identified risks over time

!MEASURE 3.1No control
“Approaches, personnel, and documentation are in place to regularly identify and track existing, unanticipated, and emergent AI risks based on factors such as intended and actual performance in deployed contexts.”
!MEASURE 3.2No control
“Risk tracking approaches are considered for settings where AI risks are difficult to assess using currently available measurement techniques or where metrics are not yet available.”
!MEASURE 3.3No control
“Feedback processes for end users and impacted communities to report problems and appeal system outcomes are established and integrated into AI system evaluation metrics.”

MEASURE 4 · Feedback about how well measurement works

!MEASURE 4.1No control
“Measurement approaches for identifying AI risks are connected to deployment context(s) and informed through consultation with domain experts and other end users. Approaches are documented.”
!MEASURE 4.2No control
“Measurement results regarding AI system trustworthiness in deployment context(s) and across the AI lifecycle are informed by input from domain experts and relevant AI actors to validate whether the system is performing consistently as intended. Results are documented.”
!MEASURE 4.3No control
“Measurable performance improvements or declines based on consultations with relevant AI actors, including affected communities, and field data about context-relevant risks and trustworthiness characteristics are identified and documented.”

MANAGE

MANAGE 1 · Risks prioritized, responded to and managed

!MANAGE 1.1No control
“A determination is made as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed.”
!MANAGE 1.2No control
“Treatment of documented AI risks is prioritized based on impact, likelihood, and available resources or methods.”
!MANAGE 1.3No control
“Responses to the AI risks deemed high priority, as identified by the MAP function, are developed, planned, and documented. Risk response options can include mitigating, transferring, avoiding, or accepting.”
−MANAGE 1.4Partial fit · 1 control
“Negative residual risks (defined as the sum of all unmitigated risks) to both downstream acquirers of AI systems and end users are documented.”

Partly met. GV-01 covers part of it. No control fits it fully.

  • Partial fit−GV-01 Stated risk tolerance and a forward-looking risk registerResidual scores are ordinal judgments, not a sum of unmitigated risk.

MANAGE 2 · Maximizing benefits, minimizing negative impacts

−MANAGE 2.1Partial fit · 1 control
“Resources required to manage AI risks are taken into account – along with viable non-AI alternative systems, approaches, or methods – to reduce the magnitude or likelihood of potential impacts.”

Partly met. OI-04 covers part of it. No control fits it fully.

  • Partial fit!OI-04 Spend controlWeighs cost per run; non-AI alternatives aren't considered.
!MANAGE 2.2No control
“Mechanisms are in place and applied to sustain the value of deployed AI systems.”
!MANAGE 2.3No control
“Procedures are followed to respond to and recover from a previously unknown risk when it is identified.”
−MANAGE 2.4Partial fit · 1 control
“Mechanisms are in place and applied, and responsibilities are assigned and understood, to supersede, disengage, or deactivate AI systems that demonstrate performance or outcomes inconsistent with intended use.”

Partly met. AC-01 covers part of it. No control fits it fully.

  • Partial fit✓AC-01 Destructive-command blocksBlocks specific destructive actions; it can't disengage an agent.

MANAGE 3 · Third-party risks and benefits

−MANAGE 3.1Partial fit · 1 control
“AI risks and benefits from third-party resources are regularly monitored, and risk controls are applied and documented.”

Partly met. OI-05 covers part of it. No control fits it fully.

  • Partial fit−OI-05 Vendor change watch and portabilityMonitors one vendor's tool interface.
!MANAGE 3.2No control
“Pre-trained models which are used for development are monitored as part of AI system regular monitoring and maintenance.”

MANAGE 4 · Treatment, response, recovery and communication

−MANAGE 4.1Partial fit · 6 controls
“Post-deployment AI system monitoring plans are implemented, including mechanisms for capturing and evaluating input from users and other relevant AI actors, appeal and override, decommissioning, incident response, recovery, and change management.”

Partly met. GV-02, CM-01, CM-02, CM-03, OP-01 and OP-05 cover part of it. No control fits it fully.

  • Partial fit−GV-02 Written change, back-out, and retirement proceduresCovers change, recovery and retirement; no monitoring plan or user input.
  • Partial fit✓CM-01 Spec-and-review build gateChange management only.
  • Partial fit✓CM-02 Hold window on rule changesA review window over rule changes, with no sign-off.
  • Partial fit✓CM-03 Versioned change recordKeeps the change record; it doesn't manage change.
  • Partial fit✓OP-01 Job monitoring and self-healing by exceptionMonitoring and recovery for scheduled jobs only.
  • Partial fit−OP-05 Register snapshots and offsite copyRecovery only.
!MANAGE 4.2No control
“Measurable activities for continual improvements are integrated into AI system updates and include regular engagement with interested parties, including relevant AI actors.”
−MANAGE 4.3Full fit · 1 control
“Incidents and errors are communicated to relevant AI actors, including affected communities. Processes for tracking, responding to, and recovering from incidents and errors are followed and documented.”

Partly met. OP-04 fits it fully, but OP-04 is only partly proven.

  • Full fit−OP-04 Root cause and corrective action

GenAI Profile (NIST AI 600-1)

12 requirements
  • ✓0 met
  • −5 partly met
  • !7 not met

5 of 12 risks have a control · 8 of 21 controls map here. 0 of the 5 have a full fit.

!CBRN Information or CapabilitiesNo control
−ConfabulationPartial fit · 1 control

Partly met. OI-01 covers part of it. No control fits it fully.

  • Partial fit✓OI-01 Unverified completion-claim blockCatches unverified completion claims, not other false statements.
!Dangerous, Violent, or Hateful ContentNo control
−Data PrivacyPartial fit · 2 controls

Partly met. GV-03 and OI-02 cover part of it. No control fits it fully.

  • Partial fit−GV-03 Privacy impact assessmentAssesses the exposure; prevention sits in OI-02.
  • Partial fit−OI-02 Public surfaces are built, never strippedCovers what reaches public pages.
!Environmental ImpactsNo control
!Harmful Bias and HomogenizationNo control
−Human-AI ConfigurationPartial fit · 3 controls

Partly met. CM-01, AC-02 and OI-03 cover part of it. No control fits it fully.

  • Partial fit✓CM-01 Spec-and-review build gateStops the agent certifying its own change; the reviewer is also an agent.
  • Partial fit−AC-02 Human approval for bulk and outward-facing actionsEnforced for bulk changes; model-applied for sends and spend.
  • Partial fit!OI-03 Decision-origin labelsUnenforced labels.
!Information IntegrityNo control
−Information SecurityPartial fit · 1 control

Partly met. AC-03 covers part of it. No control fits it fully.

  • Partial fit−AC-03 Inbound channel restrictionCloses the remote channel to strangers; injection through content stays open.
!Intellectual PropertyNo control
!Obscene, Degrading, and/or Abusive ContentNo control
−Value Chain and Component IntegrationPartial fit · 1 control

Partly met. OI-05 covers part of it. No control fits it fully.

  • Partial fit−OI-05 Vendor change watch and portabilityOne vendor's interface is watched; the rest rely on portability.

ISO/IEC 42001

16 requirements
  • ✓3 met
  • −8 partly met
  • !5 not met

11 of 16 clauses and objectives have a control · 17 of 21 controls map here. 4 of the 11 have a full fit.

Management system clauses

!Cl. 4 Context of the organizationNo control
!Cl. 5 LeadershipNo control
−Cl. 6 PlanningPartial fit · 2 controls

Partly met. GV-01 and GV-04 cover part of it. No control fits it fully.

  • Partial fit−GV-01 Stated risk tolerance and a forward-looking risk registerRisk assessment and tolerance; the objectives are in GV-04.
  • Partial fit!GV-04 Stated mission and measurable objectivesObjectives only; risk planning stays with GV-01.
✓Cl. 7 SupportFull fit · 1 control

Met. CM-03 fits it fully and is enforced and tested.

  • Full fit✓CM-03 Versioned change record
✓Cl. 8 OperationFull fit · 3 controls

Met. CM-01 fits it fully and is enforced and tested.

  • Partial fit−GV-02 Written change, back-out, and retirement proceduresDocuments how changes are planned; CM-01 enforces it.
  • Full fit✓CM-01 Spec-and-review build gate
  • Partial fit✓CM-02 Hold window on rule changesDelays the offsite record, not the change.
−Cl. 9 Performance evaluationPartial fit · 1 control

Partly met. OP-03 covers part of it. No control fits it fully.

  • Partial fit−OP-03 Control wiring checksEvaluates whether controls can fire; coverage is partial.
−Cl. 10 ImprovementFull fit · 1 control

Partly met. OP-04 fits it fully, but OP-04 is only partly proven.

  • Full fit−OP-04 Root cause and corrective action

Annex A objectives

!A.2 AI policiesNo control
−A.3 Internal organizationPartial fit · 1 control

Partly met. OI-03 covers part of it. No control fits it fully.

  • Partial fit!OI-03 Decision-origin labelsAccountability for AI-initiated work, unenforced.
!A.4 Resources for AI systemsNo control
−A.5 Assessing impacts of AI systemsPartial fit · 1 control

Partly met. GV-03 covers part of it. No control fits it fully.

  • Partial fit−GV-03 Privacy impact assessmentPrivacy impacts only, not wider effects on people or society.
✓A.6 AI system life cycleFull fit · 7 controls

Met. CM-01 fits it fully and is enforced and tested.

  • Partial fit−GV-02 Written change, back-out, and retirement proceduresChange, rollback and retirement stages only.
  • Full fit✓CM-01 Spec-and-review build gate
  • Partial fit✓CM-03 Versioned change recordRecords configuration changes, not operating event logs.
  • Partial fit−CM-04 Test before promote, including a test that proves the check can failVerification before use; no validation against stated criteria.
  • Partial fit✓OP-01 Job monitoring and self-healing by exceptionOperation and monitoring of scheduled jobs only.
  • Partial fit✓OP-02 Hollow-run detectionOperation and monitoring of scheduled model runs only.
  • Partial fit✓OI-01 Unverified completion-claim blockA verification step inside operation.
−A.7 Data for AI systemsPartial fit · 1 control

Partly met. GV-03 covers part of it. No control fits it fully.

  • Partial fit−GV-03 Privacy impact assessmentInventories the data agents read; not its quality or provenance.
!A.8 Information for interested partiesNo control
−A.9 Use of AI systemsPartial fit · 2 controls

Partly met. AC-01 and AC-02 cover part of it. No control fits it fully.

  • Partial fit✓AC-01 Destructive-command blocksKeeps use within intended actions, by pattern.
  • Partial fit−AC-02 Human approval for bulk and outward-facing actionsHuman oversight of use; half of it is model-applied.
−A.10 Third-party and customer relationshipsPartial fit · 1 control

Partly met. OI-05 covers part of it. No control fits it fully.

  • Partial fit−OI-05 Vendor change watch and portabilityWatches a supplier; no supplier assessment.

COSO 2013

17 requirements
  • ✓1 met
  • −9 partly met
  • !7 not met

10 of 17 principles have a control · 21 of 21 controls map here. 5 of the 10 have a full fit.

Control environment

!P1 tone at the top: integrity and ethicsNo control
!P2 oversight independent of managementNo control
!P3 structure, reporting lines and authorityNo control
!P4 attracts and keeps competent peopleNo control
!P5 holds people accountable for their control dutiesNo control

Risk assessment

−P6 objectives clear enough to assess riskPartial fit · 1 control

Partly met. GV-04 covers part of it. No control fits it fully.

  • Partial fit!GV-04 Stated mission and measurable objectivesRegister risks were scored before these objectives existed and aren't tied to them yet.
−P7 identifies and analyzes riskFull fit · 2 controls

Partly met. GV-01 fits it fully, but GV-01 is only partly proven.

  • Full fit−GV-01 Stated risk tolerance and a forward-looking risk register
  • Partial fit−GV-03 Privacy impact assessmentOne risk domain, assessed once.
!P8 considers the potential for fraudNo control
−P9 assesses significant changePartial fit · 1 control

Partly met. OI-05 covers part of it. No control fits it fully.

  • Partial fit−OI-05 Vendor change watch and portabilityDetects one kind of external change.

Control activities

−P10 control activities that mitigate riskFull fit · 2 controls

Partly met. AC-02 fits it fully, but AC-02 is only partly proven.

  • Full fit−AC-02 Human approval for bulk and outward-facing actions
  • Partial fit!OI-04 Spend controlA policy the model applies, with no check.
Known gap: Segregation of duties between people
✓P11 general controls over technologyFull fit · 9 controls

Met. CM-01, CM-03, AC-01, OP-01 and OP-02 fit it fully and are enforced and tested.

  • Full fit✓CM-01 Spec-and-review build gate
  • Partial fit✓CM-02 Hold window on rule changesHolds the backup, not the change.
  • Full fit✓CM-03 Versioned change record
  • Full fit−CM-04 Test before promote, including a test that proves the check can fail
  • Full fit✓AC-01 Destructive-command blocks
  • Full fit−AC-03 Inbound channel restriction
  • Full fit✓OP-01 Job monitoring and self-healing by exception
  • Full fit✓OP-02 Hollow-run detection
  • Full fit−OP-05 Register snapshots and offsite copy
−P12 deployed through policy and procedureFull fit · 3 controls

Partly met. GV-02 fits it fully, but GV-02 is only partly proven.

  • Full fit−GV-02 Written change, back-out, and retirement procedures
  • Partial fit−AC-02 Human approval for bulk and outward-facing actionsThe send-and-spend half is a policy nothing checks.
  • Partial fit−OI-02 Public surfaces are built, never strippedA policy everywhere, tested on some surfaces.

Information and communication

−P13 relevant, quality informationPartial fit · 1 control

Partly met. OI-01 covers part of it. No control fits it fully.

  • Partial fit✓OI-01 Unverified completion-claim blockProtects one kind of information: claims that work is done.
−P14 internal communicationPartial fit · 1 control

Partly met. OI-03 covers part of it. No control fits it fully.

  • Partial fit!OI-03 Decision-origin labelsInternal communication, unenforced.
!P15 communication with outside partiesNo control

Monitoring

−P16 ongoing and separate evaluationsPartial fit · 2 controls

Partly met. CM-04 and OP-03 cover part of it. No control fits it fully.

  • Partial fit−CM-04 Test before promote, including a test that proves the check can failEvaluates a control before it's relied on, on change only.
  • Partial fit−OP-03 Control wiring checksRuns on change, with partial coverage.
Known gap: Scheduled operating-effectiveness testingKnown gap: Risk metrics with thresholds
−P17 evaluates and communicates deficienciesFull fit · 1 control

Partly met. OP-04 fits it fully, but OP-04 is only partly proven.

  • Full fit−OP-04 Root cause and corrective action

SOX IT general controls

4 requirements
  • ✓1 met
  • −3 partly met
  • !0 not met

4 of 4 domains have a control · 10 of 21 controls map here. 2 of the 4 have a full fit.

−Program changePartial fit · 3 controls

Partly met. GV-02, CM-01 and CM-03 cover part of it. No control fits it fully.

  • Partial fit−GV-02 Written change, back-out, and retirement proceduresThe written policy; the operating evidence comes from CM-01 and CM-03.
  • Partial fit✓CM-01 Spec-and-review build gateThe approver is a second agent, not a second person.
  • Partial fit✓CM-03 Versioned change recordThe job is the recorded author, so the record doesn't show who made each change.
Known gap: Gate blind spots
−Program developmentPartial fit · 1 control

Partly met. CM-04 covers part of it. No control fits it fully.

  • Partial fit−CM-04 Test before promote, including a test that proves the check can failThe builder runs the tests; nothing re-runs them on a schedule.
−Access to programs and dataFull fit · 2 controls

Partly met. AC-03 fits it fully, but AC-03 is only partly proven.

  • Partial fit✓AC-01 Destructive-command blocksRestricts one class of privileged function by pattern, not by permission.
  • Full fit−AC-03 Inbound channel restriction
Known gap: Segregation of duties between people
✓Computer operationsFull fit · 4 controls

Met. OP-01 and OP-02 fit it fully and are enforced and tested.

  • Full fit✓OP-01 Job monitoring and self-healing by exception
  • Full fit✓OP-02 Hollow-run detection
  • Full fit−OP-04 Root cause and corrective action
  • Partial fit−OP-05 Register snapshots and offsite copySnapshots don't survive a restart, and the offsite copy is daily.

OWASP Top 10 for LLM Applications

10 requirements
  • ✓0 met
  • −6 partly met
  • !4 not met

6 of 10 risks have a control · 8 of 21 controls map here. 0 of the 6 have a full fit.

−LLM01 Prompt InjectionPartial fit · 2 controls

Partly met. AC-02 and AC-03 cover part of it. No control fits it fully.

  • Partial fit−AC-02 Human approval for bulk and outward-facing actionsLimits what an injected instruction can do; doesn't detect it.
  • Partial fit−AC-03 Inbound channel restrictionBlocks instructions from unknown senders only.
−LLM02 Sensitive Information DisclosurePartial fit · 1 control

Partly met. OI-02 covers part of it. No control fits it fully.

  • Partial fit−OI-02 Public surfaces are built, never strippedTests published pages for identifiers; not model output generally.
−LLM03 Supply ChainPartial fit · 1 control

Partly met. OI-05 covers part of it. No control fits it fully.

  • Partial fit−OI-05 Vendor change watch and portabilityDetects interface changes; doesn't vet components.
!LLM04 Data and Model PoisoningNo control
!LLM05 Improper Output HandlingNo control
−LLM06 Excessive AgencyPartial fit · 4 controls

Partly met. CM-01, AC-01, AC-02 and AC-03 cover part of it. No control fits it fully.

  • Partial fit✓CM-01 Spec-and-review build gateLimits unreviewed changes; doesn't narrow the agent's tool permissions.
  • Partial fit✓AC-01 Destructive-command blocksPattern-based; a deletion by another route isn't caught.
  • Partial fit−AC-02 Human approval for bulk and outward-facing actionsEnforced for bulk changes; model-applied for sends.
  • Partial fit−AC-03 Inbound channel restrictionLimits what a remote trigger can run; other paths aren't covered.
!LLM07 System Prompt LeakageNo control
!LLM08 Vector and Embedding WeaknessesNo control
−LLM09 MisinformationPartial fit · 1 control

Partly met. OI-01 covers part of it. No control fits it fully.

  • Partial fit✓OI-01 Unverified completion-claim blockCompletion claims only.
−LLM10 Unbounded ConsumptionPartial fit · 1 control

Partly met. OI-04 covers part of it. No control fits it fully.

  • Partial fit!OI-04 Spend controlApproval per run; no hard spending cap.