Justina Dong

The system learns from what breaks. I learn from what I don’t understand.

How the system learns

One real failure in my lab, in four steps. Every line on screen is from a real file.

then round again

Watch it break

The system is my lab: AI jobs that run on a schedule. In one set of them, about one run in four wrote almost nothing, and every status screen still said success: 47 out of 47.

See the real note
… every status surface reports those runs as `success` — history.json is Counter({'success': 47}), 47 of 47, including a 38 B run and a `529 Overloaded` run. A field that cannot say no is not a check.

Its code, lines 6 to 9

From the notes at the top of the control’s code.

Build the control

A control is a check that runs by itself, outside the AI. This one reads what each run wrote. Under 400 bytes, about a short paragraph, means hollow: it said success and wrote almost nothing. Empty runs wrote 266 at most; the smallest working one, 593.

See the real note
hollow run -> file exists, 0-266 B (measured)
non-run -> no file at all
healthy run -> file exists, 593 B - 10 KB (measured; …)

Its code, lines 14 to 16

From the control’s own notes.

Monitor the control

The control is a scheduled job too, so each morning a dashboard marks it like the other 63: green if it ran on time with nothing wrong.

How the system is graded

JobSa26Su27Mo28Tu29We30Th1Fr2Clean
  1. Hollow-run detectorEvery 3 hours · one of the Watchers
    100%
Watchers18 jobs that watch other jobs, or the machine they run on
43%
  1. Agent self-auditWeekly · Sun 05:00
    0%
  2. Agent usage scanWeekly · Sun 03:30
    0%
  3. Broken-reference checkWeekly · Sun 03:00
    0%
  4. Core-file drift checkWeekly · Sun 06:00
    0%
  5. Docs-vs-reality auditWeekly · Sun 05:30
    0%
  6. Verdict-file sweepDaily · 03:10
    0%
  7. Weekly health passWeekly · Sun 08:07
    0%
  8. New-job checkDaily · 08:20
    29%
  9. Cloud watchdogWeekly
    43%
  10. Heartbeat watchdogEvery hour
    43%
  11. Network probeEvery 15 min
    43%
  12. Fleet surveyDaily · 09:40
    67%
  13. Lost-question sweepDaily · 03:25
    71%
  14. Cost and security auditDaily · 06:00
    100%
  15. Design-rule scanWeekly · Mon 04:30
    100%
  16. Editor format checkDaily · 08:35
    100%
  17. Message-bot health checkEvery 15 min
    100%
Builders12 jobs that build the pages and records the rest of the site reads
100%
  1. Daily commitDaily · 04:00
    100%
  2. Digest publish checkEvery 10 min
    100%
  3. Inbox pageEvery 15 min
    100%
  4. Jobs dashboardEvery 15 min
    100%
  5. Nightly graderDaily · 00:15
    100%
  6. Project datesDaily · 03:30
    100%
  7. Report cardWeekly · Mon 06:00
    100%
  8. Research publish checkEvery 10 min
    100%
  9. Session ledgerDaily · 03:00
    100%
  10. Site publisherHourly
    100%
  11. Skills usage pageEvery hour
    100%
  12. To-do pageHourly
    100%
Writers9 jobs that write something new for me to read
75%
  1. Weekly essay (switched off)Weekly · Wed
    —
  2. Lecture news bridgeWeekly
    43%
  3. Lecture readerDaily · 08:47
    43%
  4. Weekly AI synthesisWeekly
    43%
  5. AI news digestDaily · 08:30
    71%
  6. AI lab blog watcherDaily · 19:00
    100%
  7. Essay topic refreshWeekly · Wed 07:00
    100%
  8. Insight extractorDaily · 09:00
    100%
  9. Weekly essayWeekly · Wed 08:30
    100%
Upkeep9 jobs that tidy, file and keep the services up
79%
  1. File cleanupDaily · 02:30
    14%
  2. Screenshot organizerDaily · 11:00
    43%
  3. Nightly to-do closerDaily · 23:30
    57%
  4. Message-bot listenerAlways on
    100%
  5. Nightly wind-downDaily · 22:00
    100%
  6. Remote-control serviceAlways on
    100%
  7. Repeat-request detectorWeekly · Sun 23:30
    100%
  8. Unlinked-note scanDaily · 02:00
    100%
  9. Weekly inbox filingWeekly · Sun 09:00
    100%

cleanneeds a look: late, or it raised somethingfailedno reportswitched offshare of a group clean that morning

Yellow means look, not broken: on 2 October, 9 of 10 yellow jobs had run on schedule.

16 personal jobs are counted; names stay private.

Watch that break

Every control has a way around it. Size alone would miss a run that fails late. One in July wrote 2,169 bytes, then quit on an error from the AI service. So the control also looks for the error message, at any size.

See the real log line
[2026-09-09 16:25] HOLLOW|capture-close-backfill|2026-07-16-capture-close-backfill.md|2169B|api-error

Its log, line 43

From its first scan, the day it was built.

Where it keeps what it learns is in the lab: How the AI remembers →

How I learn

I say I don’t understand something.

My research agent’s essays were being scored by a script that counts words, headers and dashes. After a blind test came back, Claude said it had killed the metric. I didn’t follow, so I asked:

what do you mean, killed the metric?

16 September 2026

Claude writes it up, with the context.

Claude explains on the spot. Terms worth keeping go on a list, and the ones I approve become entries in my glossary, each with where it came up. It holds 111 terms. This is that entry, cut into cards.

“what do you mean, killed the metric?”

1 of 6
The top of my glossary page. Title: Learnings. A working glossary of what I’ve learned, each term as an article I can revisit cold.Further down the same page, the list of terms. The first is Construct validity, the proxy is not the thing, dated 16 September 2026.

My real glossary. This entry is on the next cards.

Context it came up in

Killed means thrown out: the score wasn’t measuring what we thought. A script rated how different a set of essays sounded by counting five things, like words and dashes. Every count was right. Each was only a stand-in, a proxy, for the real question: does a reader hear different writers?

The analogy that landed

Decide whether two people look alike by comparing height, weight, shoe size and hair length. Every measurement correct, repeatable, cheap. None of them measures the face.

What it is

Construct validity asks whether a measure captures the thing it claims to measure. That is a separate question from whether it is reliable (the same answer twice) or computed right. A metric can be flawless on both and empty on this one.

Watch for

I was the judge. I couldn’t see which essay was which, but by the fourth pass I knew what the study was asking. If your judge can guess the result you are hoping for, the judge is a stand-in too, and nothing real is underneath.

Leverage

Ask of any AI score: what is this number a stand-in for, and what did you check it against? Then add one case you predict will fail, and see whether it does. A study that can only confirm is not a study.

I get graded every night.

Just after midnight each night, a scheduled AI job reads what I did with AI that day and grades it like school, A to F. A failed run retries four times that day. Every day in September has a grade.

A line chart from my report card: my overall grade for each graded day from 24 March to 30 September, on a scale from F to A.
From D and C in March to mostly B now. My real report card.

How that’s graded.

Each day gets up to nine letters: eight for things like how clear my prompts were, what it cost and whether I wrote a plan first, then one overall. On the real card each letter opens to why, and what to do differently. Nobody independent grades the grader yet.

One day on my report card, 29 September. Prompt B, Model A, Routing A, Cost B, Spec A, Capability B, Learning B, Breadth A, Overall B. The overall row is open and reads: Exemplary build discipline and specialist sequencing; one security judgment gap holds it from A.
29 September, the day I rebuilt this site.