Justina Dong

These are six terms I didn’t follow during one week in September, and each card asks you to guess before it turns over.

This page is my side of the learning: terms I didn’t follow, explained until I did. The system learns another way, by turning each failure into a check that runs in code. How the system learns →

How a term gets written up, with one real term
  1. Me

    I say I don’t understand something

    After a blind read of my essay agent’s work, I asked the short version: what do you mean, killed the metric?

  2. Claude

    Claude explains, right then

    In plain words, with a picture I can hold: judging whether two people look alike by comparing height, weight, shoe size and hair length. Every measurement is right, and none of them looks at the face. Building waits until I can read the output.

  3. Me

    It gets queued to a short list

    At the end of the session I see the terms that came up and pick the ones worth keeping. This one stayed.

  4. Claude

    Upon my approval, it becomes an entry

    Four parts: where it came up, what it is, what to watch for, and how to use it.

dashed, I do it · solid, Claude does it

This is one week, 13 to 16 September 2026, rewritten for this page: my own notes say more about where each term came up, and here that part is shorter and leaves my private life out. Six of the week’s ten entries made it.

0 of 6 turned
16 Sep 2026

Construct validity

A script counts words, headers and dashes in ten pairs of essays and calls nine of the pairs different writers. Is it measuring voice?

No. Every count was right, and none of them measured a voice. Read aloud, the pairs sounded like one writer.

I wrote a script to score how different my essay agent's pieces sounded. It counted words, headers and dashes. It called nine of ten pairs different where I heard one writer.

Five counts per essay

  • word count
  • headers
  • dashes
  • dashes per 1,000 words
  • verbal tics

9 of 10 pairs: different writers

Reading the same ten pairs

One writer. Every count was right, and none of them measured a voice.

What it is

Whether a measure captures the thing it claims to. That is a separate question from whether it gives the same answer twice, or whether the arithmetic is right. A number can pass both and still measure the wrong thing.

What to watch for

My own check had the same flaw. I was the ground truth, and by the fourth round I knew what the test was looking for, so my answers stopped counting as evidence. If the judge can guess the result you want, you have two stand-ins and nothing real underneath.

How I use it. Ask of any AI metric: what is this a stand-in for, and what did you check it against? Then add one case built to fail, and see whether it does.

Related
14 Sep 2026

“No disk” isn’t “no device”

An external drive shows nothing in the disk list. Is it plugged in?

That list can’t tell you. It only shows disks the system managed to hand over. This drive was plugged in and powered the whole time.

An external drive showed nothing in the disk list. Twice, three days apart, I wrote it down as not connected. It was plugged in and powered the whole time.

  1. USB devicepresent, powered
  2. Asked “what are you?”answered
  3. Asked “are you ready?”timed out, 10 s
  4. Disk listnothing

What it is

A drive comes up in layers: the USB device, the storage driver, then the disk the system hands you. The tool I used only reads the last layer. The drive answered when asked what it was and timed out when asked if it was ready, so its electronics worked and its mechanism didn’t.

What to watch for

The failure shows up as silence rather than an error. The list never says a drive is missing; it prints the disks it has, and I read the gap as absence.

How I use it. When a system says something isn’t there, ask which layer is doing the reporting. An AI tool should return “nothing matched” and “you don’t have access” as two different answers, not the same empty list.

Related
14 Sep 2026

Retrieval and ranking

A search hands back five good-looking results, and the right answer isn’t one of them. Which stage lost it?

Can’t tell yet. Read what the first stage retrieved: if the answer isn’t in there, the ranking never saw it.

I was reading how large video feeds decide what to show, and the design turned out to have the same shape as the search I run over my own notes.

  1. Everythingmillions
  2. 1 · Retrieval, cheapa few thousand ★
  3. 2 · Ranking, carefultop 5 ★

Ranking still returns five results, and nothing on the screen says one is missing.

An illustration, not measured numbers.

What it is

Systems that pick a few things out of millions work in two stages. Retrieval is cheap and rough: it cuts millions down to a few thousand. Ranking is expensive and careful, and it only scores what survived. The split is a budget decision, since nobody can afford to run the careful model on everything.

What to watch for

Something that never gets retrieved never gets scored, so a miss in the first stage leaves no record at all. Teams tune the ranker because that’s the part they can see, and the fix never reaches the stage that dropped the item.

How I use it. When an answer built on search is wrong, I pull what was retrieved and read it before touching the prompt. If the right passage isn’t in there, no prompt change will fix it.

Related
13 Sep 2026

Hashing: identity by content, not by name

Two photos have the same filename, and one file is twice the size of the other. Same picture?

Neither fact tells you. Hash both files and compare the bytes.

I was recovering photos from a failing drive. A recovered photo and a backed-up one had the same filename, and one file was about twice the size of the other. Neither fact told me whether they were the same picture.

Type in either box. Change one character and watch the fingerprint.

… …

Real SHA-256, first 24 of 64 characters, computed in your browser.

What it is

A filename is a label someone typed. A hash is computed from the file’s bytes, and changing one byte changes it completely. Two files with the same hash are the same file, whatever they’re called.

What to watch for

Names collide, and sizes collide at volume: 68,772 files on that drive had only 50,780 distinct sizes. A hash has the opposite blind spot. It proves identical bytes, not the same photo, so a JPEG saved again looks the same to me and hashes differently.

How I use it. Ask what an integrity check actually compares. A backup that only compares filenames will pass every time, including on a file that was quietly corrupted.

Related
13 Sep 2026

File carving

A drive was quick-formatted and now reports that it’s empty. Are the photos gone?

Still on the drive. The format rewrote the index and left the photos where they were. What’s lost is their names.

The recovery scan found thousands of photos on a drive that had been quick-formatted, a drive the computer said was empty.

  • beach.jpgblocks 1–3
  • notes.txtblock 5
  • dinner.jpgblocks 7–8

The index is gone. The blocks are all still there.

  1. FF D8 FF
  2. …
  3. …
  4. 00 00
  5. 6E 6F
  6. 00 00
  7. FF D8 FF
  8. …

Two photos back, found by how their bytes begin. Their names are not coming back.

An illustration of eight blocks, not the real drive.

What it is

A drive holds two things: an index (names, folders, dates, where each file sits) and the data itself. A quick format rewrites the index and leaves the data. Carving ignores the index and finds files by how their bytes begin. I get the photo back, but not its name or its folder.

What to watch for

Short patterns match by chance. The three bytes that open every JPEG turn up in random data about once every 16 MB, which is thousands of fake files across a drive. Matching on a longer marker that real photos carry cut the chance hits by roughly 250 times.

How I use it. Deleting usually cuts the pointer, not the data. When a system says a record is deleted, ask what overwrote it.

Related
13 Sep 2026

When a law is amended by another law

Your notes give 2 August 2026 as the start of the AI Act’s high-risk rules, copied from the official text. Still right?

Not any more. A later regulation moved it: 2 December 2027 for Annex III systems, 2 August 2028 for Annex I.

I was checking the dates in my notes on the EU AI Act against the official text. They didn’t match, and nothing had been mistyped.

  1. High-risk rules apply2 Aug 2026
  2. Annex III systems2 Dec 2027
  3. Annex I systems2 Aug 2028

This is the text a search for “AI Act” brings back.

What it is

An amending regulation is a separate law whose job is to edit another one. The original act is never reissued, so its old text keeps circulating unchanged. Regulation (EU) 2026/1744 moved the AI Act’s high-risk obligations from 2 August 2026 to 2 December 2027 for Annex III systems, and to 2 August 2028 for Annex I systems.

What to watch for

A search for the AI Act returns the original text, and the article you cite from it is authentic and out of date, because the amendment doesn’t contain the words you searched for and never comes up.

How I use it. I put a source and a checked-on date beside every date in a compliance document. An AI that answers legal questions should work from the consolidated text and show its as-of date in every answer.

Related

Checked 30 Sep 2026 against Regulation (EU) 2026/1744, Official Journal 24 July 2026, Article 1, the amendment to Article 113(c).