Construct validity
A script counts words, headers and dashes in ten pairs of essays and calls nine of the pairs different writers. Is it measuring voice?
No. Every count was right, and none of them measured a voice. Read aloud, the pairs sounded like one writer.
I wrote a script to score how different my essay agent's pieces sounded. It counted words, headers and dashes. It called nine of ten pairs different where I heard one writer.
Five counts per essay
- word count
- headers
- dashes
- dashes per 1,000 words
- verbal tics
9 of 10 pairs: different writers
Reading the same ten pairs
One writer. Every count was right, and none of them measured a voice.
What it is
Whether a measure captures the thing it claims to. That is a separate question from whether it gives the same answer twice, or whether the arithmetic is right. A number can pass both and still measure the wrong thing.
What to watch for
My own check had the same flaw. I was the ground truth, and by the fourth round I knew what the test was looking for, so my answers stopped counting as evidence. If the judge can guess the result you want, you have two stand-ins and nothing real underneath.
How I use it. Ask of any AI metric: what is this a stand-in for, and what did you check it against? Then add one case built to fail, and see whether it does.