Field note

Grading the grader

I use a model judge to decide which of my own paragraphs survive. So I graded 30 of its calls against my own, blind, on a protocol fixed before I saw the data.

A judge that cuts your prose is an instrument, and an instrument you have never scored is a rumor. I run a rubric-driven model judge over my own drafts: it reads a passage and returns one of four verdicts, from keep it to cut this to a paragraph. I have been letting it shape published writing. At some point the honest move is to point the measuring stick at the measuring stick.

So: a calibration receipt. Aggregate numbers only, no manuscript text, companion to the OPERANT note, which does the same thing for agent judgment.

The method, fixed before the data

Thirty sections were sampled deterministically (seed 20260717) from three corpora: a published book (12 sections), an unpublished draft (12), and published essays (6). Each was graded independently on the rubric's four-point ordinal scale by me and by two blind model judges running the same rubric dimension.

The metrics were written down first, in a method file committed before any grade existed. There was exactly one amendment, disclosed: raw agreement was added alongside kappa, still before a single operator grade was recorded. That ordering is the whole point. A calibration study you design after seeing the results is a story about the results.

Results

MetricValue
Judge vs judge (linear-weighted κ)1.00, verdict-identical 30/30
Operator vs judge A (linear-weighted κ)0.78 (unweighted 0.78)
Operator vs judge B (linear-weighted κ)0.78 (unweighted 0.78)
Raw agreement, operator vs judges96.7% exact, 100% within one step
Length bias (Spearman, words vs verdict)judges −0.52, operator −0.42
Severity skew (judge minus operator, mean)−0.03
Family gap by corpusbook −0.08, draft 0.00, essays 0.00

Reading the numbers

0.78 lands in the pre-registered "substantial" band, which sounds like a hedge until you look at what is underneath it. Twenty-nine of thirty verdicts matched exactly. All thirty matched within one step. The marginals are heavily skewed toward keep, because this is a post-developmental-edit corpus, and prevalence skew is precisely the condition under which kappa punishes an agreeing pair of raters hardest. Kappa is reported with raw agreement for exactly that reason, and that decision was made in advance, not after 0.78 came back.

The entire operator-versus-judge gap is one call: a single flag on a published section I still stand behind as written. Not a pattern. A disagreement.

Three other things did not happen, and their absence is the finding. No family favoritism: perfect agreement on the unpublished draft and on the essays, so the judge is not softer on work that is not yet public. No judge-specific length bias: all three raters, me included, trend the same direction against long passages, with the judges amplifying mildly rather than inventing the effect. No severity drift: the judges were harsher than me by one step, once, in thirty.

And the two judges agreed with each other on every single passage while quoting different verbatim anchors to justify it. That is the reproducibility result hiding inside the calibration result: the rubric is mechanically repeatable, not a mood.

Limits, stated in the method rather than discovered afterward

What I get for the trouble is narrow and worth having: when this judge tells me to cut something, I now know roughly how often I would have agreed, and that the answer is "almost always." That is not the judge being right. It is the judge being calibrated to me, which is the only claim a receipt like this can honestly make.

Adjacent reading

← Field notes · RSS