Grading the grader
I use a model judge to decide which of my own paragraphs survive. So I graded 30 of its calls against my own, blind, on a protocol fixed before I saw the data.
A judge that cuts your prose is an instrument, and an instrument you have never scored is a rumor. I run a rubric-driven model judge over my own drafts: it reads a passage and returns one of four verdicts, from keep it to cut this to a paragraph. I have been letting it shape published writing. At some point the honest move is to point the measuring stick at the measuring stick.
So: a calibration receipt. Aggregate numbers only, no manuscript text, companion to the OPERANT note, which does the same thing for agent judgment.
The method, fixed before the data
Thirty sections were sampled deterministically (seed 20260717) from three corpora: a published book (12 sections), an unpublished draft (12), and published essays (6). Each was graded independently on the rubric's four-point ordinal scale by me and by two blind model judges running the same rubric dimension.
The metrics were written down first, in a method file committed before any grade existed. There was exactly one amendment, disclosed: raw agreement was added alongside kappa, still before a single operator grade was recorded. That ordering is the whole point. A calibration study you design after seeing the results is a story about the results.
Results
| Metric | Value |
|---|---|
| Judge vs judge (linear-weighted κ) | 1.00, verdict-identical 30/30 |
| Operator vs judge A (linear-weighted κ) | 0.78 (unweighted 0.78) |
| Operator vs judge B (linear-weighted κ) | 0.78 (unweighted 0.78) |
| Raw agreement, operator vs judges | 96.7% exact, 100% within one step |
| Length bias (Spearman, words vs verdict) | judges −0.52, operator −0.42 |
| Severity skew (judge minus operator, mean) | −0.03 |
| Family gap by corpus | book −0.08, draft 0.00, essays 0.00 |
Reading the numbers
0.78 lands in the pre-registered "substantial" band, which sounds like a hedge until you look at what is underneath it. Twenty-nine of thirty verdicts matched exactly. All thirty matched within one step. The marginals are heavily skewed toward keep, because this is a post-developmental-edit corpus, and prevalence skew is precisely the condition under which kappa punishes an agreeing pair of raters hardest. Kappa is reported with raw agreement for exactly that reason, and that decision was made in advance, not after 0.78 came back.
Three other things did not happen, and their absence is the finding. No family favoritism: perfect agreement on the unpublished draft and on the essays, so the judge is not softer on work that is not yet public. No judge-specific length bias: all three raters, me included, trend the same direction against long passages, with the judges amplifying mildly rather than inventing the effect. No severity drift: the judges were harsher than me by one step, once, in thirty.
And the two judges agreed with each other on every single passage while quoting different verbatim anchors to justify it. That is the reproducibility result hiding inside the calibration result: the rubric is mechanically repeatable, not a mood.
Limits, stated in the method rather than discovered afterward
- Single operator. Sources were unnamed on the grading sheet as a blinding aid, not a guarantee: I may still recognize my own prose.
- The judges ran concurrently in one repository under instructed non-access rather than true process isolation. Their independence is evidenced, not assumed, by those differing verbatim anchors under identical verdicts.
- The rubric's seven-gate vocabulary was collapsed to a four-point ordinal scale, and cross-unit gates such as duplicate placement were judged at passage grain. That understates what the rubric does across a whole draft.
- Prevalence-skewed marginals, as above. Read the pair of numbers, not either one alone.
What I get for the trouble is narrow and worth having: when this judge tells me to cut something, I now know roughly how often I would have agreed, and that the answer is "almost always." That is not the judge being right. It is the judge being calibrated to me, which is the only claim a receipt like this can honestly make.