I pointed an eval harness at my own reflection
Listen to this piece · 5 min
0:00 / 5:55 · resumes at 0:00
Read by a synthetic voice (Kokoro).
Somewhere right now a recruiter is asking an AI engine about a candidate before opening the résumé. The candidate will never see the answer. They cannot A/B it, cannot check its sources, and nobody sends a report. I wanted to know what that answer looked like for me — and, more usefully, whether it could be changed on purpose.
So I built the same kind of harness I build at work and pointed it at my own name. Four engines — ChatGPT, Claude, Perplexity, Gemini — each asked 83 questions about me, 63 of them with search on, all of them repeated. Every answer tagged against a canon of verified facts. Every citation logged. Then a month of fixes to the pages those engines read, and the whole battery run again.
I expected a before-and-after chart. I got three things I did not expect, and each of them only surfaced because I counted instead of scored.
The headline split cleanly by engine. Perplexity went from citing my site on 13 of 63 search probes to 63 of 63. Claude went from 0 to 0. Same fixes, same month, opposite outcomes — and one variable explains it: index access. A Bing Webmaster submission reached the index Perplexity reads. There is no equivalent door into Anthropic's.
But Claude's zero was not silence. The junk that used to feed its answers — name-etymology pages, a job-board scrape — dropped from 79 citations to 11. And a data broker rose from 0 to 74 to take their place. Claude's third most-cited source about me, after a month of cleanup, was a page I have never controlled and had asked to be removed.
That is the pattern worth naming. An engine that cannot find the authoritative page does not say "I couldn't find much." It assembles an answer from whatever it can find, with the same fluency it would use for the truth. Clearing the junk did not create room for the right source; it created a vacuum, and the vacuum filled. If you fix your surfaces without fixing the index, you have rearranged the substitutes.
While labelling answers by hand I kept meeting the same move. An engine would find my site, read a number on it — a metric from a project, a result from a harness — and then decline to stand behind it. Not because it was wrong, but because I was the one who had published it. The claim was treated as testimony rather than evidence.
The engines are right to do this, which is what makes it uncomfortable. Self-attestation is weak evidence. A harness result on my own site is a claim; the same result in a repository with commits, or in a piece someone else wrote about the work, is corroboration. The lesson for anyone whose work is mostly self-published is not to publish more. It is to get the claim restated somewhere you do not own — and to make the thing you own as easy to verify as possible: public code, public data, dated releases, a predictions file scored in the open.
I had a version of this belief before the study. Watching an engine apply the discount, sentence by sentence, turned it from a belief into a measurement I now want to run properly.
The harness uses two LLM judges from different model families to tag each answer with a failure category. Then it does the thing that makes a judge a judge: a blind, stratified sample of 40 answers, labelled by a human who cannot see the verdicts.
The judges failed. The Claude judge matched the human's exact tag set on 3 of 40. The Gemini judge on 12. The two judges agreed with each other on 35% of answers. Ten repetitions of an identical prompt changed the tag set 86% of the time. It was the third blind check in a row to say so.
In CrossSource, the same validation step came back 15 for 15 — and caught a real harness bug in the process. Here it came back a failure. I think the second result is the more useful one to have in public, because of what it forces. You cannot publish a failure rate on the strength of a judge that agrees with a human 8% of the time. So the interpretive categories became bands — a floor where both judges agree, a ceiling where either fires — and the headline numbers became counts: which surfaces each engine cited, of how many probes. The one category that survived as a point estimate, poisoned citations, did so because it was moved out of the judge entirely and computed against canon. It went from a 77%-agreement judgement to an exact calculation the moment it stopped being a question for a language model.
The method did not soften the claim when the judge failed. It changed what kind of number was allowed to appear in a headline. That is what validation is for.
Count before you score. The citation trail — which pages an engine actually read — is judge-free, cheap, and turned out to carry the whole story. Repeat every probe; a single-shot before/after is indistinguishable from resampling. Treat your own site as a claim, not a proof, and go get the corroboration. And when the validation step comes back ugly, publish that too. A harness that only reports the numbers its judge can be trusted with is worth more than one that reports everything.
The harness is public and entity-agnostic — canon.yaml is the only file with facts in it. Run it on yourself. I would like to know whether your engines fill the way mine did.
If you're shipping an LLM product, let's talk about what your evals miss.