mirror-eval: pointing an eval harness at what AI search says about me
Same fixes, four engines, opposite outcomes — and a zero that turned out to be the most useful number in the study. An open harness for measuring what AI search engines say about a person, built to the CrossSource method: explicit ground truth, a failure taxonomy instead of a bare score, a judge whose agreement with a human is measured rather than assumed, and a limitations section that says what the numbers cannot support.
Listen to this piece · 7 min
0:00 / 7:25 · resumes at 0:00
Read by a synthetic voice (Kokoro).
Recruiters, buyers and counterparties increasingly ask an AI engine about you before they ask you. The answer is assembled from whatever the crawlers found — a dead portfolio, a scraped aggregator, a job you left two years ago — and you cannot see it from inside your own account, cannot A/B it, and nobody sends you a report. Existing tooling measures whether you are mentioned. That is the easy half. The hard half is whether the mention is true, current, and sourced to something you control. That is a scoring problem, which makes it an eval problem — the same class of problem as asking whether a production LLM's citation actually supports its claim. So I pointed the harness at my own reflection.
- Subject
- Me.
canon.yamlholds the ground truth — true claims, stale claims, clean and poisoned sources — and is the only file with facts in it. Point it at anyone. - Engines
- ChatGPT, Claude, Perplexity, Gemini, through each provider's own API and retrieval stack. Aggregators would bolt a third-party search layer onto the model and measure a system nobody uses.
- Battery
- 83 probes per engine per wave — 63 search-mode, 20 knowledge-mode — 332 per wave. Prompts are derived from facets (scaffolding, prior, decision, output shape), not topics; every cell is repeated, because a single probe cannot tell "the fix worked" from "we resampled."
- Two waves
- Baseline 2026-08-06, lift 2026-09-04. Identical battery, canon, and composition.
- Between the waves (Aug 7–20)
- Six changes to the surfaces engines read — a Bing Webmaster submission, a homepage link to the CrossSource repo, LinkedIn and profile-page cleanup, a DOI and identifier records, a new /writing/ page. They landed as a cluster, and are attributed as one.
- Two judges from different model families
- Tag every answer against a claim-failure taxonomy; then a blind, stratified human-labelled sample of 40 decides whether any judged number can be cited.
- Two kinds of number
- Counted: which surfaces each engine actually cited — judge-free. Judged: taxonomy categories — reported as bands unless validation says otherwise.
Probes citing an owned surface, of 63 search-mode probes per engine, Wave 1 → Wave 2:
| Engine | zoebnomi.com | github.com/zoeb-nomi | Owned (site or repo) |
|---|---|---|---|
| ChatGPT | 62 → 54 | 0 → 25 | 62 → 63 |
| Claude | 0 → 0 | 0 → 0 | 0 → 0 |
| Perplexity | 13 → 63 | 0 → 0 | 13 → 63 |
| Gemini | 61 → 63 | 0 → 32 | 61 → 63 |
What filled the gap (citation counts, Wave 1 → Wave 2):
| Surface | Engine | Count |
|---|---|---|
| Four poisoned name-etymology and job-board pages | Claude | 79 → 11 |
| zoominfo.com (data broker) | Claude | 0 → 74 |
| Wrong-person pages (imdb, nomi.ai, youtube) | Perplexity | 228 → 36 |
The finding that matters: the same cluster of fixes produced a 13 → 63 lift on Perplexity and nothing at all on Claude — 0 of 63 in both waves, a true null at four weeks. One variable explains the split: index access. The Bing submission reached the index Perplexity reads; there is no equivalent path into Anthropic's. And Claude's zero is not silence. Its poisoned citations fell 79 → 11, and a data broker rose 0 → 74 to fill the vacuum. An engine that cannot find the authoritative page does not abstain. It substitutes.
CrossSource's judge passed its blind check at 15/15. This one did not. Against 40 blind human labels, the Claude judge matched the exact tag set on 3 (8%) and the Gemini judge on 12 (30%); the two judges agreed with each other on 117 of 332 answers (35%). Ten repetitions of an identical prompt changed the tag set 86% of the time. It was the third consecutive blind check to say the same thing.
So no interpretive category on this page is a rate. Bands only — and every headline number is a count with a denominator. The one exception is poisoned_citation, 27% → 19%, which is computed against canon rather than judged.
It became a point estimate the moment it stopped being asked of a language model.
The method survives the judge failing. When validation says the judge cannot be trusted, you do not soften the claim — you change what kind of number you publish. Instrument what can be counted; demote what must be interpreted; separate index-driven effects from model-driven ones so you know which fix owns the failure. It is the same discipline as CrossSource, and the opposite outcome, which is why the two belong together.
Engine versions are not verifiably frozen between waves (the citation trail is index-driven and robust to that; the taxonomy is not). The last fixes had about two weeks of recrawl, not four. The six interventions were a cluster, not isolated tests. Human validation covers 40 of 227 eligible items, stratified — not a full audit. Canon had a coverage gap: several "fabricated" numbers were true, published figures missing from canon.yaml. Verbatim engine answers are withheld because cited surfaces can include other people's public posts.
PREDICTIONS.md was committed before Wave 1 ran and scored publicly, misses first: two of nine held. The misses changed the roadmap more than the hits did.
v1 is closed with this wave. A Claude-only mini-wave on the frozen instrument runs once a real crawler-access change has had time to land. v2 moves from chat engines to the agentic stacks that actually screen and source people — paired designs, counted metrics, human-gated labels.
Read the code, run it on yourself → Release v1.0 → Predictions →
Read the code, run it on yourself →