Three harnesses, one method.
Three open eval harnesses, side by side: legal RAG citations, what AI search says about a person, and LLM résumé screeners. Each row is one harness — the question it answers, the number it stands on, how it was run.
-
CrossSource
When the RAG cites a court opinion, does that opinion support the claim?
[judged]
0.994
citation precision, strict prompt (baseline 0.981) · n=25 · judge 15/15 vs human
- 22 opinions · 25 questions
- LLM judge, blind-validated 15/15
- BM25 top-5 · baseline vs strict
-
mirror-eval
What do AI search engines say about me, and is it true and sourced to something I control?
[counted]
13 → 63
Perplexity citations, wave 1 → wave 2, of 63 probes · Claude 0 → 0
- 4 engines · 63 search probes each
- judges failed 3/40, 12/40 → counted instead
- 2 waves · Aug → Sep 2026
-
screener-eval
Do employer names and evidence links change what an LLM résumé screener scores?
[null]
< 1 pt
score shift when every employer is swapped · 885 calls · 21 postings
- 21 postings · 2 screeners · 5 reps
- no judge in the loop
- 95% CIs straddle zero
One judge passed and caught a bug. One judge failed and the study survived on counts. One needed no judge at all. Same method three times — that is the point.
| Harness | Question | Judge | Counted metric | Sample | Verdict |
|---|---|---|---|---|---|
| CrossSource | When the RAG cites a court opinion, does that opinion support the claim? | LLM judge, blind-validated against a human: 100% (15/15) | Claim-level citation precision, 0.981 → 0.994; recall 0.760 in both configurations | 22 opinions · 25 questions | Judge passed and caught a harness bug; prompting buys precision, retrieval owns recall |
| mirror-eval | What do AI search engines say about me, and is it true and sourced to something I control? | Two LLM judges from different model families — failed blind validation, 3/40 and 12/40 | Probes citing an owned surface, of 63 search-mode probes per engine — counted, judge-free | 4 engines · 63 search probes each · 2 waves | Perplexity 13 → 63, Claude 0 → 0; index access explains the split |
| screener-eval | Do employer names and evidence links change what an LLM résumé screener scores? | None — nothing is judged by a model | Paired delta in fit score (B − A), 10,000-resample bootstrap 95% CI | 1 résumé · 21 postings · 2 screeners · 885 calls | Employer swap and link deletion each moved the score under a point; which screener read it moved it 22 |
Build the judge. Validate it blind against a human. Count what you can. Report the rest as bands.
If you're shipping an LLM product, let's talk about what your evals miss.