screener-eval: I ran my résumé through an LLM screener 885 times

Same résumé, same 21 job descriptions, two LLM screeners, five repeats each. Swapping every employer on the résumé for a fictional one moved the fit score by less than a point. Deleting every link moved it by less than a point. Which screener read it moved it by 22.

Before a person reads a résumé, a parser extracts it and, increasingly, a language model scores it. Nobody outside the vendor sees the score, and the advice industry around it — brand names up top, links to proof, keywords — is advice nobody has measured. This is the measurement: one real résumé, the currently open US AI product roles at large tech and frontier labs, two cheap screening models of the kind a vendor would actually run, and a paired design so the delta is the unit. The same eval discipline as CrossSource and mirror-eval, pointed at the one system where applications actually die.

Résumé
Mine, v4.2 (résumé A). Two variants, each generated by a script from A: B-institution replaces every employer and institution name with a fictional unknown of matching description; B-links deletes every URL (portfolio, GitHub, LinkedIn, repo).
Job descriptions
21 open, US-based AI Product Manager postings retrieved 2026-09-10 — 12 at frontier labs (OpenAI 6, Anthropic 3, Google DeepMind 3), 9 at large tech (Google 2, Meta 2, Amazon 2, NVIDIA 2, Microsoft 1). Director, Principal and Head titles excluded. Posting text is used locally and never published; the repo ships URLs and hashes.
Screeners
claude-haiku-4-5 and gpt-5-mini — the cheap tier, because that is what screening at volume runs on. The screening prompt is a rubric from published research on LLM résumé screening; no vendor publishes theirs, so this is a proxy and is labelled as one. Each call returns strict JSON: a 0–100 fit score, advance / hold / reject, reasons, objections, discounted claims.
Design
Paired. For every job description, A and B are scored by the same model in the same run window, order randomised; the unit of analysis is the per-JD delta (B − A). Conditions are interleaved within one window, not run on different days.
Reps
A 10-rep pilot on one posting put the run-to-run standard deviation at 3.7 (Haiku) and 4.4 (gpt-5-mini) points and recommended 12 reps; 5 were run, which resolves effects of about 3 points and up. 885 scored calls in total, 0 parse errors, under $4 in API calls.
Statistics
Mean paired delta with a 10,000-resample bootstrap 95% confidence interval and a two-sided sign test. Nothing here is judged by a model.
Parser test (separate, deterministic)
The résumé PDF through three plain text extractors and two open-source résumé parsers, field-level errors counted against a hand-written canon of 29 fields.

Paired delta in fit score, B − A, 21 job descriptions, 5 reps:

Table 1 — paired delta by lever and screener
Lever ScreenerMean delta95% CISign test p
Every employer → fictional unknown claude-haiku-4-5+0.07−3.05 to +2.761.00
Every employer → fictional unknown gpt-5-mini−1.14−3.03 to +0.840.19
Every link deleted claude-haiku-4-5−0.90−2.90 to +1.031.00
Every link deleted gpt-5-mini+0.41−1.82 to +2.711.00

What did move the score — the same résumé, the same posting, condition A only:

Table 2 — condition A, by screener
Measure claude-haiku-4-5gpt-5-mini
Mean fit score 44.864.7
Most common verdict reject (136 of 233)hold (157 of 232)
Distinct scores given, across 233 calls 1030
Postings where every rep returned the identical score 17 of 42 cells0 of 42
Run-to-run SD, same posting, same résumé 3.04.8
Postings where the screener flipped its own verdict across reps 6 of 2112 of 21

Between the two screeners, on the same résumé and the same posting: mean gap 22.3 points, largest 51.3; gpt-5-mini scored higher on 19 of 21 postings; the two agreed on the majority verdict for 6 of 21 postings; correlation between their per-posting scores 0.46.

The finding that matters: the two levers résumé advice is built on did nothing measurable. Replacing every employer with a company that does not exist, and deleting every link to proof, each moved the score by less than a point, with confidence intervals that straddle zero on both screeners. The choice of screener moved it by 22 points on average and flipped the verdict on 15 of 21 postings. And the screener's own noise — a 3-to-5-point wobble on identical input, a verdict that flips against itself on a third to a half of postings — is larger than either lever. A single-shot "ATS score" from any tool is a sample from that wobble.

The screener you get is the variable.

The employer swap replaced companies a US screener has not heard of with companies that do not exist. It says nothing about swapping in a famous name; that is the prestige test, and it needs a different résumé than mine to run. The link deletion removed URL text a language model cannot follow anyway; it tests whether the presence of links signals anything to the screener, and it did not. Both nulls are about this résumé, these 21 postings, this rubric, and these two cheap models. They are counted honestly and they are narrow.

Before any model sees a résumé, an extractor does. Field-level extraction against a 29-field canon, résumé v4.1 → v4.2:

Table 3 — field-level extraction, v4.1 → v4.2
Parser v4.1 (ok / missing / wrong / garbled)v4.2
pdftotext 26 / 3 / 0 / 029 / 0 / 0 / 0
PyMuPDF 26 / 3 / 0 / 029 / 0 / 0 / 0
pdfminer.six 26 / 3 / 0 / 029 / 0 / 0 / 0
resumix (npm) 7 / 9 / 12 / 19 / 8 / 11 / 1
resume-parser (npm) 2 / 4 / 1 / 222 / 2 / 1 / 24

The three fields every plain extractor missed on v4.1 were the same three: LinkedIn, portfolio, GitHub. The links were hyperlink annotations under plain words, with no URL text on the page — invisible to anything that reads text. Printing the addresses fixed it: 29 of 29. The two résumé-specific parsers were worse than plain text extraction on both versions, and one of them read the two-column header as my name.

Paired, repeated, counted. The experiment cost under $4 and a morning, and it replaces two pieces of received wisdom with two measured nulls and one measured effect nobody talks about: the screener you get is the variable. The same discipline as the other two harnesses — decide what is countable, count it, state what the count cannot support — with no judge in the loop at all.

One candidate's résumé. Twenty-one postings from eight employers, so shared boilerplate limits independence. Five reps against a pilot recommendation of twelve; the intervals are wide enough that effects under about three points are not excluded. The rubric is a published-research proxy, not any vendor's system. The institution-swap run was stopped after five complete reps and a partial sixth, which is included where both conditions of a pair completed. Cheap-tier models only; a frontier-tier screener may behave differently. Early in the pilot, gpt-5-mini returned empty content under a small output budget because its reasoning tokens consumed it — fixed before any counted run, disclosed because a screening vendor could make the same mistake.

A prestige arm (well-known employer names swapped in), a JD-vocabulary-alignment arm, and the same design on a second candidate's résumé — each is one config change on the public harness.

Read the code, run it on your own résumé →

Read the code, run it on your own résumé →