885 screenings, one résumé, and the two levers that did not move

Every piece of résumé advice I have ever been given comes down to two levers. Put the recognisable names where the eye lands first. Link to the proof. I had never seen either one measured against the thing that actually reads a résumé now, so I measured them.

The setup is small on purpose. One résumé — mine. Twenty-one job descriptions, every open US AI product manager role I could find at the frontier labs and large tech. Two cheap screening models, the tier a vendor would run at volume, given a scoring rubric from published research and asked for a 0–100 fit score and a verdict. Then the two levers, one at a time: a copy of the résumé with every employer replaced by a company that does not exist, and a copy with every link deleted. Same job description, both versions, same model, same hour, five times over. 885 scored calls, under four dollars.

Swapping every employer for a fiction moved the score by 0.07 points on one screener and −1.1 on the other. Deleting every link: −0.9 and +0.4. All four confidence intervals straddle zero. On this résumé, against these postings, the two levers did nothing a counting method could see.

What did move the score was the screener. On the same résumé and the same posting, the two models disagreed by 22 points on average and by 51 at the extreme. One rejected me 58% of the time; the other put me on hold 68% of the time. They agreed on the verdict for six postings out of twenty-one. One of them handed out only ten distinct scores across 233 calls, and gave exactly 28 to eighty-four of them. That is not a scale. That is a rubric leaking through.

And under all of it, the noise. Ask the same model the same question about the same résumé five times and the score wobbles by three to five points. On a third to a half of the postings, the screener flipped its own verdict between reps. Every free "ATS score" tool I have seen returns one number from one call. That number is a sample from the wobble.

I want to be careful about what the nulls mean. The employer swap replaced companies a US screener has not heard of with companies nobody has heard of; it says nothing about whether a famous name would help, and I cannot run that test on my own résumé. The deleted links were text a language model cannot follow anyway, so the test was whether the presence of links signals anything, and it did not. These are narrow, honest results about one candidate and two cheap models.

The wider lesson is not about résumés. It is that the interesting variable in an LLM-mediated system is rarely the one people are optimising. I spent a week making sure my evidence links were printed as text after finding that every extractor dropped them as invisible annotations — that fix was real, 26 of 29 fields to 29 of 29, and it is the one change in this study I would tell anyone to make. But once the text is in, the score belongs to the screener, and the screener belongs to whoever chose it.

The harness is public. Point it at your own résumé; the only thing you need to change is the PDF.

If you're shipping an LLM product, let's talk about what your evals miss.