Writing

Notes on eval methodology, LLM-as-judge reliability, and RAG output quality — drawn from CrossSource and production eval work. And, occasionally, an industry thesis written so it can be wrong.