Work
Four companies, two promotions, one through-line: quality you can measure.
Open work
Open work
Three public evaluation harnesses, built to one method.
CrossSource — citation accuracy in legal RAG Case studyRepo
mirror-eval — what AI search engines say about a person Case studyRepov1.0
screener-eval — what LLM résumé screeners actually reward Case studyRepo
Instead
AI-native tax research & planning platform; first new entrant to clear IRS e-filing approval alongside incumbents.
-
Eval loop
Own end-to-end output quality for the production tax-research LLM: a citation-accuracy evaluation loop (flag → classify by failure mode → weekly review) and a four-dimension rubric (answerability, accuracy, citation quality, actionability) that benchmarks model configurations and catches regressions before release — ~95% citation accuracy held on the eval set.
-
Benchmarking
Built the competitive model-benchmarking program: the platform scored against three competing platforms across 50+ scenarios over 5 rounds; designed a "pipeline-collapse" eval measuring chain coherence and self-QC honesty across a full multi-step tax workflow.
-
Failure mode
Fixed a systemic "knowledge–citation gap" failure mode — plausible-but-wrong citations, a false-positive problem — by attaching document identity and sub-type to every cited source. Citation precision rose; outputs became traceable and audit-ready.
-
Latency
Cut workflow latency ~90% (10–15 minutes → under a minute) by scoping incremental edits to changed inputs instead of reprocessing the full document set.
-
Corpus
Architected and QA'd the 270K+-record RAG corpus across 100+ legal source types: ingestion-pipeline PRDs, MongoDB schemas, a Python fetchability harness across 378 sources, and remediation of ~139K scraped documents.
-
Research
Led legal-support research across 160+ tax strategies (a Legal Support Matrix built from Tax Court cases, IRS rulings, and audit guidance) and shipped Source Explorer — command-palette search with cross-type related sources.
Public record
What's public. The eval numbers above are from production; the reproducible version is CrossSource.
How the AI cites tax law — Research product page
instead.com
CrossSource — open harness, MIT
github.com
The judge caught a bug — essay
zoebnomi.com
Multiplier
Global employment platform enabling compliant hiring across 150+ countries.
-
Owned the Value-Added Services vertical (procurement, BGV, ITSM, partnerships): $100K incremental revenue and a 47% efficiency gain from new SOPs; launched the PosterElite compliance partnership end-to-end in 21 days.
Public record
What's public. The PosterElite launch was not announced publicly; the claim stands, unlinked.
Background screening for global hires — Veremark partner page
usemultiplier.com
Initiate background verification through Multiplier — help centre
help.usemultiplier.com
IT asset & equipment support at onboarding — product news
updated Aug 2025
Keka HR
India's leading HR technology platform; $25M+ ARR CoreHR suite.
Built the Background Verification module zero-to-one on the Checkr API — Keka's first US-market product. From nothing to 28 US enterprise clients and $2.7M MRR, it drove the company's first US expansion and was a full course in integration edge cases, compliance constraints, and enterprise onboarding.
-
Also
6 SSO integrations (Azure AD / Google Workspace) → 27% adoption lift and $121K upsell; led the Exit Module revamp → offboarding time down 28%, CSAT up 80%.
Public record
What shipped, on Keka's own surfaces. None of these name me; the profile piece does.
Background Verification module — product video
youtube.com
Keka × Checkr — marketplace listingCheckr's integration guide
keka.com · help.checkr.com
Employee Exit module revamp — announcementexit-process guide
help.keka.com
help.keka.com
Verification partners — SpringVerifyOnGridHelloVerify
keka.com
Profile — "Meet Zoeb Nomi: the rabbit-hole explorer and CoreHR maestro"
Jan 2025
Hurix Digital
-
Led agile transformation of LMS development (43% faster delivery); drove 31% growth in customer conversions via an end-to-end video-learning rollout.
Mechanical engineering → enterprise product → AI product quality.
If you're shipping an LLM product, let's talk about what your evals miss.