Lakshya

Artificial Intelligence

Applied track Named 59% · titled <2%

6 · AI Evaluation & Quality

The largest gap between what is demanded and what is studied. Read this section even if you are targeting another archetype.

What the job actually is

Deciding whether a non-deterministic system got better or worse, with enough rigour that a team can ship on the answer. Golden sets, regression suites, human labelling protocols, LLM-as-judge and its many biases, and the statistics of drawing conclusions from a few hundred examples. It is measurement science applied to a system that gives a different answer every time you ask.

The arbitrage is stark: 59% of job descriptions name evaluation, and under 2% have it as a job title. Almost nobody prepares for it, and it is assessed in every applied-track loop. Getting good at this raises your performance in archetypes 1, 2 and 3 simultaneously — which makes it the highest leverage thing on this page.

Titles this hides behind

Data Scientist — EvaluationsAI Quality Engineer SDET — AI EvaluationApplied Scientist, Evaluation Research Engineer, Evals

Prepare like this

  • Build a real eval harness for something you have shipped. A golden set, a scoring function, a regression run, a report. Fifty well-chosen examples beat five thousand scraped ones.
  • Learn the LLM-as-judge failure modes by name: position bias, verbosity bias, self-preference, and score compression at the top of the scale. Then learn the mitigations — swapping order, pairwise comparison, calibrating against human labels.
  • Always carry a control. A number without a baseline is not a result. If your agent resolves 71% of tickets, the interview question is 71% against what — the previous version, a human, or a trivial keyword rule that would have got 65%. This single habit distinguishes strong candidates more reliably than any technical depth.
  • Know your statistical power. On 200 examples, a 3-point difference is usually noise. Be able to say so, with the arithmetic.
  • Design a human labelling protocol including the disagreement rate you would accept and what you do when annotators disagree.

Be ready for

  • "How do you evaluate something with no single correct answer?" Rubrics, pairwise preference, and task-completion proxies — plus honesty about what each misses.
  • "Your LLM judge agrees with humans 85% of the time. Is that good enough to ship on?" Depends entirely on where the 15% falls. Strong answers ask about the error distribution.
  • "How do you stop your golden set from going stale?"
  • "You improved the eval score and users complained more. Explain." The eval measured the wrong thing — and whether you can say that plainly is the test.

Compensation

United States — base$160k – $260k Underpriced relative to demand today, which is exactly what an emerging title looks like.
India — total₹25L – ₹55L Estimated. Sarvam and Postman both post evaluation-specific roles.

What this role tests

Themes, and where to learn them

These chapters are shared across every role that tests them, so preparation here compounds rather than being spent once.

Measuring what resists measurementNamed in 59% of AI postings and 42% of platform postings. Almost nobody studies it deliberately.READ THE CHAPTER →The controlOne habit separates strong candidates from plausible ones more reliably than any technical depth.READ THE CHAPTER →

Practice questions across all themes →  ·  Back to Artificial Intelligence →