Home › AI Interview Prep › AI Evaluation Engineer
Growing Fast

AI Evaluation Engineer Interview Questions and Answers

Thirty scenarios on eval design, LLM-as-judge, regression testing, human review and safety evaluation. Write your own answer first, by typing or speaking, then open the model answer to compare structure and reasoning.

30 scenariosWhat each question testsRed-flag answersNo sign-up, private
Written by Ayush Bisht · Reviewed by Sanjay Saini
Last updated 2026-09-30
AI Evaluation Engineer interview questions

Answer in your own words before opening a model answer, by typing or by pressing Speak your answer. Your text is saved in this browser only and is never uploaded to us. Voice input uses your browser's speech service to turn speech into text; in Chrome that audio is processed by Google.

Scenario 1Eval DesignPractitioner

How do you build an eval set for a new LLM feature?

What the interviewer is testing: Whether your eval set covers real, edge and adversarial cases.

Model answer, red flags and follow-up

A strong answer

I collect real inputs from logs, support tickets and domain experts, then add synthetic cases to cover rare, edge and adversarial situations. Each item gets an expert-labelled answer or a rubric, and the set is versioned like code. I tag items by scenario so I can report results by segment. A held-out portion stays untouched so prompts are not tuned to the tests. The set is never finished: production failures flow back in, and I review coverage regularly against how users actually use the feature.

Answers that lose you the room

  • Uses only easy or synthetic examples
  • Keeps no held-out set
  • Has no labelled answers or rubrics

Expect this follow-up: How do you know your eval set covers real usage?

Scenario 2LLM-as-JudgeAdvanced

What are the risks of using an LLM as a judge, and how do you mitigate them?

What the interviewer is testing: Whether you calibrate judges against humans and known biases.

Model answer, red flags and follow-up

A strong answer

LLM judges show position bias, length bias and a preference for their own style, and they can be inconsistent between runs. I mitigate by writing clear rubrics with score definitions and examples, randomising answer order, and using a judge from a different model family than the one being graded. Most importantly I calibrate against human labels: I measure agreement on a sample, and I keep auditing samples over time because judge behaviour drifts when models update. A judge I have not validated against humans is an opinion, not a metric.

Answers that lose you the room

  • Trusts the judge without calibration
  • Ignores position and length bias
  • Uses the same model to generate and judge

Expect this follow-up: Judge and human agree only 65% of the time. What do you do?

Scenario 3MetricsPractitioner

How do you choose metrics for a summarisation or Q&A system?

What the interviewer is testing: Whether metrics map to the failures that matter.

Model answer, red flags and follow-up

A strong answer

I start from the failure that matters for the product. For summarisation that might be factual errors and missed key points; for Q&A, wrong or unsupported answers. I define metrics for each dimension, such as factuality, completeness, relevance and tone, and combine automatic checks, rubric-based judging and human review for the harder calls. Results are reported by segment, because an average can hide a group of users who are badly served. I retire metrics that no longer predict user satisfaction.

Answers that lose you the room

  • Reports one average metric
  • Uses BLEU or ROUGE alone
  • Ignores factuality

Expect this follow-up: Summaries score well but users say they miss key points. What's wrong?

Scenario 4RegressionPractitioner

How do you stop prompt or model changes from silently breaking quality?

What the interviewer is testing: Whether evals gate releases in CI with real thresholds.

Model answer, red flags and follow-up

A strong answer

I put the eval suite in CI so every prompt, model or retrieval change is scored automatically against the baseline. Each critical metric and segment has a threshold, and the comparison uses confidence intervals so noise is not mistaken for a change. A change that regresses a critical case blocks the release, with an explicit override that needs an owner's sign-off. I also monitor production, because some regressions only appear on real traffic. Silent breakage happens when nobody is measuring, so I make measurement automatic.

Answers that lose you the room

  • Tests manually before release
  • Sets no thresholds or baselines
  • Ignores segment-level regressions

Expect this follow-up: The overall score rose but one critical segment dropped. Do you ship?

Scenario 5Human ReviewPractitioner

How do you design a reliable human annotation process?

What the interviewer is testing: Whether annotation is reliable and agreement is measured.

Model answer, red flags and follow-up

A strong answer

I write guidelines with clear definitions and worked examples, train annotators and run a pilot round before scaling. A subset of items is labelled by several raters so I can measure inter-rater agreement, using a statistic that corrects for chance. Disagreements are reviewed together, and they usually reveal ambiguous rubric wording that I then fix. I monitor annotator quality over time with hidden gold items, and I keep the guidelines versioned. Low agreement means the task definition is unclear, not that raters are careless.

Answers that lose you the room

  • Skips guidelines and training
  • Uses a single annotator
  • Never measures agreement

Expect this follow-up: Raters keep disagreeing on one criterion. What do you fix?

Scenario 6Online vs OfflinePractitioner

How do offline evals and online metrics work together?

What the interviewer is testing: Whether offline scores are validated against real outcomes.

Model answer, red flags and follow-up

A strong answer

Offline evals are cheap, fast and repeatable, so they gate releases. Online metrics such as acceptance rate, edits, retries, complaints and task completion show how the feature behaves with real users. The two must connect: I check whether offline scores actually predict online outcomes, and if they do not, I fix the offline eval. Production failures and disagreements are fed back as new offline cases. Each side covers the other's blind spot: offline misses real usage, and online is slow and noisy.

Answers that lose you the room

  • Relies on offline scores alone
  • Ignores online signals
  • Never checks that offline predicts online

Expect this follow-up: Offline scores improved but users are less satisfied. Why?

Scenario 7Safety EvalsAdvanced

How would you evaluate a model for harmful or biased outputs?

What the interviewer is testing: Whether you test harm and group gaps systematically.

Model answer, red flags and follow-up

A strong answer

I define harm categories and user groups relevant to the product, build test prompts for each, and mix expert-written and automated adversarial attacks. I measure violation and refusal rates, including over-refusal of legitimate requests, and compare performance across demographic groups and languages to find gaps. After mitigations I re-test to confirm they worked without creating new problems. Results feed explicit release criteria with thresholds and owners. A one-time safety review is not enough, because models and prompts change.

Answers that lose you the room

  • Tests only obvious harmful prompts
  • Ignores group-level gaps
  • Doesn't re-test after fixes

Expect this follow-up: How do you decide the residual risk is acceptable?

Scenario 8BenchmarksFoundation

A vendor claims state-of-the-art benchmark scores. Do you trust them?

What the interviewer is testing: Whether you distrust claims that skip your own task.

Model answer, red flags and follow-up

A strong answer

Not on their own. Public benchmarks can be contaminated by training data, saturate quickly, and often measure something different from my task. A vendor's number also comes from their prompts and settings, not mine. I run the candidate model on my own eval set with my prompts, and compare quality alongside cost, latency, context limits and data terms. The benchmark can help build a shortlist, but the decision rests on evidence from my data. If the vendor will not let me test, that is itself informative.

Answers that lose you the room

  • Accepts the claim as fact
  • Ignores contamination
  • Doesn't test cost and latency

Expect this follow-up: The model wins on the vendor's benchmark and loses on yours. What do you report?

Scenario 9Golden DatasetsPractitioner

How do you keep a golden dataset useful over time?

What the interviewer is testing: Whether datasets evolve with production failures.

Model answer, red flags and follow-up

A strong answer

A golden dataset decays if it is frozen. I version it, add new failures from production, retire cases that no longer reflect how the product is used, and review labels periodically since standards and facts change. I track coverage across use cases and user segments so gaps are visible. I keep a protected held-out slice to avoid overfitting. Each change to the dataset is logged so score changes can be explained by a code change or a dataset change, and not confused.

Answers that lose you the room

  • Freezes the dataset
  • Never adds production failures
  • Doesn't review labels

Expect this follow-up: How often would you refresh it, and what triggers a change?

Scenario 10Synthetic DataPractitioner

When is synthetic data appropriate for evals?

What the interviewer is testing: Whether you use synthetic data without fooling yourself.

Model answer, red flags and follow-up

A strong answer

Synthetic data is useful for expanding coverage of rare or adversarial cases and for bootstrapping before real traffic exists. I validate samples by hand, mix them with real data and label them so results can be split. I avoid using the same model to both generate and grade the data, because that inflates scores. Synthetic items tend to be cleaner and more predictable than real ones, so I never rely on them alone to judge readiness. As real data arrives it takes over.

Answers that lose you the room

  • Uses synthetic data only
  • Lets the same model generate and grade
  • Doesn't validate samples by hand

Expect this follow-up: Where could synthetic data mislead you?

Scenario 11RAG EvalsPractitioner

How do you evaluate a RAG system?

What the interviewer is testing: Whether you evaluate retrieval and generation separately.

Model answer, red flags and follow-up

A strong answer

I evaluate retrieval and generation separately. For retrieval I measure recall, precision and ranking quality against labelled relevant passages. For generation I measure faithfulness to the retrieved sources, answer relevance and completeness, and I test unanswerable questions to check the system says so. Separating the two shows whether to fix search or the prompt. A wrong answer can come from missing context or from misusing good context, and the fix is different. I also track citation accuracy where the product shows sources.

Answers that lose you the room

  • Evaluates only final answers
  • Doesn't separate retrieval from generation
  • Ignores faithfulness to sources

Expect this follow-up: Recall is high but answers are wrong. Where do you look next?

Scenario 12Agent EvalsAdvanced

How do you evaluate multi-step agent trajectories?

What the interviewer is testing: Whether you score trajectories, not just outcomes.

Model answer, red flags and follow-up

A strong answer

I judge the outcome and the path. Outcome checks ask whether the task was completed correctly. Path checks look at tool choices, unnecessary or repeated steps, policy or safety violations, recovery from errors, cost and latency. I use sandboxed environments with mocked or resettable tools so runs are repeatable, and I run each scenario several times because behaviour varies. Trajectory-level scoring can use rubrics or judges calibrated against humans. Only checking the final message misses agents that get the right answer by risky means.

Answers that lose you the room

  • Judges only the final outcome
  • Ignores unnecessary steps and unsafe actions
  • Runs on live systems

Expect this follow-up: Two agents succeed, one in 5 steps and one in 40. How do you compare them?

Scenario 13Statistical RigourAdvanced

LLM eval scores fluctuate between runs. How do you draw reliable conclusions?

What the interviewer is testing: Whether you draw conclusions with proper statistics.

Model answer, red flags and follow-up

A strong answer

Because outputs vary, I use enough samples for the effect size I care about and repeat runs to estimate noise. I report confidence intervals, control randomness where I can with fixed seeds and settings, and use paired comparisons, since the same items run under both versions reduce variance. I avoid concluding anything from small differences, and I check by segment, correcting for the number of comparisons I make. If a change is within the noise, I say so, and I either collect more data or treat it as no difference.

Answers that lose you the room

  • Concludes from small score differences
  • Runs each case once
  • Reports no confidence intervals

Expect this follow-up: How many runs would you do to trust a 2-point gain?

Scenario 14Scoring MethodsPractitioner

When do you use pairwise comparison instead of absolute scoring?

What the interviewer is testing: Whether you match the scoring method to the decision.

Model answer, red flags and follow-up

A strong answer

Pairwise comparison asks which of two answers is better, and people and judges are more consistent at that than at assigning absolute scores, especially for subjective qualities like helpfulness or tone. It is the right tool for comparing two versions. Absolute rubrics suit pass or fail thresholds, safety criteria and tracking a metric over time. I often use both: pairwise for choosing between candidates and rubric scores for gating and trends. I randomise order in pairwise tests to remove position bias.

Answers that lose you the room

  • Uses absolute scores for everything
  • Doesn't know when pairwise helps
  • Ignores rater consistency

Expect this follow-up: Which method would you use to track quality over six months, and why?

Scenario 15Rubric DesignFoundation

What makes a good evaluation rubric?

What the interviewer is testing: Whether criteria are specific, observable and separable.

Model answer, red flags and follow-up

A strong answer

A good rubric has specific, observable criteria, clear definitions for each score level with examples, and separates dimensions such as accuracy, completeness and tone instead of blending them into one number. I keep the scale small so raters can distinguish the levels. Then I test it: several raters score the same items, I measure agreement, and I revise wording where they disagree. A vague rubric like 'is it good' produces noisy scores that cannot support decisions.

Answers that lose you the room

  • Uses vague criteria like 'good quality'
  • Bundles accuracy and tone together
  • Never tests with several raters

Expect this follow-up: Write one criterion with score definitions for 'helpfulness'.

Scenario 16Eval CostPractitioner

Your eval suite is slow and expensive. How do you fix it?

What the interviewer is testing: Whether you tier evals to stay fast and affordable.

Model answer, red flags and follow-up

A strong answer

I tier the suite. A small, fast smoke set runs on every change, the fuller suite runs nightly or before release, and the most expensive checks, such as human review, run only at milestones. I cache results for unchanged inputs, sample smartly instead of running everything, remove redundant cases and use cheaper judges once I have calibrated them against a stronger one. I track the cost of each eval so I can see where the money goes. The aim is fast feedback for developers without giving up coverage.

Answers that lose you the room

  • Runs the full suite on every change
  • Cuts cases at random
  • Uses cheaper judges without calibration

Expect this follow-up: Which tests go in the fast tier, and why?

Scenario 17ContaminationAdvanced

How do you detect and avoid benchmark contamination?

What the interviewer is testing: Whether you protect evals from leakage into training.

Model answer, red flags and follow-up

A strong answer

Contamination means the model saw the test items during training, which inflates scores. I keep private held-out sets that never appear publicly, use fresh or paraphrased items, and check for verbatim overlap with known corpora where I can. I prefer task-specific evals built from my own data over public leaderboards, and I look for suspicious signs such as very high scores on old items but lower ones on new but similar items. I treat any public benchmark result with caution.

Answers that lose you the room

  • Trusts public leaderboards
  • Keeps no private held-out set
  • Doesn't check for overlap

Expect this follow-up: How would you suspect a model has seen your test data?

Scenario 18Multilingual EvalsPractitioner

How do you evaluate quality across languages?

What the interviewer is testing: Whether you evaluate each language with native reviewers.

Model answer, red flags and follow-up

A strong answer

I build test sets per language, written or reviewed by native speakers, using real user text including code-mixed and transliterated input. I compare scores by language and never assume English results transfer. I check for translation artefacts in the data, cultural fit and tone, and I include language-specific failure modes such as wrong script or mixed languages in the reply. Judge models can be weaker in some languages, so I calibrate them against native-speaker ratings before trusting them.

Answers that lose you the room

  • Assumes English results transfer
  • Uses machine-translated tests only
  • Has no native reviewers

Expect this follow-up: Scores are lower in Tamil. How do you tell if it's the model or the test?

Scenario 19No Ground TruthAdvanced

How do you evaluate open-ended tasks with no single right answer?

What the interviewer is testing: Whether you can evaluate open-ended output credibly.

Model answer, red flags and follow-up

A strong answer

When there is no single right answer, I define what good looks like through rubrics covering properties such as factuality against sources, coverage of key points, clarity and safety. I combine expert review, pairwise preference comparisons and automated property checks, and calibrate any LLM judge against human ratings, tracking agreement over time. I accept that scores are noisy and report distributions, not a single number. The goal is not certainty but decisions that are better informed than opinion.

Answers that lose you the room

  • Insists on one right answer
  • Relies on exact match
  • Never checks the judge against humans

Expect this follow-up: How do you evaluate a creative-writing assistant?

Scenario 20Error AnalysisPractitioner

How do you do error analysis on failing cases?

What the interviewer is testing: Whether you turn failures into a prioritised taxonomy.

Model answer, red flags and follow-up

A strong answer

I sample failures, read them carefully and group them into a taxonomy such as retrieval miss, misread instruction, hallucination, formatting or tool error. Then I count frequency and severity for each category and trace root causes. This gives a prioritised fix list: the biggest and most severe category first. I review the taxonomy as the product changes, and I keep examples of each category so the team shares one language. Error analysis is usually the highest-value hour in an eval process.

Answers that lose you the room

  • Reads a few failures anecdotally
  • Has no failure taxonomy
  • Doesn't rank by frequency and severity

Expect this follow-up: You have 200 failures. How do you start?

Scenario 21ReportingFoundation

How do you report eval results to executives?

What the interviewer is testing: Whether results become a decision, not a score dump.

Model answer, red flags and follow-up

A strong answer

Executives need a decision, not a data dump. I give a one-page view: headline metrics against agreed thresholds, the trend over time, the top failure categories with examples, the remaining risk, and a clear recommendation to ship, hold or fix. I explain what the numbers mean in business terms and state the uncertainty honestly. Detail is available in an appendix for those who want it. If the recommendation is not obvious from the first paragraph, the report is too complicated.

Answers that lose you the room

  • Sends raw score dumps
  • Gives no thresholds or recommendation
  • Hides failure categories

Expect this follow-up: What one sentence would you say to an executive about release readiness?

Scenario 22Red TeamingAdvanced

How do you structure a red-teaming exercise?

What the interviewer is testing: Whether red teaming is structured and feeds release gates.

Model answer, red flags and follow-up

A strong answer

I start by defining harm categories and threat models relevant to the product, including realistic attackers and misuse. I combine expert human red-teamers, who are creative, with automated attack generation, which gives scale. Every attempt is recorded with its outcome so I can report success rates. Findings are triaged, fixed and retested, and successful attacks are added to the eval suite as permanent regression tests. The exercise feeds release gates and is repeated when the model, prompt or tools change.

Answers that lose you the room

  • Runs a one-off jailbreak session
  • Has no threat model or harm categories
  • Lets findings skip the release gates

Expect this follow-up: How do you turn red-team findings into permanent tests?

Scenario 23CalibrationAdvanced

How do you check whether a model's confidence can be trusted?

What the interviewer is testing: Whether you test if confidence tracks accuracy.

Model answer, red flags and follow-up

A strong answer

I compare the model's confidence, whether stated or derived from token probabilities or self-consistency, with actual accuracy. Calibration plots and error rates by confidence band show whether a 90% confident answer is right about 90% of the time. If confidence is well calibrated, I can use it for routing, abstention and human escalation. If not, I recalibrate or use other signals, such as retrieval agreement. Verbalised confidence from LLMs is often overconfident, so I never assume it is reliable without checking.

Answers that lose you the room

  • Takes model-stated confidence at face value
  • Uses no calibration plots
  • Doesn't use confidence for routing

Expect this follow-up: The model says 95% confident but is right 70% of the time. What do you do?

Scenario 24A/B TestingPractitioner

How do you A/B test an LLM feature?

What the interviewer is testing: Whether you run experiments with guardrail metrics.

Model answer, red flags and follow-up

A strong answer

I define a primary metric and guardrail metrics such as cost, latency and safety before starting. Users are randomly assigned, the test runs long enough to reach significance and cover weekly patterns, and I watch for novelty effects that fade. I also read a sample of actual outputs from both arms, because a metric like engagement can improve while quality falls. I decide in advance what result would make me ship, hold or stop, so I do not rationalise afterwards.

Answers that lose you the room

  • Ends the test early on a good day
  • Has no guardrail metrics
  • Ignores cost and novelty effects

Expect this follow-up: The variant wins on clicks but complaints rise. What do you do?

Scenario 25Eval CulturePractitioner

How do you build an eval-driven culture in a team?

What the interviewer is testing: Whether evals become habit, not a launch event.

Model answer, red flags and follow-up

A strong answer

I make evals easy to run so nobody has an excuse to skip them: one command, fast feedback, clear output. They run in CI, appear on team dashboards and are part of the definition of done. Every production bug becomes a test case so the same failure does not return. I celebrate catches, where an eval prevented a bad release, and I share failure examples openly. Culture changes when people see evals saving them time and embarrassment, not when they are mandated.

Answers that lose you the room

  • Runs evals only before big launches
  • Makes evals hard to run
  • Doesn't turn production bugs into tests

Expect this follow-up: How would you make evals part of the definition of done?

Scenario 26Judge DesignAdvanced

Your LLM judge and your human reviewers disagree on 30% of cases. What do you do?

What the interviewer is testing: Whether you diagnose disagreement rather than simply trusting one side.

Model answer, red flags and follow-up

A strong answer

I would not assume either side is right. I sample the disagreements and read them, looking for patterns: the judge favouring longer answers, humans and the judge reading the rubric differently, or genuinely ambiguous items. If the rubric is unclear, I fix it and re-label. If the judge has a systematic bias, I adjust the prompt, add examples or change the judge model. Then I re-measure agreement on a fresh sample. Until agreement is acceptable, I use the judge only for trends, not for gating releases.

Answers that lose you the room

  • Trusts the judge because it is cheaper
  • Discards the human labels
  • Never reads the disagreeing cases

Expect this follow-up: How much agreement would you require before letting the judge gate a release?

Scenario 27Production MonitoringPractitioner

How do you monitor quality in production when you have no ground-truth labels?

What the interviewer is testing: Whether you use proxy signals, sampling and drift detection sensibly.

Model answer, red flags and follow-up

A strong answer

I combine several proxy signals: user feedback, edits, retries and abandonment, automated checks for format, grounding and policy, and an LLM judge scoring a sample of traffic. Regular human review of a random sample keeps everything honest. I watch trends and distributions for drift in input types, output length or refusal rates, and alert on changes. Flagged and low-scoring cases feed the eval set. No single proxy is trustworthy alone, but together they catch most problems early.

Answers that lose you the room

  • Waits for customers to complain
  • Relies on a single proxy metric
  • Reviews no real samples

Expect this follow-up: Thumbs-down rates are flat but you suspect quality has fallen. How do you check?

Scenario 28Eval DataPractitioner

How do you decide how large your eval set needs to be?

What the interviewer is testing: Whether you connect sample size to the decision and effect size.

Model answer, red flags and follow-up

A strong answer

It depends on the decision. To detect a small improvement with confidence I need many items; to catch big failures a smaller set is enough. I estimate the noise from repeated runs, decide the smallest difference I care about, and calculate the sample size needed. I also need enough items per segment to report on each one. Quality matters as much as quantity: a few hundred well-labelled, representative cases beat thousands of poor ones. I grow the set where uncertainty is highest.

Answers that lose you the room

  • Picks a round number arbitrarily
  • Ignores per-segment counts
  • Values size over label quality

Expect this follow-up: Your set has 200 items and two versions differ by 2 points. What do you conclude?

Scenario 29ToolingFoundation

What would you look for when choosing an evaluation framework or platform?

What the interviewer is testing: Whether you choose tools by workflow fit rather than popularity.

Model answer, red flags and follow-up

A strong answer

I look at how well it fits our workflow: running in CI, versioning datasets and prompts, supporting custom metrics and judges, tracing production requests and comparing runs side by side. I check data privacy, cost, export options to avoid lock-in, and how easily non-engineers such as reviewers can use it. I prototype on one real feature before committing. Sometimes a small in-house harness is enough, and a heavy platform adds overhead without value.

Answers that lose you the room

  • Picks the most popular tool by default
  • Ignores data privacy and export
  • Never prototypes on a real feature

Expect this follow-up: Would you build or buy for a team of five engineers? Why?

Scenario 30BehaviouralAdvanced

Tell me about a time an evaluation result changed a decision the team had already made.

What the interviewer is testing: Whether you can influence decisions with evidence and handle pushback.

Model answer, red flags and follow-up

A strong answer

A strong answer describes a specific case where the eval contradicted expectation, such as a new model that looked better in demos but scored worse on an important segment. I would explain how I checked the result was reliable, presented it clearly with examples and uncertainty, and dealt with disagreement calmly. The outcome might be delaying a launch or choosing a different model. I would also say what I learned about making eval evidence persuasive: concrete examples land better than tables.

Answers that lose you the room

  • Cannot give a specific case
  • Presents it as personal victory
  • Ignores the pushback

Expect this follow-up: How did you make sure the result was not just noise?

0 of 30 attempted