Home › AI Interview Prep › LLMOps Engineer
In Demand

LLMOps Engineer Interview Questions and Answers

Thirty scenarios on deployment, monitoring, cost and latency control, reliability, self-hosting and incident response. Write your own answer first, by typing or speaking, then open the model answer to compare structure and reasoning.

30 scenariosWhat each question testsRed-flag answersNo sign-up, private
Written by Ayush Bisht · Reviewed by Sanjay Saini
Last updated 2026-09-30
LLMOps Engineer interview questions

Answer in your own words before opening a model answer, by typing or by pressing Speak your answer. Your text is saved in this browser only and is never uploaded to us. Voice input uses your browser's speech service to turn speech into text; in Chrome that audio is processed by Google.

Scenario 1DeploymentPractitioner

How do you deploy and version LLM applications safely?

What the interviewer is testing: Whether you version prompts, models and configs together.

Model answer, red flags and follow-up

A strong answer

I version prompts, model identifiers and configuration together as one release unit, stored in source control, and infrastructure is defined as code. Changes move through environments with eval gates, then reach production through a canary or shadow release with a small share of traffic. Rollback must be instant and tested, ideally a config switch, not a rebuild. Every release records which prompt, model and settings were live, so any behaviour can be traced to a version. Deploying straight to production is how silent quality regressions happen.

Answers that lose you the room

  • Versions code but not prompts and models
  • Deploys straight to production
  • Has no rollback path

Expect this follow-up: How do you deploy a prompt change without redeploying the app?

Scenario 2MonitoringPractitioner

What do you monitor in production LLM systems?

What the interviewer is testing: Whether you monitor quality and cost, not just uptime.

Model answer, red flags and follow-up

A strong answer

I monitor the operational basics: latency percentiles, error rates, throughput, token usage and cost. On top of that I watch quality signals: a sampled evaluation of outputs, guardrail hits, refusal rates, user feedback and drift in the types of inputs arriving. Each has agreed thresholds and alerts that reach the right on-call person. Dashboards show trends by feature and model version. The gap in many teams is quality monitoring, since a service can be perfectly healthy technically while producing poor answers.

Answers that lose you the room

  • Monitors uptime only
  • Doesn't sample quality
  • Sets alerts with no thresholds

Expect this follow-up: Everything is green but users complain. What is missing?

Scenario 3Cost ControlPractitioner

How do you control LLM costs at scale?

What the interviewer is testing: Whether you control cost without cutting quality.

Model answer, red flags and follow-up

A strong answer

I attack cost from several angles. Caching removes repeated work, prompt trimming and history caps reduce tokens, routing sends easy requests to smaller models, and batching handles non-urgent work at lower prices. I set budgets and rate limits per team and feature, and track cost per successful task, since a cheap failing request is waste. Regular reviews of the top spenders find surprises early. Every optimisation is checked against the eval set so savings do not silently cost quality.

Answers that lose you the room

  • Cuts quality to save cost
  • Tracks no cost per task
  • Sets no per-team budgets

Expect this follow-up: Which cost lever do you pull first, and why?

Scenario 4LatencyPractitioner

How do you reduce latency for a chat application?

What the interviewer is testing: Whether you reduce latency with measured techniques.

Model answer, red flags and follow-up

A strong answer

I measure where time goes first, then act on the biggest part. Streaming makes responses feel faster, and prompt caching and shorter context cut processing time. I run independent steps, such as retrieval and classification, in parallel, use smaller or faster models for simple requests, and deploy close to users. Output length often dominates, so I constrain it. I set a latency budget for each step and track p95 and p99, because the slow tail is what users remember.

Answers that lose you the room

  • Only switches to a smaller model
  • Ignores streaming and caching
  • Measures average latency

Expect this follow-up: p95 is bad but the average is fine. What do you check?

Scenario 5ReliabilityAdvanced

How do you handle provider outages and rate limits?

What the interviewer is testing: Whether you design for provider outages and limits.

Model answer, red flags and follow-up

A strong answer

I design for failure from the start. Every call has timeouts and retries with exponential backoff and jitter, and circuit breakers stop hammering a failing provider. Behind a model abstraction I keep a fallback provider or model, and queue non-urgent work. When everything fails, the system degrades gracefully, for example by showing a cached answer or a non-AI path. I respect rate limits with client-side throttling, and I rehearse failover regularly, because untested failover tends not to work when needed.

Answers that lose you the room

  • Uses a single provider with no fallback
  • Retries with no backoff
  • Never rehearses failover

Expect this follow-up: The fallback model behaves differently. How do you manage that?

Scenario 6Self-hostingPractitioner

When would you self-host an open model instead of using an API?

What the interviewer is testing: Whether you weigh self-hosting on total cost.

Model answer, red flags and follow-up

A strong answer

I consider self-hosting when data must stay within our control, when volume is high and steady enough to keep GPUs busy, when I need customisation such as fine-tuning, or when latency needs cannot be met by an API. I compare the full cost, including GPUs, engineering time, on-call, upgrades and idle capacity, against API pricing. At low or spiky volume, APIs are usually cheaper and simpler. I also weigh model quality, since the best hosted models may outperform open ones for my task.

Answers that lose you the room

  • Assumes self-hosting is always cheaper
  • Ignores GPU and on-call costs
  • Has no volume estimate

Expect this follow-up: At what volume does self-hosting break even?

Scenario 7Quality DriftAdvanced

Quality dropped and no code changed. What do you investigate?

What the interviewer is testing: Whether you diagnose drift when no code changed.

Model answer, red flags and follow-up

A strong answer

Since no code changed, I look outside our repository. The provider may have updated the model behind an alias, the mix of user inputs may have shifted, the retrieval index may be stale or a data pipeline may have broken. Seasonal or event-driven effects also matter. I compare current results with the eval set and traces from before the drop, segment by input type and time, and check provider change logs and index freshness. To prevent recurrence, I pin model versions and monitor quality continuously.

Answers that lose you the room

  • Says nothing changed so it can't be us
  • Ignores provider updates
  • Has no eval set to compare against

Expect this follow-up: How would you confirm the provider changed the model?

Scenario 8IncidentPractitioner

A prompt change caused a spike of bad outputs at 2am. How do you respond?

What the interviewer is testing: Whether you roll back first and learn afterwards.

Model answer, red flags and follow-up

A strong answer

First I stop the harm: roll back to the last known good version, then confirm recovery in the metrics. I communicate status to stakeholders and keep a timeline. Once stable, I capture the failing examples, find the cause and add those examples to the eval suite so the release gate would catch them next time. I tighten the release process, for example with canary stages or required eval passes for prompt changes. The follow-up is a blameless review focused on how the system allowed the change through.

Answers that lose you the room

  • Debugs live before rolling back
  • Skips the blameless review
  • Doesn't turn failures into tests

Expect this follow-up: What release gate would have caught this?

Scenario 9Prompt RegistryPractitioner

Why use a prompt registry, and what should it support?

What the interviewer is testing: Whether prompts get ownership, versions and audit.

Model answer, red flags and follow-up

A strong answer

A prompt registry treats prompts as managed assets. It gives each prompt versions, owners, environments, diffs and linked eval results, and supports instant rollback. Where appropriate it decouples prompt changes from application releases, so a fix does not require a full deployment, while access controls, review and audit history keep this safe. It also helps non-engineers such as product or domain experts propose changes through a controlled process. Without it, prompts get edited in place and nobody knows what was running when.

Answers that lose you the room

  • Keeps prompts scattered across code
  • Has no owners or audit history
  • Has no environment separation

Expect this follow-up: Who should be allowed to change a production prompt?

Scenario 10CI/CDPractitioner

What does CI/CD look like for an LLM application?

What the interviewer is testing: Whether evals gate releases in your pipeline.

Model answer, red flags and follow-up

A strong answer

It looks like conventional CI/CD with an extra layer. Code goes through linting and unit tests, then the eval suite runs on prompt, model and retrieval changes, with safety checks. Deployment is staged, with canary monitoring of quality and cost, and automatic rollback if thresholds are breached. The evals play the role of tests, and they gate each release. Because LLM outputs vary, I use statistical thresholds and confidence intervals instead of exact matches, and I keep the pipeline fast enough that developers do not bypass it.

Answers that lose you the room

  • Skips evals in the pipeline
  • Tests only the code
  • Has no automatic rollback

Expect this follow-up: What fails the build in your pipeline?

Scenario 11Model GatewayPractitioner

What does a model gateway provide?

What the interviewer is testing: Whether you centralise access, routing and policy.

Model answer, red flags and follow-up

A strong answer

A model gateway gives applications a single interface to multiple providers and models. It centralises authentication, rate limits, routing, retries, caching, logging, cost tracking and policy enforcement, so those features are not re-implemented in every service. It reduces lock-in because switching providers becomes a configuration change, and it gives security and finance a single point of visibility. The trade-off is another critical component to keep reliable and low-latency, so I treat it as production infrastructure with its own SLOs.

Answers that lose you the room

  • Calls providers directly from each app
  • Has no central logging
  • Enforces no policy

Expect this follow-up: What would you put in the gateway first?

Scenario 12TracingPractitioner

What should LLM tracing capture?

What the interviewer is testing: Whether traces make production bugs debuggable.

Model answer, red flags and follow-up

A strong answer

Each trace should capture the prompt and its version, the retrieved context, the model and parameters, tool calls and results, latency, token counts and cost, and any user feedback, all linked by a trace ID. This makes debugging possible, since I can reconstruct exactly what the model saw. It also lets me build evaluation sets from real production data and analyse cost by feature. I handle privacy through redaction and retention limits, and sample where volume is very high.

Answers that lose you the room

  • Logs only final responses
  • Uses no trace IDs
  • Stores raw personal data

Expect this follow-up: How do you debug one bad conversation from production?

Scenario 13Semantic CachingAdvanced

When is semantic caching a good idea?

What the interviewer is testing: Whether you cache safely and avoid wrong shared answers.

Model answer, red flags and follow-up

A strong answer

Semantic caching suits high-volume, repetitive questions where near-duplicates should get the same answer, such as FAQ-style support. I set the similarity threshold carefully and test it, because a threshold that is too loose returns wrong answers confidently. I avoid caching personalised, time-sensitive or sensitive responses, scope the cache by tenant and permissions, and expire entries. I measure the hit rate and the quality of cached answers to confirm the saving is worth the risk.

Answers that lose you the room

  • Caches every response
  • Sets loose similarity thresholds
  • Caches personalised answers

Expect this follow-up: A cached answer is wrong for a different user. How did that happen?

Scenario 14GPU AutoscalingAdvanced

How do you autoscale GPU inference?

What the interviewer is testing: Whether you scale GPUs on the right signals.

Model answer, red flags and follow-up

A strong answer

I scale on signals that reflect real load, such as queue depth, request latency and GPU utilisation, not just CPU. GPUs take minutes to start and models take time to load, so I keep some warm capacity and scale ahead of predictable peaks. Batching and mixed instance types improve utilisation, and I set cost ceilings to avoid runaway spend. I load-test with realistic prompt and output lengths to find true limits, since token counts vary widely and throughput depends on them.

Answers that lose you the room

  • Scales on CPU only
  • Ignores cold starts
  • Sets no cost ceiling

Expect this follow-up: Traffic spikes at 9am daily. How do you handle cold starts?

Scenario 15Inference OptimisationAdvanced

How can you make self-hosted inference cheaper and faster?

What the interviewer is testing: Whether you optimise inference while validating quality.

Model answer, red flags and follow-up

A strong answer

The main levers are quantisation to reduce memory and speed up inference, continuous batching to raise throughput, KV-cache reuse for shared prefixes, and an optimised serving engine such as vLLM. I right-size hardware, and consider smaller or distilled models where quality allows. Speculative decoding can cut latency for some workloads. Each change is validated against the eval set, because optimisations can degrade quality in subtle ways. I measure cost per thousand tokens and latency, not just raw speed.

Answers that lose you the room

  • Buys bigger GPUs first
  • Doesn't validate quality after quantisation
  • Ignores batching

Expect this follow-up: How do you check quality didn't drop after quantising?

Scenario 16Multi-tenancyAdvanced

How do you isolate tenants in a shared LLM platform?

What the interviewer is testing: Whether tenants are isolated and that is tested.

Model answer, red flags and follow-up

A strong answer

I isolate tenants at every layer: separate data stores or indexes, or strict filtering enforced in the retrieval layer, per-tenant keys, quotas and rate limits, and separate logs. Caches must not be shared across tenants, and prompts must never include another tenant's data. I test isolation explicitly with cross-tenant probes, including prompt-injection attempts to extract other tenants' data. A noisy-neighbour problem is also possible, so fair-use limits protect performance for everyone.

Answers that lose you the room

  • Shares indexes and caches across tenants
  • Sets no per-tenant quotas
  • Never tests isolation

Expect this follow-up: How do you prove one tenant can't see another's data?

Scenario 17SecretsFoundation

How do you manage API keys and secrets for LLM services?

What the interviewer is testing: Whether you handle keys and secrets safely.

Model answer, red flags and follow-up

A strong answer

API keys live in a secrets manager, never in code, prompts or logs. They are scoped by service and environment, rotated regularly, and issued with the minimum permissions. I monitor usage for anomalies such as sudden spikes, which can indicate a leak, and set spending limits on keys. Developers get separate keys from production. Where possible I use short-lived credentials or workload identity, and have a tested procedure for revoking and replacing a compromised key quickly.

Answers that lose you the room

  • Keeps keys in code or config files
  • Never rotates them
  • Puts secrets in prompts or logs

Expect this follow-up: A key leaks in a log. What do you do?

Scenario 18Logging & PrivacyPractitioner

How do you balance detailed logging with privacy?

What the interviewer is testing: Whether you balance debuggability with privacy.

Model answer, red flags and follow-up

A strong answer

I log what is needed for debugging and evaluation, and no more. Sensitive fields are redacted or hashed at the point of logging, retention is short, and access is restricted and audited. A separate, consented or sanitised sample set can be kept for review and evals. Privacy requirements from legal and the data protection team decide what may be stored and for how long. The trade-off is real: less logging makes debugging harder, so I invest in good redaction rather than turning logs off.

Answers that lose you the room

  • Logs everything forever
  • Logs nothing to protect privacy
  • Sets no access restrictions

Expect this follow-up: Debugging needs full prompts. How do you allow that safely?

Scenario 19Index RefreshPractitioner

How do you keep a RAG index fresh?

What the interviewer is testing: Whether you keep retrieval indexes fresh and clean.

Model answer, red flags and follow-up

A strong answer

I trigger incremental ingestion when sources change, and schedule periodic full checks for anything missed. Deleted or updated documents must be removed or re-embedded so stale answers disappear. I version indexes so I can roll back a bad refresh, and I monitor freshness, ingestion failures and index size. After each refresh I run a set of retrieval tests to confirm quality did not drop. Stale data is a quiet failure, so I alert on lag between the source and the index.

Answers that lose you the room

  • Re-indexes everything manually
  • Ignores deleted documents
  • Doesn't monitor freshness

Expect this follow-up: A document is deleted at the source. How does it leave the index?

Scenario 20Feature FlagsPractitioner

How do feature flags help with LLM releases?

What the interviewer is testing: Whether you decouple deployment from release.

Model answer, red flags and follow-up

A strong answer

Feature flags let me separate deployment from release. I can expose a new prompt or model to a small share of traffic or specific segments, compare metrics with the control, and switch it off instantly if something goes wrong without redeploying. They also support experiments and gradual rollout by tenant. I keep flags well managed, with owners and expiry dates, so they do not accumulate into untested combinations of settings.

Answers that lose you the room

  • Releases to everyone at once
  • Has no way to switch off instantly
  • Ties release to deployment

Expect this follow-up: How would you run a 5% rollout of a new model?

Scenario 21SLOsPractitioner

How do you define SLOs for an LLM service?

What the interviewer is testing: Whether SLOs include quality, not only availability.

Model answer, red flags and follow-up

A strong answer

I define SLOs that reflect the user's experience: availability, latency percentiles, error rate and quality indicators from sampled evals. Each has an objective and an error budget that guides how much risk we can take with releases. Because quality is harder to measure than uptime, I make the quality SLO explicit and track it. I review SLOs after incidents and adjust them when they do not match what users care about. Alerts are based on budget burn, not on every blip.

Answers that lose you the room

  • Sets availability targets only
  • Adds no quality indicators
  • Has no error budgets

Expect this follow-up: How do you turn a quality dip into an SLO breach?

Scenario 22Cost AttributionPractitioner

How do you attribute LLM costs to teams and features?

What the interviewer is testing: Whether costs are attributable by team and feature.

Model answer, red flags and follow-up

A strong answer

I tag every request with team, feature and environment through the gateway, then aggregate token and infrastructure costs in dashboards. Each team gets a budget with alerts, and anomalies are reviewed weekly. Shared costs, such as a common index or platform overhead, are allocated using a documented rule. Making cost visible changes behaviour: teams notice wasteful prompts when they can see the bill. I also report cost per successful outcome, so cheap-but-failing features do not look efficient.

Answers that lose you the room

  • Tracks total spend only
  • Uses no tags by team or feature
  • Reviews costs annually

Expect this follow-up: One feature's cost spikes. How fast can you find it?

Scenario 23Model RegistryAdvanced

How do you manage fine-tuned models across their lifecycle?

What the interviewer is testing: Whether models trace back to data, code and evals.

Model answer, red flags and follow-up

A strong answer

I use a model registry that records each model's lineage: training data version, code, hyperparameters and evaluation results. Models move through stages, such as staging and production, with approvals and automatic checks, and I can reproduce any model from its record. Rollback to a previous version is straightforward. Every deployed model traces to its evidence, which supports audit and debugging. I also track dependencies on base models, so a base-model deprecation triggers a plan.

Answers that lose you the room

  • Stores models with no lineage
  • Cannot reproduce training
  • Requires no approvals

Expect this follow-up: An auditor asks how model v3 was produced. What do you show?

Scenario 24RunbooksFoundation

What belongs in an on-call runbook for an LLM service?

What the interviewer is testing: Whether on-call runbooks are concrete and maintained.

Model answer, red flags and follow-up

A strong answer

It lists the symptoms and the dashboards to check, how to confirm provider status, and step-by-step actions for rollback and failover. It covers known failure modes and their fixes, escalation contacts, and templates for communicating with users and stakeholders. It is written to be usable at 3am by someone who did not build the system. I update it after every incident and test it in game days, because an out-of-date runbook is worse than none.

Answers that lose you the room

  • Writes runbooks with generic steps
  • Leaves out rollback and failover steps
  • Never updates them after incidents

Expect this follow-up: What is on the first page of your runbook?

Scenario 25StagingPractitioner

How do you make staging environments realistic for LLM apps?

What the interviewer is testing: Whether staging exposes real-world problems.

Model answer, red flags and follow-up

A strong answer

I make staging as close to production as practical: the same configuration, the same model versions, representative traffic patterns and real provider rate limits. Data is sanitised or synthetic but realistic in shape and difficulty. I run the eval suite there and load-test at expected volumes. Mocked model responses hide exactly the problems I need to find, such as latency, rate limiting and unexpected outputs, so I use real calls with cost controls. A staging environment that behaves differently gives false confidence.

Answers that lose you the room

  • Uses mocked responses only
  • Uses unrealistic traffic and data
  • Skips real provider limits

Expect this follow-up: What issue only shows up in realistic staging?

Scenario 26ObservabilityAdvanced

How would you detect that an LLM application has started producing lower-quality answers before customers complain?

What the interviewer is testing: Whether you monitor quality proactively, not only uptime.

Model answer, red flags and follow-up

A strong answer

I combine several signals. An LLM judge or rule-based checks score a sample of live traffic against rubrics, and a small share is reviewed by humans to keep the judge honest. I watch proxy signals such as edit rates, retries, abandonment and thumbs-down, and I track drift in input topics, output length and refusal rates. Alerts fire on shifts beyond normal variation. Detection is only useful if flagged cases feed the eval set and an owner investigates.

Answers that lose you the room

  • Monitors only uptime and latency
  • Waits for customer complaints
  • Uses a judge that was never validated

Expect this follow-up: The judge score is stable but complaints rise. What do you check?

Scenario 27Rate LimitsPractitioner

Your traffic doubles during a marketing campaign and you hit provider rate limits. What do you do in the moment and afterwards?

What the interviewer is testing: Whether you can manage capacity and degrade gracefully under load.

Model answer, red flags and follow-up

A strong answer

In the moment I shed or queue low-priority work, apply client-side throttling, turn on caching, route eligible traffic to a secondary provider or smaller model, and communicate status. Afterwards I ask what warning we missed: capacity was not requested in advance, there was no load test and there was no priority tiering. I request higher limits, set quotas per feature, add autoscaling queues and rehearse the scenario. Marketing should tell engineering about campaigns in advance.

Answers that lose you the room

  • Only asks the provider for more quota
  • Has no priority between traffic types
  • Never load-tests the peak

Expect this follow-up: Which requests would you drop first, and how would you decide?

Scenario 28DeploymentPractitioner

How do you roll out a new model version when you cannot fully predict how it will behave?

What the interviewer is testing: Whether you use staged exposure and comparison, not blind switches.

Model answer, red flags and follow-up

A strong answer

I run the eval suite first, including regression and safety tests, then shadow the new model on live traffic without showing its outputs so I can compare results, cost and latency. Next comes a canary at a few percent, watching quality, guardrail hits and user signals, and then a gradual ramp with flags for instant rollback. Prompts may need adjusting for behavioural differences. I keep the old model available until the new one has proven itself over a full traffic cycle.

Answers that lose you the room

  • Switches all traffic at once
  • Relies only on public benchmarks
  • Removes the old model immediately

Expect this follow-up: The canary looks fine on average but worse for one customer. What now?

Scenario 29Data GovernancePractitioner

How do you handle user data that ends up in prompts, logs and traces in your LLM platform?

What the interviewer is testing: Whether you treat observability data as sensitive data.

Model answer, red flags and follow-up

A strong answer

I classify what may appear in prompts and traces, then apply redaction at the point of capture, encryption, short retention and role-based access. Debug access is audited. Data that must be kept for evaluation is sampled, minimised and, where required, consented. I honour deletion requests across logs, caches and indexes. I make sure third-party observability tools are covered by the same data terms. Traces are enormously useful, but they are also a copy of user data that needs protecting.

Answers that lose you the room

  • Logs full prompts forever
  • Gives all engineers access to traces
  • Forgets caches when handling deletion

Expect this follow-up: A user requests deletion. Where might their data still exist in your platform?

Scenario 30BehaviouralAdvanced

Tell me about an outage or incident you handled on an ML or LLM system. What did you change afterwards?

What the interviewer is testing: Whether you respond calmly and turn incidents into lasting improvements.

Model answer, red flags and follow-up

A strong answer

A strong answer gives a concrete incident with its impact and timeline: how it was detected, what I did to stabilise it and how I communicated. Then the root cause, often a combination of factors instead of one mistake, and the fixes: better alerts, release gates, runbooks or capacity planning. I would describe the blameless review and one specific change that measurably reduced the risk of a repeat. I would be honest about what I would do differently.

Answers that lose you the room

  • Blames one person
  • Cannot describe any lasting change
  • Focuses only on the heroics

Expect this follow-up: How did you know your fix worked?

0 of 30 attempted