AI Engineer (Prompt & Context Engineering) Interview Questions and Answers
Thirty scenarios on prompt and context engineering, RAG, structured output, model selection and cost for production LLM features. Write your own answer first, by typing or speaking, then open the model answer to compare structure and reasoning.
30 scenariosWhat each question testsRed-flag answersNo sign-up, private
Answer in your own words before opening a model answer, by typing or by pressing Speak your answer. Your text is saved in this browser only and is never uploaded to us. Voice input uses your browser's speech service to turn speech into text; in Chrome that audio is processed by Google.
Voice input is not supported in this browser. Try Chrome, Edge or Safari, or use your keyboard's dictation (Windows + H on Windows, the microphone key on your phone keyboard).
No scenarios match that combination. Choose another topic or level.
Scenario 1Prompt EngineeringPractitioner
How do you structure a production prompt so it stays reliable as requirements change?
What the interviewer is testing: Whether prompts are structured, versioned and eval-gated.
Model answer, red flags and follow-up
A strong answer
I split the prompt into role, task, constraints, context and output format so each part can change without disturbing the others, and I keep it in version control with a changelog. Examples come from real inputs, not invented ones. Before any change I run a fixed eval set of common and edge cases and compare against the previous version. Small targeted edits beat rewrites because I can attribute a regression to one change. When a requirement shifts, I add the new cases to the eval set first, watch them fail, then adjust the prompt until they pass without breaking the older ones.
Answers that lose you the room
Writes one long unstructured prompt
Doesn't version prompts
Changes prompts without running evals
Expect this follow-up: A prompt fix helps one case and breaks another. How do you prevent that?
Scenario 2Context EngineeringPractitioner
What is context engineering and how does it differ from writing prompts?
What the interviewer is testing: Whether you curate what the model sees, not just what you say.
Model answer, red flags and follow-up
A strong answer
Prompting is the instruction; context engineering is deciding everything the model sees when it reads that instruction: retrieved documents, memory, tool results, conversation history, and the order and size of each. I rank, filter and compress context because irrelevant tokens cost money, add latency and pull attention away from what matters. I also decide what goes near the start or end of the window, since position affects recall. In practice most quality gains in production systems come from better context, not cleverer wording, and I measure that with evals rather than intuition.
Answers that lose you the room
Treats it as just better prompt wording
Sends all retrieved text unfiltered
Ignores context order and size
Expect this follow-up: What would you cut first when the context is too long?
Scenario 3RAGPractitioner
Your RAG answers are wrong even though the right document exists. How do you debug it?
What the interviewer is testing: Whether you separate retrieval failures from generation failures.
Model answer, red flags and follow-up
A strong answer
I split the problem in two: retrieval and generation. First I check whether the right chunk was retrieved and where it ranked. If it was missing, I look at chunking, embedding quality, hybrid search, metadata filters and query rewriting. If it was retrieved but ranked low, I add a reranker. If it was in the context and the answer is still wrong, the problem is generation: prompt clarity, context order, or too much noise. I keep retrieval metrics such as recall at k so I can tell which side broke before I touch the prompt.
Answers that lose you the room
Changes the prompt before checking retrieval
Blames the model by default
Has no retrieval metrics
Expect this follow-up: The right chunk is retrieved but ignored. What next?
Scenario 4HallucinationsPractitioner
How do you reduce hallucinations in a customer-facing assistant?
What the interviewer is testing: Whether you reduce hallucination with grounding and measurement.
Model answer, red flags and follow-up
A strong answer
I ground answers in retrieved sources and require citations, so every claim can be checked. I explicitly allow 'I don't know' and reward it in evals, because a model that must always answer will invent. I constrain output format, add a verification step for high-risk claims, and route low-confidence or high-stakes cases to a human. Then I measure: a labelled eval set gives me a hallucination rate I can track across releases. A stricter prompt helps a little, but only measurement tells me whether the assistant is actually safer.
Answers that lose you the room
Says a stricter prompt solves it
Never allows 'I don't know'
Tracks no hallucination rate on an eval set
Expect this follow-up: How do you decide when to route to a human?
Scenario 5Structured OutputFoundation
How do you get reliable structured output from an LLM?
What the interviewer is testing: Whether you enforce schemas instead of asking politely.
Model answer, red flags and follow-up
A strong answer
I stop asking politely and enforce a schema. Native structured output or schema-constrained decoding guarantees valid shape, and I still validate values against the schema in code, because valid JSON can hold wrong content. On failure I retry once with the validation error included so the model can correct itself, then fall back to a safe default or a human. I keep schemas simple, with enums where possible, and log every failure. Recurring failures usually point to an ambiguous field description or an over-complex schema, which I fix at the source.
Answers that lose you the room
Asks the model politely for JSON
Does no schema validation
Retries blindly without the error
Expect this follow-up: The schema is complex and fails often. What do you change?
Scenario 6Model SelectionFoundation
How do you choose a model for a new feature?
What the interviewer is testing: Whether you choose models by your own evals and constraints.
Model answer, red flags and follow-up
A strong answer
I start from requirements: quality bar, latency, cost per request, context length, data privacy and language coverage. I shortlist two or three models and compare them on my own eval set, not public benchmarks, because benchmarks rarely match my task. I pick the cheapest model that clears the quality bar with some margin, and I keep a second model as a fallback behind a thin abstraction so switching is cheap. I also plan to re-run the comparison periodically, since new models change the cost and quality picture every few months.
Answers that lose you the room
Picks the newest or biggest model
Relies on public benchmarks
Has no fallback model
Expect this follow-up: The cheaper model is 3% worse. How do you decide?
Scenario 7Long ContextPractitioner
When would you use a long context window instead of retrieval?
What the interviewer is testing: Whether you weigh long context against retrieval honestly.
Model answer, red flags and follow-up
A strong answer
Long context suits small, self-contained material and one-off analysis, like reading a single contract. Retrieval wins when the corpus is large, changes often, or needs citations, and it is cheaper per query because I send only what is relevant. I test both on the real task, because models can lose facts buried in the middle of a long prompt. Cost matters too: sending a full corpus on every request adds up quickly. Often the best design is hybrid: retrieve broadly, then place a generous but curated set of passages in a long window.
Answers that lose you the room
Says long context replaces retrieval
Ignores cost per query
Doesn't test middle-of-context recall
Expect this follow-up: At what corpus size would you switch to retrieval?
Scenario 8CostPractitioner
Your token bill doubled after a launch. What do you check first?
What the interviewer is testing: Whether you diagnose cost from tokens and traffic, not guesses.
Model answer, red flags and follow-up
A strong answer
I start with the numbers, not guesses. I break the bill down by feature: tokens per request, request volume, retries, and how much context grew. Common causes are ballooning conversation history, an over-retrieved context, retry storms and a new feature routed to an expensive model. Fixes follow the cause: trim prompts, cap history, cache repeated prefixes, route easy tasks to a smaller model and set budgets with alerts per feature. I avoid cutting quality first; I only accept a quality trade-off after the eval set shows it is small.
Answers that lose you the room
Blames traffic alone
Has no per-feature cost visibility
Cuts quality first
Expect this follow-up: Context per request grew 40%. Where would you look?
Scenario 9Few-shot PromptingFoundation
When do few-shot examples help, and when do they hurt?
What the interviewer is testing: Whether you know when examples help or bias output.
Model answer, red flags and follow-up
A strong answer
Few-shot examples help when the format or edge-case behaviour is hard to describe in words, such as tone, labelling rules or output structure. They hurt when examples are unrepresentative, when the model over-copies their surface pattern, or when they consume tokens that context could use better. I choose diverse examples from real data, include an edge case, and vary them so the model learns the rule, not the sample. Then I test with and without them on the eval set. If the gain is small, I drop them and save the tokens.
Answers that lose you the room
Adds many examples by default
Uses unrepresentative examples
Never tests with and without
Expect this follow-up: The model copies your example too literally. What do you do?
Scenario 10Prompt InjectionAdvanced
How do you defend an LLM app against prompt injection?
What the interviewer is testing: Whether you layer defences and treat inputs as untrusted.
Model answer, red flags and follow-up
A strong answer
I assume all retrieved and user-supplied text is untrusted. I separate instructions from data with clear delimiters, but I know that alone is weak. The real protections are architectural: least-privilege tool permissions, confirmation for sensitive actions, no secrets in the prompt, output filtering, and never letting retrieved text trigger actions unchecked. I add an input and output classifier as another layer and log suspicious attempts. Before launch I run adversarial tests, including indirect injection through documents. No single defence works, so I layer them and plan for the case where one fails.
Answers that lose you the room
Relies on 'ignore malicious instructions' in the prompt
Trusts retrieved text as instructions
Gives tools broad permissions
Expect this follow-up: How would you test your defences before launch?
Scenario 11ChunkingPractitioner
How do you decide on a chunking strategy for documents?
What the interviewer is testing: Whether chunking follows structure and is tuned on metrics.
Model answer, red flags and follow-up
A strong answer
I start with the document's own structure: headings, sections, lists and tables, so chunks hold complete ideas. Then I tune size and overlap against retrieval metrics on real user questions, since the best setting depends on the content and the questions. Each chunk carries metadata such as source, section title, date and access level, which supports filtering and citation. Tables and code need special handling so rows and columns stay together. I treat chunking as an experiment: change one parameter, re-run retrieval evals, and keep what measurably improves recall.
Answers that lose you the room
Uses one fixed chunk size everywhere
Ignores headings and tables
Has no retrieval metrics on real questions
Expect this follow-up: A table gets split across chunks. How do you handle it?
Scenario 12EmbeddingsPractitioner
How do you choose an embedding model and vector store?
What the interviewer is testing: Whether you test embeddings on your own data and languages.
Model answer, red flags and follow-up
A strong answer
I compare candidate embedding models on my own retrieval eval, using real queries and documents in the languages my users write in. I weigh quality against dimensions, storage, latency, cost and licensing. For the vector store I look at scale, metadata filtering, hybrid search support, update patterns, and whether it fits our operations, such as managed versus self-hosted. Leaderboards are a starting shortlist, not a decision. I also plan for re-embedding, since changing the model later means reprocessing the whole corpus.
Answers that lose you the room
Picks the top leaderboard model
Doesn't test on their own data
Ignores filtering and hybrid support
Expect this follow-up: Retrieval is good in English but poor in Hindi. What do you check?
Scenario 13Hybrid SearchPractitioner
Why combine keyword and vector search?
What the interviewer is testing: Whether you know where vector search alone fails.
Model answer, red flags and follow-up
A strong answer
Vector search captures meaning but often misses exact tokens such as order IDs, product codes, names and rare terms. Keyword search does the opposite: precise on exact terms, weak on paraphrase. Combining both, then reranking the merged list, gives better recall and precision than either alone. I confirm the benefit on a set of real queries rather than assuming it, and I tune the fusion weights. For structured identifiers I sometimes add a direct lookup path, because no retrieval trick beats an exact match.
Answers that lose you the room
Says vectors alone are enough
Doesn't know about exact-term failures
Adds no reranker
Expect this follow-up: Users search by order ID and get nothing. Why, and what's the fix?
Scenario 14Function CallingPractitioner
How do you make function calling reliable?
What the interviewer is testing: Whether tool calls are validated and evaluated.
Model answer, red flags and follow-up
A strong answer
Reliability starts with the tool definitions: clear names, precise descriptions and strict typed schemas with enums for closed choices. I validate every argument in code before executing, and return an informative error to the model so it can correct itself, with a hard cap on retries. Fewer, well-separated tools are chosen more accurately than many overlapping ones. I evaluate tool-selection and argument accuracy on realistic tasks, including ambiguous ones, and I log every call so failures can be traced and turned into new test cases.
Answers that lose you the room
Writes vague function descriptions
Does no argument validation
Allows unlimited retries
Expect this follow-up: The model calls a function with an invalid argument. What happens next?
Scenario 15Prompt EvaluationPractitioner
How do you evaluate whether a prompt change is actually better?
What the interviewer is testing: Whether you prove improvement statistically and by segment.
Model answer, red flags and follow-up
A strong answer
I run the old and new prompt on the same eval set and compare metrics overall and by segment, because an average can hide a regression in an important group. I read a sample of failures by hand rather than trusting a score alone, and I check whether the difference is bigger than run-to-run noise. If it holds up, I release to a small share of traffic first and watch real-world metrics before going wide. Judging by a handful of examples is how teams ship regressions with confidence.
Answers that lose you the room
Judges by a few examples
Reports only averages
Ships without a canary
Expect this follow-up: The new prompt is better overall but worse for one segment. What do you do?
Scenario 16CachingPractitioner
Where can caching reduce cost and latency in an LLM app?
What the interviewer is testing: Whether you cut cost without leaking or staling answers.
Model answer, red flags and follow-up
A strong answer
Prefix or prompt caching cuts cost and latency when a long system prompt or shared document is reused across requests. Response caching helps for identical queries. Embedding caching avoids recomputing vectors for unchanged text, and semantic caching can serve near-duplicate questions when the answers are safe to reuse. The risks are staleness and privacy: cached answers can go out of date, and a cache shared between users must never leak one person's data to another. I set expiry, scope caches by tenant, and measure hit rate to confirm the saving is real.
Answers that lose you the room
Caches everything
Ignores privacy and staleness
Doesn't know prefix caching
Expect this follow-up: What would you never put in a semantic cache?
Scenario 17Fine-tuningAdvanced
When is fine-tuning worth it over prompting?
What the interviewer is testing: Whether you justify fine-tuning with evidence of a gap.
Model answer, red flags and follow-up
A strong answer
Fine-tuning earns its place when prompting has a proven gap: consistent style or format at scale, lower latency and cost from a smaller model, or behaviour prompts cannot reach. It needs quality labelled data and an eval set to prove it worked. I always start with prompting, retrieval and evals, because they are faster to iterate and cheaper to undo. Fine-tuning also adds a maintenance burden, since I must retrain when the base model or the data changes. I fine-tune to close a measured gap, not to feel sophisticated.
Answers that lose you the room
Fine-tunes before trying prompting
Has no quality labelled data
Cannot state the gap fine-tuning would close
Expect this follow-up: What evidence would convince you fine-tuning is needed?
Scenario 18MultilingualPractitioner
How do you handle Indian languages in an LLM feature?
What the interviewer is testing: Whether you test quality per language, including code-mixing.
Model answer, red flags and follow-up
A strong answer
I test quality per language on real user text, not translated English, because performance often drops for lower-resource languages. Users mix languages and write Hindi in Roman script, so I test code-mixing and transliteration explicitly. I tune prompts and retrieval per language, check that the embedding model actually handles them, and involve native speakers in evaluation. I also watch token cost, since some scripts tokenise into many more tokens. If a language falls below the quality bar, I limit the feature or add human review rather than ship something unreliable.
Answers that lose you the room
Assumes English quality carries over
Ignores code-mixing and transliteration
Has no native-speaker evaluation
Expect this follow-up: Users type Hindi in Roman script. What breaks and how do you handle it?
Scenario 19Streaming UXFoundation
How does streaming change the user experience of an LLM feature?
What the interviewer is testing: Whether you understand streaming's effect on perceived latency.
Model answer, red flags and follow-up
A strong answer
Streaming shows the first words within a second or so, which sharply improves perceived latency even though total time is similar. The cost is complexity. Partial output can be wrong or later retracted, structured output cannot be parsed until it is complete, and safety checks may need the whole response. I handle cancellation so abandoned requests stop spending tokens, and I design the UI to handle a response that changes or stops midway. Where a check needs the full text, I may stream a draft and validate before enabling actions.
Answers that lose you the room
Says streaming makes it faster overall
Ignores validation of partial output
Doesn't handle cancellation
Expect this follow-up: Safety checks need the full response. How do you stream anyway?
Scenario 20DeterminismFoundation
How do you get more consistent outputs from an LLM?
What the interviewer is testing: Whether you accept variability and design validation around it.
Model answer, red flags and follow-up
A strong answer
I lower temperature, tighten the prompt, use structured output and fixed examples, and use a seed where the provider supports it. I also accept that full determinism is not guaranteed: batching, hardware and model updates can shift outputs even at temperature zero. So instead of assuming identical outputs, I design tolerance into the feature: validation, tests that check properties rather than exact strings, and monitoring for drift. If exact repeatability matters, such as for audit, I store the output rather than trying to regenerate it.
Answers that lose you the room
Says temperature zero makes it deterministic
Adds no validation of outputs
Ignores structured output
Expect this follow-up: The output still varies at temperature zero. Why?
Scenario 21PII HandlingAdvanced
How do you handle personal data in prompts?
What the interviewer is testing: Whether you minimise and control personal data in prompts.
Model answer, red flags and follow-up
A strong answer
I minimise what I send: only the fields the task needs. Identifiers are redacted or tokenised before the call and restored afterwards where required. I choose providers whose data terms match our obligations, such as no training on our data and suitable retention, and I restrict what is logged and for how long. I map the data flows end to end and review them with security and legal. I also make sure evals and debugging tools do not become a side door that copies personal data into less protected places.
Answers that lose you the room
Sends raw personal data to any provider
Logs full prompts indefinitely
Has no data-flow map
Expect this follow-up: Which fields would you redact before the model call?
Scenario 22Conversation MemoryPractitioner
How do you manage long conversations within context limits?
What the interviewer is testing: Whether summaries preserve the facts that matter.
Model answer, red flags and follow-up
A strong answer
I keep recent turns verbatim, summarise older ones, and store durable facts such as the user's name, preferences and open tasks in a separate memory that I retrieve when relevant. The risk is that summaries quietly drop something important, so I test that key facts survive summarisation across many turns. I also cap history to control cost and latency. For sensitive facts I make memory visible and correctable to the user. The goal is that the conversation feels continuous without paying to resend everything each time.
Answers that lose you the room
Sends the entire history every turn
Loses key facts when summarising
Never tests summarisation
Expect this follow-up: The user's name is forgotten after summarisation. How do you prevent it?
Scenario 23Ambiguous RequirementsFoundation
Product gives you a vague requirement for an AI feature. What do you do?
What the interviewer is testing: Whether you turn vague asks into testable examples.
Model answer, red flags and follow-up
A strong answer
I turn the vague ask into something testable. I talk to the requester about the user problem and what success would look like, then build a quick prototype and run it on real examples. Showing actual outputs to stakeholders is far more effective than debating a spec, because people recognise good and bad results when they see them. I collect those reactions into an eval set, which becomes the definition of done. That way the requirement becomes concrete through iteration instead of waiting for a perfect document.
Answers that lose you the room
Starts building without clarifying
Waits for a perfect spec
Doesn't show concrete outputs
Expect this follow-up: The stakeholder still can't say what 'good' means. What do you do?
Scenario 24DebuggingPractitioner
Outputs are inconsistent across users. How do you investigate?
What the interviewer is testing: Whether you debug by comparing traces and changing one variable.
Model answer, red flags and follow-up
A strong answer
I compare traces across affected and unaffected cases: the input, the retrieved context, the prompt version, the model and its settings, and the tool results. I look for what differs, segment failures by user type, language or input length, and form a hypothesis. Then I change one variable at a time and re-run the evals so I know what fixed it. Changing several things at once makes the result unexplainable. Good tracing from the start makes this fast, which is why I insist on logging the full chain of context per request.
Answers that lose you the room
Changes several variables at once
Doesn't compare traces
Has no hypothesis
Expect this follow-up: What is the first thing you record in a trace so you can compare users?
Scenario 25Model DeprecationPractitioner
A provider deprecates the model you depend on. How do you prepare?
What the interviewer is testing: Whether you plan migration before the deadline.
Model answer, red flags and follow-up
A strong answer
I treat deprecation as a scheduled event, not a surprise. I track provider notices, keep a model abstraction layer, and maintain an eval suite so I can test a replacement quickly on my own tasks. As soon as a successor is available, I compare it, adapt prompts where behaviour differs, and roll it out gradually with canary traffic while watching quality and cost. I keep the old model available until the new one has proven itself, and I never move all traffic at once. Starting early turns a deadline into a routine migration.
Answers that lose you the room
Notices only at the deadline
Has no abstraction or eval suite
Switches all traffic at once
Expect this follow-up: The replacement model behaves differently on your prompts. What do you do?
Scenario 26LatencyPractitioner
Your AI feature has good answers but a p95 latency of twelve seconds. How do you bring it down?
What the interviewer is testing: Whether you diagnose latency by stage and know the levers beyond a faster model.
Model answer, red flags and follow-up
A strong answer
I measure where the time goes: retrieval, prompt size, time to first token and output length. Output tokens usually dominate, so I shorten responses, stream them and cap length. I trim context, use prefix caching, and run independent calls such as retrieval and classification in parallel. Simple requests go to a smaller, faster model and only hard ones to the larger one. I set a latency budget per stage and watch p95, not the average, because users feel the slow tail.
Answers that lose you the room
Only proposes a faster model
Measures average latency, not p95
Has no per-stage timing
Expect this follow-up: Streaming is on and users still complain. What else could be slow?
Scenario 27RetrievalPractitioner
When is a reranker worth the extra latency and cost?
What the interviewer is testing: Whether you add components because measurement shows a gain.
Model answer, red flags and follow-up
A strong answer
A reranker helps when first-stage retrieval finds the right passage but ranks it too low, which is common with large corpora and ambiguous queries. I measure it: compare recall and answer quality at the top few results with and without reranking on real queries. If the right chunk is already first most of the time, the extra hop is waste. I keep the candidate list small to limit latency, and I consider a lighter reranker if speed matters. It earns its place only if the gain is visible in the eval.
Answers that lose you the room
Adds a reranker without measuring
Ignores the added latency
Reranks a huge candidate list
Expect this follow-up: Your reranker adds 400 ms. Product says it is too slow. What are your options?
Scenario 28Evaluation DataAdvanced
You are building a RAG assistant and have no labelled data. How do you create an eval set?
What the interviewer is testing: Whether you can bootstrap evaluation honestly from real material.
Model answer, red flags and follow-up
A strong answer
I start with real questions from support tickets, search logs or subject experts, since they reflect real usage. Where those are thin, I generate synthetic questions from the documents, then have an expert review and correct them, because unreviewed synthetic data flatters the system. Each item gets an expected answer or a rubric and the source passage. I include hard cases: ambiguous, multi-document and unanswerable questions. I keep a held-out part untouched for final checks and grow the set from production failures.
Answers that lose you the room
Uses only unreviewed synthetic questions
Tunes on the same set they report on
Includes no unanswerable questions
Expect this follow-up: How do you stop the synthetic questions from being too easy?
Scenario 29Model SelectionAdvanced
When would you use a reasoning model instead of a standard model, and what does it cost you?
What the interviewer is testing: Whether you match model type to task difficulty and account for the trade-offs.
Model answer, red flags and follow-up
A strong answer
Reasoning models help on multi-step problems such as planning, complex analysis or code with many constraints, where extra thinking improves correctness. They cost more, respond more slowly and consume extra tokens for the hidden reasoning. For simple extraction, classification or lookup they add cost with little benefit. I test both on my eval set by task type and route accordingly, sending only hard requests to the reasoning model. I also check that prompts written for standard models do not over-constrain the reasoning one.
Answers that lose you the room
Uses the reasoning model for everything
Ignores latency and token cost
Never compares against a standard model on their eval
Expect this follow-up: How would you decide which requests get routed to the reasoning model?
Scenario 30BehaviouralAdvanced
Tell me about an AI feature you built that did not work in production. What did you do?
What the interviewer is testing: Whether you own failures, learn from them and change your process.
Model answer, red flags and follow-up
A strong answer
A strong answer names a specific feature and what went wrong in concrete terms, such as answers that looked fine in demos but failed on real user phrasing. I would explain how I detected it, through user feedback, monitoring or an eval gap, and what I did first to limit harm. Then the root cause: usually an eval set that did not resemble real traffic. The lasting change is process: adding production samples to the eval set and gating releases on it. I would say plainly what I got wrong.
Answers that lose you the room
Blames the model or the users
Describes a success dressed up as a failure
Cannot say what changed afterwards
Expect this follow-up: What would you have needed to see before launch to catch it?
0 of 30 attempted
Keep going
Practise the neighbouring roles, or see every question in one place.