AI Engineering Interview Questions
Scenario practice for the fastest-growing AI engineering roles. Every question shows what the interviewer is actually testing, a model answer, the answers that lose you the room, and the follow-up you should expect.
Choose your role
Each simulator drills the situations interviewers actually test, not definitions.

AI Engineer (Prompt & Context Engineering)
Prompt and context engineering, RAG, structured output, model selection and cost for production LLM features.
Start practice (30 scenarios) →
Agentic AI Engineer
Agent design, tools and MCP, guardrails, memory, multi-agent systems and observability.
Start practice (30 scenarios) →
AI Evaluation Engineer
Eval design, LLM-as-judge, regression testing, human review and safety evaluation.
Start practice (30 scenarios) →
Forward Deployed Engineer
Customer scoping, messy data, trust, security reviews and taking pilots to production.
Start practice (30 scenarios) →
Responsible AI / AI Governance Engineer
Risk assessment, regulation, bias testing, transparency, privacy and incident response.
Start practice (30 scenarios) →
LLMOps Engineer
Deployment, monitoring, cost and latency control, reliability, self-hosting and incident response.
Start practice (30 scenarios) →
AI Solutions Architect
Enterprise GenAI architecture, build vs buy, security, multi-model design, scale and stakeholder trade-offs.
Start practice (30 scenarios) →How it works
Three steps from cold to interview-ready.
All 210 interview questions
Every scenario across the 7 roles in one list. Choose a question to jump straight to it, then write your own answer before you open the model answer.
Showing all 210 questions
No questions match. Try a different word, role or level.
AI Engineer (Prompt & Context Engineering) 30 questions
- How do you structure a production prompt so it stays reliable as requirements change?
- What is context engineering and how does it differ from writing prompts?
- Your RAG answers are wrong even though the right document exists. How do you debug it?
- How do you reduce hallucinations in a customer-facing assistant?
- How do you get reliable structured output from an LLM?
- How do you choose a model for a new feature?
- When would you use a long context window instead of retrieval?
- Your token bill doubled after a launch. What do you check first?
- When do few-shot examples help, and when do they hurt?
- How do you defend an LLM app against prompt injection?
- How do you decide on a chunking strategy for documents?
- How do you choose an embedding model and vector store?
- Why combine keyword and vector search?
- How do you make function calling reliable?
- How do you evaluate whether a prompt change is actually better?
- Where can caching reduce cost and latency in an LLM app?
- When is fine-tuning worth it over prompting?
- How do you handle Indian languages in an LLM feature?
- How does streaming change the user experience of an LLM feature?
- How do you get more consistent outputs from an LLM?
- How do you handle personal data in prompts?
- How do you manage long conversations within context limits?
- Product gives you a vague requirement for an AI feature. What do you do?
- Outputs are inconsistent across users. How do you investigate?
- A provider deprecates the model you depend on. How do you prepare?
- Your AI feature has good answers but a p95 latency of twelve seconds. How do you bring it down?
- When is a reranker worth the extra latency and cost?
- You are building a RAG assistant and have no labelled data. How do you create an eval set?
- When would you use a reasoning model instead of a standard model, and what does it cost you?
- Tell me about an AI feature you built that did not work in production. What did you do?
Practise all AI Engineer (Prompt & Context Engineering) scenarios with model answers
Agentic AI Engineer 30 questions
- When should you use an agent instead of a fixed workflow?
- How do you design tools so an agent uses them correctly?
- What problem does the Model Context Protocol solve?
- How do you stop an agent from taking harmful actions?
- How do you handle memory in a long-running agent?
- An agent keeps looping and burning tokens. What do you do?
- When are multi-agent systems worth the extra complexity?
- How do you debug a wrong agent decision after the fact?
- How do agents plan, and when do you choose plan-then-execute over ReAct?
- How do you design human-in-the-loop approvals for an agent?
- How can tool outputs be used to attack an agent, and what do you do?
- How do you evaluate an agent end to end?
- How do you keep agent session costs predictable?
- How do you make long-running agents resumable?
- How should an agent handle tool failures?
- How do you choose an agent framework?
- How do you handle tasks that take hours or many steps?
- How do you give agents access to user accounts safely?
- How do you safely let an agent execute code?
- Agent context grows with every step. How do you manage it?
- What is agent-to-agent communication and when do you need it?
- Which metrics show whether an agent is production-ready?
- How do you test agents whose behaviour varies run to run?
- How should an agent handle an ambiguous goal?
- How do you roll out a new agent to users safely?
- How do you write the system prompt for an agent so its behaviour stays predictable?
- Your agent passes its evals but users say it is unreliable. What could explain the gap?
- An agent can read email and also send it. What risks does that combination create, and how do you reduce them?
- A user asks the agent to complete a task that will take 20 minutes. How do you design the experience?
- Tell me about an agent behaviour that surprised you in testing. How did you respond?
Practise all Agentic AI Engineer scenarios with model answers
AI Evaluation Engineer 30 questions
- How do you build an eval set for a new LLM feature?
- What are the risks of using an LLM as a judge, and how do you mitigate them?
- How do you choose metrics for a summarisation or Q&A system?
- How do you stop prompt or model changes from silently breaking quality?
- How do you design a reliable human annotation process?
- How do offline evals and online metrics work together?
- How would you evaluate a model for harmful or biased outputs?
- A vendor claims state-of-the-art benchmark scores. Do you trust them?
- How do you keep a golden dataset useful over time?
- When is synthetic data appropriate for evals?
- How do you evaluate a RAG system?
- How do you evaluate multi-step agent trajectories?
- LLM eval scores fluctuate between runs. How do you draw reliable conclusions?
- When do you use pairwise comparison instead of absolute scoring?
- What makes a good evaluation rubric?
- Your eval suite is slow and expensive. How do you fix it?
- How do you detect and avoid benchmark contamination?
- How do you evaluate quality across languages?
- How do you evaluate open-ended tasks with no single right answer?
- How do you do error analysis on failing cases?
- How do you report eval results to executives?
- How do you structure a red-teaming exercise?
- How do you check whether a model's confidence can be trusted?
- How do you A/B test an LLM feature?
- How do you build an eval-driven culture in a team?
- Your LLM judge and your human reviewers disagree on 30% of cases. What do you do?
- How do you monitor quality in production when you have no ground-truth labels?
- How do you decide how large your eval set needs to be?
- What would you look for when choosing an evaluation framework or platform?
- Tell me about a time an evaluation result changed a decision the team had already made.
Practise all AI Evaluation Engineer scenarios with model answers
Forward Deployed Engineer 30 questions
- What does a forward deployed engineer do that a regular engineer doesn't?
- A customer asks for 'an AI agent for everything'. How do you scope it?
- The customer's data is messy and locked in legacy systems. What do you do?
- How do you build trust with sceptical stakeholders at a customer?
- How do you turn customer-specific work into product improvements?
- The customer's security team blocks your deployment. How do you respond?
- How do you take a successful demo to production?
- The customer keeps adding requests mid-pilot. What do you do?
- How do you run discovery with a new customer?
- How do you design a demo that convinces a customer?
- Your executive sponsor leaves mid-project. What do you do?
- Users aren't using the solution you delivered. What do you do?
- How do you integrate with legacy ERP or CRM systems?
- How do you set expectations about AI accuracy with a customer?
- How do you hand over a solution to the customer's team?
- How do you show ROI to a customer's leadership?
- A pilot didn't meet its success criteria. How do you handle it?
- You support several customers with urgent requests. How do you prioritise?
- How do you work with sales without overpromising?
- Two customer departments want conflicting behaviour from the system. What do you do?
- A customer bug needs a product fix that engineering hasn't prioritised. How do you escalate?
- How do you train a customer's engineers to work with the solution?
- A customer requires data to stay in a specific region. How do you respond?
- How do you balance fast prototyping with technical debt?
- How do you deliver bad news to a customer?
- A customer wants a fine-tuned model, but you think retrieval plus prompting would work. How do you handle the disagreement?
- How do you explain a technical limitation to a non-technical executive?
- You are two weeks from a customer go-live and evals show the system misses the agreed accuracy threshold. What do you do?
- The customer wants to send confidential documents to a hosted model API. What questions do you ask?
- Tell me about a deployment that went wrong at a customer site. What did you do and what did you change?
Practise all Forward Deployed Engineer scenarios with model answers
Responsible AI / AI Governance Engineer 30 questions
- How do you assess the risk of a new AI use case?
- How do you keep systems aligned with rules like the EU AI Act and India's DPDP Act?
- How do you test an AI system for bias?
- What documentation should accompany a deployed model?
- How do you prevent sensitive data leaking through an LLM application?
- How do you design meaningful human oversight?
- An AI system produced a harmful output publicly. What is your response?
- The business wants to launch despite unresolved fairness findings. What do you do?
- Why keep an inventory of AI systems, and what goes in it?
- How do you run an AI impact assessment?
- How do you approach explainability for different audiences?
- Fairness metrics can conflict. How do you choose?
- How do you assess data provenance, consent and copyright for training or retrieval?
- What do you check before adopting a third-party model or AI vendor?
- How do you design safety testing before launch?
- How do you turn a responsible-AI policy into guardrails engineers can use?
- What should be logged to support AI audits?
- How do you set up an AI governance operating model?
- How do you embed governance into the delivery pipeline?
- How would you write a generative-AI acceptable-use policy for employees?
- Employees are using unapproved AI tools. How do you respond?
- What extra safeguards are needed when users may be children or vulnerable?
- How do you monitor fairness and safety after launch?
- How do you handle a data deletion request when data may be in a model?
- How do you manage differing regulations across countries?
- What are the main governance risks specific to generative AI compared with traditional ML?
- How does governance change when an AI system can take actions, not only produce text?
- A team says documentation slows them down. How do you keep it useful and lightweight?
- A provider updates its model and outputs change for your regulated use case. How do you govern that?
- Tell me about a time you had to say no to a launch or slow one down for governance reasons. How did you handle it?
Practise all Responsible AI / AI Governance Engineer scenarios with model answers
LLMOps Engineer 30 questions
- How do you deploy and version LLM applications safely?
- What do you monitor in production LLM systems?
- How do you control LLM costs at scale?
- How do you reduce latency for a chat application?
- How do you handle provider outages and rate limits?
- When would you self-host an open model instead of using an API?
- Quality dropped and no code changed. What do you investigate?
- A prompt change caused a spike of bad outputs at 2am. How do you respond?
- Why use a prompt registry, and what should it support?
- What does CI/CD look like for an LLM application?
- What does a model gateway provide?
- What should LLM tracing capture?
- When is semantic caching a good idea?
- How do you autoscale GPU inference?
- How can you make self-hosted inference cheaper and faster?
- How do you isolate tenants in a shared LLM platform?
- How do you manage API keys and secrets for LLM services?
- How do you balance detailed logging with privacy?
- How do you keep a RAG index fresh?
- How do feature flags help with LLM releases?
- How do you define SLOs for an LLM service?
- How do you attribute LLM costs to teams and features?
- How do you manage fine-tuned models across their lifecycle?
- What belongs in an on-call runbook for an LLM service?
- How do you make staging environments realistic for LLM apps?
- How would you detect that an LLM application has started producing lower-quality answers before customers complain?
- Your traffic doubles during a marketing campaign and you hit provider rate limits. What do you do in the moment and afterwards?
- How do you roll out a new model version when you cannot fully predict how it will behave?
- How do you handle user data that ends up in prompts, logs and traces in your LLM platform?
- Tell me about an outage or incident you handled on an ML or LLM system. What did you change afterwards?
AI Solutions Architect 30 questions
- How do you design an enterprise architecture for generative AI?
- How do you decide between build, buy and partner for AI capabilities?
- How do you advise a client choosing between RAG and fine-tuning?
- What are the key security risks in LLM architectures?
- How do you design for multiple models and vendors?
- How do you plan for scale and cost in an AI platform?
- How do you explain trade-offs to non-technical executives?
- A bank wants AI on top of legacy systems under strict compliance. What is your approach?
- What does a production RAG reference architecture include?
- How do you prepare enterprise data for AI use?
- How do you choose between cloud AI services, private cloud and on-premises?
- How do you design for tight latency requirements?
- How do you add AI features to a multi-tenant SaaS product?
- How do you make an AI system resilient?
- Which integration patterns work well for AI in enterprise systems?
- What architectural concerns are specific to enterprise agents?
- How do you estimate total cost of ownership for an AI solution?
- What typically breaks between a PoC and production?
- How do you run a fair vendor evaluation?
- How should identity and access work in an AI architecture?
- How do you migrate an application to a new model safely?
- A vendor promises big accuracy gains. How do you validate the claim?
- When does on-device or edge AI make sense?
- What kinds of technical debt are unique to AI systems?
- How do you choose between a workflow and an agentic architecture for a client?
- How do you design an AI platform so that individual teams can build quickly while central risk requirements are still met?
- A client wants to use customer conversations to improve their AI assistant. How do you advise them?
- How do you build evaluation into the architecture, not treat it as a later testing phase?
- A client asks you to choose their first generative AI use case. How do you decide?
- Tell me about an architecture decision you got wrong. What happened and what did you learn?
Practise all AI Solutions Architect scenarios with model answers
Frequently asked questions
Is the interview prep free?
Yes. Every role simulator is free with no sign-up. You write your own answer to each scenario, then open the model answer to compare.
How many questions are there per role?
Thirty scenario-based questions per role, 210 across 7 roles. Each one shows what the interviewer is testing, a model answer, the answers that lose you the room, and the follow-up question you should expect.
How does the simulator work?
Pick a role, then work through the scenarios in any order. Type your answer or press Speak your answer to dictate it, then open the model answer under that question to compare. A progress bar tracks how many you have attempted.
Are the model answers the only correct ones?
No. They show the structure and reasoning interviewers reward. Use them to sharpen your own answer rather than to memorise a script.
Is my typed answer stored anywhere?
No. Your answers stay in your own browser using local storage and are never uploaded to us. Clearing your browser data removes them. If you use the Speak your answer button, your browser's speech service converts your voice to text; in Chrome that audio is processed by Google.
Which roles are covered?
AI Engineer (Prompt & Context Engineering), Agentic AI Engineer, AI Evaluation Engineer, Forward Deployed Engineer, Responsible AI / AI Governance Engineer, LLMOps Engineer and AI Solutions Architect, with further roles being added.