Back to blog

AI Engineer Interview in 2026: RAG, Evaluation, and System Design

Hello HaWkers, knowing model names is no longer enough to explain how an AI application works. The open repository AI Engineering Interview Questions Company Wise, created in September 2026, organizes questions about information retrieval, inference, agents, evaluation, and security. When I checked GitHub for this article, it had 1,562 stars. That signals interest in the subject; it does not prove that every question appeared in a real hiring process.

If an interviewer asks you to design an assistant that answers questions using internal documents, where would you start? Here you will build a verifiable answer: define the problem, create a small set of cases, measure retrieval, inspect failures, and defend architecture decisions. The framework is useful for practicing in English and adapting to an interview in Portuguese, Spanish, or French.

Start with the job the system must do

Before choosing a model, describe the user, the available information, and an acceptable outcome. Imagine a support team consulting internal policies. The question “Can I cancel after renewal?” sounds simple, but the answer depends on the policy version, country, contract type, and the requester's permissions. Ignore those details and a fluent answer can still be wrong.

In an interview, say which documents enter the index, how often they change, and who may access them. Then define the output: a short answer, a reference to the supporting passage, and a refusal when evidence is insufficient. This shows an understanding of the product and its risks. OpenAI's evaluation documentation emphasizes representative cases and clear criteria when testing an application. The NIST AI Risk Management Framework also helps you think about context, measurement, and follow-up without treating a list of principles as an automatic safety guarantee.

A useful question to ask back is: “Which error costs more here: failing to answer or answering with an outdated policy?” The answer changes the confidence threshold, the need for human review, and even the user experience. Ask before promising accuracy. It is worth more than reciting a generic architecture.

Turn isolated questions into test cases

Many people prepare for interviews by reading a huge list of questions and memorizing answers. For AI Engineer roles, it is more useful to turn a few of them into small experiments. If the topic is RAG, build a set containing each question, its expected document, and an acceptable answer. Include an unanswerable case: a trustworthy system needs to say when it could not find the information.

The example below uses only Python's standard library. Save it as casos.py and run python casos.py. The data is fictional and intended for practice; it does not represent any company's rules. The initial validation prevents a case from missing its question or reference, a common mistake in evaluation spreadsheets assembled in a hurry.

# Fictional cases for practicing a technical interview.
casos = [
    {"pergunta": "Qual é o prazo de cancelamento?", "fonte": "politica_2026", "resposta": "Consulte o contrato vigente."},
    {"pergunta": "Como redefinir a senha?", "fonte": "acesso", "resposta": "Use a página de recuperação."},
    {"pergunta": "Qual é o preço do plano futuro?", "fonte": None, "resposta": None},
]

for numero, caso in enumerate(casos, start=1):
    assert caso["pergunta"], f"Caso {numero} sem pergunta"
    assert (caso["fonte"] is None) == (caso["resposta"] is None), f"Caso {numero} incoerente"
    print(numero, "respondível" if caso["fonte"] else "sem evidência")

That set is still far too small to estimate real performance. In an interview, explain how you would expand it: frequent questions, ambiguous wording, contradictory documents, policy updates, and attempts to access someone else's data. Keep development examples separate from the final evaluation set so you do not tune the system while looking at the test answers. If the domain is regulated, request a subject matter expert's review.

The repository behind this topic is a useful catalog of themes, but treat its descriptions of hiring processes as the authors' accounts. Always verify the job opening and the company's contact. Public question lists change; the ability to build a reproducible test remains useful even when an interview format changes.

Explain RAG as a sequence of decisions

RAG combines passage retrieval with answer generation. The minimal design includes ingestion, document splitting, indexing, search, passage selection, and generation with source references. Yet each stage can lose information. Splitting a contract in the middle of an exception can remove the condition that makes a clause valid. Finding the right document but retrieving an old version is another failure.

Show that you distinguish retrieval quality from answer quality. If the right passage never reaches the model, changing the prompt will not solve the main problem. A simple practice metric is recall@k: in how many cases does the expected document appear among the first k results? This code takes simulated results and counts hits without relying on an external service.

# Check whether the expected source appeared in the first three results.
avaliacoes = [
    {"esperado": "politica_2026", "recuperados": ["faq", "politica_2026", "arquivo"]},
    {"esperado": "acesso", "recuperados": ["acesso", "suporte"]},
    {"esperado": "faturamento", "recuperados": ["vendas", "faq"]},
]

k = 3
acertos = sum(item["esperado"] in item["recuperados"][:k] for item in avaliacoes)
recall = acertos / len(avaliacoes)
print(f"recall@{k}: {recall:.1%} ({acertos}/{len(avaliacoes)})")

The result measures only this fictional sample. It does not show that the generated answer is correct or that the cited source supports the final statement. Say so explicitly. Then propose filters for version, language, and permission; compare lexical, vector, and hybrid search against the same cases; and inspect errors manually before switching tools. A mature interview answer also states when to shut down or revise the system.

Evaluate the answer, refusal, and source

An answer can cite the right document and still invent a conclusion. Evaluation therefore needs to examine at least three things: whether it answers the question, whether the passage supports the content, and whether the system refuses when it has no basis. Do not reduce everything to one score without understanding which errors it hides. For an internal policy, a false claim may be more serious than a cautious refusal.

The following exercise shows a very limited deterministic checker. It accepts only answers whose text appears literally in an allowed passage. In production, legitimate paraphrases call for more sophisticated evaluation and human review. Here the strict rule makes the contract visible and open to discussion in an interview.

# Teaching example: the answer must appear in the authorized passage.
def verificar_resposta(resposta, trechos):
    if not resposta.strip():
        return "recusa"
    texto_permitido = " ".join(trechos).casefold()
    return "sustentada" if resposta.casefold() in texto_permitido else "revisar"

trechos = ["Use a página de recuperação para redefinir a senha."]
print(verificar_resposta("Use a página de recuperação", trechos))
print(verificar_resposta("A senha é enviada por SMS", trechos))
print(verificar_resposta("", []))

Do not present this code as a universal hallucination detector. It fails with paraphrases, negations, and long texts. A serious solution combines automated criteria, samples reviewed by people, incident investigation, and monitoring after release. In the conversation, explain what you would measure by segment: language, document type, version, and question class. An overall average may conceal a serious failure in a small group.

Also record the origin of every claim the system presents. If a document is updated, you need to reproduce which version supported an earlier answer. Without that trail, a complaint turns into an argument based on memory. This requirement connects evaluation to architecture: the document identifier, version, indexing time, and passages used need to travel with the answer.

Defend the system design with explicit limits

On the interview whiteboard or in a document, draw two paths. The first updates the documents: it receives a version, validates metadata, applies access control, and publishes a new index. The second serves a question: it authenticates the person, retrieves only what they may see, produces an answer with references, and records the operation. Explain how you prevent a partial update from mixing old and new passages.

Then discuss the budget: time limit, cost per request, context size, and ability to handle spikes. Do not invent a latency target if the prompt did not provide one. Ask for the service objective and describe how you would measure it. A queue, cache, or smaller model might help, but each choice has a cost: a cache needs invalidation when a policy changes; a smaller model needs evaluation on difficult cases; a queue might be unsuitable for interactive support.

A minimal trace helps compare alternatives. The code below times a local function and records fictional identifiers without logging sensitive text. In a real system, identifiers must follow the organization's privacy and retention policies.

# Time the operation and record only the necessary metadata.
from time import perf_counter

def responder(pergunta):
    return "Consulte a política vigente" if pergunta else "Pergunta vazia"

inicio = perf_counter()
resposta = responder("Posso cancelar?")
duracao_ms = (perf_counter() - inicio) * 1000
evento = {"documento": "politica_2026", "modelo": "exemplo_local", "duracao_ms": round(duracao_ms, 2)}
print(resposta, evento)

This example does not measure a network call or predict production performance. It demonstrates the discipline: measure the real operation, record the version used, and compare results under equivalent conditions. In the interview, spell out what happens when search fails, the model takes too long, or a citation points to a removed document. A clear error path often reveals more engineering maturity than a diagram packed with boxes.

Security is part of the technical answer

Retrieved documents are data, not trusted instructions. A passage may contain a malicious sentence telling the model to ignore rules or send information to another address. The OWASP Top 10 for LLM Applications covers risks such as prompt injection and sensitive information exposure. Use it to structure security questions without assuming a text filter eliminates every risk.

Describe concrete boundaries: authorization before retrieval, tools with minimal permissions, confirmation for sensitive actions, isolation between customers' data, and tests with adversarial inputs. If the application only answers questions, do not grant it write access. If it can take action, record who authorized it, what action was requested, and what result came back. A responsible answer clearly distinguishes reading, suggesting, and executing.

This also connects to career development. Someone just starting out can show maturity without building a complete product. A small project with test cases, documented failures, and justified decisions teaches more than a flawless demo built around one prompt. The article on how to stand out when applying for junior roles in 2026 explains how to present evidence of your own work. Here the evidence is simple: show what you tested, where it failed, and what you would do next.

A practical plan for your next interview

Set aside one session to understand the role and list the tasks it actually requires. “AI Engineer” may mean an LLM product, inference infrastructure, evaluation, or data integration. Sort the topics in the job description into three groups: I can explain and demonstrate this; I can explain it but need practice; I cannot defend it yet. This keeps you from spending all your study time on new topics irrelevant to that hiring process.

In the next session, pick a small problem and build the cases. First write the expected manual answer and its source, then implement a simple search and measure where it fails. In another session, present your architecture aloud: input, permissions, updates, search, generation, evaluation, and failure recovery. Ask someone to interrupt with “What if the document is outdated?” or “How do you know the answer is correct?” Practicing objections prepares you better than memorizing acronyms.

Finish with a decision sheet: known assumptions, open questions, rejected alternatives, and missing metrics. If the interviewer changes the scenario, you can adapt the design instead of defending a fixed recipe. Be honest about the examples' limits. The code in this article is educational; a real application requires authorization, representative data, monitoring, and tests proportional to the risk.

Outlook: less memorization, more evidence

Public question lists help map the territory, but they cannot replace an explanation that holds up when a new case appears. The strongest signal in a technical interview is the ability to say what the system should do, how you would verify its result, and under which conditions it should not answer. RAG, evaluation, and system design belong in the same conversation because an error at one stage affects the others.

Bring a small project you can open, run, and critique. If the role calls for another technology, keep the method: define the problem, show a test, present trade-offs, and make uncertainty visible. That demonstrates engineering practice, including when the right answer is to ask for more context before choosing a tool.

Let's go! 🦅

📚 Want to Keep Up With What Is Coming?

This article covered AI Engineer interviews, but the ecosystem changes every week, and not every development becomes an article here.

On X, I share what I am testing, behind-the-scenes notes from projects, and developments that appear before they turn into posts.

Follow Me There

👉 Follow @jeffbruchado on X

💡 Daily content about development, careers, and the tools I actually use

Comments (0)

This article has no comments yet 😢. Be the first! 🚀🦅

Add comments