OpenAI Interview: Design an Enterprise RAG Assistant
A full solution to the OpenAI enterprise RAG question: connectors and permission sync, chunking, hybrid retrieval, reranking, permission-aware search, tenant isolation, grounded generation with citations, multi-turn context, caching, and evaluation.
By The Forward Deployed editorial teamReviewed
Part of the OpenAI system design question bank. The question is representative of the round. The analysis and solution are this site's own.
Problem statement
The product connects to each customer's content systems, indexes their documents, and answers employees' questions in a chat with citations to the source passages. Each company is a tenant. Inside a tenant, each employee may only see what they could open in the source system. Building connectors for every possible tool and training the base model are out of scope.
Clarifying questions
- How many tenants, and how big is the largest? For practice: 1,000 companies; the largest has 50 million documents and 100,000 employees.
- What sources? Wikis, document drives, ticket systems, and chat archives, each with its own permission model.
- How fresh must answers be? New and edited content within minutes. Permission removals faster than that.
- What latency? First token within about two seconds; a full answer streams over several seconds.
- Multi-turn? Yes. Follow-up questions refer to earlier turns.
- What must an answer include? Citations to the passages it used, and a clear "I could not find this" when the sources do not contain the answer.
- Query volume? For practice: 10% of employees ask 5 questions a day.
What makes enterprise RAG hard
A tutorial RAG pipeline is five steps: chunk, embed, store, retrieve, generate. It works on a folder of public documents. Enterprise data breaks it in four ways.
Permissions are per document and change constantly. The pipeline must never show a passage from a document the user cannot open, and it must stop showing it within minutes when access is removed.
Tenants share infrastructure. One company's data must never appear in another's answers, whatever the bug.
Content is messy and huge. Tables, code, slide decks, and ticket threads do not chunk well, and a 50-million-document tenant must be indexed without stalling everyone else.
And quality is invisible without measurement. A fluent wrong answer with a real-looking citation is worse than no answer, and nobody notices until a customer does.
So the driving tension is recall versus control. Retrieval wants to search as widely as possible; the enterprise needs every result filtered, isolated, fresh, and verified.
flowchart LR U([Employee]):::user --> Q[Query service]:::svc Q --> R[Retrieve<br/>permission-filtered]:::svc --> RR[Rerank]:::svc --> G[Generate with citations]:::svc --> U C[Connectors]:::svc --> I[Ingest: chunk, embed, index]:::svc --> IDX[(Per-tenant index)]:::store R --> IDX classDef user fill:#e6efec,stroke:#315e55,color:#171717; classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
Key idea. Retrieval quality is table stakes. The enterprise parts, permissions, isolation, freshness, and evaluation, are what the round grades.
Key concepts
Retrieval-augmented generation
RAG answers a question by retrieving relevant passages and putting them in the model's prompt. The model writes from the passages instead of from memory, so answers can cite sources and reflect private, current data.
Dense and sparse retrieval
Dense retrieval embeds text into vectors so that similar meanings land close together; it finds "time off policy" for a question about vacation. Sparse retrieval, such as BM25, scores exact term overlap; it finds error code E4012 and product names that embeddings blur. Hybrid retrieval runs both and merges the results.
Approximate nearest neighbor search
Comparing a query vector with hundreds of millions of vectors exactly is too slow. ANN indexes such as HNSW trade a small loss in recall for search in milliseconds.
Reranking
A cross-encoder reads the query and one passage together and outputs a relevance score. It is much more accurate than comparing two separate embeddings and much too slow for a whole corpus. So retrieve about 100 candidates cheaply, then rerank them to the best 10.
Grounding
A grounded answer makes only claims supported by the retrieved passages. Grounding is enforced by the prompt, checked after generation, and measured by evaluation.
Key idea. Hybrid retrieval for recall, reranking for precision, grounding for trust.
- Requirements
Before reading on. List the requirements. Then name the property that must never break and the constraint that drives the design.
1.1 Functional requirements
- Connect to each tenant's content systems and sync documents and permissions.
- Answer questions in a chat, streaming, with citations to passages.
- Support follow-up questions that depend on earlier turns.
- Say when the sources do not contain the answer.
- Let admins see usage, disable sources, and remove content.
1.2 Non-functional requirements
- Access correctness. Zero passages shown to users without access. Removals apply within minutes.
- Tenant isolation. No data crosses tenants, even through bugs in one layer.
- Freshness. Edits searchable within minutes.
- Latency. First token within about two seconds at p95.
- Answer quality. Measured, not assumed, with a release gate.
1.3 The constraint versus the property
Access correctness is the property. One leaked salary document ends the customer relationship. Index scale is the constraint. Hundreds of millions of chunks per large tenant drive the index design, the ingestion pipeline, and the cost.
- Back-of-the-envelope estimation
2.1 Index size for the largest tenant
50 million documents × about 10 chunks each = 500 million chunks. A 1,024-dimension embedding in 4-byte floats is 4 KB, so raw vectors take 500 M × 4 KB ≈ 2 TB. Quantizing to 1 byte per dimension cuts that to about 512 GB, which a sharded index holds in memory across a handful of machines. Across all tenants, assume the total is several times the largest.
2.2 Ingestion
The initial backfill for the largest tenant is 500 million chunks to embed. If an embedding service processes 5,000 chunks per second per GPU, one GPU needs 100,000 seconds, about 28 hours. Twenty GPUs finish in under 1.5 hours. After the backfill, edits are a trickle: if 1% of documents change per day, that is 500,000 documents, about 6 per second.
2.3 Query load
100,000 employees × 10% × 5 questions = 50,000 questions a day for the largest tenant, and perhaps 5 million a day across all tenants: about 58 per second on average and a few hundred at peak. Each question makes one or two retrieval calls, one rerank over about 100 passages, and one generation.
2.4 Tokens per answer
Ten passages of 400 tokens, plus instructions and history, is about 5,000 input tokens per question. At 5 million questions a day, that is 25 billion input tokens a day. Input tokens dominate cost, which is why reranking to 10 passages, not 30, matters.
Key idea. Hundreds of millions of vectors per large tenant, a backfill that needs GPU parallelism, a modest query rate, and input tokens that dominate cost.
- API design
Before reading on. Should the client send the chat history with every question, or should the server keep it?
Keep it on the server. The server needs history to rewrite follow-up questions, it must apply the same permission checks to any documents mentioned in history, and admins may need to audit or delete conversations. The client sends a conversation ID.
3.1 Ask
POST /v1/conversations/:id/messages
{text}
-> text/event-stream
answer_delta {text}
citation {n, doc_id, title, url, snippet}
done {message_id, answered: true|false}
POST /v1/conversations -> {conversation_id}
GET /v1/conversations/:id -> messages with citations
POST /v1/messages/:id/feedback {rating: up|down, reason?}3.2 Admin and connectors
POST /v1/sources {type: wiki|drive|tickets, credentials_ref, scope}
GET /v1/sources/:id -> {status, docs_indexed, last_sync, errors}
DELETE /v1/sources/:id -> removes all content from the index
POST /v1/documents/:id/purge
- Data model
document (tenant_id, doc_id, source_id, title, url, mime, updated_at,
acl_groups[], content_hash, deleted_at)
chunk (tenant_id, chunk_id, doc_id, section_path, text, token_count,
acl_groups[], updated_at)
vector_index per tenant or namespace:
chunk_id -> embedding, filterable fields: acl_groups, source, updated_at
keyword_index per tenant: inverted index on chunk text, same filter fields
group_member (tenant_id, group_id, user_id) -- synced from identity provider
conversation (tenant_id, conversation_id, user_id, created_at)
message (conversation_id, message_id, role, text, citations[], created_at)
eval_case (tenant_type, question, relevant_doc_ids[], reference_answer)Chunks copy the document's ACL groups so the index can filter on them without a join.
- High-level design
5.1 The tutorial pipeline
flowchart LR D[Documents]:::store --> CH[Chunk + embed]:::svc --> V[(Vector DB)]:::store Q([Question]):::user --> E[Embed] --> V --> P[Top 5 chunks]:::svc --> L[LLM]:::svc --> A([Answer]):::user classDef user fill:#e6efec,stroke:#315e55,color:#171717; classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
It answers questions about public docs. It shows any user any document, mixes tenants, misses exact terms, forgets the previous turn, and has no way to know if it is right.
5.2 Fix 1: connectors that sync permissions
Each connector pulls documents and their access lists from the source system, and group memberships from the identity provider. Every chunk carries its document's ACL groups.
flowchart LR SRC[Source systems]:::store --> CN[Connector]:::new IDP[Identity provider]:::store --> CN CN --> DOC[(Documents + ACLs)]:::store CN --> GM[(Group members)]:::store classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717; classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;
5.3 Fix 2: filter at retrieval, per tenant
The query service expands the user's groups and passes them as a filter into the vector and keyword searches, inside the tenant's own index. Forbidden passages never leave the index.
5.4 Fix 3: hybrid retrieval and reranking
Run dense and keyword retrieval in parallel, merge with reciprocal rank fusion, and rerank the top 100 with a cross-encoder to the best 10.
flowchart LR Q[Standalone query]:::svc --> D[Dense search<br/>top 100]:::new Q --> K[Keyword search<br/>top 100]:::new D --> F[Rank fusion]:::new K --> F F --> X[Cross-encoder rerank<br/>to top 10]:::new classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;
5.5 Fix 4: conversation-aware queries and grounded answers
A small model rewrites each follow-up into a standalone query using recent turns. Generation uses numbered passages and must cite them. A checker validates citations before the answer completes.
5.6 Fix 5: evaluation as a release gate
A golden set per tenant type and a permission test set run on every change to chunking, retrieval, prompts, or models. A regression blocks the release.
5.7 The composed design
sequenceDiagram
autonumber
actor U as Employee
participant A as API
participant Q as Query service
participant ACL as Group cache
participant IX as Tenant indexes
participant RR as Reranker
participant LLM as Model
U->>A: follow-up question
A->>Q: question + conversation id
Q->>Q: rewrite into standalone query using history
Q->>ACL: user's groups (cached, short TTL)
par dense and keyword
Q->>IX: vector search, filter tenant + groups
Q->>IX: keyword search, filter tenant + groups
end
Q->>Q: rank fusion
Q->>RR: rerank top 100
RR-->>Q: top 10
Q->>Q: re-check access for top 10 docs (source API)
Q->>LLM: numbered passages + question + instructions
LLM-->>A: streamed answer with [n] citations
A->>A: validate citations
A-->>U: answer + citationsKey idea. Sync permissions with content, filter inside the tenant's index, retrieve hybrid, rerank, rewrite follow-ups, cite and verify, and gate releases on evaluation.
- Deep dives
6.1 Permission-aware retrieval
Before reading on. Why is filtering the results after retrieval a bug, and not just a slower way to do the same thing?
Suppose the user asks about compensation bands. The top 100 passages by similarity are mostly from HR documents the user cannot open. A post-filter removes 95 of them and leaves 5 weak matches. The user gets a poor answer while better, permitted passages ranked 101 to 300 were never considered. Post-filtering hurts recall exactly on the questions where permissions matter most.
Pre-filtering passes the user's groups into the ANN search, so the index only returns permitted chunks. Most vector databases support filtered search; with a very restrictive filter, some switch to a brute-force scan over the matching subset, which is fine because that subset is small.
Group expansion is the next detail. A user belongs to groups, and groups nest. Expand the full set at query time from a cache refreshed from the identity provider every few minutes, and invalidate a user's entry when the provider reports a change.
Sync lag leaves a window: access is removed at 10:00, the connector syncs at 10:05. For the handful of documents actually cited in an answer, re-check access against the source system's API at answer time. It costs one call per cited document and closes the window where it matters.
What separates answers: permissions
WeakFilters in the prompt or after retrieval
Asks the model not to reveal restricted content, or filters after retrieval, which leaks or loses recall.
GoodFilters inside the search
Stores ACL groups on chunks and filters in the vector and keyword queries.
StrongFilters, expands, and re-checks
Pre-filters in both indexes, expands nested groups from a fresh cache, re-checks cited documents at answer time, and tests for leaks with a permission test set gated at zero.
6.2 Tenant isolation
Before reading on. A bug in the query service drops the tenant filter. What stops one company's data from appearing in another's answers?
Defense in depth. Large tenants get their own physical index, so there is nothing else to find. Small tenants share an index, but the retrieval service adds the tenant filter itself from the authenticated request, never from client input, and the index rejects queries without one. Each tenant's documents are encrypted with the tenant's own key, and the ingestion path for tenant A cannot obtain tenant B's key. Caches key on tenant ID first. Logs and evaluation data are partitioned by tenant too.
Say out loud that prompts are not an isolation boundary. Nothing a model is told stops it from repeating what is in its context.
6.3 Chunking real enterprise content
Before reading on. A wiki page has a long table of pricing tiers. How do you chunk it so a question about one tier retrieves the right row with its header?
Split by structure first: headings, sections, list items, table rows, code blocks. For tables, keep the header with every group of rows, so a chunk of rows 40 to 50 still says what the columns mean. For ticket threads, chunk per message and keep the ticket title and status as metadata. For slides, one slide per chunk with the deck title.
Then cap sizes, a few hundred tokens with a small overlap. Store each chunk's section path, such as "Pricing > Enterprise > Limits", and prepend it to the chunk text before embedding. That short context makes retrieval much better for chunks that do not name their own topic.
Keep a link to the parent section. If a retrieved chunk is too thin to answer from, the generator can pull in the neighboring chunks.
6.4 Grounded generation and citation checks
The prompt gives numbered passages with titles and asks the model to answer only from them, cite passage numbers after each claim, and say "I could not find this in your sources" when they do not contain the answer.
After generation, check citations mechanically: every cited number must exist in the retrieved set. An answer that cites nothing for a factual claim, or cites a number that was not provided, is regenerated or returned with a warning. For high-risk tenants, sample answers and run a support check: for each sentence and its cited passage, does the passage support the claim? A small model can run this check at scale, calibrated against human labels.
6.5 Multi-turn conversations
"What about contractors?" cannot be searched alone. A small, fast model rewrites it into "What is the parental leave policy for contractors?" using the last few turns. Retrieval runs on the rewrite. Generation still sees the conversation, so the answer reads naturally.
Bound the context. Keep the last few turns verbatim and summarize older ones. Do not carry old retrieved passages forward; retrieve fresh ones for each turn, so permission checks and freshness apply every time.
6.6 Caching without leaking
Query embeddings are safe to cache: they contain no document content. Retrieval results and answers are not safe to share between users, because two users asking the same question may have different permissions. If you cache them, key the cache on tenant ID plus a hash of the user's full group set, and keep the time-to-live short so permission changes take effect. Semantic caching across users is a trap in this product.
6.7 Evaluation
Before reading on. The team wants to switch embedding models. How do you decide?
With numbers from a golden set. Build one per kind of tenant: real questions, the documents that answer them, and reference answers, labeled by people who know the content. Measure:
| Layer | Metric | What it catches |
|---|---|---|
| Retrieval | Recall at 10, recall at 100 | The right passage never reached the model |
| Reranking | nDCG at 10 | The right passage was buried |
| Generation | Faithfulness to cited passages | Claims the sources do not support |
| Generation | Answer correctness vs reference | Wrong answers that cite real passages |
| Access | Leaks on the permission test set | Must be exactly zero |
| Online | Thumbs down rate, no-answer rate, source click rate | Real-world drift |
An LLM judge can score correctness and faithfulness at scale, if you first check its agreement with human labels. Run the whole suite on every change and block releases that regress. The embedding switch ships only if recall improves and nothing else drops.
What separates answers: evaluation
WeakLooks right in a demo
No golden set, no metrics, and no release gate.
GoodOffline metrics per layer
Measures retrieval recall and answer correctness on a labeled set.
StrongEvaluation as infrastructure
Separates retrieval, ranking, generation, and access metrics, calibrates the LLM judge, gates releases, and watches online signals for drift.
6.8 Scaling ingestion
Separate the ingestion pipeline from query serving. Connectors write changed documents to a queue per tenant. Parser and chunker workers scale on queue depth. Embedding runs in batches on GPUs. Index writers apply updates in bulk. A large tenant's backfill runs at lower priority with its own quota, so it cannot delay small tenants' incremental updates.
Deduplicate on content hash: many enterprises store the same file in several places. Deletions propagate as tombstones through the same pipeline and remove chunks from both indexes.
6.9 Freshness and deletion
Before reading on. An employee deletes a document with a customer's personal data at 10:00. At 10:02 someone asks a question that the document would answer. What do they see?
Nothing from it, if the pipeline is built for deletion. Connectors learn of changes in two ways: webhooks from the source system when it offers them, and periodic incremental polling when it does not. A deletion arrives as a tombstone event and jumps the queue ahead of ordinary edits. The index writer removes the document's chunks from both the vector and keyword indexes, and the document store marks it deleted.
The window between the delete and the index update is the risk. Close it the same way as permissions: before showing a citation, check the cited document's deleted_at and access in the document store, which updates first. A deleted document never appears in an answer even if its chunks are still in the index for a few seconds.
Edits follow the same path at normal priority. Re-chunk and re-embed only the changed document. If the content hash did not change, such as a metadata-only edit, update metadata and skip the embedding cost.
Set targets and measure them: for practice, 99% of deletions reflected in the index within one minute, and edits within five. Track the lag per connector, because a stalled connector silently serves stale answers.
6.10 Latency budget
Break the two-second target to first token into parts and assign each a budget:
| Step | Budget (p95) | How |
|---|---|---|
| Auth, group lookup | 20 ms | Cached group expansion |
| Query rewrite | 250 ms | Small, fast model, short prompt |
| Dense + keyword retrieval | 80 ms | Run in parallel; filtered ANN in memory |
| Rerank top 100 | 120 ms | Batched cross-encoder on GPU |
| Access re-check for top 10 | 100 ms | Parallel calls to source APIs, cached briefly |
| Model time to first token | 1,000 ms | 5,000-token prompt; prompt caching on instructions |
| Slack | 430 ms | Network, queueing |
If the rewrite step blows the budget, skip it on first-turn questions, which are already standalone. If the access re-check is slow for some sources, start generation in parallel and hold back citations until the check passes.
What separates answers: freshness and latency
WeakNo freshness story
Treats indexing as a batch job and has no latency breakdown.
GoodIncremental updates and a budget
Processes edits and deletions incrementally and breaks latency into steps.
StrongDeletion-safe and budgeted
Prioritizes tombstones, re-checks cited documents against the document store, sets and measures freshness targets per connector, and assigns a latency budget to every step with fallbacks.
- Variants
7.1 Agentic retrieval
For complex questions, let the model issue several searches, read results, and search again. It improves hard multi-part questions and costs more latency and tokens. Keep every tool call permission-filtered the same way, and cap the number of steps.
7.2 Structured data
Questions like "how many P1 tickets last week" need a database query, not passages. Route them to a text-to-SQL tool over a read-only replica with row-level security, and cite the query.
7.3 On-premises deployment
Some customers require the index and model inside their network. The same components deploy into their environment; the vendor runs only the control plane, and no content leaves.
7.4 At ten times the tenants
With 10,000 tenants, most are small. Put small tenants into shared indexes by region with a mandatory tenant filter, and give the largest few hundred their own. Connector scheduling becomes a fairness problem: a queue per tenant with a share of connector workers, so one tenant's resync cannot delay everyone's edits. Evaluation also scales: maintain golden sets per industry, not per tenant, and let large tenants add their own.
- The transferable pattern
Enterprise RAG is search with an authorization layer and a verifier. Retrieval finds candidates, authorization prunes them before anything else sees them, generation writes from what remains, and verification checks the output against the inputs. The same shape applies to any system that mixes private data with a model: code assistants over private repositories, support bots over ticket history, and agents with tool access.
Review: the 30-second answer
- Sync permissions with content. ACL groups on every chunk.
- Filter inside the search, per tenant. Never after retrieval, never in the prompt.
- Hybrid retrieval, then rerank. Dense plus keyword, fused, cross-encoder to top 10.
- Rewrite follow-ups; cite and verify. Numbered passages, mechanical citation checks, sampled support checks.
- Evaluate every layer and gate releases. Including a permission test set that must show zero leaks.
Quiz
+Why is post-filtering by permission a recall bug?
If most top-ranked passages are forbidden, filtering them afterward leaves few, weak results, while permitted passages ranked lower were never retrieved. Pre-filtering searches only the permitted set, so the best permitted passages come back.
+Why use keyword search next to embeddings?
Embeddings capture meaning but blur exact strings like error codes, product names, and ticket numbers. Keyword search matches them exactly. Fusing both gives better recall than either alone.
+Why rerank only the top 100 candidates?
A cross-encoder reads the query and passage together, which is accurate but slow. Running it on millions of chunks is impossible; running it on 100 cheap candidates costs tens of milliseconds and sharply improves the top 10.
+Why is a shared semantic answer cache dangerous here?
Two users can ask the same question with different permissions. A cached answer built from one user's permitted documents could reveal content to a user without access.
+What does a permission test set check, and what is its target?
It asks questions as users who lack access to specific documents and checks whether any of those documents appear in retrieval or citations. The target is exactly zero leaks, and it gates every release.
+How does the system avoid citing a document deleted seconds ago?
Deletions travel ahead of edits as tombstones, and before citing any document the answer path checks the document store, which records the deletion first. A document marked deleted is never cited, even if its chunks remain in the index briefly.
+Why assign a latency budget to each step?
A single end-to-end target does not show where time goes. Per-step budgets show which step to optimize or skip when the total runs over.
Sources and further reading
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks is the paper that introduced RAG.
- Efficient and robust approximate nearest neighbor search using HNSW graphs describes the index behind most vector search.
- Reciprocal Rank Fusion outperforms Condorcet and individual rank learning methods introduces the fusion method.
- RAG and vector search, Evaluations, and Security and compliance cover each layer from the deployment side.
- Embedding training and reranking are worked through in embedding and search design.
