OpenAI System Design Interview Questions
Reported OpenAI system design interview questions, from enterprise RAG and GPU video pipelines to payments, chess, and Slack, with what each round probes.
By The Forward Deployed editorial teamReviewed
This bank covers the system-design pool reported for software, ML, and infrastructure roles at OpenAI. The forward-deployed loop has its own shape, covered in the OpenAI FDE interview guide. Use this page when your loop includes a system-design round, or when you are weighing an engineering role at OpenAI next to the FDE role. Each question below links to a full analysis and solution. Answer the question out loud before you open it.
The questions fall into two families. The first is AI products: retrieval, streaming, GPU scheduling, and training data. The second is the classic distributed-systems pool, where the interviewer grades depth on one or two components. Answer each question out loud, on a timer, with the protocol from AI System Design Cases.
AI product and ML system questions
Design an enterprise assistant over a company's internal documents
A company uploads its wikis, tickets, and shared drives. Employees ask questions and get answers with sources. Design the whole system for many tenants, with permissions on each document.
The round goes deep into retrieval: chunking, dense versus sparse search, the reranking stage, and how each answer stays tied to its sources. Permissions are the trap. A user must never see a passage from a document they cannot open, so enforce access at retrieval time and filter before generation. Candidates also lose points when they design for a single turn and forget conversation history, or when they offer no evaluation plan.
Prepare with RAG and vector search, Evaluations, and Security and compliance.
Read the full analysis and solution.
Explain how you would train and use a text-embedding model for search
An oral round. An interviewer with a search background asks how an embedding model learns, then follows the retrieval flow from query to ranked results.
Be ready to write a contrastive loss from memory and explain each term. The follow-ups are about negatives. Why are in-batch negatives free? Why does a bigger batch help? When do you mine hard negatives because the in-batch ones are too easy? Then compare pointwise, pairwise, and listwise reranking. Have the listwise tradeoffs ready before the question arrives; a vague answer there is a reported place to lose the round.
Read the full analysis and solution.
Design a chat interface where all history stays in the browser
Users log in, send messages, and watch answers stream in token by token. The server stores no chat history. A page refresh starts a new conversation.
This is a frontend and full-stack round. The interviewer wants a reasoned choice between server-sent events and WebSockets, and a plan for a stream that drops halfway. Show how you render new tokens without re-rendering the whole transcript. Then cover token-based login, browser storage limits, and retries when the model API fails. Cost and latency covers streaming and time to first token.
Read the full analysis and solution.
Design a model playground with adjustable parameters and saved presets
Users type a prompt, tune settings such as temperature, and see the completion stream in. They can save a prompt with its settings as a preset and run it again later. You are the only developer, so every technical choice is yours to defend.
Start with the problem the product solves. Then the data model: users, presets, and the settings each preset freezes. The rest is the data flow between browser, backend, and model API, drawn on the shared board. Show where state lives in the UI and what happens when a user edits one preset in two tabs. Prompting and structured output helps you argue which settings belong in a preset.
Read the full analysis and solution.
Design a real-time AI feature for a very large audience
Pick the transport, then defend it. The question asks for streaming answers at scale, with admission control, queueing, overload behavior, and cost limits around the model call.
Connect each backend decision to what the user sees. When you shed load, what does the screen show? When a stream fails, how does the client recover, and who pays for the lost tokens? Draw the request lifecycle from submit through admission, model execution, streaming, cancellation, and error recovery. Then redo the design at ten times the traffic. Bring one production story with a failure, the debugging path, and the tradeoff you chose. Pair Cost and latency with Observability.
Read the full analysis and solution.
Design a video-generation service on a limited GPU pool
A client submits a prompt and gets a video minutes later. GPUs are scarce, their count changes, and workers can disappear in the middle of a job.
The API must be asynchronous: return a durable job ID at once, and expose status, progress, and cancel. Most of the round is scheduling and failure. Leases and heartbeats send a dead worker's job back to the queue. Checkpoints bound the work a preemption destroys. Admission control gives an honest ETA when the queue is long. Do the capacity math out loud, because jobs per second times minutes per job gives the workers you need at peak. The metadata side is small. GPU capacity dominates, so spend your time there. See Deployment and Cost and latency.
Read the full analysis and solution.
Mine new, useful examples from a huge unlabeled corpus
An ML design question with a thin spec. Find novel information in a very large unlabeled dataset, or find every image that contains a given object. One variant uses medical data and asks for high-signal training examples.
The spec is thin on purpose. Your first job is to define novel and useful in a way you can measure. Then build the filtering and retrieval pipeline around that definition: embeddings, deduplication, and a small labeled sample to check precision. That sample is the same move as a golden set in Evaluations.
Read the full analysis and solution.
Distributed-systems questions
The rest of the pool is familiar ground. The interviewer usually picks one or two components and goes deep. A template that gives every box equal time runs out of clock.
Design payments for a coffee shop, or whatever the question actually says
The title says payments. The question often narrows to one vertical: in-store ordering, an online checkout, or the path from a merchant to its payment provider. Read it twice before you draw. A common failure is a polished generic payment design delivered to a narrower question.
Wherever it lands, the same core gets probed. Idempotency keys on every request, so a retry never charges twice. A state machine from authorized to captured to settled, with holds that expire. A double-entry ledger. Webhook handlers that are safe to replay, because providers retry. The coffee-shop version adds a hold at order time, a final charge that can differ from the hold because of tips, and a nightly settlement batch. Name the index that makes the batch query cheap, and give the batch its own idempotency token. Keep strong consistency for money. Write the real table schemas with keys and indexes; vague storage answers lose points.
Read the full analysis and solution.
Design an online chess platform
Matchmaking by rating, moves in real time over WebSockets, and a clock the server owns. The timer on screen is only a display.
The clock is the interesting part. Store the time of the last move and each player's remaining time, and compute the rest on demand; writing every tick to a database is the classic mistake. Handle reconnects, resignations, draw offers, and duplicate move submissions with move numbers that make each move idempotent. Matchmaking widens the rating window the longer a player waits.
Read the full analysis and solution.
Design a browser-based cloud IDE
Users edit files, run commands, and see terminal output live, with nothing installed locally. One rotation replaces the browser with SSH access to a fleet of hosts.
Isolation comes first: one user can never reach another user's files or processes. Next is the sandbox lifecycle, with fast starts from a warm pool, idle shutdown, and a decision about what survives a restart. Terminal output must stream fast enough that typing feels local. Ask early whether sessions last minutes or hours, because the answer changes the lifecycle design.
Read the full analysis and solution.
Design a multi-tenant CI/CD system
A git push triggers a workflow defined in the repository. Jobs run in order, in containers, and each job must run exactly once, even across crashes. Users watch status and logs live.
Exactly-once is the whole round. Say plainly that the queue delivers at least once. The effect becomes exactly-once through leases, fencing tokens, and idempotent job-state transitions. Then show how the system detects a crashed worker and reassigns its job without a second run. Tenant isolation covers compute, secrets, and fair scheduling, so one busy tenant cannot starve the others.
Read the full analysis and solution.
Design a webhook delivery system
Customers register callback URLs for event types. The system posts each event to the right URL at very high volume, and every event must arrive at least once.
Design the API and schema first. Then the retry path: exponential backoff with jitter, a dead-letter queue, and a delivery log customers can query. Isolation is the follow-up that matters. One customer's slow or dead endpoint must not delay every other delivery, so partition the queues or cap concurrency per endpoint. Sign each payload so receivers can verify it, and include an event ID so they can deduplicate.
Read the full analysis and solution.
Design Slack
Direct messages and channels, several devices per user, notifications, and workspace isolation. Start with direct messages, then build up to channels.
Write the message and a per-recipient inbox entry first, then publish for live delivery. The durable write guarantees delivery; the pub/sub layer supplies speed. Large channels change the fan-out. Past some size, publish once per channel and let each gateway deliver to its local members. Cover multiple devices before you are asked, with a per-device cursor that finds gaps on reconnect. Be ready to compare Redis pub/sub with Kafka as the transport, and to explain why user data and message data live in separate services.
Read the full analysis and solution.
Design a nearby-places search
Find points of interest near a user, anywhere in the world, under heavy read traffic and a tight latency budget. Variants ask for the exact K nearest results, a sharded index, or fast updates to business details. Geo indexing is the core. Explain geohash or a quadtree precisely, compare the two, and handle the case where the nearest result sits in a neighboring cell. Then shard the index and cache the busy areas.
Read the full analysis and solution.
Design a distributed crossword solver
Given a large grid with about a hundred slots and a dictionary of about a million words, find one valid fill. The interviewer is not looking for a clever algorithm. Show quickly that one machine cannot search the space. Then design a job system: split the search tree into tasks, prune dead branches early with constraint checks, move work away from stuck workers, and stop every worker once one finds a solution.
Read the full analysis and solution.
Design Google Calendar
Users create events and view them by day, week, month, or year, on several devices. The two hard parts are fast range queries for each view and near-real-time sync. Recurring events are the depth test: store the rule, expand it on read, and treat an edit to one occurrence as an exception.
Read the full analysis and solution.
Design a URL shortener
The warm-up of the pool. Cover key generation, with a counter and base62 encoding or a hash with collision handling. Then the read-heavy redirect path and a cache for hot links. Depth comes from global reads and abuse controls.
Read the full analysis and solution.
What carries across the pool
Two habits score in almost every question above. Do the capacity arithmetic out loud before you pick components. Write concrete schemas and APIs, with keys, indexes, and request shapes. The AI questions add a third habit: say how you would know the system works, with an evaluation you could actually run. To see how the FDE loop tests these skills with a customer attached, go back to the OpenAI FDE interview guide.
Frequently asked questions
What system design questions does OpenAI ask?
Reported questions span AI products and classic distributed systems. The AI side includes an enterprise document assistant built on retrieval, embedding and reranking design, a streaming chat interface, a model playground, and a video-generation pipeline on scarce GPUs. The classic side includes payments with holds and settlement, online chess, a cloud IDE, multi-tenant CI/CD, webhook delivery, Slack, and nearby-places search. The pool rotates, so confirm the round format with your recruiter.
Does the OpenAI FDE loop use these questions?
The FDE loop has its own shape: a take-home on OpenAI's APIs, a technical deep dive, and an open customer solution-design round. The questions on this page are reported for the engineering system-design rounds. They still build the skills the FDE design round tests. See the OpenAI FDE interview guide.
How should I practice OpenAI system design questions?
Pick one question, set a timer, and answer out loud. Do the capacity math first, write the schema and API, then go deep on the component the question makes hardest. The protocol and the AI interviewer on AI System Design Cases work for these questions too.
