The Forward Deployed

OpenAI Interview: Design a Real-Time AI Feature at Scale

A full solution to the OpenAI real-time AI feature question: picking a concrete feature, the request lifecycle, admission and queueing, load shedding, caching, cost control, frontend recovery, observability, a debugging story, and 10x scale.

By Reviewed

Part of the OpenAI system design question bank. The question is representative of the round. The analysis and solution are this site's own.

Problem statement

The question is broad on purpose. It tests whether you can make it concrete. Pick one feature early and design it end to end. This solution uses a Summarize button in a document editor: a user clicks it and a summary of the open document streams into a side panel.

Clarifying questions

  • What is the feature? Propose one: document summaries in an editor with 50 million daily users. Confirm the interviewer is happy with it.
  • Where does the model run? Assume a third-party model API with a fixed quota of tokens per minute, and the option to self-host later.
  • How fast must it feel? First words within a second or two; the full summary within about ten seconds.
  • How fresh must summaries be? They must match the current document version.
  • What about cost? There is a monthly budget. Ask for its size, or state the cost formula and treat the budget as an input.
  • Who can use it? All users, with lower limits for the free tier.

What makes a real-time AI feature hard

Model capacity is the scarce resource, and it behaves unlike a normal backend. A model call takes seconds, costs real money per token, and runs against a quota that does not grow when traffic does. A traffic spike that a web tier absorbs with autoscaling turns, for a model-backed feature, into queues, timeouts, and a bill.

Meanwhile users see everything. A slow first token feels broken. A stream that dies halfway feels worse than an error. A request that waits 60 seconds and then fails wastes both the user's time and the tokens it consumed.

So the driving tension is user experience versus scarce, expensive capacity. Every design choice either protects the capacity or spends it on something the user will notice.

flowchart LR
  U([User clicks Summarize]):::user --> GW[Gateway]:::svc
  GW --> C{Cache hit?}:::svc
  C -->|yes| S[Stream cached summary]:::store
  C -->|no| A{Admit?}:::svc
  A -->|no| R[Fast, clear rejection]:::bad
  A -->|yes| M[Model call, streamed]:::store
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef bad fill:#fbe9e4,stroke:#c4492d,color:#171717;
Key idea. Treat model tokens per second as the scarce resource. Protect it with caching and admission, and spend it only on requests the user will still be waiting for.

Key concepts

Admission control

Deciding at the door whether a request can be served in time. If not, reject it at once with a clear message. It keeps queues short, so admitted requests stay fast.

Load shedding and degradation

Under overload, drop or cheapen the least valuable work first: free-tier requests, long outputs, or the large model. Degradation keeps the feature useful at lower quality instead of failing for everyone.

Little's law

Concurrency equals arrival rate times duration. Streams that last ten seconds at 300 per second mean 3,000 open streams. It sizes connection capacity and model concurrency.

Prompt caching

Many model providers cache a repeated prompt prefix and bill it at a lower rate, with faster processing. A fixed instruction block at the start of every prompt benefits.

Distributed tracing

One trace ID follows a request across the browser, gateway, queue, and model call. Each step records a timed span. It turns "summaries are slow" into "queue wait is fine, model time doubled."

Key idea. Admit what you can serve, shed the least valuable work first, size by Little's law, cache prefixes and results, and trace every request.

  1. Requirements

Before reading on. Write the requirements for the Summarize feature. Which single user-facing number would you promise?

1.1 Functional requirements

  • Summarize the current document version, streamed into a panel.
  • Return an existing summary instantly when the document has not changed.
  • Cancel when the user closes the panel.
  • Retry on failure.
  • Enforce per-user and per-tier limits.

1.2 Non-functional requirements

  • Time to first token under 1.5 seconds at p95 for admitted requests.
  • No long hangs. A request either starts streaming within a few seconds or is rejected clearly.
  • Cost within budget, with alerts before overruns.
  • Observability. Any slow or failed request can be explained from its trace.

1.3 The constraint versus the property

A fast, honest experience is the property. Users forgive "busy, try again in a minute"; they do not forgive a spinner that runs for a minute. Model capacity and cost are the constraint. They are fixed in the short term and expensive to grow.

  1. Back-of-the-envelope estimation

2.1 Request rate

50 million daily users; 10% use Summarize once a day: 5 million requests a day, about 58 per second on average. Peaks at 5 times the average: about 290 per second.

2.2 Tokens

1,500 input tokens (document plus instructions) and 300 output tokens per summary.

  • Peak output: 290 × 300 = 87,000 tokens per second.
  • Peak input: 290 × 1,500 = 435,000 tokens per second.
  • Daily: 7.5 billion input and 1.5 billion output tokens.

2.3 Cost

Daily cost = input tokens × input price + output tokens × output price. With output often priced several times higher than input per token, both terms matter here. Write the formula even without prices. Then note the two biggest levers: the cache hit rate, which removes whole requests, and output length, which scales the expensive term.

2.4 Concurrency

If a summary streams for 6 seconds, open streams at peak = 290 × 6 ≈ 1,740. Gateways handle that easily; the model quota is what binds.

2.5 Cache potential

Popular documents are read and summarized by many people. If 30% of requests hit an existing summary for the same document version, peak model load drops to about 200 requests per second, and cost drops by 30%.

Key idea. Tokens per second at peak is the capacity number; cost is a formula with two levers, cache hit rate and output length.

  1. API design

Before reading on. The user clicks Summarize twice quickly. How many model calls should happen?

One. Deduplicate on the document version: the second request joins the first request's stream, or gets the cached result when it finishes.

3.1 Summaries

POST /v1/documents/:doc_id/summaries
  {doc_version, length: short|medium}
  -> 200 text/event-stream
       event: meta    {summary_id, cached: false, model: "fast"}
       event: delta   {text}
       event: done    {usage}
  -> 200 (cached)  same stream, emitted at once
  -> 429 {reason: "user_limit", retry_after}
  -> 503 {reason: "busy", retry_after}
  -> 413 {reason: "document_too_long"}

3.2 Why SSE

The answer flows one way after one request, over plain HTTP, with built-in reconnection semantics. WebSockets would add connection state for no benefit here.

  1. Data model

summary_cache   key: (doc_id, doc_version, prompt_version, length)
                value: summary_text, model, created_at     TTL: 30 days
inflight        key: same as cache -> stream id (dedupe concurrent requests)
user_limits     (user_id, tier, tokens_today, requests_this_minute)
usage_events    (ts, user_id, doc_id, model, input_tokens, output_tokens,
                 cached, latency_ms, outcome)               -> analytics
prompt_versions (id, template, created_at, active)

The prompt version is part of the cache key. Changing the prompt must not serve summaries made with the old one.

  1. High-level design

5.1 Call the model on every click

flowchart LR
  U([User]):::user --> API[API]:::svc --> M[Model]:::store
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;

Every click costs a full model call, including repeats. A spike sends more requests than the quota allows; the provider returns errors or slows down, and users wait on timeouts.

5.2 Fix 1: cache by document version

Summaries for the same document version, prompt version, and length are identical in purpose. Cache them. Deduplicate requests in flight.

5.3 Fix 2: admission control and a bounded queue

A token bucket per user caps individual use. A global concurrency limit matches the model quota. Requests beyond it wait in a short queue with a deadline; if the estimated wait exceeds it, reject at once.

flowchart LR
  R[Request]:::svc --> PU{User bucket ok?}:::new
  PU -->|no| X1[429]:::bad
  PU -->|yes| G{Global slots free?}:::new
  G -->|yes| M[Model call]:::store
  G -->|no| Q[[Queue with deadline]]:::new
  Q -->|slot frees in time| M
  Q -->|deadline passes| X2[503 busy]:::bad
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef bad fill:#fbe9e4,stroke:#c4492d,color:#171717;
  classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;

5.4 Fix 3: degrade under overload

When the queue grows, switch free-tier requests to a smaller model and shorter summaries, then shed free-tier requests if needed. Paid users keep the best experience longest.

5.5 Fix 4: cancellation and tracing

When the browser closes the stream, the gateway cancels the upstream call. Every request carries a trace ID through every hop.

5.6 The composed design

sequenceDiagram
  autonumber
  actor U as User
  participant B as Browser
  participant G as Gateway
  participant C as Cache
  participant A as Admission
  participant M as Model API
  U->>B: click Summarize
  B->>G: POST summaries {doc_version} (trace id)
  G->>C: lookup (doc, version, prompt, length)
  alt hit
    C-->>G: summary
    G-->>B: stream cached text
  else miss
    G->>A: admit? (user bucket, global slots, tier)
    alt rejected
      A-->>G: busy, retry_after
      G-->>B: 503 with message
    else admitted
      G->>M: stream request (prompt version N)
      M-->>G: tokens
      G-->>B: delta events
      G->>C: store summary
    end
  end
  opt user closes panel
    B->>G: disconnect
    G->>M: cancel
  end
Key idea. Cache first, then admit, then degrade, then cancel what nobody waits for, and trace all of it.

  1. Deep dives

6.1 Admission and overload

Before reading on. Traffic triples for twenty minutes after a product launch. The model quota does not change. Walk through what users see.

Minute 0: the cache absorbs repeats of popular documents. Misses go to admission. Global slots fill.

Minute 2: the queue grows. The gateway estimates wait as queue length divided by recent completion rate. For admitted requests the estimate stays under the deadline, so they stream normally.

Minute 3: the estimate passes the free-tier threshold. Free users get the smaller model with a short summary, which costs a fraction of the tokens, so more requests fit through the same quota.

Minute 5: still over. Free users now get "Summaries are busy; try again in a minute" at once. Paid users still get full summaries.

Minute 20: traffic falls; the thresholds relax in reverse order.

At no point does anyone wait a minute and then fail. That is the property, kept under an unchanging constraint.

What separates answers: overload

WeakAutoscales the API tier

Adds servers, which does nothing when the model quota is the limit.

GoodRate limits and a queue

Limits per user, caps concurrency, and queues with a timeout.

StrongAdmission by estimated wait, with tiered degradation

Rejects early when the wait would exceed the deadline, degrades free traffic to a cheaper model before shedding it, and explains what each tier sees minute by minute.

6.2 Caching and deduplication

Cache by document version and prompt version. An edit to the document changes its version and misses the cache, which is correct. A prompt change invalidates everything, which is also correct, and costs a burst of misses; roll prompt versions gradually to spread it.

Two users who click Summarize on a popular document within seconds should share one model call. Keep an in-flight map from cache key to stream. The second request subscribes to the first stream and replays the tokens so far, then follows live.

Do not use semantic similarity caching across documents. Two similar documents can differ in the one sentence that matters.

6.3 Cost control

  • Cap output. A maximum length per tier; "short" summaries by default.
  • Route by size. Short documents to a smaller model; long ones to the larger model, or summarize in chunks with the small model and combine.
  • Prompt caching. Put the fixed instructions first so the provider can cache them.
  • Budgets and alerts. Track cost per hour against a daily budget; alert at 70% of budget by time of day, not at 100%.
  • Measure per feature. Cost per summary and per active user, so product decisions see the cost.

Cost drifts quietly. The usual cause is a prompt change that grows input tokens. Track tokens per request as a metric with an alert.

6.4 The frontend

Show the panel at once with a placeholder. Render tokens as they arrive, batched per animation frame. Keep partial markdown safe to render. If the stream drops, keep the partial text, mark it interrupted, and offer Retry. Map errors to plain language:

CaseMessage
429 user limit"You have reached today's summary limit."
503 busy"Summaries are busy right now. Try again in a minute."
413 too long"This document is too long to summarize in one pass."
Stream droppedPartial text, marked, with Retry

Abort the request when the user closes the panel or navigates away.

6.5 Observability and a debugging story

Before reading on. Users say summaries got slow yesterday. Walk through how you find the cause.

Track, at p50 and p95: time to first token, total time, queue wait, and model time. Also cache hit rate, rejection rate by reason, cancel rate, error rate by type, input and output tokens per request, and cost per hour. Log prompts and outputs only with sampling and redaction, because they contain user content.

The investigation: open the latency dashboard. Time to first token doubled yesterday at 14:00. Queue wait is flat, so it is not overload. Model time doubled. Input tokens per request doubled at the same minute. The deploy log shows a prompt change at 13:58 that added the document's full revision history to the prompt. Roll back the prompt version; the metrics recover within minutes. Add an alert on input tokens per request.

Tell that story concretely. The interviewer asks for a production story because it shows whether your monitoring would have found the cause.

What separates answers: operations

WeakLogs errors

Has error logs and an uptime check, and no way to explain latency or cost changes.

GoodLatency and error metrics

Tracks p95 latency and error rates, with dashboards.

StrongTraces with a debugging path

Breaks latency into queue wait and model time with trace IDs, tracks tokens per request and cost, and shows how those metrics would localize a real regression.

6.6 At ten times the traffic

  • Capacity. Raise provider quotas ahead of time; add a second provider or region behind a router with health checks; consider self-hosting a smaller model for the free tier.
  • Precompute. Summarize the most-read documents when they change, in background batch jobs that use off-peak capacity.
  • Cache. Hit rates rise with traffic, because popularity is skewed; a larger cache pays off more.
  • Limits. The same admission rules, with tighter free-tier limits.
  • Cost. Ten times the tokens is ten times the bill unless caching and routing improve. Show the cost model to the product owner before launch.

6.7 Long documents

Before reading on. A user clicks Summarize on a 300-page document, about 150,000 tokens. The model's context is 128,000 tokens. What happens?

Three options, ordered by quality and cost.

  • Reject clearly. Return 413 with "This document is too long to summarize in one pass." Cheap and honest, and bad for the users who need it most.
  • Map-reduce. Split the document by section into chunks that fit, summarize each chunk with the small model in parallel, then summarize the summaries with the larger model. It handles any length. It loses cross-section connections and costs roughly one pass over the whole document plus a final pass.
  • Incremental. Keep summaries per section cached by section content hash. When the document changes, re-summarize only changed sections and rebuild the top summary. For long, frequently edited documents this is by far the cheapest.

For practice: 150,000 tokens in 10 chunks of 15,000, each producing a 500-token summary, then a final pass over 5,000 tokens. Input is about 155,000 tokens, similar to one pass, and the 10 chunk calls run in parallel, so time to first token for the final summary is one chunk call plus the final call's first token. Stream a "summarizing sections" progress state while the chunks run.

6.8 Quality and safety of the output

A fast, cheap summary that is wrong is worse than none. Build a small evaluation set of documents with reference summaries, and score new prompt versions and models on faithfulness (no claims absent from the document) and coverage (main points present), using a calibrated model judge. Gate prompt releases on it, the same way code is gated on tests.

For safety, the document is untrusted input. Text inside it such as "ignore previous instructions" must not change the feature's behavior. Keep instructions in the system message, wrap the document in clear delimiters, and never give this feature tools that can act on the user's data.

What separates answers: content edge cases

WeakAssumes documents fit

No plan for long documents or for bad summaries.

GoodHandles length

Uses map-reduce or chunking for long documents.

StrongHandles length, quality, and injection

Caches section summaries for incremental updates, gates prompt releases on an evaluation set, and treats document text as untrusted input.

  1. Variants

7.1 Autocomplete while typing

A latency-critical variant: suggestions within a few hundred milliseconds. Use a small, fast model, debounce keystrokes, cancel stale requests aggressively, and cache by prefix.

7.2 Voice features

Audio both ways needs WebSockets or WebRTC, barge-in handling, and much tighter latency budgets per step.

7.3 Background features

Features that can wait, such as nightly digests, use batch APIs at lower cost and never compete with interactive traffic for quota.

7.4 Multiple providers

Relying on one model provider makes its outages yours. Put a router in front with health checks and error-rate tracking per provider. Keep prompts portable, evaluate each provider on the same evaluation set, and fail over only to providers that pass. Cache keys must include the provider and model, because outputs differ.

  1. The transferable pattern

A real-time AI feature is a scarce-resource front door. Cache what repeats, admit what fits, degrade before failing, cancel what nobody waits for, and trace everything. The same front door protects any expensive backend: payment processors, search clusters, GPU pools.

Review: the 30-second answer

  • Make it concrete. One feature, with numbers: 290 requests per second, 87,000 output tokens per second at peak.
  • Cache by content version, dedupe in flight. The cheapest token is the one never generated.
  • Admit by estimated wait. Short queues, fast rejections, degrade free traffic first.
  • Cancel on disconnect; cap output. Cost follows output length and hit rate.
  • Trace every request. Split latency into queue and model time, and track tokens per request.

Quiz

+Why doesn't autoscaling the API servers fix overload for this feature?

The binding limit is model capacity, a quota of tokens per minute. More API servers only send more requests into the same quota, which lengthens queues and timeouts.

+Why include the prompt version in the cache key?

A new prompt produces different summaries. Without the version in the key, users would keep seeing summaries made with the old prompt after a change.

+What does admission by estimated wait achieve?

Requests that would wait longer than the deadline are rejected immediately instead of failing after a long wait. Admitted requests then stay fast, and users get honest feedback.

+Why cancel the model call when the user closes the panel?

Nobody will read the rest of the output, but it still costs tokens and capacity. Cancelling frees the slot for someone who is waiting.

+In the debugging story, which metric pointed to the cause?

Input tokens per request doubled at the moment latency doubled, while queue wait stayed flat. That pointed to a prompt change, not to overload.

+How can a feature summarize a document longer than the model's context?

Split it into chunks that fit, summarize the chunks in parallel, then summarize the summaries. Caching per-section summaries by content hash makes later edits cheap.

+Why is the document treated as untrusted input?

Its text can contain instructions meant to change the model's behavior. Keeping instructions in the system message, delimiting the document, and giving the feature no tools limit the damage.

Sources and further reading

NextVideo Generation Pipeline