The Forward Deployed

OpenAI Interview: Design a ChatGPT-Style Chat Interface

A full solution to the OpenAI chat frontend question: a key-holding proxy, login and sessions, SSE versus WebSockets, parsing a streamed response, rendering tokens efficiently, browser state, context limits, error recovery, and scale.

By Reviewed

Part of the OpenAI system design question bank. The question is representative of the round. The analysis and solution are this site's own.

Problem statement

The product is a web app with a login, a message box, and a transcript that fills in as the model writes. The interviewer usually has a frontend background and expects depth on streaming, rendering, state, and failure handling in the browser. The no-storage rule narrows the backend, but it does not remove it.

Clarifying questions

  • Who logs in, and how? Assume company single sign-on or a standard identity provider using OpenID Connect.
  • Which model API? A chat completion API that supports streaming.
  • What happens on refresh? The conversation clears. Confirm that losing the chat on an accidental reload is acceptable.
  • How long can a conversation get? Long enough to hit the model's context limit. Ask what to do then.
  • Markdown and code in answers? Yes, rendered.
  • How many users? For practice: 1 million daily users, 100,000 at peak at once, 20 messages each per day.

What makes a streaming chat UI hard

Three things are harder than they look.

The API key cannot ship to the browser. Anything in client code is public. So even with no stored history, the design needs a backend that authenticates users and calls the model with a secret key.

Streaming is a long-lived response. The server sends a slow trickle of small events for tens of seconds. Every hop, from the load balancer to the browser, must pass partial data through without buffering it, and every hop's timeout must allow it.

Rendering can become the bottleneck. Tokens arrive dozens of times a second. If each one re-renders the whole transcript and re-parses all its markdown, the page stutters and typing lags, even though the network is fine.

So the driving tension is responsiveness versus correctness under partial data. The UI must show text as soon as it arrives, while the text is incomplete, the markdown is half-formed, and the stream may break at any moment.

flowchart LR
  B([Browser SPA]):::user -->|POST /api/chat, full history| P[Chat proxy<br/>auth, limits, key]:::svc
  P -->|stream=true| M[Model API]:::store
  M -->|tokens| P -->|text/event-stream| B
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
Key idea. Stateless is not backendless. A thin proxy holds the key and relays the stream; the browser owns the conversation and must render partial, fragile data smoothly.

Key concepts

Server-sent events

SSE is a one-way stream from server to browser over ordinary HTTP. The response has content type text/event-stream, and each event is a few field: value lines followed by a blank line. It passes through proxies, works with HTTP/2 multiplexing, and needs nothing special on the server.

WebSockets

A WebSocket upgrades an HTTP connection into a two-way channel. Either side can send at any time. It suits voice, live interruption, and collaboration, and it needs its own routing, heartbeats, and reconnect logic.

Fetch streaming

The browser's EventSource API only makes GET requests with no body. A chat request needs a POST with the history in the body. So read the stream with fetch and the response's ReadableStream, and parse SSE lines yourself.

Render batching

The browser paints about 60 times a second. Updating state more often than that wastes work. Collect tokens in a buffer and flush it once per animation frame.

Key idea. SSE over fetch for one-way streams, WebSockets only when the client must talk mid-stream, and render at frame rate, not token rate.

  1. Requirements

Before reading on. What must the product do, and which quality would users notice first if it failed?

1.1 Functional requirements

  • Log in and log out.
  • Send a message and see the answer stream in.
  • Stop a response in progress.
  • Retry a failed or interrupted response.
  • Render markdown and code with copy buttons.
  • Clear the conversation on refresh.

1.2 Non-functional requirements

  • Time to first token under about one second after send, most of which is the model.
  • Smooth rendering. No dropped frames while streaming; typing stays responsive.
  • Security. No API key in the browser, no session token readable by scripts, no script injection through model output.
  • Resilience. A dropped stream never loses the text already shown and always offers a retry.

1.3 The constraint versus the property

A responsive, trustworthy UI is the property. Users judge the product by how fast text appears and whether it ever freezes or loses work. The no-server-state rule is the constraint. It moves conversation state, context management, and recovery into the browser.

  1. Back-of-the-envelope estimation

2.1 Requests

1 million users × 20 messages = 20 million chat requests a day, about 230 per second on average. Peaks at 5 times give about 1,200 per second.

2.2 Concurrent streams

If an answer streams for 15 seconds on average, the proxy holds about 1,200 × 15 = 18,000 open streams at peak (Little's law: arrival rate times duration). A proxy host with async I/O holds thousands of idle connections easily, so a handful of hosts suffices, with headroom.

2.3 Request size

Because the browser sends the whole history each time, requests grow. A 40-turn conversation at 300 tokens per turn is 12,000 tokens, about 48 KB of text. Uploading 48 KB per message is fine; the model's context limit arrives long before bandwidth matters.

2.4 Token rate in the browser

A model streaming 50 to 100 tokens per second delivers a token every 10 to 20 ms. Rendering per token would mean 50 to 100 React updates a second on a growing transcript. Batching to frames caps it at 60 and usually far fewer.

Key idea. Capacity is measured in open streams, not requests per second, and the browser's render loop is as much a bottleneck as the network.

  1. API design

Before reading on. Should the stream be a GET with EventSource or a POST with fetch? Why?

A POST with fetch. The request carries the full message history, which does not fit in a URL, and it should not be cached or logged by intermediaries the way GET URLs often are.

3.1 Chat

POST /api/chat
Cookie: session=...            (HTTP-only, Secure, SameSite=Lax)
{
  "messages": [{"role": "user", "content": "..."}, ...],
  "model": "default",
  "client_request_id": "uuid"
}

200 text/event-stream
  event: delta      data: {"text": "Hel"}
  event: delta      data: {"text": "lo"}
  event: done       data: {"finish_reason": "stop", "usage": {...}}
  event: error      data: {"code": "upstream_timeout", "retryable": true}

401  session expired            -> client redirects to login
413  history exceeds context    -> client trims and retries
429  rate limited, Retry-After

3.2 Auth

GET  /auth/login      -> redirect to identity provider (code flow + PKCE)
GET  /auth/callback   -> sets session cookie, redirects to app
POST /auth/logout     -> clears cookie
GET  /api/me          -> {user_id, name}

  1. Data model

The server stores users' sessions only. The browser holds the conversation.

// browser, in memory
type Message = {
  id: string;
  role: "user" | "assistant";
  text: string;
  status: "streaming" | "done" | "error" | "stopped";
};

type ChatState = {
  messages: Message[];
  inFlight?: { controller: AbortController; messageId: string };
  draft: string;
};

// server
session (session_id, user_id, expires_at)      // or a signed, encrypted cookie
usage   (user_id, day, requests, tokens)       // for rate limits, no content

  1. High-level design

5.1 The browser calls the model directly

flowchart LR
  B([Browser<br/>API key in JS]):::bad --> M[Model API]:::store
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef bad fill:#fbe9e4,stroke:#c4492d,color:#171717;

It streams, and it publishes the API key to anyone who opens developer tools. There is no login and no per-user limit.

5.2 Fix 1: a stateless proxy

A small backend holds the key, checks the user's session, applies rate limits, and forwards the request with streaming on. It relays each event as it arrives and stores no content.

flowchart LR
  B([Browser]):::user --> P[Chat proxy]:::new --> M[Model API]:::store
  P --> RL[(Rate-limit counters)]:::new
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;

5.3 Fix 2: login with an HTTP-only session

The proxy runs a standard login flow and sets a session cookie that scripts cannot read. Every chat request carries it automatically.

5.4 Fix 3: a streaming parser and a frame-rate renderer

The client reads the response stream, parses SSE events, appends text to a buffer, and flushes to state once per frame. Finished messages are memoized.

5.5 Fix 4: cancel, retry, and context trimming

A Stop button aborts the fetch; the proxy sees the closed connection and cancels the upstream call. Failed streams keep their partial text and show Retry. Before each send, the client trims history to fit the context window.

5.6 The composed design

sequenceDiagram
  autonumber
  actor U as User
  participant B as Browser
  participant P as Proxy
  participant M as Model API
  U->>B: type message, press Enter
  B->>B: append user message, trim history to fit
  B->>P: POST /api/chat (cookie, full history)
  P->>P: check session, rate limit
  P->>M: stream=true, API key
  loop tokens
    M-->>P: chunk
    P-->>B: event: delta
    B->>B: buffer, flush once per frame
  end
  M-->>P: done
  P-->>B: event: done
  B->>B: mark message done
  opt user clicks Stop
    B->>P: abort (connection closes)
    P->>M: cancel upstream request
  end
Key idea. A proxy for secrets and limits, a cookie session for identity, a frame-rate renderer for smoothness, and abort, retry, and trimming for recovery.

  1. Deep dives

6.1 SSE or WebSockets

Before reading on. The interviewer asks why you did not use WebSockets. Defend SSE, then name the case where you would switch.

The traffic shape decides it. A chat turn is one request followed by a one-way stream of tokens. SSE matches that exactly: plain HTTP, no connection upgrade, no custom protocol, automatic support in proxies and CDNs, and each request stands alone, which makes load balancing and retries simple.

WebSockets add state: a long-lived connection per user, sticky routing, heartbeats, reconnect logic, and a message protocol of your own. That cost buys two-way messaging during a stream. You need it for voice, where audio flows both ways; for interrupting the model with new input mid-answer; or for multi-user sessions where other people's messages arrive unprompted. None of those is in scope here.

What separates answers: transport

WeakPicks WebSockets because chat is real time

Chooses the heavier tool without noticing the traffic is one-way per turn.

GoodPicks SSE for one-way streams

Explains that SSE fits the request-then-stream shape and runs over plain HTTP.

StrongKnows the details and the switch point

Reads SSE with fetch because EventSource cannot POST, handles buffering proxies, and names voice, interruption, and multi-user sessions as reasons to move to WebSockets.

6.2 Parsing the stream

async function streamChat(messages: Msg[], onText: (t: string) => void, signal: AbortSignal) {
  const res = await fetch("/api/chat", {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ messages }),
    signal,
  });
  if (!res.ok || !res.body) throw new HttpError(res.status);

  const reader = res.body.pipeThrough(new TextDecoderStream()).getReader();
  let buffer = "";
  for (;;) {
    const { value, done } = await reader.read();
    if (done) break;
    buffer += value;
    const events = buffer.split("\n\n");
    buffer = events.pop() ?? "";          // keep a partial event for the next chunk
    for (const raw of events) {
      const ev = parseEvent(raw);         // reads "event:" and "data:" lines
      if (ev.type === "delta") onText(ev.data.text);
      if (ev.type === "error") throw new StreamError(ev.data);
      if (ev.type === "done") return ev.data;
    }
  }
  throw new StreamError({ code: "ended_early", retryable: true });
}

Two details matter. A network chunk can end in the middle of an event, so keep the partial tail in the buffer. And a stream that ends without a done event is an error, not a success: the answer was cut off.

6.3 Rendering tokens without jank

Before reading on. The answer is 2,000 tokens of markdown with code blocks. What does naive React do on each token, and what do you do instead?

Naively, each token calls setState with a new messages array. React re-renders the transcript, and the markdown renderer re-parses the entire message on every token. The work per token grows with the message length, so the stream slows down as it gets longer.

Fixes, in order:

  • Batch to frames. onText appends to a ref. A requestAnimationFrame loop flushes the ref into state at most once per frame.
  • Isolate the streaming message. Finished messages are memoized components keyed by ID, so they never re-render. Only the last message updates.
  • Parse incrementally. Split the streaming text into completed blocks and the current block. Render completed blocks once and re-parse only the tail.
  • Tolerate broken syntax. An open code fence or a half-written table must render as text until it closes, not flicker between layouts.
  • Virtualize long transcripts. Only messages near the viewport mount.
  • Keep scroll sane. Auto-scroll only if the user is already at the bottom; if they scrolled up to read, leave them there.

Sanitize the rendered HTML. Model output is untrusted input, and a response containing a script tag or a malicious link must not execute.

6.4 State in the browser and the context limit

Before reading on. The conversation reaches the model's context limit. What should the client do?

Because the server stores nothing, the browser must manage context. Count tokens with a client-side tokenizer, or estimate by characters. Before each send, if the history plus a reserve for the answer exceeds the limit, trim: drop the oldest turns first, keeping the system prompt and the most recent turns. Show a small notice that earlier messages are no longer in the model's view. A stronger option asks the proxy to summarize old turns into one short message, which keeps more meaning in less space.

On refresh, state clears, as specified. If the product later wants chats to survive an accidental reload, sessionStorage keeps them per tab, with a limit of a few megabytes. Warn about that limit and trim code blocks first when saving.

6.5 Authentication without server state

The only server state is the session. Use the OpenID Connect authorization code flow with PKCE. The proxy exchanges the code for tokens and sets a session cookie marked HTTP-only, Secure, and SameSite. Scripts cannot read it, so a script-injection bug cannot steal it. Tokens in localStorage would be readable by any script on the page.

The session can live in a small store, or in the cookie itself, encrypted and signed, which keeps the proxy fully stateless. On a 401 mid-conversation, keep the conversation in memory, refresh the session silently if possible, and resend. Losing the user's conversation to a login redirect is a bad experience.

SameSite cookies block most cross-site request forgery. For extra safety, require a custom header on /api/chat, which browsers will not send cross-site without a CORS preflight.

6.6 Errors and recovery

FailureWhat the user seesWhat the client does
Network drops mid-streamPartial answer, marked interruptedOffer Retry, which resends the same history
429 rate limited"Too many requests, try again in N seconds"Disable Send until Retry-After
5xx or upstream timeout"Something went wrong" with RetryRetry up to a few times with backoff and jitter
Stream ends without donePartial answer, marked cut offSame as a drop
Browser offlineSend disabled, offline bannerListen for the online event
413 context too longAutomatic trim, then sendTrim oldest turns and retry once

Stopping is a user action, not an error. Abort the fetch; mark the message stopped; keep its text. The proxy cancels upstream when the connection closes, so tokens nobody reads are not billed.

What separates answers: failures

WeakAssumes the stream completes

No plan for a stream that dies halfway; errors blank the message.

GoodKeeps partial text and offers retry

Detects drops and missing done events, backs off on 429 and 5xx.

StrongDesigns every failure's UX

Maps each failure to a message and an action, cancels upstream on stop, trims context automatically, and never loses what the user already saw.

6.7 Operating the proxy

Streams hold connections for tens of seconds, so plan by concurrent connections. Raise idle timeouts on the load balancer and proxy, and disable response buffering for the chat route, or tokens arrive in bursts. Send an SSE comment line every 15 seconds during long pauses so intermediaries do not time out.

Rate limit per user in the proxy with a token bucket, and cap concurrent streams per user. Record usage counts, not content. Measure time to first token and tokens per second from the browser, since that is what users feel, and alert on error rates by type.

6.8 Component architecture

Before reading on. Sketch the component tree and say where state lives. Which components re-render on each token?
<App>
  <AuthProvider>                 session state, login redirect
    <ChatProvider>               messages, inFlight, send(), stop(), retry()
      <Header />                 static
      <Transcript>               virtualized list
        <MessageView id=... />   memoized by id + status; finished messages never re-render
        <StreamingMessage />     subscribes to the token buffer only
      </Transcript>
      <Composer />               draft text, local state; Enter to send
    </ChatProvider>
  </AuthProvider>
</App>

State lives in one store owned by ChatProvider. The streaming message reads from a separate buffer that updates once per frame, so only StreamingMessage re-renders during a stream. When the stream finishes, the buffer's text is committed to the message list once, and the message becomes a memoized MessageView.

The composer keeps its draft in local state, so typing never touches the transcript. That separation keeps input responsive during a fast stream, which is the most common complaint in chat UIs.

6.9 Accessibility and keyboard

Streaming text is hard for screen readers: announcing every token is noise. Mark the transcript as a log region that announces completed messages, and announce "response finished" at the end. Enter sends, Shift+Enter adds a line, Escape stops a stream. Focus returns to the composer after sending. Code blocks get a copy button reachable by keyboard.

These details rarely decide a round alone, and interviewers with frontend backgrounds notice them. Mentioning them briefly signals that you have shipped a real UI.

What separates answers: frontend structure

WeakOne big component

Keeps all state in one component that re-renders on every token and every keystroke.

GoodSplit state

Separates the streaming buffer and the composer draft from the message list.

StrongDesigned for scale and access

Memoizes finished messages, virtualizes the transcript, isolates the streaming component, and handles screen readers and keyboard flows.

  1. Variants

7.1 Server-side history

If the interviewer lifts the no-storage rule, add a conversations service with messages stored per user, a sidebar of past chats, and sync across devices. The browser then sends only the new message and a conversation ID.

7.2 Voice

Voice needs audio in both directions and interruption, so switch to WebSockets or WebRTC. Stream audio chunks up, stream audio and text down, and handle barge-in: the user talks while the model speaks, and the client cancels playback and the upstream generation.

7.3 File uploads

Upload files directly to storage with a presigned URL, then send a file reference in the chat request. Show upload progress separately from the answer stream.

7.4 At ten times the users

With 10 million daily users, the proxy holds about 180,000 open streams at peak. Scale the proxy horizontally behind a layer-4 or HTTP/2-aware load balancer, and watch memory per connection. Upstream quota becomes the limit: add per-user concurrency caps, a queue with honest wait messages, and degradation to a faster model for free users. None of the browser design changes.

  1. The transferable pattern

A streaming chat UI is a stateless relay plus a stateful client renderer. The server guards secrets and limits and passes bytes; the client owns state, recovery, and presentation. The same split applies to any streaming product: live logs, progress feeds, collaborative cursors. Decide which side owns state, then design the other side to be dumb and replaceable.

Review: the 30-second answer

  • A thin proxy holds the key. Auth, per-user limits, stream relay, no content stored.
  • SSE read with fetch. POST bodies, plain HTTP, WebSockets only for two-way needs.
  • Render at frame rate. Buffer tokens, memoize finished messages, parse incrementally, sanitize.
  • The browser owns context. Count tokens and trim before each send.
  • Every failure has a UX. Partial text kept, Retry offered, Stop cancels upstream.

Quiz

+Why does a "no backend storage" chat still need a backend?

The model API key must stay secret, and anything shipped to the browser is public. A proxy holds the key, authenticates users, and applies rate limits, while storing no conversation content.

+Why read the stream with fetch instead of EventSource?

EventSource only makes GET requests without a body. A chat request needs a POST with the message history in the body.

+Why does naive token-by-token rendering slow down as the answer grows?

Each token triggers a re-render and a full markdown re-parse of the growing message, so the work per token grows with the message length. Batching per frame and parsing only the tail fixes it.

+Why keep the session in an HTTP-only cookie?

Scripts cannot read an HTTP-only cookie, so a cross-site scripting bug cannot steal the session. Tokens in localStorage are readable by any script on the page.

+What should happen when the user clicks Stop?

The client aborts the fetch and keeps the text so far, marked stopped. The proxy sees the connection close and cancels the upstream request, so no more tokens are generated or billed.

+Why keep the composer's draft in local component state?

So typing updates only the composer. If the draft lived in the shared chat store, every keystroke would re-render the transcript, and typing would lag during a stream.

+How should a screen reader experience a streaming answer?

The transcript is a log region that announces completed messages, not each token, plus a short announcement when the response finishes.

Sources and further reading

NextModel Playground