The Forward Deployed

OpenAI Interview: Design a Model Playground with Presets

A full solution to the OpenAI playground question: the product problem, stack choices, presets data model, API, streaming completions, frontend state for prompt and completion ranges, undo, concurrency, and export to code.

By Reviewed

Part of the OpenAI system design question bank. The question is representative of the round. The analysis and solution are this site's own.

Problem statement

The page has a large editor, a settings panel, and a Submit button. The completion appears after the prompt, highlighted, as the model writes it. A developer iterates, saves presets they like, and loads them later. The round is usually a whiteboard discussion with a reference mockup of the finished page.

Clarifying questions

  • Who are the users? Developers with API accounts, experimenting before they write code.
  • What settings? Model, temperature, maximum tokens, top-p, stop sequences, and frequency and presence penalties.
  • Can presets be shared? Assume private by default, with a read-only share link as a stretch goal.
  • Are runs saved? Assume usage is recorded for billing and limits, and run output is not stored by default.
  • Scale? For practice: 50,000 daily users, 40 runs and 3 saves each per day.
  • Mobile? Desktop first.

What makes a playground hard

The backend is small. The difficulty sits in two places.

The editor holds two kinds of text in one document. The user's prompt and the model's completion sit side by side, the completion must stay highlighted, and the user can then edit anywhere, including inside the completion, and submit again. The UI must track which characters came from the model while text keeps changing.

And the product exists to move experiments into code. A preset that forgets one setting, such as a stop sequence, produces different results in the developer's own code, and they lose trust in the tool.

So the driving tension is free-form editing versus precise state. The user edits freely; the system must always know what the prompt is, what the model wrote, and exactly which settings produced it.

flowchart LR
  B([Browser: editor + settings]):::user --> API[Playground backend]:::svc
  API --> M[Model API]:::store
  API --> DB[(Presets, usage)]:::store
  M -->|tokens| API -->|SSE| B
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
Key idea. Start from the problem: shorten the loop from idea to working API call. That goal decides what state to track and what a preset must freeze.

Key concepts

Completion versus chat

A completion model continues a text. The prompt is plain text, and the output is the most likely continuation. That is why the UI shows the completion appended to the prompt, in one document.

Sampling settings

Temperature scales randomness: 0 is nearly deterministic, higher values vary more. Top-p limits sampling to the smallest set of tokens whose probabilities add up to p. Maximum tokens caps output length. Stop sequences end generation when the model writes them. Penalties discourage repetition. Every one of them changes the output, so every one belongs in a preset.

Text ranges

A range is a start and end offset in the document. Tracking the completion as a range, and updating it as the user edits, is how the editor knows what to highlight.

Optimistic concurrency

Each preset has a version number. A save sends the version it loaded. If the server's version moved on, the save is rejected and the client resolves the conflict.

Key idea. Completion models continue text, so the UI is one document with a tracked range. Every sampling setting changes output, so a preset must freeze all of them.

  1. Requirements

Before reading on. Write the problem statement in one sentence, then the requirements. Which failure would make developers stop using the tool?

1.1 Functional requirements

  • Enter a prompt and receive a streamed completion appended after it, highlighted.
  • Adjust model and sampling settings, validated per model.
  • Stop a completion in progress.
  • Undo the last completion.
  • Save the prompt and settings as a named preset; list, load, rename, delete presets.
  • Export the current state as an API call in curl or SDK code.

1.2 Non-functional requirements

  • Fast first token. The completion starts within about a second.
  • Smooth editing. Typing never lags, even while a completion streams.
  • Faithful presets. Loading a preset reproduces exactly the settings that were saved.
  • Security. No API key in the browser; per-user rate limits.

1.3 The constraint versus the property

Faithful presets are the property. The tool's value is carrying an experiment into code. Editor complexity is the constraint. Tracking prompt and completion in one freely edited document shapes the frontend design.

  1. Back-of-the-envelope estimation

  • Runs: 50,000 × 40 = 2 million a day, about 23 per second on average, perhaps 100 at peak.
  • Open streams at peak: 100 per second × 8 seconds average = 800. Trivial for a small proxy fleet.
  • Saves: 150,000 a day, under 2 per second. One Postgres instance, with a replica for failover.
  • Preset size: a prompt of a few KB plus a few hundred bytes of settings. A million presets is a few GB.

Say plainly that scale is not the problem, and spend the time on state and correctness.

  1. API design

Before reading on. The user loads a preset, edits it, and meanwhile saves the same preset from another tab. What should the second save do?

Reject it with a conflict, and let the user choose: overwrite, or save as a new preset. Silent overwrite would destroy the first tab's work.

POST   /api/completions
  {model, prompt, params: {temperature, max_tokens, top_p, stop[], ...}}
  -> text/event-stream: delta {text} ... done {finish_reason, usage}

GET    /api/models              -> [{id, context_window, param_ranges}]

GET    /api/presets             -> [{id, name, updated_at}]
POST   /api/presets             {name, prompt, model, params} -> {id, version}
GET    /api/presets/:id         -> {id, name, prompt, model, params, version}
PUT    /api/presets/:id         {name, prompt, model, params, version}
                                -> 200 {version+1} | 409 {current}
DELETE /api/presets/:id
POST   /api/presets/:id/share   -> {share_token}   (read-only copy link)
GET    /api/export?format=curl|python|node  (body: current state) -> code

GET /api/models drives the settings panel: ranges and defaults come from the server, so a new model does not need a frontend release.

  1. Data model

users    (id, email, org_id, created_at)
presets  (id, owner_id, name, prompt_text, model, params_json,
          version, created_at, updated_at, deleted_at)
          index (owner_id, updated_at desc)
shares   (token, preset_id, created_at, revoked_at)
usage    (user_id, day, runs, prompt_tokens, completion_tokens)

Settings live in one JSON column validated against a schema per model. New settings appear often, and a column per setting would need a migration each time.

  1. High-level design

5.1 A static page calling the model

The page calls the model API directly with a key in the JavaScript bundle. It works for a demo and leaks the key to every visitor.

5.2 Fix 1: a backend proxy

A small backend holds the key, authenticates users, rate limits them, validates settings, streams completions, and records usage.

flowchart LR
  B([Browser]):::user --> P[Backend: auth, validate, stream]:::new --> M[Model API]:::store
  P --> U[(Usage)]:::new
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;

5.3 Fix 2: presets with versions

A presets table with a version column. Saves send the loaded version; stale saves get 409.

5.4 Fix 3: an editor that tracks the completion range

The frontend keeps the text and the completion's range in state, updates the range on every edit, and highlights it.

5.5 The composed design

sequenceDiagram
  autonumber
  actor U as Developer
  participant E as Editor state
  participant B as Backend
  participant M as Model API
  U->>E: click Submit
  E->>E: snapshot text for undo, completion range starts at end of text
  E->>B: POST /api/completions {model, prompt=text, params}
  B->>B: auth, rate limit, validate params for model
  B->>M: stream request
  loop tokens
    M-->>B: chunk
    B-->>E: delta
    E->>E: append to text, extend range (batched per frame)
  end
  B-->>E: done + usage
  U->>E: click Save preset
  E->>B: PUT /api/presets/:id {prompt, params, version}
  B-->>E: 200 or 409

  1. Deep dives

6.1 The editor state

Before reading on. The completion is highlighted. The user types three characters in the middle of it. What happens to the highlight, and what is the prompt for the next Submit?

State:

type EditorState = {
  text: string;                               // prompt + accepted completions
  completion: { start: number; end: number } | null;
  model: string;
  params: Params;
  status: "idle" | "streaming" | "error";
  preset: { id: string; version: number; dirty: boolean } | null;
  undo: Snapshot[];                           // text + completion range before each run
};

On each edit, transform the range. An edit fully before the range shifts it. An edit after it does nothing. An edit inside it splits the model's text from the user's: the simplest rule ends the highlight at the edit, so everything after the edit counts as user text. Editors with a mark or decoration API do this automatically: the completion is a mark, and marks move with edits.

The next Submit sends the whole text as the prompt. Once a completion is accepted into the text, it is part of the prompt. That is how completion playgrounds work: iterate by editing the continuation and continuing again.

Undo restores the snapshot from before the last run: text and range together.

What separates answers: editor state

WeakStores one string

Cannot highlight the completion after edits or undo a run cleanly.

GoodTracks the completion range

Keeps a range, shifts it on edits, and snapshots before each run for undo.

StrongUses the editor's mark model

Represents the completion as a mark that moves with edits, defines what happens on edits inside it, and batches streaming updates so typing never lags.

6.2 Streaming into an editor the user can touch

While the completion streams, the user may scroll, select, or even type. Two options:

  • Lock the editor during streaming. It is simple and predictable. Show a Stop button.
  • Allow edits outside the completion's range. Append tokens at the range's end, not at the cursor, and transform the range as usual.

Recommend locking for v1: the run takes seconds, and edits during it are rare. Batch token appends per animation frame so a fast stream never blocks input handling.

6.3 Settings validation and model changes

Each model has its own ranges: context window, maximum tokens, supported settings. The backend returns them from /api/models, and the panel renders from that data. When the user switches models, clamp settings to the new ranges and show what changed. Before submit, count prompt tokens on the client, and warn if prompt plus maximum tokens exceeds the context window. The backend validates again, because the client can be wrong.

6.4 Presets that reproduce

Before reading on. A developer exports a preset to code and gets different outputs than in the playground. List the causes.

Missing settings: a stop sequence, top-p, or a penalty left at a playground default that the code does not set. A different model version: the playground used an alias that has since moved. Randomness: temperature above zero gives different outputs by design.

So a preset stores every setting explicitly, including defaults, and pins the exact model version. The export writes every setting into the code. If temperature is above zero, the export includes a comment that outputs vary run to run, and shows the seed parameter if the API supports one.

6.5 Concurrency

Each save sends the version it loaded, and the server updates with WHERE id = ? AND version = ?. If no row matches, return 409 with the current preset. The client shows "This preset changed elsewhere" with options to overwrite, save as new, or reload. Autosave is tempting and multiplies conflicts; save explicitly, with a dirty indicator and a warning before leaving the page.

6.6 Export to code

The export renders the current state, not the saved preset, into curl, Python, or Node code with every setting explicit. The API key appears as an environment variable reference. This feature is small, and it is the product's reason to exist: an experiment that works becomes code in one click.

6.7 Layout and theme

The interviewer may ask about layout. A wide editor on the left, settings on the right, Submit and token count below the editor, and presets in a dropdown at the top. Completion text highlighted with a background color that meets contrast guidelines in both light and dark themes. Keyboard shortcuts: submit with Ctrl+Enter, stop with Escape.

6.8 Rate limits, usage, and cost display

Before reading on. A user sets maximum tokens to 4,000 and clicks Submit 20 times in a minute. What should the backend do, and what should the user see?

Limit in tokens per minute, per user and per organization, because one request can cost a hundred times another. At admission, reserve prompt tokens plus maximum tokens against the user's budget; refund what was not used when the stream ends. If the reservation does not fit, return 429 with the time until the budget refills, and show it next to the Submit button.

Show cost before and after each run. Before: estimated prompt tokens from the client-side count, and the maximum output. After: actual usage from the done event. Developers use the playground to learn what their production calls will cost, so this is a feature, not a nicety.

Record usage per run for billing, with the model, token counts, and the user. Aggregate per organization per day for dashboards.

6.9 The whiteboard walkthrough

The round is diagram-driven. A clear order for the board:

  1. The problem, in one sentence, top left.
  2. Three boxes: browser, backend, model API. Label the key on the backend.
  3. The two main flows as numbered arrows: run (streamed) and save preset.
  4. The data model: users, presets, runs, usage.
  5. The frontend state object, next to the browser box.
  6. Then the deep dives the interviewer picks: editor ranges, concurrency, export.

Talk while drawing, and state every decision as a choice with a reason: Postgres because the data is small and relational; SSE because the stream is one-way; one JSON column for settings because models change often.

flowchart LR
  subgraph Browser
    ED[Editor state:<br/>text, completion range, params, preset]:::user
  end
  subgraph Backend
    API[API: auth, limits, validation]:::svc
    PG[(Postgres: users, presets, runs, usage)]:::store
  end
  ED -->|1. POST /completions| API
  API -->|2. stream| MA[Model API]:::store
  MA -->|3. tokens| API -->|4. SSE| ED
  ED -->|5. PUT /presets with version| API --> PG
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;

What separates answers: ownership

WeakLists technologies

Names a stack without reasons and never states the problem.

GoodJustifies each choice

States the problem first and gives one reason per technical choice.

StrongDesigns for the developer's workflow

Adds token-based limits with visible cost, reproducible presets, export to code, and a whiteboard flow the interviewer can follow step by step.

  1. Variants

7.1 Chat models

For chat models, the editor becomes a list of messages with roles. A preset stores the system message and example turns. The streaming and preset logic stays the same.

7.2 Comparing settings

Run the same prompt with two settings side by side. Each run is an independent stream; the UI shows two panes, and a preset can store both.

7.3 Saving runs

If runs are saved, store the output with the exact prompt and settings, so a developer can compare runs later. Mind storage and privacy: prompts may contain customer data, so give users a clear retention setting.

7.4 At ten times the users

At 500,000 daily users, runs reach about 230 per second on average. The backend scales horizontally; Postgres gains read replicas for preset lists. The real limit is model API quota per organization, which needs fair sharing between users of the same organization and a queue with visible wait times at peaks.

  1. The transferable pattern

A playground is a precise state machine behind a free-form editor, plus a proxy that guards the expensive resource. The same pattern appears in query consoles, notebook tools, and API explorers. Pin down exactly what the output depends on, store all of it, and make the path from experiment to production one step.

Review: the 30-second answer

  • Problem first. Shorten the loop from idea to working API call.
  • A small backend. Holds the key, validates settings per model, streams, meters.
  • Editor state with a completion range. Marks move with edits; snapshots enable undo.
  • Presets freeze everything. Every setting explicit, model version pinned, versioned saves.
  • Export to code. The feature that justifies the product.

Quiz

+Why does the completion appear appended to the prompt in one editor?

A completion model continues the given text. Showing the continuation in place makes it natural to edit the result and continue again, which is how people iterate.

+How does the editor know what to highlight after the user edits?

It tracks the completion as a range or mark and transforms it on every edit: shift it for edits before, leave it for edits after, and split or end it for edits inside.

+Why store every setting in a preset, including defaults?

Because defaults can differ between the playground and the developer's code, and they can change over time. Storing every value explicitly makes the preset reproduce the same behavior anywhere.

+What does the version column on presets prevent?

Lost updates. A save only succeeds if the preset still has the version the editor loaded; otherwise the server returns a conflict instead of silently overwriting another tab's work.

+Why fetch model settings ranges from the backend?

So a new model or changed limits only needs a backend change. The settings panel renders from data instead of hard-coded ranges.

+Why reserve maximum tokens at admission for rate limiting?

The backend cannot know the output length in advance. Reserving the maximum bounds the cost, and refunding unused tokens afterward keeps the limit fair.

+Why show the estimated cost before a run?

Developers use the playground to learn what their production calls will cost. Showing cost before and after each run makes that visible.

Sources and further reading

NextReal-Time AI Feature at Scale