OpenAI Interview: Design a Cloud IDE
A full solution to the OpenAI cloud IDE question: sandbox isolation for untrusted code, capacity math that shows memory binds, warm pools, the workspace lifecycle, file persistence and snapshots, terminal streaming, sharing, and the SSH variant.
By The Forward Deployed editorial teamReviewed
Part of the OpenAI system design question bank. The question is representative of the round. The analysis and solution are this site's own.
Problem statement
A user opens a project and gets an editor, a file tree, and a terminal within seconds. Commands run on a machine in the cloud, and output streams back as fast as a local terminal. Installed packages last for the session; files last forever. A current rotation of this question drops the browser: users connect over SSH to workspaces on a host fleet, and a scheduler manages the hosts.
Clarifying questions
- Who are the users? Anyone with an account, which means some of them will run hostile code.
- Session length? Interactive coding sessions, with a cap such as 12 hours and an idle timeout.
- What persists? Files always. System packages and running processes only during the session, unless the user pays for full persistence.
- Resources per workspace? For practice: 1 vCPU and 2 GB of memory by default.
- Scale? 100,000 workspaces running at peak.
- Collaboration? Shared view and edit access, without live co-editing of the same file.
- Latency? Workspace ready in under 5 seconds; terminal echo under 100 ms.
What makes a cloud IDE hard
Every user runs arbitrary code on your machines. Some will try to escape the sandbox, read other users' files, mine cryptocurrency, or attack your internal network. The sandbox is a security boundary, not a convenience.
Users also expect it to feel local. A new workspace must be ready in seconds, not the minutes a fresh virtual machine and toolchain take. The terminal must echo keystrokes fast enough that typing does not feel remote.
And resources are expensive. 100,000 running workspaces at 2 GB each is 200 TB of memory. Most of them are idle at any moment, so the fleet must pack them tightly without letting one user starve another.
So the driving tension is isolation versus density and speed. Stronger isolation costs startup time and memory; tighter packing and warm reuse risk leaking between users.
flowchart LR B([Browser: editor, files, terminal]):::user <-->|WebSocket| GW[Gateway]:::svc GW --> CP[Control plane:<br/>create, start, stop, share]:::svc GW <--> HA[Host agent]:::svc <--> SB[Sandbox microVM<br/>user code runs here]:::store SB <--> VOL[(Workspace volume)]:::store classDef user fill:#e6efec,stroke:#315e55,color:#171717; classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
Key idea. The sandbox is the product and the security boundary at once. Design it for hostile users first, then make it fast with warm pools and dense with overcommit.
Key concepts
Containers, microVMs, and user-space kernels
A container isolates processes with kernel namespaces and cgroups, but all containers on a host share one kernel. A kernel bug can let code escape. A microVM runs a tiny virtual machine with its own kernel; escaping requires breaking the hypervisor, a much smaller target, at a cost of some memory and startup time. A user-space kernel intercepts system calls in a separate process and implements them itself, which shrinks the host kernel's exposure without full virtualization.
cgroups
Linux control groups cap CPU, memory, process count, and I/O per sandbox. They stop a fork bomb or a memory leak from hurting neighbors.
Pseudo-terminals
A PTY is the kernel device that makes a program believe it is talking to a terminal. The shell writes to it; something reads the bytes and sends them to the browser, where a terminal emulator renders them.
Warm pools
Sandboxes started in advance, idle, waiting to be claimed. Claiming one skips boot and image setup.
Overcommit
Promising more CPU in total than the host has, because most workspaces are idle. Safe for CPU, which can be shared by time slicing. Dangerous for memory, which cannot.
Key idea. MicroVMs for hostile code, cgroups for fairness, PTYs for terminals, warm pools for speed, and overcommit only where it is safe.
- Requirements
Before reading on. List the requirements. Which property do you protect first, and what resource will bind the fleet?
1.1 Functional requirements
- Create a workspace from a template or repository; open it in the browser.
- Edit files; create and delete folders.
- Run commands with live output; stop running processes.
- Install packages that persist for the session.
- Share a workspace as viewer or editor.
- Stop idle workspaces and resume them later with files intact.
1.2 Non-functional requirements
- Isolation. No user can read another's files or affect another's processes, even with hostile code.
- Startup under 5 seconds at p95.
- Terminal latency under 100 ms for echo at p95.
- Durability of files: at most a few minutes of changes lost on a host failure.
- Availability 99.9% for the control plane.
1.3 The constraint versus the property
Isolation is the property. A sandbox escape is a breach of every user's data. Memory is the constraint. Section 2 shows it binds before CPU, so it sets placement, pricing, and idle policy.
- Back-of-the-envelope estimation
2.1 Fleet size
100,000 workspaces × 1 vCPU × 2 GB.
- CPU: most workspaces idle; overcommit 4 times. On 64-vCPU hosts: 100,000 / (64 × 4) ≈ 391 hosts.
- Memory: cannot be safely overcommitted. On 256 GB hosts, reserving 16 GB for the host itself: 100,000 × 2 GB / 240 GB ≈ 834 hosts.
Memory binds at more than twice the CPU count. That fact drives decisions: price by memory, pack by memory, and stop idle workspaces aggressively to free it.
2.2 Warm pool
If 20,000 workspaces start per hour at peak, that is about 5.6 per second. With a refill time of 30 seconds for a new sandbox, the pool needs about 5.6 × 30 ≈ 170 ready sandboxes per base image, plus headroom for bursts. A few popular images cover most starts; rare images start cold.
2.3 Storage
If the average workspace holds 500 MB of user files, 1 million workspaces hold 500 TB. Package installs live in a session layer and are discarded, so they do not count.
2.4 Terminal traffic
Most terminals are idle. A busy build may print 1 MB per second. Batching output into 10 ms frames keeps message counts low and adds no noticeable delay.
Key idea. Memory binds at about 830 hosts versus 390 for CPU. Design, price, and schedule around memory.
- API design
Before reading on. What should the browser talk to directly: the control plane, the host, or the sandbox?
The control plane, over HTTPS, for lifecycle actions. The gateway, over one WebSocket, for terminal and file traffic, which the gateway forwards to the right host. The browser never reaches a host or sandbox directly, so hosts can stay on a private network.
3.1 Control plane
POST /v1/workspaces {template | repo_url, size} -> {workspace_id}
POST /v1/workspaces/:id/start -> {status: "starting", connect_url, token}
POST /v1/workspaces/:id/stop
GET /v1/workspaces/:id -> {status, host?, last_active_at, size}
POST /v1/workspaces/:id/shares {user_id, role: viewer|editor}
DELETE /v1/workspaces/:id/shares/:user_id3.2 Session channel (WebSocket via gateway)
fs.list {path} fs.read {path} fs.write {path, content, base_hash}
fs.mkdir {path} fs.delete {path}
proc.start {cmd, cwd} -> {pid, pty_id}
pty.input {pty_id, bytes} pty.output {pty_id, seq, bytes} (server -> client)
pty.resize {pty_id, cols, rows}
proc.signal {pid, signal: "INT" | "KILL"}fs.write carries the hash of the version the editor loaded, so a write from a stale tab is rejected instead of overwriting.
- Data model
workspaces (id, owner_id, template, size, status: stopped|starting|running|stopping,
host_id?, sandbox_id?, volume_id, last_active_at, created_at)
shares (workspace_id, user_id, role)
hosts (id, zone, cpu_total, mem_total, mem_reserved, status, last_heartbeat)
sandboxes (id, host_id, image, state: warm|claimed|draining, workspace_id?)
volumes (id, workspace_id, snapshot_ref, size, last_snapshot_at)
sessions (token, workspace_id, user_id, role, expires_at)
- High-level design
5.1 One big shared server
All users get accounts on one large Linux machine. It starts instantly and isolates nothing: users see each other's processes, one heavy build slows everyone, and a kernel exploit owns the machine.
5.2 Fix 1: a container per workspace
Each workspace runs in its own container with cgroup limits. Processes and files separate, and one user's load is capped. The kernel is still shared with every other container on the host.
5.3 Fix 2: microVMs for untrusted code
Run each workspace in a microVM with its own kernel. A kernel exploit now stays inside the VM. Block the VM's network from internal services and the cloud metadata endpoint.
5.4 Fix 3: a warm pool
Booting a microVM and preparing an image takes tens of seconds. Keep booted sandboxes ready per base image. Starting a workspace claims one and mounts the user's volume: about a second.
flowchart LR REQ[Start workspace]:::svc --> SCH[Scheduler]:::svc SCH -->|claim| WP[(Warm pool<br/>image: python-3.12)]:::new WP --> SB[Sandbox]:::store VOL[(User volume)]:::store -->|mount| SB PM[Pool manager]:::new -->|refill| WP classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717; classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;
5.5 Fix 4: persistent volumes and snapshots
User files live on a volume attached to the sandbox, with snapshots to object storage every few minutes and on stop. A host failure loses minutes at most.
5.6 Fix 5: a gateway for terminals and files
The browser keeps one WebSocket to a gateway, which routes to the host agent that owns the sandbox. The agent bridges the PTY and file operations.
5.7 The composed design
sequenceDiagram autonumber actor U as User participant CP as Control plane participant S as Scheduler participant H as Host agent participant SB as Sandbox participant G as Gateway U->>CP: start workspace CP->>S: place (size 2 GB, image python) S->>H: claim warm sandbox on host with free memory H->>SB: mount volume from latest snapshot, start agent CP-->>U: connect_url + session token U->>G: WebSocket (token) G->>H: route to sandbox U->>G: pty.input "pytest\n" G->>H: forward H->>SB: write to PTY SB-->>H: output bytes H-->>G: pty.output (batched 10 ms) G-->>U: render in terminal
Key idea. MicroVMs for isolation, cgroups for fairness, warm pools for speed, snapshotted volumes for files, and a gateway that routes each session to its host.
- Deep dives
6.1 Isolation in depth
Before reading on. A user runs code designed to reach other workspaces. List every path they might try and what blocks each.
| Attack path | Defense |
|---|---|
| Kernel exploit from inside the sandbox | MicroVM or user-space kernel; host kernel not exposed |
| Reading another user's volume | Each volume attached only to its own VM; no shared file system |
| Network to internal services | VM network allows only the internet, through an egress proxy; internal ranges blocked |
| Cloud metadata endpoint for credentials | Blocked at the VM network layer |
| Exhausting host resources | cgroup and VM limits on CPU, memory, disk, process count |
| Leftover data in reused sandboxes | Warm sandboxes are never reused across users; destroyed after one session |
| Abuse such as mining or spam | Egress rate limits, CPU usage anomaly detection, account signals |
Warm pools deserve a sentence: a warm sandbox is fresh and unused. It is claimed once and destroyed after the session. Reusing a sandbox between users would leak files, processes, or cached credentials.
What separates answers: isolation
WeakContainers and trust
Runs user code in plain containers on a shared kernel and never mentions the network.
GoodStronger sandboxes and limits
Uses microVMs or a user-space kernel, cgroup limits, and blocks internal networks.
StrongThreat-modeled
Walks each attack path, blocks the metadata endpoint, never reuses sandboxes across users, and adds abuse detection for mining and spam.
6.2 The workspace lifecycle
stateDiagram-v2 [*] --> stopped stopped --> starting: user opens starting --> running: sandbox claimed, volume mounted running --> stopping: idle 30 min, user stops, or 12 h cap stopping --> stopped: final snapshot saved, sandbox destroyed running --> recovering: host lost recovering --> starting: restore from last snapshot
Idle detection counts no terminal input, no file writes, and no running foreground processes for 30 minutes. Before stopping, warn the user in the UI. On stop, take a final snapshot and destroy the sandbox, which frees its memory, the binding resource.
6.3 Files: volumes, snapshots, and packages
Before reading on. A user installs 800 MB of packages, then the workspace stops. What survives, and why?
Split the file system into layers. The base image is read-only and shared. The user's project directory is a persistent volume, snapshotted to object storage. Everything else, including system packages, lives in a session layer written over the base image and discarded on stop.
So the project files survive. The 800 MB of system packages do not, unless they were installed into the project directory, such as a local virtual environment. That keeps storage small and restarts clean. Offer a setup script in the project that reinstalls dependencies on start, and a paid option that snapshots the whole disk.
Snapshots are incremental: only changed blocks since the last snapshot are uploaded. Every few minutes during activity, and always on stop.
6.4 Terminal streaming
Before reading on. A build prints two million lines. The user's browser is on a slow connection. What should happen?
The path is PTY, host agent, gateway, WebSocket, browser. Keep it short, with no queue in between. The agent reads PTY output and batches it into frames every 10 ms or 16 KB, whichever comes first, tagged with a sequence number.
When the client cannot keep up, do not block the process. Blocking the PTY would pause the user's build because their Wi-Fi is slow. Instead, the agent keeps a bounded buffer. When it overflows, the agent drops old output and sends a marker, and the terminal emulator shows the latest screen. The full log can still be written to a file if the user wants it.
Keystrokes travel the other way. Echo depends on a round trip to the sandbox, so place sandboxes in the region nearest the user.
On reconnect, the client sends the last sequence number it rendered. The agent keeps recent scrollback, such as the last 10,000 lines, and replays from there. Running processes are not affected by a browser disconnect.
What separates answers: terminals
WeakOutput through a message queue
Routes terminal bytes through a queue or database, adding latency and cost.
GoodDirect streaming path
Streams PTY output through the gateway over WebSocket with batching.
StrongFlow control and reconnect
Drops old output instead of blocking the process when clients lag, sequence-numbers frames, replays scrollback on reconnect, and places sandboxes near users for echo latency.
6.5 Scheduling and placement
Place workspaces by free memory, not CPU. Spread workspaces of one organization across hosts, so one host failure does not take down a whole team. Keep warm pools per zone. When a host drains for maintenance, stop its workspaces with a warning and restart them elsewhere from their snapshots.
6.6 Sharing
Sharing is an ACL on the workspace. An editor gets a session token that allows file writes and terminal input. A viewer's token allows file reads and a read-only copy of the terminal stream. The gateway checks the role on every message type. Removing a share revokes the user's tokens and closes their active sessions.
Concurrent edits to the same file by two editors are handled by the base_hash check: the second save gets a conflict and sees a diff. Live co-editing would need operational transforms or CRDTs, and is out of scope unless asked.
6.7 Previewing a web server inside the workspace
Before reading on. A user runs a web app on port 3000 inside their workspace and wants to open it in a browser tab. How do you expose it safely?
Give each workspace port a unique URL, such as https://3000-ws8f2k.preview.example.dev. A preview proxy receives the request, checks that the user's session may access that workspace, looks up the workspace's host, and forwards the request through the host agent to the port inside the sandbox.
Three safety details:
- A separate domain. Previews run untrusted code that serves pages to the user's browser. Put them on a different registrable domain from the IDE, so a malicious preview cannot read the IDE's cookies.
- Authentication by default. Previews are private unless the user explicitly shares a port publicly.
- No direct routes. The proxy talks only to host agents; sandboxes never accept connections from outside.
WebSocket traffic from the preview, such as hot reload, goes through the same proxy.
6.8 Cost and idle policy
Memory binds, so idle workspaces cost money for nothing. For practice: if the average workspace runs 3 hours a day but is active only 1 hour, a 30-minute idle timeout cuts running time to about 1.5 hours, halving memory cost. Suspending to disk, where the VM's memory is saved and restored in seconds, keeps the user's processes alive across short breaks while freeing memory.
Tie the policy to plans. Free workspaces stop after 30 minutes idle and have smaller limits; paid ones get longer timeouts and optional always-on. Show users the timeout and a warning before it fires.
What separates answers: exposure and cost
WeakOpens ports to the internet
Lets sandboxes accept traffic directly, and keeps idle workspaces running.
GoodAuthenticated preview proxy
Routes previews through a proxy with per-workspace auth, and stops idle workspaces.
StrongIsolated previews and memory-aware policy
Serves previews from a separate domain, keeps sandboxes unreachable, suspends idle workspaces to disk, and ties timeouts to plans with the cost arithmetic.
- Variants
7.1 SSH to a host fleet
Users connect over SSH instead of a browser. An SSH gateway authenticates the user's key, looks up the workspace, starts it if needed, and forwards the session to the sandbox. Port forwarding lets users reach web servers running inside their workspace. The scheduler, sandbox, and volume design are unchanged.
7.2 Long-running jobs
If sessions can run for days, workspaces behave like servers. Idle timeouts give way to explicit quotas, and host maintenance needs live migration or scheduled restarts with notice.
7.3 GPU workspaces
GPUs are scarce and cannot be overcommitted. Give GPU workspaces a separate pool, shorter idle timeouts, and a queue when capacity runs out.
7.4 At ten times the users
At a million running workspaces, memory reaches about 2 PB and hosts about 8,000. Run the fleet in cells of a few hundred hosts, each with its own scheduler and warm pools, and route new workspaces to cells with capacity. Suspend-to-disk and aggressive idle stops become major cost levers, and image caching per cell keeps start times low.
- The transferable pattern
A cloud IDE is multi-tenant compute for untrusted code: a strong sandbox, a warm pool to hide startup, a persistent layer separated from a disposable one, and a direct streaming path for interaction. The same pattern runs online judges, notebook services, CI runners, and AI agent sandboxes.
Review: the 30-second answer
- MicroVMs for hostile code. Block internal networks and metadata; never reuse sandboxes across users.
- Memory binds. About 830 hosts versus 390 for CPU; place and price by memory, stop idle workspaces.
- Warm pools per image. Start in about a second.
- Persistent project volume, disposable session layer. Incremental snapshots every few minutes.
- Direct PTY streaming. Batched frames, drop instead of block, replay on reconnect.
Quiz
+Why are plain containers not enough for a public cloud IDE?
All containers on a host share one kernel. A kernel vulnerability lets code escape to the host and reach other users. MicroVMs or user-space kernels remove that shared-kernel exposure.
+Why does memory, not CPU, set the fleet size?
Idle workspaces use little CPU, so CPU can be overcommitted several times. Memory cannot be safely overcommitted, so 100,000 workspaces at 2 GB need about 200 TB of real memory.
+Why must a warm sandbox never be reused for another user?
A used sandbox can hold the previous user's files, processes, or credentials in memory or on disk. Destroying it after one session guarantees nothing leaks.
+What should happen when a client cannot keep up with terminal output?
Drop old output and show the latest screen. Blocking the PTY would pause the user's process because of their network speed.
+Why do installed system packages disappear when a workspace stops?
They live in a disposable session layer over the base image. Only the project volume persists, which keeps storage small and makes every restart clean.
+Why serve workspace previews from a separate domain?
Previews run untrusted code that serves pages in the user's browser. A separate domain keeps that code from reading the IDE's cookies or acting as the IDE.
+How much can an idle timeout save when memory binds?
If workspaces run 3 hours a day but are active only 1, a 30-minute timeout cuts running time to about 1.5 hours, roughly halving memory cost.
Sources and further reading
- Firecracker is an open-source microVM monitor built for multi-tenant, short-lived workloads.
- gVisor documents a user-space kernel that intercepts system calls for container isolation.
- Linux cgroups v2 explains resource limits per group of processes.
- Sandboxed execution and tenant isolation return in multi-tenant CI/CD. Security and compliance covers isolation boundaries in customer environments.
