The Forward Deployed

OpenAI Interview: Design a Multi-Tenant CI/CD System

A full solution to the OpenAI CI/CD question: what exactly-once can honestly mean, idempotent triggers, workflow parsing, job state transitions, leases with fencing, fair scheduling across tenants, image caching, isolation and secrets, and live logs.

By Reviewed

Part of the OpenAI system design question bank. The question is representative of the round. The analysis and solution are this site's own.

Problem statement

A developer pushes code. The git host calls a webhook with the repository ID and commit hash. The service reads a YAML workflow at that commit, creates a run with a linear chain of jobs, schedules each job onto a worker that runs it in a container, streams logs to the UI, and records results. The service is multi-tenant: many companies share the worker fleet.

Clarifying questions

  • Workflow shape? A linear chain of jobs; job 2 starts only after job 1 succeeds. Parallel jobs are a follow-up.
  • What runs? Arbitrary customer code: builds, tests, and deploys.
  • What does exactly once mean here? Ask. Usually: each job's result is recorded once, and a crash never leaves two copies of a job both acting.
  • Scale? For practice: 5,000 tenants, 200,000 runs a day, 5 jobs per run, 6 minutes per job on average.
  • Secrets? Jobs need tenant secrets such as deploy keys.
  • Log retention? 90 days.

What makes CI/CD hard

Three problems carry the round.

Exactly-once is impossible as usually stated. A worker can finish a deploy and crash before it reports success. The system cannot tell that from a crash before the deploy started. Any design that claims a queue delivers exactly once is wrong, and the interviewer knows it.

Customer code is hostile by default. A build step can try to read other tenants' secrets, escape its container, or attack the network.

And tenants compete. One company pushing 500 commits in an hour must not freeze everyone else's builds.

So the driving tension is throughput versus safety guarantees. Workers must run as many jobs as possible across tenants, while each job's effect happens once, each tenant stays isolated, and fairness holds.

flowchart LR
  GH[Git host]:::user -->|webhook: repo, commit| TR[Trigger service]:::svc
  TR --> WF[Workflow service:<br/>parse YAML, create run + jobs]:::svc --> DB[(Runs, jobs)]:::store
  SCH[Scheduler: fair queues]:::svc --> DB
  W[Workers]:::svc -->|lease jobs| SCH
  W --> LOGS[Log service]:::svc --> UI([Web UI]):::user
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
Key idea. Say early that delivery is at least once. Exactly-once is an effect you build with leases, fencing, and idempotent steps.

Key concepts

At-least-once plus idempotency

A message queue can guarantee a job is delivered at least once. If processing the same job twice has the same effect as once, the combination behaves like exactly once. That is the only honest way to get it.

Leases and fencing tokens

A worker leases a job for a limited time and renews it with heartbeats. Each lease increments an attempt number, the fencing token. Every write about the job carries it, and stale attempts are rejected.

Idempotent external effects

A deploy to a customer's cluster happens outside your system. Pass an idempotency key, such as run ID plus job name, to the target, or check its state first, so a repeated deploy of the same commit does nothing.

Weighted fair queuing

Keep a queue per tenant and pick the next job so that each tenant's share of workers matches its weight, regardless of how many jobs it has queued.

Key idea. At-least-once delivery, fenced leases, idempotent effects, and fair queues per tenant.

  1. Requirements

Before reading on. Write the requirements, then define exactly-once in one honest sentence.

1.1 Functional requirements

  • Receive push webhooks; create one run per repository, commit, and workflow.
  • Parse the workflow file at that commit.
  • Run jobs in order, each in a container, passing results forward.
  • Show live status and logs; keep logs for 90 days.
  • Cancel runs; retry failed jobs manually.
  • Inject tenant secrets into jobs.

1.2 Non-functional requirements

  • Exactly-once effect per job: one recorded result per job, and no two workers acting on the same job at once.
  • Fault tolerance: worker, scheduler, and database failures do not lose runs.
  • Isolation: no tenant can read another's code, secrets, or logs.
  • Fairness: one tenant's burst cannot starve others.
  • Horizontal scale: add workers to add capacity.

1.3 The constraint versus the property

Correct, isolated execution is the property. Worker capacity is the constraint. It is shared across tenants, it is the main cost, and it forces fair scheduling and caching.

  1. Back-of-the-envelope estimation

  • Jobs per day: 200,000 runs × 5 = 1 million.
  • Average arrival: 1 M / 86,400 ≈ 11.6 jobs per second; peak during working hours about 3 times: 35 per second.
  • Concurrent jobs at peak (Little's law): 35 × 360 s = 12,600.
  • Workers: at 4 jobs per worker machine, about 3,150 machines at peak, fewer at night; autoscale.
  • Logs: 1 MB per job average = 1 TB a day; 90 TB over 90 days, compressed perhaps a quarter of that.
  • Scheduler writes: a few state changes per job plus heartbeats every 10 seconds from 12,600 jobs = about 1,300 writes per second. One relational database handles it; shard by tenant later.
Key idea. About 12,600 jobs running at peak. The scheduler's write load is modest; worker capacity and image pulls dominate.

  1. API design

Before reading on. The git host sends the same webhook twice for one push. How many runs are created?

One. The run table has a unique key on repository, commit, and workflow name. The second insert does nothing and returns the existing run.

3.1 External API

POST /v1/hooks/push           (signed by the git host)
  {repo_id, commit_sha, ref}
  -> 202 {run_id}             (existing run_id if already created)

GET  /v1/runs/:id             -> {status, jobs: [{name, status, started_at, ...}]}
POST /v1/runs/:id/cancel
POST /v1/runs/:id/jobs/:name/retry
GET  /v1/runs/:id/jobs/:name/logs?from_seq=   (WebSocket upgrade for live tail)

3.2 Worker API

POST /internal/lease      {worker_id, labels: [linux, 4cpu], cached_images[]}
  -> {job_id, attempt, spec, secrets_token, lease_expires_at} | 204
POST /internal/heartbeat  {job_id, attempt}  -> {lease_expires_at, cancel}
POST /internal/logs       {job_id, attempt, seq, bytes}
POST /internal/finish     {job_id, attempt, exit_code, outputs}

Every worker call carries the attempt. Any call with a stale attempt returns 409.

  1. Data model

runs  (id, tenant_id, repo_id, commit_sha, workflow_name, status, created_at)
      unique (repo_id, commit_sha, workflow_name)
jobs  (id, run_id, tenant_id, seq, name, image, commands, needs_secrets[],
       status: waiting | ready | leased | running | succeeded | failed | cancelled,
       attempt int, max_attempts, worker_id, lease_expires_at,
       exit_code, outputs_json, log_ref, queued_at, started_at, finished_at)
      index (tenant_id, status, queued_at)     -- per-tenant ready queues
      index (status, lease_expires_at)         -- sweeper
tenants (id, plan, weight, max_concurrent_jobs)
secrets (tenant_id, name, ciphertext, key_id)   -- encrypted per tenant
log_chunks (job_id, attempt, seq, object_ref)

  1. High-level design

5.1 A cron job that runs builds on one machine

A single server receives webhooks and runs builds in a loop. It works for one team. A crash loses the run in progress, builds from different companies share a machine, and one heavy tenant blocks the queue.

5.2 Fix 1: a durable run and job model

Webhooks create runs and jobs in a database, idempotently. The first job becomes ready; the rest wait. The run survives any service crash.

5.3 Fix 2: workers pull jobs with leases

Workers lease ready jobs with a conditional update, heartbeat while running, and finish with their attempt number. A sweeper returns expired leases to ready.

flowchart LR
  W[Worker]:::svc -->|lease| S[Scheduler]:::new
  S -->|UPDATE jobs SET status='leased', attempt=attempt+1, worker=W<br/>WHERE id=? AND status='ready'| DB[(Jobs)]:::store
  W -->|heartbeat with attempt| S
  W -->|finish with attempt| S
  SW[Sweeper]:::new -->|leased + expired -> ready| DB
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;

5.4 Fix 3: fair queues per tenant

The scheduler picks the next job by weighted fair share across tenants, capped by each tenant's concurrency limit.

5.5 Fix 4: sandboxed execution and scoped secrets

Each job runs in a fresh microVM or sandboxed container. Secrets are fetched at start with a short-lived token scoped to that job.

5.6 Fix 5: a log service

Workers stream log chunks with sequence numbers; the log service stores them and publishes them to a live channel per job.

5.7 The composed design

sequenceDiagram
  autonumber
  participant GH as Git host
  participant T as Trigger
  participant D as Database
  participant S as Scheduler
  participant W as Worker
  participant L as Log service
  participant U as UI
  GH->>T: push webhook (repo, sha)
  T->>D: insert run (unique repo+sha+workflow), jobs, job 1 ready
  W->>S: lease (labels, cached images)
  S->>D: pick fairly across tenants, lease job 1, attempt 1
  S-->>W: spec, secrets token
  W->>W: start sandbox, pull image (cache), run commands
  loop while running
    W->>S: heartbeat (attempt 1)
    W->>L: log chunk (seq n)
    L-->>U: live tail
  end
  W->>S: finish (attempt 1, exit 0)
  S->>D: job 1 succeeded AND job 2 ready (one transaction, attempt checked)
Key idea. Idempotent triggers, a durable job chain, fenced leases, fair scheduling, sandboxes with scoped secrets, and a sequenced log stream.

  1. Deep dives

6.1 Exactly-once, honestly

Before reading on. A worker runs a deploy job, the deploy succeeds, and the worker crashes before calling finish. What happens next, and what guarantees can you still offer?

The lease expires. The sweeper returns the job to ready. Another worker leases it with attempt 2 and runs the deploy again. Nothing in the CI system can know the first deploy happened; that knowledge died with the worker.

So offer three guarantees, and be precise:

  1. One owner at a time. Leases plus fencing mean at most one worker can change a job's state, and a worker whose lease expired can no longer record anything.
  2. One recorded result. The finish write succeeds only for the current attempt, and moves the job to a final state once.
  3. Idempotent effects are the job's responsibility, with help. The system gives each job a stable idempotency key, such as run ID plus job name, the same across attempts. Deploy steps pass it to the target, or check whether that commit is already deployed, so the second run is a no-op.

State the limit plainly: the system runs a job at least once and records it exactly once; side effects outside the system are exactly once only if they are idempotent.

What separates answers: exactly once

WeakClaims the queue guarantees it

Says the queue provides exactly-once delivery and moves on.

GoodLeases with retries

Uses leases and retries crashed jobs, and notes that jobs may run twice.

StrongPrecise guarantees

Separates single ownership (fencing), single recorded result (conditional finish), and idempotent external effects (stable keys across attempts), and names what cannot be guaranteed.

6.2 State transitions without races

Each job moves through states with conditional updates only:

ready -> leased      WHERE status = 'ready'                       attempt += 1
leased -> running    WHERE attempt = :a AND status = 'leased'
running -> succeeded WHERE attempt = :a AND status = 'running'
running -> failed    WHERE attempt = :a AND status = 'running'
leased|running -> ready (sweeper)  WHERE lease_expires_at < now()

When job N succeeds, making job N+1 ready happens in the same transaction. If the transaction fails, neither change happens, and the retry is safe. A run cannot have two jobs ready at once, and it cannot skip a job.

Cancellation sets a flag on the run. Ready and waiting jobs move to cancelled at once; running jobs see cancel: true in their next heartbeat and stop.

6.3 Fair scheduling

Before reading on. Tenant A queues 2,000 jobs. Tenant B queues 5. How long does B wait?

With one global FIFO queue, B waits behind all 2,000 of A's jobs, perhaps an hour. With fair queuing, B waits roughly one scheduling cycle.

Keep a ready queue per tenant. When a worker asks for work, choose the tenant with the lowest ratio of running jobs to weight, among tenants with ready jobs that fit the worker's labels. Take that tenant's oldest ready job. Each tenant also has a concurrency cap from its plan, so a large tenant cannot take the whole fleet even when others are idle, unless you allow borrowing idle capacity.

flowchart LR
  QA[[Tenant A: 2,000 ready<br/>running 40, weight 4]]:::svc --> P{Pick min running/weight}:::new
  QB[[Tenant B: 5 ready<br/>running 0, weight 1]]:::svc --> P
  QC[[Tenant C: 30 ready<br/>running 12, weight 2]]:::svc --> P
  P -->|B: 0/1 is lowest| W[Worker]:::store
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;

6.4 Isolation and secrets

Before reading on. A malicious workflow tries to read another tenant's secrets. Walk through every layer that stops it.
  • Execution. Each job runs in a fresh microVM or a user-space-kernel sandbox, destroyed after the job. Plain containers share the host kernel with other tenants' jobs.
  • Dedicated capacity. Large or regulated tenants get their own worker pools.
  • Secrets. Stored encrypted per tenant. A job receives a short-lived token that can fetch only the secrets its workflow declares, for its own tenant, only while the job's lease is valid.
  • Logs. The log service masks known secret values in output before storing it.
  • Network. Jobs cannot reach the CI system's internal services or the cloud metadata endpoint.
  • Artifacts and caches. Keyed by tenant; a cache from tenant A is never mounted for tenant B.

Pull-request builds from forks are a special case: they run code from outside the tenant. Run them without secrets by default.

6.5 Image caching

Pulling a 2 GB image for every job wastes minutes and bandwidth. Each worker keeps a local cache of recently used images. Workers report their cached images when they ask for work, and the scheduler prefers a job whose image is already cached, within the fairness rules. A registry mirror in each zone serves cache misses fast. Popular base images are pre-pulled onto new workers at boot.

6.6 Live logs

Before reading on. A user opens a job's log page halfway through a run. How do they see the earlier output and then follow new output?

The worker sends output in chunks with increasing sequence numbers. The log service appends each chunk to object storage under the job and attempt, and publishes it to a live channel for that job.

The UI opens a WebSocket and asks for logs from sequence 0. The service replays stored chunks, then switches to the live channel. If the socket drops, the UI reconnects with the last sequence it rendered and receives only what it missed. A retried job has a new attempt number, so its logs appear separately and never mix with the failed attempt's.

After the job ends, chunks are compacted into one compressed file per attempt, kept for 90 days.

6.7 Failures

FailureRecovery
Worker crashLease expires; job retried with a new attempt
Worker network partitionIts later writes carry a stale attempt and are rejected
Scheduler crashState in the database; other instances continue
Webhook lostA reconciler polls recent commits and creates missing runs
Job fails every attemptStop at max attempts; run marked failed with the error
Database failoverLeases are longer than failover time, so running jobs are not requeued

6.8 Caching build dependencies

Before reading on. Every job downloads 800 MB of dependencies before it runs tests. How do you cut that without leaking between tenants?

Cache by content key, per tenant. The workflow declares a cache key, typically a hash of the lockfile, and a path. At job start, the worker fetches the cache for (tenant, key) from object storage and unpacks it; at job end, if the key was new, it uploads the path. A lockfile change creates a new key, so stale dependencies never mix in.

Scope matters. Caches from a pull request branch must not poison the main branch's cache, so reads fall back from branch to main, but writes stay on the branch. Caches never cross tenants.

For practice: 1 million jobs a day × 800 MB is 800 TB of downloads a day. With a 90% cache hit rate served from a regional object store, public registry traffic drops to 80 TB, and jobs start minutes sooner.

6.9 Walking a run end to end

A concrete walk the interviewer can follow:

  1. 10:00:00 Push arrives. Run 881 created; jobs build, test, deploy; build is ready.
  2. 10:00:01 Tenant T has 3 of its 10 slots in use; the scheduler leases build to worker W7, attempt 1.
  3. 10:00:05 W7 starts a microVM, restores the dependency cache, pulls the image from its local cache.
  4. 10:02:40 Build succeeds. In one transaction: build succeeded (attempt 1 checked), test ready.
  5. 10:02:41 Worker W12 leases test, attempt 1.
  6. 10:04:10 W12's host loses power. Heartbeats stop.
  7. 10:04:40 Lease expires; sweeper sets test ready.
  8. 10:04:41 W3 leases test, attempt 2; logs for attempt 2 start fresh in the UI.
  9. 10:06:30 Test succeeds; deploy ready.
  10. 10:06:31 Deploy runs with idempotency key run-881/deploy. The target sees a new key and deploys.
  11. 10:07:10 Deploy succeeds. Run 881 succeeded.

If W12 had come back at 10:05 and called finish with attempt 1, the update would have matched no row.

What separates answers: concreteness

WeakAbstract boxes

Describes components but never walks a run through them.

GoodA happy-path walk

Walks a run from push to success.

StrongA walk with failure and caching

Walks a run through a worker loss, a retry with a new attempt, cache restores, and an idempotent deploy, with the fencing check shown at the moment it matters.

  1. Variants

7.1 Parallel jobs

For a graph of jobs, a job becomes ready when all its dependencies succeed. Track remaining dependencies per job and decrement them in the finish transaction.

7.2 Self-hosted runners

Tenants run their own workers inside their networks. Workers long-poll the scheduler for their tenant's jobs only. The same lease and fencing protocol applies.

7.3 Deploy gates

A deploy job waits for approval. Model it as a job state awaiting_approval that a person moves to ready, with the approval recorded for audit.

7.4 At ten times the jobs

At 10 million jobs a day, peak concurrency passes 120,000. Shard the scheduler by tenant, each shard owning a set of tenants' queues and jobs, with workers pulling from several shards. Log storage reaches about 10 TB a day before compression. Regional worker pools with local image and dependency caches keep start times low.

  1. The transferable pattern

CI/CD is a leased job queue over untrusted code, with fairness across tenants. The pieces transfer to any shared execution platform: idempotent intake, conditional state transitions, fenced leases, per-tenant fair queues, disposable sandboxes, scoped secrets, and sequenced log streams.

Review: the 30-second answer

  • Idempotent triggers. Unique run per repository, commit, and workflow.
  • At-least-once runs, exactly-once records. Leases with fencing; conditional finish; stable idempotency keys for external effects.
  • Fair queues per tenant. Weighted share with concurrency caps.
  • Fresh sandboxes and scoped secrets. No shared kernel, masked logs, no internal network.
  • Sequenced logs. Replay then follow; reconnect from the last sequence.

Quiz

+Why can't the system guarantee a deploy happens exactly once?

A worker can complete the deploy and crash before recording it. The system cannot distinguish that from a crash before the deploy, so it must run the job again. Only an idempotent deploy step makes the repeat harmless.

+What does the attempt number protect?

It fences each lease. A worker whose lease expired still holds the old attempt number, so its heartbeats and finish calls are rejected, and it cannot overwrite the current owner's result.

+Why make job N+1 ready in the same transaction that marks job N succeeded?

So a crash between the two cannot leave the run stuck with job N done and job N+1 never scheduled, or schedule job N+1 twice.

+How does fair queuing help a small tenant?

The scheduler picks the tenant with the lowest ratio of running jobs to weight. A tenant with nothing running gets the next free worker, regardless of how many jobs a large tenant has queued.

+How does a user who opens the log page mid-run see everything?

The log service replays stored chunks from sequence zero, then switches to the live channel. On reconnect, the client asks from its last sequence number.

+How do dependency caches avoid leaking between tenants and branches?

Caches are keyed by tenant and a hash of the lockfile. Branch builds can read the main branch's cache but write only their own, and no cache is ever read across tenants.

+In the walkthrough, why does the returning worker's finish call fail?

Its lease expired and another worker took the job with attempt 2. The returning worker's call carries attempt 1, so the conditional update matches no row.

Sources and further reading

NextWebhook Delivery System