The Forward Deployed

OpenAI Interview: Design a Webhook Delivery System

A full solution to the OpenAI webhook question: API and schema, capacity math, fan-out, per-endpoint lanes, concurrency caps and circuit breakers, retries with backoff and jitter, dead letters and replay, signing, SSRF protection, and observability.

By Reviewed

Part of the OpenAI system design question bank. The question is representative of the round. The analysis and solution are this site's own.

Problem statement

Customers register endpoints, such as https://example.com/hooks, and choose event types, such as invoice.paid. When the platform produces an event, the service delivers it to every subscribed endpoint, retries failures for days, and lets customers see and replay deliveries. The endpoints belong to customers, so they are slow, down, or misconfigured all the time.

Clarifying questions

  • Volume? 1 billion events a day, with peaks around three times the average.
  • Fan-out? Most events go to one endpoint; some go to several.
  • Delivery guarantee? At least once. Duplicates are acceptable; losses are not.
  • Ordering? Ask. Assume no global ordering, with per-object ordering hints.
  • Retry window? About three days, then give up and notify the customer.
  • Payload size? Up to 256 KB, 2 KB on average.
  • Security? Customers must be able to verify that requests came from us.

What makes webhook delivery hard

You deliver to endpoints you do not control. At any moment, thousands of them are timing out, returning errors, or pointing at servers that no longer exist. A design that works when endpoints are healthy falls over when a few large customers' endpoints start timing out.

The core failure mode is head-of-line blocking. If all deliveries share one queue and one pool of senders, a slow endpoint holds senders for its full timeout on every attempt. A few slow endpoints with high volume can occupy every sender, and every other customer's webhooks stall behind them.

So the driving tension is reliability for each endpoint versus fairness across all of them. The system must keep retrying broken endpoints for days, without letting them slow down healthy ones.

flowchart LR
  P[Producers]:::user --> EL[[Event log]]:::svc --> FO[Fan-out]:::svc
  FO --> Q1[[Lane: endpoint A]]:::store
  FO --> Q2[[Lane: endpoint B]]:::store
  FO --> Q3[[Lane: endpoint C]]:::store
  Q1 --> S[Senders]:::svc
  Q2 --> S
  Q3 --> S
  S -->|POST| EP([Customer endpoints]):::user
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
Key idea. Isolate endpoints from each other. A broken endpoint should cost the system a bounded, small share of capacity, however much traffic it has.

Key concepts

At-least-once delivery

The sender retries until it gets a success response or gives up. An endpoint may process a request and fail to answer in time, so it can receive the same event twice. Receivers must deduplicate by event ID.

Exponential backoff with jitter

After each failure, wait longer before the next try: 1 minute, 5, 30, 2 hours, and so on. Jitter adds randomness to each wait, so retries from many events do not all arrive at the moment an endpoint recovers.

Circuit breakers

After repeated failures, stop sending to an endpoint for a while. It protects the endpoint while it recovers, and it protects your senders from wasting time on requests that will fail.

HMAC signatures

The sender computes a keyed hash of the payload and a timestamp with a secret shared with the customer, and sends it in a header. The receiver recomputes it to verify the request is authentic and recent.

Server-side request forgery (SSRF)

If customers can enter any URL, an attacker can register http://169.254.169.254/ or an internal hostname and make your servers call your own internal systems. Webhook senders are a classic SSRF target.

Key idea. Retry with backoff and jitter, isolate with lanes and breakers, prove authenticity with signatures, and defend your own network from customer-supplied URLs.

  1. Requirements

Before reading on. List the requirements. Then name the one failure mode you would design against first.

1.1 Functional requirements

  • Register, update, disable, and delete endpoints with event type filters.
  • Deliver each event to every matching endpoint with an HTTP POST.
  • Retry failures with backoff for up to three days.
  • Record every attempt; let customers query delivery history.
  • Let customers retry one delivery or replay a time range.
  • Sign every request; support secret rotation.

1.2 Non-functional requirements

  • Durability. An accepted event is never lost.
  • Isolation. One slow or failing endpoint does not delay other customers.
  • Latency. Healthy endpoints receive events within seconds.
  • Scale. 1 billion events a day, 35,000 per second at peak.
  • Security. Signed requests; no SSRF into internal networks.

1.3 The constraint versus the property

Durable, at-least-once delivery is the property. Unreliable customer endpoints are the constraint. They drive the lane design, concurrency caps, breakers, and retry scheduling.

  1. Back-of-the-envelope estimation

  • Average events: 1 billion / 86,400 ≈ 11,600 per second; peak about 35,000 per second.
  • With average fan-out of 1.2, peak deliveries are about 42,000 per second.
  • Senders: if a healthy delivery takes 200 ms, one sender slot handles 5 per second, so 42,000 per second needs about 8,400 concurrent slots. Async HTTP clients hold thousands of slots per machine, so tens of machines.
  • The slow-endpoint problem in numbers: an endpoint timing out at 10 seconds with 1,000 events per second would, without caps, occupy 10,000 slots, more than the whole healthy fleet.
  • Payload storage: 1 billion × 2 KB = 2 TB a day. Keep payloads for the 3-day retry window plus a replay window of 30 days: about 60 TB.
  • Attempt log: with about 1.3 attempts per delivery on average and 200 bytes per attempt record, about 300 GB a day.
Key idea. 42,000 deliveries per second is ordinary. The dangerous number is one slow endpoint's demand on sender slots, which can exceed the whole fleet.

  1. API design

Before reading on. Should the event payload include the full object, or just an ID the customer fetches?

Both have uses. A full payload saves the customer a round trip and works when your API is down. A thin payload with an ID avoids sending stale or sensitive data and keeps payloads small. A common answer: send the full object with a version number, and document that the API is the source of truth for the latest state.

3.1 Customer API

POST   /v1/webhook_endpoints
  {url, event_types: ["invoice.paid", "invoice.failed"], description}
  -> 201 {id, secret}            // shown once
GET    /v1/webhook_endpoints/:id
PATCH  /v1/webhook_endpoints/:id  {url?, event_types?, disabled?}
DELETE /v1/webhook_endpoints/:id
POST   /v1/webhook_endpoints/:id/rotate_secret  -> {secret, old_secret_expires_at}

GET  /v1/webhook_endpoints/:id/deliveries?status=failed&limit=50&cursor=...
GET  /v1/deliveries/:id          -> {event, attempts: [{at, status_code, latency_ms, error}]}
POST /v1/deliveries/:id/retry
POST /v1/webhook_endpoints/:id/replay  {from, to, event_types?}

3.2 The request customers receive

POST https://example.com/hooks
Content-Type: application/json
Webhook-Id: evt_01J9...                     // stable across retries
Webhook-Timestamp: 1790000000
Webhook-Signature: v1,<base64 HMAC-SHA256(secret, id.timestamp.body)>

{"id": "evt_01J9...", "type": "invoice.paid", "created": 1790000000,
 "data": {"object": {...}, "object_version": 12}}

A 2xx response within 10 seconds counts as success. Anything else, including a redirect, counts as failure.

  1. Data model

endpoints (
  id, customer_id, url, event_types text[], status,        -- active | disabled
  secret_ref, old_secret_ref, old_secret_expires_at,
  max_concurrency int DEFAULT 10,
  created_at
)
events (
  id, customer_id, type, payload_ref, created_at           -- payload in object storage
)
deliveries (
  id, event_id, endpoint_id, status,                       -- pending | delivered | retrying | dead
  attempt_count, next_attempt_at, last_status_code, last_error,
  created_at, delivered_at
)
  index (endpoint_id, created_at desc)                     -- history page
  index (status, next_attempt_at)                          -- retry scheduler
attempts (delivery_id, attempt_no, started_at, status_code, latency_ms, error)

At this volume, deliveries and attempts go to a partitioned store: a distributed key-value or wide-column database, partitioned by endpoint and time, with a time-to-live for retention.

  1. High-level design

5.1 Send inline from the producer

The service that creates the invoice calls the customer's URL directly. A slow endpoint slows the invoice API, a failure is lost, and nothing retries.

5.2 Fix 1: a durable event log and a fan-out step

Producers append events to a durable log. A fan-out worker reads each event, finds the subscribed endpoints, and creates one delivery per endpoint. The producer's work ends when the event is in the log.

5.3 Fix 2: lanes per endpoint

Put each delivery into a lane for its endpoint, such as a partition keyed by endpoint ID. Senders take work across lanes, never more than the endpoint's concurrency cap from any one lane.

flowchart LR
  FO[Fan-out]:::svc --> LA[[Lane A: healthy]]:::new
  FO --> LB[[Lane B: timing out]]:::new
  FO --> LC[[Lane C: healthy]]:::new
  LA -->|up to 10 in flight| SND[Sender pool]:::svc
  LB -->|up to 10 in flight, breaker open| SND
  LC -->|up to 10 in flight| SND
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;

5.4 Fix 3: a retry scheduler

Failed deliveries leave the hot lane and go to a scheduler with their next attempt time. When the time comes, the delivery goes back into its lane.

5.5 Fix 4: dead letters, replay, and a dashboard

After the final attempt, a delivery becomes dead. The customer sees it in the dashboard, and can retry it or replay a range.

5.6 The composed design

sequenceDiagram
  autonumber
  participant P as Producer
  participant L as Event log
  participant F as Fan-out
  participant Q as Endpoint lane
  participant S as Sender
  participant C as Customer endpoint
  participant R as Retry scheduler
  P->>L: append invoice.paid (event id)
  L->>F: consume
  F->>F: lookup subscribed endpoints (cached)
  F->>Q: delivery per endpoint
  S->>Q: take (respecting concurrency cap, breaker)
  S->>S: resolve URL, check IP not private, sign
  S->>C: POST (10 s timeout)
  alt 2xx
    C-->>S: 200
    S->>S: mark delivered, record attempt
  else error or timeout
    S->>R: schedule retry at now + backoff + jitter
    R-->>Q: re-enqueue when due
  end
Key idea. Log first, fan out to per-endpoint lanes, cap concurrency per endpoint, schedule retries off the hot path, and end with dead letters the customer can replay.

  1. Deep dives

6.1 Isolating slow endpoints

Before reading on. Customer X's endpoint starts timing out at 10 seconds. X receives 1,000 events per second. Walk through what your system does.
  1. Concurrency cap. Lane X can have at most 10 requests in flight. X's timeouts now use 10 slots, not 10,000. Everyone else is unaffected.
  2. Timeouts. A 10-second timeout caps each slot's waste. Some systems use shorter connect timeouts, such as 3 seconds, and a longer read timeout.
  3. Circuit breaker. After, say, 20 consecutive failures, the breaker opens. For the next minute, lane X sends nothing; deliveries go straight to the retry scheduler. Then a single probe request tests the endpoint. On success, the breaker closes and delivery resumes gradually.
  4. Backlog. X's lane grows by 1,000 events per second while the breaker is open. The lane lives in durable storage, so it can hold hours of backlog.
  5. Recovery. When X recovers, drain the backlog at a controlled rate so the recovering endpoint is not flattened.
  6. Notification. After a set failure period, email X's owners. After days, disable the endpoint.

What separates answers: isolation

WeakOne shared queue

All deliveries share a queue and a sender pool, so one slow customer stalls everyone.

GoodPer-endpoint caps

Limits concurrent requests per endpoint and uses short timeouts.

StrongLanes, breakers, and controlled recovery

Adds circuit breakers with probes, durable per-endpoint backlogs, rate-limited draining after recovery, and automatic notification and disabling for long failures.

6.2 Retry scheduling

A schedule such as: 1 minute, 5 minutes, 30 minutes, 2 hours, then every 5 hours until about 3 days. Add jitter of about 20% to each wait, so a batch of events that failed together does not retry together.

Retries must not sit in the hot lane, where they would block fresh events. Keep them in a scheduler: a table indexed by next attempt time, polled every second for due items, or a delay queue that releases messages at a given time. When due, the delivery re-enters its endpoint's lane.

Treat responses differently. Retry on timeouts, connection errors, 5xx, and 429 (honoring Retry-After). Do not retry on 400, 401, 404, or 410 more than a couple of times; those rarely fix themselves, and 410 Gone should disable the endpoint.

6.3 Ordering and duplicates

Before reading on. A customer complains they received invoice.updated before invoice.created. Is that a bug?

No; it is expected with retries. If created failed and was retried after updated succeeded, the customer sees them out of order. Guaranteeing order would mean blocking every later event for an object behind a failed earlier one, which trades one broken delivery for many.

Instead, document that order is not guaranteed and give customers the tools to handle it: the event creation time, and an object_version on the payload so they can ignore updates older than what they have. Some customers can simply fetch the latest object state on any event.

Duplicates are also expected. The event ID stays the same across retries, and receivers deduplicate on it. Publish that contract clearly.

6.4 Signing and secret rotation

Sign every request with HMAC-SHA256 over the event ID, the timestamp, and the raw body, using the endpoint's secret. The receiver recomputes the signature and rejects requests whose timestamp is more than five minutes old, which stops replay of captured requests.

Rotation: when a customer rotates, issue a new secret and keep the old one valid for 24 hours. During that window, send two signatures in the header, one per secret. The customer updates their server at any point in the window without dropping events.

6.5 SSRF protection

Before reading on. A customer registers http://metadata.internal/latest/credentials as their URL. What stops your sender from fetching it?

Several checks, because each alone can be bypassed:

  • Require HTTPS and public DNS names at registration.
  • Resolve and check at send time, not only at registration. DNS can change: an attacker can point a name at a public IP during registration and a private one later.
  • Block private, loopback, link-local, and metadata ranges for every resolved address, IPv4 and IPv6.
  • Connect to the IP you checked, not by re-resolving the name, to prevent DNS rebinding between check and connect.
  • Do not follow redirects, or check every hop the same way.
  • Run senders in an isolated network that has no route to internal services anyway, as defense in depth.

6.6 Observability

Show each customer their endpoint health: success rate, latency, recent errors, and backlog. It answers most support questions before they are asked. Internally, track lane lag per endpoint, sender slot use, breaker state counts, dead letters per hour, and the age of the oldest undelivered event. Alert on lag, because lag is what customers feel.

6.7 Choosing the queue technology

Before reading on. Would you build lanes with Kafka partitions, a database table, or a message broker with per-queue semantics? What breaks with each?

Kafka partitioned by endpoint. Durable and fast. But a partition is ordered: a message stuck at the head blocks every message behind it in that partition. Since many endpoints share each partition, one slow endpoint delays its partition-mates. Kafka works as the event log feeding fan-out, and less well as the lanes themselves.

A database table of deliveries. Senders pick due rows per endpoint with skip-locked queries and per-endpoint concurrency counters. Flexible and easy to inspect; at 42,000 deliveries per second, the write load needs a partitioned database and careful indexing.

A queue per endpoint in a broker that supports many queues and per-message acknowledgment. It maps directly to lanes, but millions of endpoints mean millions of queues, which many brokers handle poorly.

A common production shape: Kafka for the event log; a partitioned key-value or relational store for deliveries and retry state; and a scheduler that hands out due deliveries per endpoint within concurrency caps. Say the tradeoff; the interviewer wants to see you know why a single ordered partition is a problem.

6.8 A worked day for a failing endpoint

For practice: endpoint E receives 50 events per second. At 09:00 its server starts returning 500.

  • 09:00 to 09:01: 20 consecutive failures open E's breaker. Senders stop calling E. New deliveries for E go to the retry scheduler with a 1-minute delay.
  • 09:01 onward: every minute, a single probe request tests E. Each probe fails.
  • 09:30: E's backlog is 90,000 events, stored durably. Other endpoints are unaffected.
  • 10:00: an email alerts E's owner; the dashboard shows the failure and backlog.
  • 11:15: E recovers; a probe succeeds; the breaker closes.
  • 11:15 onward: the backlog of about 400,000 events drains at a capped rate, such as 200 per second, above E's normal 50 per second, finishing in about 45 minutes.
  • Result: no events lost; the delay was the outage plus the drain; nobody else noticed.

What separates answers: mechanics

WeakNo view of the queue

Names a queue without considering head-of-line blocking or backlogs.

GoodAware of ordered partitions

Explains why a shared ordered partition blocks healthy endpoints behind a slow one.

StrongWalks an outage end to end

Picks storage for deliveries and retries with reasons, and walks an endpoint failure through breaker, probes, backlog, alerting, and a rate-limited drain.

  1. Variants

7.1 Strict ordering per object

If some customers truly need order, offer an opt-in mode that orders deliveries per object ID: an object's lane blocks behind its first failure. Make the cost clear: one failing event holds back every later event for that object.

7.2 Batching

High-volume customers may prefer batches: up to 100 events per request, flushed every second. Fewer requests, and each failure affects more events.

7.3 Pull instead of push

Offer an events API customers can poll with a cursor as a fallback. It removes the endpoint reliability problem for customers who prefer it.

7.4 At ten times the events

At 10 billion events a day, peak deliveries reach about 420,000 per second. The design holds, with more senders and a larger delivery store. Cost shifts to outbound bandwidth and the attempt log: sample successful attempt records after a few days and keep failures in full. Offer batching to the largest customers, which cuts request count by up to a hundredfold for them.

  1. The transferable pattern

Webhook delivery is reliable delivery to unreliable consumers. The same design applies to push notifications, email sending, and any integration that calls partner systems: a durable log, isolated lanes per destination, caps and breakers, retries with backoff and jitter off the hot path, dead letters with replay, and a strict egress policy.

Review: the 30-second answer

  • Log, then fan out to per-endpoint lanes. Producers never wait on customers.
  • Cap concurrency per endpoint; add breakers. A slow endpoint costs 10 slots, not the fleet.
  • Retries with backoff and jitter in a scheduler. About three days, then dead letters and replay.
  • At least once, unordered, with stable event IDs. Receivers deduplicate and use object versions.
  • Sign with HMAC and timestamps; block SSRF. Check resolved IPs at send time and do not follow redirects.

Quiz

+How can one customer's slow endpoint stall every other customer?

If all deliveries share one sender pool, each slow request holds a sender for the full timeout. A high-volume slow endpoint can occupy every sender, so all other deliveries wait.

+Why add jitter to retry delays?

Events that failed together would otherwise retry together, hitting the recovering endpoint with a burst at the same moment. Jitter spreads the retries out.

+Why check a URL's resolved IP at send time instead of only at registration?

DNS can change after registration. An attacker can register a name that points to a public address, then repoint it to an internal address.

+Why not guarantee ordered delivery by default?

Ordering would force every later event for an object to wait behind a failed earlier one. Most customers are better served by unordered delivery with timestamps and object versions.

+How does secret rotation avoid dropped events?

The old secret stays valid for a window, and requests carry signatures for both secrets. The customer switches over at any time within the window.

+Why are Kafka partitions a poor fit for per-endpoint lanes?

A partition is ordered, and many endpoints share each partition. A stuck message at the head delays every message behind it, including other endpoints' deliveries.

+After an endpoint recovers, why drain its backlog at a capped rate?

The endpoint is fragile right after recovery. Sending the whole backlog at full speed can knock it over again; a capped rate above its normal traffic clears the backlog without overwhelming it.

Sources and further reading

NextDesign Slack