OpenAI Interview: Design a Webhook Delivery System
A full solution to the OpenAI webhook question: API and schema, capacity math, fan-out, per-endpoint lanes, concurrency caps and circuit breakers, retries with backoff and jitter, dead letters and replay, signing, SSRF protection, and observability.
By The Forward Deployed editorial teamReviewed
Part of the OpenAI system design question bank. The question is representative of the round. The analysis and solution are this site's own.
Problem statement
Customers register endpoints, such as https://example.com/hooks, and choose event types, such as invoice.paid. When the platform produces an event, the service delivers it to every subscribed endpoint, retries failures for days, and lets customers see and replay deliveries. The endpoints belong to customers, so they are slow, down, or misconfigured all the time.
Clarifying questions
- Volume? 1 billion events a day, with peaks around three times the average.
- Fan-out? Most events go to one endpoint; some go to several.
- Delivery guarantee? At least once. Duplicates are acceptable; losses are not.
- Ordering? Ask. Assume no global ordering, with per-object ordering hints.
- Retry window? About three days, then give up and notify the customer.
- Payload size? Up to 256 KB, 2 KB on average.
- Security? Customers must be able to verify that requests came from us.
What makes webhook delivery hard
You deliver to endpoints you do not control. At any moment, thousands of them are timing out, returning errors, or pointing at servers that no longer exist. A design that works when endpoints are healthy falls over when a few large customers' endpoints start timing out.
The core failure mode is head-of-line blocking. If all deliveries share one queue and one pool of senders, a slow endpoint holds senders for its full timeout on every attempt. A few slow endpoints with high volume can occupy every sender, and every other customer's webhooks stall behind them.
So the driving tension is reliability for each endpoint versus fairness across all of them. The system must keep retrying broken endpoints for days, without letting them slow down healthy ones.
flowchart LR P[Producers]:::user --> EL[[Event log]]:::svc --> FO[Fan-out]:::svc FO --> Q1[[Lane: endpoint A]]:::store FO --> Q2[[Lane: endpoint B]]:::store FO --> Q3[[Lane: endpoint C]]:::store Q1 --> S[Senders]:::svc Q2 --> S Q3 --> S S -->|POST| EP([Customer endpoints]):::user classDef user fill:#e6efec,stroke:#315e55,color:#171717; classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
Key idea. Isolate endpoints from each other. A broken endpoint should cost the system a bounded, small share of capacity, however much traffic it has.
Key concepts
At-least-once delivery
The sender retries until it gets a success response or gives up. An endpoint may process a request and fail to answer in time, so it can receive the same event twice. Receivers must deduplicate by event ID.
Exponential backoff with jitter
After each failure, wait longer before the next try: 1 minute, 5, 30, 2 hours, and so on. Jitter adds randomness to each wait, so retries from many events do not all arrive at the moment an endpoint recovers.
Circuit breakers
After repeated failures, stop sending to an endpoint for a while. It protects the endpoint while it recovers, and it protects your senders from wasting time on requests that will fail.
HMAC signatures
The sender computes a keyed hash of the payload and a timestamp with a secret shared with the customer, and sends it in a header. The receiver recomputes it to verify the request is authentic and recent.
Server-side request forgery (SSRF)
If customers can enter any URL, an attacker can register http://169.254.169.254/ or an internal hostname and make your servers call your own internal systems. Webhook senders are a classic SSRF target.
Key idea. Retry with backoff and jitter, isolate with lanes and breakers, prove authenticity with signatures, and defend your own network from customer-supplied URLs.
- Requirements
Before reading on. List the requirements. Then name the one failure mode you would design against first.
1.1 Functional requirements
- Register, update, disable, and delete endpoints with event type filters.
- Deliver each event to every matching endpoint with an HTTP POST.
- Retry failures with backoff for up to three days.
- Record every attempt; let customers query delivery history.
- Let customers retry one delivery or replay a time range.
- Sign every request; support secret rotation.
1.2 Non-functional requirements
- Durability. An accepted event is never lost.
- Isolation. One slow or failing endpoint does not delay other customers.
- Latency. Healthy endpoints receive events within seconds.
- Scale. 1 billion events a day, 35,000 per second at peak.
- Security. Signed requests; no SSRF into internal networks.
1.3 The constraint versus the property
Durable, at-least-once delivery is the property. Unreliable customer endpoints are the constraint. They drive the lane design, concurrency caps, breakers, and retry scheduling.
- Back-of-the-envelope estimation
- Average events: 1 billion / 86,400 ≈ 11,600 per second; peak about 35,000 per second.
- With average fan-out of 1.2, peak deliveries are about 42,000 per second.
- Senders: if a healthy delivery takes 200 ms, one sender slot handles 5 per second, so 42,000 per second needs about 8,400 concurrent slots. Async HTTP clients hold thousands of slots per machine, so tens of machines.
- The slow-endpoint problem in numbers: an endpoint timing out at 10 seconds with 1,000 events per second would, without caps, occupy 10,000 slots, more than the whole healthy fleet.
- Payload storage: 1 billion × 2 KB = 2 TB a day. Keep payloads for the 3-day retry window plus a replay window of 30 days: about 60 TB.
- Attempt log: with about 1.3 attempts per delivery on average and 200 bytes per attempt record, about 300 GB a day.
Key idea. 42,000 deliveries per second is ordinary. The dangerous number is one slow endpoint's demand on sender slots, which can exceed the whole fleet.
- API design
Before reading on. Should the event payload include the full object, or just an ID the customer fetches?
Both have uses. A full payload saves the customer a round trip and works when your API is down. A thin payload with an ID avoids sending stale or sensitive data and keeps payloads small. A common answer: send the full object with a version number, and document that the API is the source of truth for the latest state.
3.1 Customer API
POST /v1/webhook_endpoints
{url, event_types: ["invoice.paid", "invoice.failed"], description}
-> 201 {id, secret} // shown once
GET /v1/webhook_endpoints/:id
PATCH /v1/webhook_endpoints/:id {url?, event_types?, disabled?}
DELETE /v1/webhook_endpoints/:id
POST /v1/webhook_endpoints/:id/rotate_secret -> {secret, old_secret_expires_at}
GET /v1/webhook_endpoints/:id/deliveries?status=failed&limit=50&cursor=...
GET /v1/deliveries/:id -> {event, attempts: [{at, status_code, latency_ms, error}]}
POST /v1/deliveries/:id/retry
POST /v1/webhook_endpoints/:id/replay {from, to, event_types?}3.2 The request customers receive
POST https://example.com/hooks
Content-Type: application/json
Webhook-Id: evt_01J9... // stable across retries
Webhook-Timestamp: 1790000000
Webhook-Signature: v1,<base64 HMAC-SHA256(secret, id.timestamp.body)>
{"id": "evt_01J9...", "type": "invoice.paid", "created": 1790000000,
"data": {"object": {...}, "object_version": 12}}A 2xx response within 10 seconds counts as success. Anything else, including a redirect, counts as failure.
- Data model
endpoints ( id, customer_id, url, event_types text[], status, -- active | disabled secret_ref, old_secret_ref, old_secret_expires_at, max_concurrency int DEFAULT 10, created_at ) events ( id, customer_id, type, payload_ref, created_at -- payload in object storage ) deliveries ( id, event_id, endpoint_id, status, -- pending | delivered | retrying | dead attempt_count, next_attempt_at, last_status_code, last_error, created_at, delivered_at ) index (endpoint_id, created_at desc) -- history page index (status, next_attempt_at) -- retry scheduler attempts (delivery_id, attempt_no, started_at, status_code, latency_ms, error)
At this volume, deliveries and attempts go to a partitioned store: a distributed key-value or wide-column database, partitioned by endpoint and time, with a time-to-live for retention.
- High-level design
5.1 Send inline from the producer
The service that creates the invoice calls the customer's URL directly. A slow endpoint slows the invoice API, a failure is lost, and nothing retries.
5.2 Fix 1: a durable event log and a fan-out step
Producers append events to a durable log. A fan-out worker reads each event, finds the subscribed endpoints, and creates one delivery per endpoint. The producer's work ends when the event is in the log.
5.3 Fix 2: lanes per endpoint
Put each delivery into a lane for its endpoint, such as a partition keyed by endpoint ID. Senders take work across lanes, never more than the endpoint's concurrency cap from any one lane.
flowchart LR FO[Fan-out]:::svc --> LA[[Lane A: healthy]]:::new FO --> LB[[Lane B: timing out]]:::new FO --> LC[[Lane C: healthy]]:::new LA -->|up to 10 in flight| SND[Sender pool]:::svc LB -->|up to 10 in flight, breaker open| SND LC -->|up to 10 in flight| SND classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;
5.4 Fix 3: a retry scheduler
Failed deliveries leave the hot lane and go to a scheduler with their next attempt time. When the time comes, the delivery goes back into its lane.
5.5 Fix 4: dead letters, replay, and a dashboard
After the final attempt, a delivery becomes dead. The customer sees it in the dashboard, and can retry it or replay a range.
5.6 The composed design
sequenceDiagram
autonumber
participant P as Producer
participant L as Event log
participant F as Fan-out
participant Q as Endpoint lane
participant S as Sender
participant C as Customer endpoint
participant R as Retry scheduler
P->>L: append invoice.paid (event id)
L->>F: consume
F->>F: lookup subscribed endpoints (cached)
F->>Q: delivery per endpoint
S->>Q: take (respecting concurrency cap, breaker)
S->>S: resolve URL, check IP not private, sign
S->>C: POST (10 s timeout)
alt 2xx
C-->>S: 200
S->>S: mark delivered, record attempt
else error or timeout
S->>R: schedule retry at now + backoff + jitter
R-->>Q: re-enqueue when due
endKey idea. Log first, fan out to per-endpoint lanes, cap concurrency per endpoint, schedule retries off the hot path, and end with dead letters the customer can replay.
- Deep dives
6.1 Isolating slow endpoints
Before reading on. Customer X's endpoint starts timing out at 10 seconds. X receives 1,000 events per second. Walk through what your system does.
- Concurrency cap. Lane X can have at most 10 requests in flight. X's timeouts now use 10 slots, not 10,000. Everyone else is unaffected.
- Timeouts. A 10-second timeout caps each slot's waste. Some systems use shorter connect timeouts, such as 3 seconds, and a longer read timeout.
- Circuit breaker. After, say, 20 consecutive failures, the breaker opens. For the next minute, lane X sends nothing; deliveries go straight to the retry scheduler. Then a single probe request tests the endpoint. On success, the breaker closes and delivery resumes gradually.
- Backlog. X's lane grows by 1,000 events per second while the breaker is open. The lane lives in durable storage, so it can hold hours of backlog.
- Recovery. When X recovers, drain the backlog at a controlled rate so the recovering endpoint is not flattened.
- Notification. After a set failure period, email X's owners. After days, disable the endpoint.
What separates answers: isolation
WeakOne shared queue
All deliveries share a queue and a sender pool, so one slow customer stalls everyone.
GoodPer-endpoint caps
Limits concurrent requests per endpoint and uses short timeouts.
StrongLanes, breakers, and controlled recovery
Adds circuit breakers with probes, durable per-endpoint backlogs, rate-limited draining after recovery, and automatic notification and disabling for long failures.
6.2 Retry scheduling
A schedule such as: 1 minute, 5 minutes, 30 minutes, 2 hours, then every 5 hours until about 3 days. Add jitter of about 20% to each wait, so a batch of events that failed together does not retry together.
Retries must not sit in the hot lane, where they would block fresh events. Keep them in a scheduler: a table indexed by next attempt time, polled every second for due items, or a delay queue that releases messages at a given time. When due, the delivery re-enters its endpoint's lane.
Treat responses differently. Retry on timeouts, connection errors, 5xx, and 429 (honoring Retry-After). Do not retry on 400, 401, 404, or 410 more than a couple of times; those rarely fix themselves, and 410 Gone should disable the endpoint.
6.3 Ordering and duplicates
Before reading on. A customer complains they receivedinvoice.updatedbeforeinvoice.created. Is that a bug?
No; it is expected with retries. If created failed and was retried after updated succeeded, the customer sees them out of order. Guaranteeing order would mean blocking every later event for an object behind a failed earlier one, which trades one broken delivery for many.
Instead, document that order is not guaranteed and give customers the tools to handle it: the event creation time, and an object_version on the payload so they can ignore updates older than what they have. Some customers can simply fetch the latest object state on any event.
Duplicates are also expected. The event ID stays the same across retries, and receivers deduplicate on it. Publish that contract clearly.
6.4 Signing and secret rotation
Sign every request with HMAC-SHA256 over the event ID, the timestamp, and the raw body, using the endpoint's secret. The receiver recomputes the signature and rejects requests whose timestamp is more than five minutes old, which stops replay of captured requests.
Rotation: when a customer rotates, issue a new secret and keep the old one valid for 24 hours. During that window, send two signatures in the header, one per secret. The customer updates their server at any point in the window without dropping events.
6.5 SSRF protection
Before reading on. A customer registers http://metadata.internal/latest/credentials as their URL. What stops your sender from fetching it?Several checks, because each alone can be bypassed:
- Require HTTPS and public DNS names at registration.
- Resolve and check at send time, not only at registration. DNS can change: an attacker can point a name at a public IP during registration and a private one later.
- Block private, loopback, link-local, and metadata ranges for every resolved address, IPv4 and IPv6.
- Connect to the IP you checked, not by re-resolving the name, to prevent DNS rebinding between check and connect.
- Do not follow redirects, or check every hop the same way.
- Run senders in an isolated network that has no route to internal services anyway, as defense in depth.
6.6 Observability
Show each customer their endpoint health: success rate, latency, recent errors, and backlog. It answers most support questions before they are asked. Internally, track lane lag per endpoint, sender slot use, breaker state counts, dead letters per hour, and the age of the oldest undelivered event. Alert on lag, because lag is what customers feel.
6.7 Choosing the queue technology
Before reading on. Would you build lanes with Kafka partitions, a database table, or a message broker with per-queue semantics? What breaks with each?
Kafka partitioned by endpoint. Durable and fast. But a partition is ordered: a message stuck at the head blocks every message behind it in that partition. Since many endpoints share each partition, one slow endpoint delays its partition-mates. Kafka works as the event log feeding fan-out, and less well as the lanes themselves.
A database table of deliveries. Senders pick due rows per endpoint with skip-locked queries and per-endpoint concurrency counters. Flexible and easy to inspect; at 42,000 deliveries per second, the write load needs a partitioned database and careful indexing.
A queue per endpoint in a broker that supports many queues and per-message acknowledgment. It maps directly to lanes, but millions of endpoints mean millions of queues, which many brokers handle poorly.
A common production shape: Kafka for the event log; a partitioned key-value or relational store for deliveries and retry state; and a scheduler that hands out due deliveries per endpoint within concurrency caps. Say the tradeoff; the interviewer wants to see you know why a single ordered partition is a problem.
6.8 A worked day for a failing endpoint
For practice: endpoint E receives 50 events per second. At 09:00 its server starts returning 500.
- 09:00 to 09:01: 20 consecutive failures open E's breaker. Senders stop calling E. New deliveries for E go to the retry scheduler with a 1-minute delay.
- 09:01 onward: every minute, a single probe request tests E. Each probe fails.
- 09:30: E's backlog is 90,000 events, stored durably. Other endpoints are unaffected.
- 10:00: an email alerts E's owner; the dashboard shows the failure and backlog.
- 11:15: E recovers; a probe succeeds; the breaker closes.
- 11:15 onward: the backlog of about 400,000 events drains at a capped rate, such as 200 per second, above E's normal 50 per second, finishing in about 45 minutes.
- Result: no events lost; the delay was the outage plus the drain; nobody else noticed.
What separates answers: mechanics
WeakNo view of the queue
Names a queue without considering head-of-line blocking or backlogs.
GoodAware of ordered partitions
Explains why a shared ordered partition blocks healthy endpoints behind a slow one.
StrongWalks an outage end to end
Picks storage for deliveries and retries with reasons, and walks an endpoint failure through breaker, probes, backlog, alerting, and a rate-limited drain.
- Variants
7.1 Strict ordering per object
If some customers truly need order, offer an opt-in mode that orders deliveries per object ID: an object's lane blocks behind its first failure. Make the cost clear: one failing event holds back every later event for that object.
7.2 Batching
High-volume customers may prefer batches: up to 100 events per request, flushed every second. Fewer requests, and each failure affects more events.
7.3 Pull instead of push
Offer an events API customers can poll with a cursor as a fallback. It removes the endpoint reliability problem for customers who prefer it.
7.4 At ten times the events
At 10 billion events a day, peak deliveries reach about 420,000 per second. The design holds, with more senders and a larger delivery store. Cost shifts to outbound bandwidth and the attempt log: sample successful attempt records after a few days and keep failures in full. Offer batching to the largest customers, which cuts request count by up to a hundredfold for them.
- The transferable pattern
Webhook delivery is reliable delivery to unreliable consumers. The same design applies to push notifications, email sending, and any integration that calls partner systems: a durable log, isolated lanes per destination, caps and breakers, retries with backoff and jitter off the hot path, dead letters with replay, and a strict egress policy.
Review: the 30-second answer
- Log, then fan out to per-endpoint lanes. Producers never wait on customers.
- Cap concurrency per endpoint; add breakers. A slow endpoint costs 10 slots, not the fleet.
- Retries with backoff and jitter in a scheduler. About three days, then dead letters and replay.
- At least once, unordered, with stable event IDs. Receivers deduplicate and use object versions.
- Sign with HMAC and timestamps; block SSRF. Check resolved IPs at send time and do not follow redirects.
Quiz
+How can one customer's slow endpoint stall every other customer?
If all deliveries share one sender pool, each slow request holds a sender for the full timeout. A high-volume slow endpoint can occupy every sender, so all other deliveries wait.
+Why add jitter to retry delays?
Events that failed together would otherwise retry together, hitting the recovering endpoint with a burst at the same moment. Jitter spreads the retries out.
+Why check a URL's resolved IP at send time instead of only at registration?
DNS can change after registration. An attacker can register a name that points to a public address, then repoint it to an internal address.
+Why not guarantee ordered delivery by default?
Ordering would force every later event for an object to wait behind a failed earlier one. Most customers are better served by unordered delivery with timestamps and object versions.
+How does secret rotation avoid dropped events?
The old secret stays valid for a window, and requests carry signatures for both secrets. The customer switches over at any time within the window.
+Why are Kafka partitions a poor fit for per-endpoint lanes?
A partition is ordered, and many endpoints share each partition. A stuck message at the head delays every message behind it, including other endpoints' deliveries.
+After an endpoint recovers, why drain its backlog at a capped rate?
The endpoint is fragile right after recovery. Sending the whole backlog at full speed can knock it over again; a capped rate above its normal traffic clears the backlog without overwhelming it.
Sources and further reading
- Standard Webhooks specifies headers, signatures, and retry conventions for webhook senders.
- OWASP Server-Side Request Forgery Prevention Cheat Sheet covers the network checks senders need.
- AWS Architecture Blog: Exponential Backoff and Jitter compares jitter strategies.
- The receiving side of webhooks, with idempotent handlers, appears in coffee-shop payments.
