OpenAI Interview: Design Payments for a Coffee Shop
A full solution to the OpenAI payment question: reading the actual prompt, holds and captures, idempotency keys, the payment state machine, a double-entry ledger, provider webhooks, nightly settlement with the right index, reconciliation, resilience, and sharding.
By The Forward Deployed editorial teamReviewed
Part of the OpenAI system design question bank. The question is representative of the round. The analysis and solution are this site's own.
Problem statement
The title may say "payments", and the question is a specific flow. Candidates are reported to fail after delivering a polished, generic payment-system design to a narrower question. Read the prompt twice, restate it, and design exactly that flow: authorize a hold, capture a final amount, settle nightly, and reconcile.
Fraud scoring, loyalty points, and inventory are out of scope unless the interviewer adds them.
Clarifying questions
- Card present or online? Assume a tap-to-pay terminal in the store, plus mobile orders ahead.
- Can the final amount exceed the hold? Yes, up to a limit such as 20% over, for tips.
- How long does a hold last? The provider releases it after several days if not captured. Assume orders are captured within an hour.
- Partial payments or split tenders? No, one card per order.
- How many providers? Two, for resilience and cost.
- Scale? For practice: 10,000 stores, 300 orders per store per day, peaks at the morning rush.
- Money in what form? Integer minor units, such as cents, never floating point.
What makes payments hard
Money moves across systems you do not control, over networks that fail at the worst moment.
A request can time out after the provider charged the card, so the client retries and the customer pays twice. A provider's result can arrive minutes later by webhook, and it can arrive twice. A nightly batch can fail halfway, and rerunning it must not settle anything twice. Every one of those is a real-money error that ends up in a support ticket or a chargeback.
So the driving tension is availability versus exactly-once money movement. The till must never block the line, and money must never move twice or disappear. The answer is idempotency at every boundary, an explicit state machine, and a ledger that can prove every balance.
flowchart LR POS([Store terminal]):::user --> PS[Payment service]:::svc PS --> PSP[Payment provider]:::store PSP -.webhook.-> PS PS --> DB[(Payments + ledger)]:::store B[Nightly settlement]:::svc --> DB B --> PSP classDef user fill:#e6efec,stroke:#315e55,color:#171717; classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
Key idea. Every step can be retried, so every step must be idempotent. The state machine says what may happen next; the ledger proves what did.
Key concepts
Authorization, capture, settlement
Authorization asks the card issuer to reserve an amount: a hold. No money moves yet. Capture tells the provider to charge a final amount against the hold. Settlement is when the provider actually moves captured funds to the merchant, usually in daily batches. A hold that is never captured expires and the reservation disappears.
Idempotency keys
A client-generated key attached to a request. The server records the key with the result. A retry with the same key returns the stored result instead of repeating the action.
State machines
A payment is always in exactly one state, such as authorized or captured, and only certain transitions are legal. Enforcing them in the database stops impossible sequences, such as capturing twice.
Double-entry ledger
Every money movement writes at least two entries that balance: debits equal credits. Entries are never changed; mistakes are fixed with reversing entries. Any balance is the sum of its entries, so it can always be recomputed and audited.
Reconciliation
Comparing your records with the provider's report of what actually happened, line by line, and investigating every difference.
Key idea. Authorize, capture, settle; make each step idempotent; enforce legal transitions; record everything as balanced, immutable entries; reconcile daily.
- Requirements
Before reading on. Restate the prompt in one sentence. Then list requirements, and name the property that must never break.
1.1 Functional requirements
- Place a hold when a customer orders.
- Capture a final amount, possibly larger by a tip, when the order is ready.
- Void a hold for cancelled orders.
- Handle provider results that arrive later by webhook.
- Settle all captured payments nightly, per provider.
- Reconcile with provider reports and flag differences.
- Report daily totals per store.
1.2 Non-functional requirements
- No double charges and no lost payments. Ever.
- Fast authorization. Tap to approval in about a second at p95; the line must move.
- Strong consistency for payment state and balances.
- Auditability. Every cent traceable to entries and provider references.
- Availability. A provider outage must not stop sales.
1.3 The constraint versus the property
Exactly-once money movement is the property. The authorization latency budget is the constraint. It forces the tap path to stay synchronous and short, with everything else pushed to asynchronous work.
- Back-of-the-envelope estimation
- Orders: 10,000 stores × 300 = 3 million a day.
- Average rate: 3 M / 86,400 ≈ 35 per second.
- Morning rush: if 40% of orders fall in 3 hours, that is 1.2 M / 10,800 s ≈ 110 per second, with spikes to a few hundred.
- Each order makes two provider calls (hold and capture) and writes about five rows (payment, idempotency record, three or four ledger entries).
- Rows per day: about 15 million; about 5 billion a year. Payment rows are small, a few hundred bytes, so storage is a few TB a year including indexes.
- Nightly batch: about 3 million captured payments, grouped into files per provider.
This fits one well-provisioned relational database with a replica for years. Say so, and then describe how you would shard when the chain grows.
- API design
Before reading on. The terminal sends a capture, and the network drops before the reply. What does the terminal do next, and what does the server do?
The terminal retries with the same idempotency key. The server finds the key, sees the capture already happened, and returns the stored result. The customer is charged once.
3.1 Holds and captures
POST /v1/holds
Idempotency-Key: 7f2c...
{order_id, store_id, card_token, amount: 450, currency: "USD"}
-> 201 {hold_id, status: "authorized", expires_at}
-> 402 {status: "declined", reason}
POST /v1/holds/:hold_id/capture
Idempotency-Key: 91ab...
{final_amount: 525} // 450 + 75 tip, within the allowed margin
-> 200 {payment_id, status: "captured"}
-> 409 {status: "already_voided" | "expired"}
POST /v1/holds/:hold_id/void
Idempotency-Key: c3d0...
-> 200 {status: "voided"}
GET /v1/payments/:id -> {status, amounts, provider_ref, history[]}3.2 Provider webhooks
POST /v1/webhooks/:provider
Headers: signature, timestamp
{event_id, type: "authorization.succeeded" | "capture.failed" | ..., provider_ref, ...}
-> 200 always once verified and recorded, even if already processedAmounts are integers in minor units. The API rejects decimals.
- Data model
CREATE TABLE payments (
id uuid PRIMARY KEY,
order_id uuid NOT NULL UNIQUE,
store_id uuid NOT NULL,
provider text NOT NULL,
provider_ref text UNIQUE,
status text NOT NULL CHECK (status IN
('pending','authorized','captured','settling','settled',
'voided','expired','failed')),
hold_amount bigint NOT NULL CHECK (hold_amount > 0),
final_amount bigint,
currency char(3) NOT NULL,
batch_id uuid,
authorized_at timestamptz,
captured_at timestamptz,
updated_at timestamptz NOT NULL,
version int NOT NULL DEFAULT 0
);
CREATE INDEX payments_settle_idx ON payments (status, provider, captured_at);
CREATE INDEX payments_store_day_idx ON payments (store_id, captured_at);
CREATE TABLE idempotency_keys (
key text PRIMARY KEY,
request_hash text NOT NULL,
response_code int,
response_body jsonb,
created_at timestamptz NOT NULL
);
CREATE TABLE ledger_entries (
id bigserial PRIMARY KEY,
txn_id uuid NOT NULL, -- groups the balanced entries
payment_id uuid NOT NULL,
account text NOT NULL, -- customer_receivable, store_revenue, tips_payable, ...
direction char(1) NOT NULL CHECK (direction IN ('D','C')),
amount bigint NOT NULL CHECK (amount > 0),
currency char(3) NOT NULL,
created_at timestamptz NOT NULL
);
CREATE TABLE settlement_batches (
id uuid PRIMARY KEY,
provider text NOT NULL,
business_date date NOT NULL,
status text NOT NULL, -- building | submitted | settled | failed
file_ref text,
payment_count int,
total_amount bigint,
UNIQUE (provider, business_date)
);
CREATE TABLE webhook_events (
provider text, event_id text, received_at timestamptz,
PRIMARY KEY (provider, event_id)
);The settlement index on (status, provider, captured_at) is what makes the nightly query a range scan instead of a full-table scan. Interviewers look for it.
- High-level design
5.1 One call that charges the card
The terminal calls "charge $5.25" when the order is ready. It cannot handle tips added after the hold, a timeout leaves the charge unknown, and a retry charges twice.
5.2 Fix 1: hold, then capture, with a state machine
Split the flow into authorize and capture, and give each payment an explicit state. Transitions happen only with conditional updates.
stateDiagram-v2 [*] --> pending pending --> authorized: provider approves pending --> failed: declined or error authorized --> captured: capture final amount authorized --> voided: order cancelled authorized --> expired: hold lapses captured --> settling: included in nightly batch settling --> settled: provider confirms settling --> captured: batch failed, retry tomorrow
5.3 Fix 2: idempotency keys at every write
Each hold, capture, and void carries a key. The key, a hash of the request, and the response are stored in the same transaction as the state change.
5.4 Fix 3: a double-entry ledger
Every state change that moves money writes balanced entries.
5.5 Fix 4: webhooks and a sweeper
Provider results that arrive late update the payment through the same state machine. A sweeper asks the provider about payments stuck in pending.
5.6 Fix 5: nightly settlement and reconciliation
A batch job builds one settlement file per provider, submits it with its own idempotency token, and a reconciliation job compares the provider's report with the ledger the next morning.
5.7 The composed design
sequenceDiagram autonumber participant T as Terminal participant P as Payment service participant D as Database participant V as Provider T->>P: POST /holds (key k1, 450) P->>D: insert idempotency k1, payment pending P->>V: authorize 450 (provider idempotency k1) V-->>P: approved, ref r9 P->>D: pending -> authorized, store response for k1 P-->>T: 201 authorized Note over T: order made, tip added T->>P: POST /holds/h1/capture (key k2, 525) P->>D: authorized -> captured (conditional), ledger entries P->>V: capture 525 on r9 V-->>P: ok P-->>T: 200 captured Note over P,V: 23:00 nightly batch per provider P->>D: captured -> settling, assign batch P->>V: settlement file (batch token) V-->>P: acknowledged, next morning, settlement report P->>D: settling -> settled, reconcile against ledger
Key idea. Split hold from capture, key every write, enforce transitions, post balanced entries, absorb late webhooks, and settle and reconcile as idempotent batches.
- Deep dives
6.1 Idempotency, end to end
Before reading on. Where exactly can a duplicate charge sneak in, and what stops each one?
| Boundary | Failure | Protection |
|---|---|---|
| Terminal to payment service | Timeout, terminal retries | Idempotency key stored with the result |
| Payment service to provider | Timeout, service retries | Pass the same key to the provider, which deduplicates too |
| Provider to webhook handler | Provider retries the webhook | Record (provider, event_id); skip if seen |
| Two captures racing | Two terminals or a double tap | Conditional update WHERE status = 'authorized' |
| Nightly batch rerun | Job crashes halfway | Batch row unique per provider and day; batch token sent to provider |
The stored idempotency record must be written in the same transaction as the state change. If the record is written first and the service crashes before the state change, a retry would return a result that never happened. If the state changes first and the record is lost, a retry repeats the action.
Also hash the request body with the key. The same key with a different amount is a client bug; reject it with 422 rather than return the old result.
What separates answers: idempotency
WeakRelies on retries being rare
Has no keys, so a timeout plus retry double-charges.
GoodKeys on the API
Stores idempotency keys for holds and captures and returns stored results on retry.
StrongKeys at every boundary
Passes keys to the provider, deduplicates webhooks by event ID, uses conditional transitions against races, gives the batch its own token, and writes key and state in one transaction.
6.2 The ledger
Before reading on. Write the ledger entries for a $4.50 order with a $0.75 tip, from capture to settlement.
At capture, a single balanced transaction:
| Account | Debit | Credit |
|---|---|---|
| customer_receivable (provider owes us) | 525 | |
| store_revenue | 450 | |
| tips_payable (owed to staff) | 75 |
At settlement, when the provider pays out:
| Account | Debit | Credit |
|---|---|---|
| cash_in_bank | 510 | |
| provider_fees | 15 | |
| customer_receivable | 525 |
Debits equal credits in each transaction. Entries are never updated or deleted. A refund writes new, reversing entries. Store revenue for a day is the sum of its credit entries; tips owed to staff are the balance of tips_payable. Because balances are derived, a bug in one report cannot silently change money: the entries remain the truth.
6.3 Late results and webhooks
The provider may approve an authorization after your request timed out, and tell you by webhook minutes later. Write pending before calling the provider, so a crash mid-call leaves a record. The webhook handler verifies the signature and timestamp, looks up the payment by provider reference, records the event ID, and applies the transition only if it is still legal. Then return 200, even for duplicates, so the provider stops retrying.
A sweeper runs every few minutes for payments in pending longer than a minute and asks the provider directly. If the provider says approved, move to authorized. If it has no record, mark failed, and the terminal asks the customer to tap again.
6.4 Nightly settlement
Before reading on. The batch job crashes after marking half the payments as settling and before sending the file. What happens when it reruns?
for each provider:
batch = INSERT INTO settlement_batches (provider, business_date, status)
VALUES (:p, :day, 'building')
ON CONFLICT (provider, business_date) DO NOTHING
RETURNING id
if no row returned: batch = SELECT existing; if status = 'submitted' skip
UPDATE payments SET status = 'settling', batch_id = :batch
WHERE status = 'captured' AND provider = :p AND captured_at < :cutoff
write file from SELECT ... WHERE batch_id = :batch -- deterministic
submit file to provider with batch id as idempotency token
UPDATE settlement_batches SET status = 'submitted' WHERE id = :batchOn rerun, the batch row already exists, so the job reuses it. The UPDATE picks up any captured payments not yet in the batch. The file is regenerated from batch_id, so it contains exactly the payments in the batch, including those marked before the crash. The provider deduplicates on the batch token, so a resubmission is harmless.
The index on (status, provider, captured_at) turns the selection into a range scan over captured rows only. Without it, the job scans every payment ever made.
6.5 Reconciliation
Each morning the provider sends a settlement report. Match it line by line against the batch, by provider reference:
- In the batch, not in the report: the provider did not settle it. Keep it
settling, retry in the next batch, alert after two days. - In the report, not in the batch: money arrived for something you did not submit. Investigate; often a manual capture or a duplicate.
- Amount differs: a partial capture, a fee, or a currency issue. Investigate.
Each mismatch goes to a queue for a person, with the payment, batch, and report line attached. Reconciliation is where silent bugs become visible.
6.6 Provider outages
Keep the tap path synchronous and short: terminal, payment service, provider, answer. Do not put a message queue in the authorization path; under load it adds latency exactly when the line is longest.
When a provider fails, a circuit breaker opens after several errors and stops sending it traffic for a short time. New holds go to the backup provider. Retries use exponential backoff with jitter. Holds already authorized with the failed provider must be captured with that provider; queue those captures and retry until it recovers. As a last resort for small amounts, some merchants accept offline approval with a stored card token and capture later, taking on the risk of a decline.
6.7 Consistency and scale
Money needs strong consistency: the state machine and the ledger live in one relational database with transactions. Eventual consistency for balances is a classic ding in this round.
When the chain grows tenfold, shard by store ID. Each store's payments, ledger entries, and daily reports stay on one shard, and the settlement job runs per shard. Sharding by payment ID spreads writes more evenly but scatters every store report across shards. Idempotency keys must live on the same shard as the payment they protect, so derive the shard from the store ID in the request.
6.8 Walking the failure cases
Before reading on. For each failure below, say what the customer experiences and what the system does.
| Failure | Customer sees | System does |
|---|---|---|
| Terminal times out on the hold request | "Processing..." then success on retry | Retry with the same key returns the stored result or completes the pending call |
| Provider approves, but the response is lost | Nothing unusual after retry | The payment is pending; the sweeper or webhook moves it to authorized; the terminal's retry returns it |
| Barista double-taps Capture | One charge | Second capture finds status captured and returns the stored response for its key, or a 409 for a new key |
| Order cancelled after the hold | No charge; hold released | Void moves authorized to voided and releases the hold with the provider |
| Order never captured (bug) | Hold disappears after days | Nightly job voids stale holds after an hour past close; alerts on counts |
| Provider down at the morning rush | Payment succeeds a little slower | Breaker opens; new holds route to the backup provider |
| Batch job crashes mid-run | Nothing | Rerun reuses the batch row; regenerates the file; provider deduplicates by token |
| Provider settles an amount you never submitted | Nothing | Reconciliation flags it for review the next morning |
This table is the kind of artifact interviewers remember. It shows every boundary has an answer.
6.9 The tip and the hold
A tip can exceed the hold. Providers typically allow capturing somewhat more than the authorized amount for card-present tips, within limits set by the card networks and the merchant's category. Design for the limit: validate the final amount against the allowed margin before capture. If a customer tips beyond it, capture the allowed maximum and charge the rest as a separate small transaction, or ask for a new authorization.
Record the tip as its own ledger line, tips_payable, so payroll can pay it out and reports can show it apart from revenue.
What separates answers: completeness
WeakOnly the happy path
Describes hold and capture with no failure cases.
GoodMain failures covered
Handles retries, webhooks, and batch reruns.
StrongEvery boundary has an answer
Walks each failure with what the customer sees and what the system does, handles tips beyond the hold, and cleans up stale holds.
- Variants
7.1 Generic payment system
If the prompt really is a generic payment platform, the same core applies at larger scale: a payment orchestrator, several providers, wallets, a ledger service, and reconciliation, with regional deployments for data residency.
7.2 Online checkout
Online orders add 3-D Secure challenges, fraud scoring before authorization, and inventory reservation that must be released if payment fails.
7.3 Refunds
A refund is a new transaction with its own idempotency key and reversing ledger entries. Partial refunds reference the original payment and cap the total at the captured amount.
7.4 At ten times the stores
At 100,000 stores, peak authorizations approach 1,100 per second with spikes several times higher. Shard by store ID across several database clusters, with a routing table from store to shard. Run settlement per shard in parallel, and reconcile per shard, then roll up. Provider rate limits become real: spread traffic across several merchant accounts per provider and region.
- The transferable pattern
Payments are an idempotent state machine over an immutable ledger, reconciled against an outside source of truth. The same shape governs order fulfillment, inventory, and any workflow that crosses systems you do not control. Make every step safe to repeat, make illegal transitions impossible, record facts as append-only entries, and compare against the outside world daily.
Review: the 30-second answer
- Read the prompt. Hold at order, capture with tip, settle nightly.
- Idempotency at every boundary. API, provider, webhooks, and the batch.
- State machine with conditional updates. No double capture, no impossible sequence.
- Double-entry ledger in minor units. Balances derived, entries immutable.
- Nightly batch with its own token and the right index; reconcile each morning. Synchronous tap path, circuit breakers, and a backup provider.
Quiz
+Why store the idempotency record in the same transaction as the state change?
If they are written separately, a crash between them either returns a result that never happened or repeats an action that did. One transaction makes them succeed or fail together.
+Why use integer minor units for money?
Floating-point numbers cannot represent most decimal fractions exactly, so sums drift by fractions of a cent. Integers in cents are exact.
+What index does the nightly settlement query need, and why?
An index on status, provider, and capture time. The job selects captured payments per provider before a cutoff; the index makes that a range scan instead of a scan of every payment ever made.
+How does the batch job stay safe when it is rerun after a crash?
The batch row is unique per provider and day, so the rerun reuses it. The file is regenerated from payments assigned to that batch, and the provider deduplicates on the batch token.
+Why keep the authorization path synchronous?
The customer is waiting at the counter. A queue in the path adds latency under load, exactly when the line is longest. Only settlement and reporting are pushed to asynchronous jobs.
+What happens when the barista taps Capture twice?
The second request either repeats the first idempotency key and gets the stored response, or uses a new key and finds the payment already captured, so the conditional update changes nothing and the API returns a conflict. The customer is charged once.
+How should the system handle a tip larger than the allowed capture margin?
Validate the final amount against the margin before capture. Capture up to the allowed amount and charge the remainder separately, or request a new authorization.
Sources and further reading
- Stripe: Idempotent requests describes idempotency keys as a provider implements them.
- Stripe: Place a hold on a payment method explains authorize-then-capture and hold expiry.
- Martin Fowler: Accounting Patterns covers entries, accounts, and reversal adjustments.
- Idempotent webhook handling returns from the sender's side in webhook delivery.
