The Forward Deployed

OpenAI Interview: Design Payments for a Coffee Shop

A full solution to the OpenAI payment question: reading the actual prompt, holds and captures, idempotency keys, the payment state machine, a double-entry ledger, provider webhooks, nightly settlement with the right index, reconciliation, resilience, and sharding.

By Reviewed

Part of the OpenAI system design question bank. The question is representative of the round. The analysis and solution are this site's own.

Problem statement

The title may say "payments", and the question is a specific flow. Candidates are reported to fail after delivering a polished, generic payment-system design to a narrower question. Read the prompt twice, restate it, and design exactly that flow: authorize a hold, capture a final amount, settle nightly, and reconcile.

Fraud scoring, loyalty points, and inventory are out of scope unless the interviewer adds them.

Clarifying questions

  • Card present or online? Assume a tap-to-pay terminal in the store, plus mobile orders ahead.
  • Can the final amount exceed the hold? Yes, up to a limit such as 20% over, for tips.
  • How long does a hold last? The provider releases it after several days if not captured. Assume orders are captured within an hour.
  • Partial payments or split tenders? No, one card per order.
  • How many providers? Two, for resilience and cost.
  • Scale? For practice: 10,000 stores, 300 orders per store per day, peaks at the morning rush.
  • Money in what form? Integer minor units, such as cents, never floating point.

What makes payments hard

Money moves across systems you do not control, over networks that fail at the worst moment.

A request can time out after the provider charged the card, so the client retries and the customer pays twice. A provider's result can arrive minutes later by webhook, and it can arrive twice. A nightly batch can fail halfway, and rerunning it must not settle anything twice. Every one of those is a real-money error that ends up in a support ticket or a chargeback.

So the driving tension is availability versus exactly-once money movement. The till must never block the line, and money must never move twice or disappear. The answer is idempotency at every boundary, an explicit state machine, and a ledger that can prove every balance.

flowchart LR
  POS([Store terminal]):::user --> PS[Payment service]:::svc
  PS --> PSP[Payment provider]:::store
  PSP -.webhook.-> PS
  PS --> DB[(Payments + ledger)]:::store
  B[Nightly settlement]:::svc --> DB
  B --> PSP
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
Key idea. Every step can be retried, so every step must be idempotent. The state machine says what may happen next; the ledger proves what did.

Key concepts

Authorization, capture, settlement

Authorization asks the card issuer to reserve an amount: a hold. No money moves yet. Capture tells the provider to charge a final amount against the hold. Settlement is when the provider actually moves captured funds to the merchant, usually in daily batches. A hold that is never captured expires and the reservation disappears.

Idempotency keys

A client-generated key attached to a request. The server records the key with the result. A retry with the same key returns the stored result instead of repeating the action.

State machines

A payment is always in exactly one state, such as authorized or captured, and only certain transitions are legal. Enforcing them in the database stops impossible sequences, such as capturing twice.

Double-entry ledger

Every money movement writes at least two entries that balance: debits equal credits. Entries are never changed; mistakes are fixed with reversing entries. Any balance is the sum of its entries, so it can always be recomputed and audited.

Reconciliation

Comparing your records with the provider's report of what actually happened, line by line, and investigating every difference.

Key idea. Authorize, capture, settle; make each step idempotent; enforce legal transitions; record everything as balanced, immutable entries; reconcile daily.

  1. Requirements

Before reading on. Restate the prompt in one sentence. Then list requirements, and name the property that must never break.

1.1 Functional requirements

  • Place a hold when a customer orders.
  • Capture a final amount, possibly larger by a tip, when the order is ready.
  • Void a hold for cancelled orders.
  • Handle provider results that arrive later by webhook.
  • Settle all captured payments nightly, per provider.
  • Reconcile with provider reports and flag differences.
  • Report daily totals per store.

1.2 Non-functional requirements

  • No double charges and no lost payments. Ever.
  • Fast authorization. Tap to approval in about a second at p95; the line must move.
  • Strong consistency for payment state and balances.
  • Auditability. Every cent traceable to entries and provider references.
  • Availability. A provider outage must not stop sales.

1.3 The constraint versus the property

Exactly-once money movement is the property. The authorization latency budget is the constraint. It forces the tap path to stay synchronous and short, with everything else pushed to asynchronous work.

  1. Back-of-the-envelope estimation

  • Orders: 10,000 stores × 300 = 3 million a day.
  • Average rate: 3 M / 86,400 ≈ 35 per second.
  • Morning rush: if 40% of orders fall in 3 hours, that is 1.2 M / 10,800 s ≈ 110 per second, with spikes to a few hundred.
  • Each order makes two provider calls (hold and capture) and writes about five rows (payment, idempotency record, three or four ledger entries).
  • Rows per day: about 15 million; about 5 billion a year. Payment rows are small, a few hundred bytes, so storage is a few TB a year including indexes.
  • Nightly batch: about 3 million captured payments, grouped into files per provider.

This fits one well-provisioned relational database with a replica for years. Say so, and then describe how you would shard when the chain grows.

  1. API design

Before reading on. The terminal sends a capture, and the network drops before the reply. What does the terminal do next, and what does the server do?

The terminal retries with the same idempotency key. The server finds the key, sees the capture already happened, and returns the stored result. The customer is charged once.

3.1 Holds and captures

POST /v1/holds
Idempotency-Key: 7f2c...
{order_id, store_id, card_token, amount: 450, currency: "USD"}
-> 201 {hold_id, status: "authorized", expires_at}
-> 402 {status: "declined", reason}

POST /v1/holds/:hold_id/capture
Idempotency-Key: 91ab...
{final_amount: 525}          // 450 + 75 tip, within the allowed margin
-> 200 {payment_id, status: "captured"}
-> 409 {status: "already_voided" | "expired"}

POST /v1/holds/:hold_id/void
Idempotency-Key: c3d0...
-> 200 {status: "voided"}

GET  /v1/payments/:id  -> {status, amounts, provider_ref, history[]}

3.2 Provider webhooks

POST /v1/webhooks/:provider
Headers: signature, timestamp
{event_id, type: "authorization.succeeded" | "capture.failed" | ..., provider_ref, ...}
-> 200 always once verified and recorded, even if already processed

Amounts are integers in minor units. The API rejects decimals.

  1. Data model

CREATE TABLE payments (
  id             uuid PRIMARY KEY,
  order_id       uuid NOT NULL UNIQUE,
  store_id       uuid NOT NULL,
  provider       text NOT NULL,
  provider_ref   text UNIQUE,
  status         text NOT NULL CHECK (status IN
                 ('pending','authorized','captured','settling','settled',
                  'voided','expired','failed')),
  hold_amount    bigint NOT NULL CHECK (hold_amount > 0),
  final_amount   bigint,
  currency       char(3) NOT NULL,
  batch_id       uuid,
  authorized_at  timestamptz,
  captured_at    timestamptz,
  updated_at     timestamptz NOT NULL,
  version        int NOT NULL DEFAULT 0
);
CREATE INDEX payments_settle_idx ON payments (status, provider, captured_at);
CREATE INDEX payments_store_day_idx ON payments (store_id, captured_at);

CREATE TABLE idempotency_keys (
  key            text PRIMARY KEY,
  request_hash   text NOT NULL,
  response_code  int,
  response_body  jsonb,
  created_at     timestamptz NOT NULL
);

CREATE TABLE ledger_entries (
  id             bigserial PRIMARY KEY,
  txn_id         uuid NOT NULL,          -- groups the balanced entries
  payment_id     uuid NOT NULL,
  account        text NOT NULL,          -- customer_receivable, store_revenue, tips_payable, ...
  direction      char(1) NOT NULL CHECK (direction IN ('D','C')),
  amount         bigint NOT NULL CHECK (amount > 0),
  currency       char(3) NOT NULL,
  created_at     timestamptz NOT NULL
);

CREATE TABLE settlement_batches (
  id             uuid PRIMARY KEY,
  provider       text NOT NULL,
  business_date  date NOT NULL,
  status         text NOT NULL,          -- building | submitted | settled | failed
  file_ref       text,
  payment_count  int,
  total_amount   bigint,
  UNIQUE (provider, business_date)
);

CREATE TABLE webhook_events (
  provider text, event_id text, received_at timestamptz,
  PRIMARY KEY (provider, event_id)
);

The settlement index on (status, provider, captured_at) is what makes the nightly query a range scan instead of a full-table scan. Interviewers look for it.

  1. High-level design

5.1 One call that charges the card

The terminal calls "charge $5.25" when the order is ready. It cannot handle tips added after the hold, a timeout leaves the charge unknown, and a retry charges twice.

5.2 Fix 1: hold, then capture, with a state machine

Split the flow into authorize and capture, and give each payment an explicit state. Transitions happen only with conditional updates.

stateDiagram-v2
  [*] --> pending
  pending --> authorized: provider approves
  pending --> failed: declined or error
  authorized --> captured: capture final amount
  authorized --> voided: order cancelled
  authorized --> expired: hold lapses
  captured --> settling: included in nightly batch
  settling --> settled: provider confirms
  settling --> captured: batch failed, retry tomorrow

5.3 Fix 2: idempotency keys at every write

Each hold, capture, and void carries a key. The key, a hash of the request, and the response are stored in the same transaction as the state change.

5.4 Fix 3: a double-entry ledger

Every state change that moves money writes balanced entries.

5.5 Fix 4: webhooks and a sweeper

Provider results that arrive late update the payment through the same state machine. A sweeper asks the provider about payments stuck in pending.

5.6 Fix 5: nightly settlement and reconciliation

A batch job builds one settlement file per provider, submits it with its own idempotency token, and a reconciliation job compares the provider's report with the ledger the next morning.

5.7 The composed design

sequenceDiagram
  autonumber
  participant T as Terminal
  participant P as Payment service
  participant D as Database
  participant V as Provider
  T->>P: POST /holds (key k1, 450)
  P->>D: insert idempotency k1, payment pending
  P->>V: authorize 450 (provider idempotency k1)
  V-->>P: approved, ref r9
  P->>D: pending -> authorized, store response for k1
  P-->>T: 201 authorized
  Note over T: order made, tip added
  T->>P: POST /holds/h1/capture (key k2, 525)
  P->>D: authorized -> captured (conditional), ledger entries
  P->>V: capture 525 on r9
  V-->>P: ok
  P-->>T: 200 captured
  Note over P,V: 23:00 nightly batch per provider
  P->>D: captured -> settling, assign batch
  P->>V: settlement file (batch token)
  V-->>P: acknowledged, next morning, settlement report
  P->>D: settling -> settled, reconcile against ledger
Key idea. Split hold from capture, key every write, enforce transitions, post balanced entries, absorb late webhooks, and settle and reconcile as idempotent batches.

  1. Deep dives

6.1 Idempotency, end to end

Before reading on. Where exactly can a duplicate charge sneak in, and what stops each one?
BoundaryFailureProtection
Terminal to payment serviceTimeout, terminal retriesIdempotency key stored with the result
Payment service to providerTimeout, service retriesPass the same key to the provider, which deduplicates too
Provider to webhook handlerProvider retries the webhookRecord (provider, event_id); skip if seen
Two captures racingTwo terminals or a double tapConditional update WHERE status = 'authorized'
Nightly batch rerunJob crashes halfwayBatch row unique per provider and day; batch token sent to provider

The stored idempotency record must be written in the same transaction as the state change. If the record is written first and the service crashes before the state change, a retry would return a result that never happened. If the state changes first and the record is lost, a retry repeats the action.

Also hash the request body with the key. The same key with a different amount is a client bug; reject it with 422 rather than return the old result.

What separates answers: idempotency

WeakRelies on retries being rare

Has no keys, so a timeout plus retry double-charges.

GoodKeys on the API

Stores idempotency keys for holds and captures and returns stored results on retry.

StrongKeys at every boundary

Passes keys to the provider, deduplicates webhooks by event ID, uses conditional transitions against races, gives the batch its own token, and writes key and state in one transaction.

6.2 The ledger

Before reading on. Write the ledger entries for a $4.50 order with a $0.75 tip, from capture to settlement.

At capture, a single balanced transaction:

AccountDebitCredit
customer_receivable (provider owes us)525
store_revenue450
tips_payable (owed to staff)75

At settlement, when the provider pays out:

AccountDebitCredit
cash_in_bank510
provider_fees15
customer_receivable525

Debits equal credits in each transaction. Entries are never updated or deleted. A refund writes new, reversing entries. Store revenue for a day is the sum of its credit entries; tips owed to staff are the balance of tips_payable. Because balances are derived, a bug in one report cannot silently change money: the entries remain the truth.

6.3 Late results and webhooks

The provider may approve an authorization after your request timed out, and tell you by webhook minutes later. Write pending before calling the provider, so a crash mid-call leaves a record. The webhook handler verifies the signature and timestamp, looks up the payment by provider reference, records the event ID, and applies the transition only if it is still legal. Then return 200, even for duplicates, so the provider stops retrying.

A sweeper runs every few minutes for payments in pending longer than a minute and asks the provider directly. If the provider says approved, move to authorized. If it has no record, mark failed, and the terminal asks the customer to tap again.

6.4 Nightly settlement

Before reading on. The batch job crashes after marking half the payments as settling and before sending the file. What happens when it reruns?
for each provider:
  batch = INSERT INTO settlement_batches (provider, business_date, status)
          VALUES (:p, :day, 'building')
          ON CONFLICT (provider, business_date) DO NOTHING
          RETURNING id
  if no row returned: batch = SELECT existing; if status = 'submitted' skip

  UPDATE payments SET status = 'settling', batch_id = :batch
   WHERE status = 'captured' AND provider = :p AND captured_at < :cutoff

  write file from SELECT ... WHERE batch_id = :batch       -- deterministic
  submit file to provider with batch id as idempotency token
  UPDATE settlement_batches SET status = 'submitted' WHERE id = :batch

On rerun, the batch row already exists, so the job reuses it. The UPDATE picks up any captured payments not yet in the batch. The file is regenerated from batch_id, so it contains exactly the payments in the batch, including those marked before the crash. The provider deduplicates on the batch token, so a resubmission is harmless.

The index on (status, provider, captured_at) turns the selection into a range scan over captured rows only. Without it, the job scans every payment ever made.

6.5 Reconciliation

Each morning the provider sends a settlement report. Match it line by line against the batch, by provider reference:

  • In the batch, not in the report: the provider did not settle it. Keep it settling, retry in the next batch, alert after two days.
  • In the report, not in the batch: money arrived for something you did not submit. Investigate; often a manual capture or a duplicate.
  • Amount differs: a partial capture, a fee, or a currency issue. Investigate.

Each mismatch goes to a queue for a person, with the payment, batch, and report line attached. Reconciliation is where silent bugs become visible.

6.6 Provider outages

Keep the tap path synchronous and short: terminal, payment service, provider, answer. Do not put a message queue in the authorization path; under load it adds latency exactly when the line is longest.

When a provider fails, a circuit breaker opens after several errors and stops sending it traffic for a short time. New holds go to the backup provider. Retries use exponential backoff with jitter. Holds already authorized with the failed provider must be captured with that provider; queue those captures and retry until it recovers. As a last resort for small amounts, some merchants accept offline approval with a stored card token and capture later, taking on the risk of a decline.

6.7 Consistency and scale

Money needs strong consistency: the state machine and the ledger live in one relational database with transactions. Eventual consistency for balances is a classic ding in this round.

When the chain grows tenfold, shard by store ID. Each store's payments, ledger entries, and daily reports stay on one shard, and the settlement job runs per shard. Sharding by payment ID spreads writes more evenly but scatters every store report across shards. Idempotency keys must live on the same shard as the payment they protect, so derive the shard from the store ID in the request.

6.8 Walking the failure cases

Before reading on. For each failure below, say what the customer experiences and what the system does.
FailureCustomer seesSystem does
Terminal times out on the hold request"Processing..." then success on retryRetry with the same key returns the stored result or completes the pending call
Provider approves, but the response is lostNothing unusual after retryThe payment is pending; the sweeper or webhook moves it to authorized; the terminal's retry returns it
Barista double-taps CaptureOne chargeSecond capture finds status captured and returns the stored response for its key, or a 409 for a new key
Order cancelled after the holdNo charge; hold releasedVoid moves authorized to voided and releases the hold with the provider
Order never captured (bug)Hold disappears after daysNightly job voids stale holds after an hour past close; alerts on counts
Provider down at the morning rushPayment succeeds a little slowerBreaker opens; new holds route to the backup provider
Batch job crashes mid-runNothingRerun reuses the batch row; regenerates the file; provider deduplicates by token
Provider settles an amount you never submittedNothingReconciliation flags it for review the next morning

This table is the kind of artifact interviewers remember. It shows every boundary has an answer.

6.9 The tip and the hold

A tip can exceed the hold. Providers typically allow capturing somewhat more than the authorized amount for card-present tips, within limits set by the card networks and the merchant's category. Design for the limit: validate the final amount against the allowed margin before capture. If a customer tips beyond it, capture the allowed maximum and charge the rest as a separate small transaction, or ask for a new authorization.

Record the tip as its own ledger line, tips_payable, so payroll can pay it out and reports can show it apart from revenue.

What separates answers: completeness

WeakOnly the happy path

Describes hold and capture with no failure cases.

GoodMain failures covered

Handles retries, webhooks, and batch reruns.

StrongEvery boundary has an answer

Walks each failure with what the customer sees and what the system does, handles tips beyond the hold, and cleans up stale holds.

  1. Variants

7.1 Generic payment system

If the prompt really is a generic payment platform, the same core applies at larger scale: a payment orchestrator, several providers, wallets, a ledger service, and reconciliation, with regional deployments for data residency.

7.2 Online checkout

Online orders add 3-D Secure challenges, fraud scoring before authorization, and inventory reservation that must be released if payment fails.

7.3 Refunds

A refund is a new transaction with its own idempotency key and reversing ledger entries. Partial refunds reference the original payment and cap the total at the captured amount.

7.4 At ten times the stores

At 100,000 stores, peak authorizations approach 1,100 per second with spikes several times higher. Shard by store ID across several database clusters, with a routing table from store to shard. Run settlement per shard in parallel, and reconcile per shard, then roll up. Provider rate limits become real: spread traffic across several merchant accounts per provider and region.

  1. The transferable pattern

Payments are an idempotent state machine over an immutable ledger, reconciled against an outside source of truth. The same shape governs order fulfillment, inventory, and any workflow that crosses systems you do not control. Make every step safe to repeat, make illegal transitions impossible, record facts as append-only entries, and compare against the outside world daily.

Review: the 30-second answer

  • Read the prompt. Hold at order, capture with tip, settle nightly.
  • Idempotency at every boundary. API, provider, webhooks, and the batch.
  • State machine with conditional updates. No double capture, no impossible sequence.
  • Double-entry ledger in minor units. Balances derived, entries immutable.
  • Nightly batch with its own token and the right index; reconcile each morning. Synchronous tap path, circuit breakers, and a backup provider.

Quiz

+Why store the idempotency record in the same transaction as the state change?

If they are written separately, a crash between them either returns a result that never happened or repeats an action that did. One transaction makes them succeed or fail together.

+Why use integer minor units for money?

Floating-point numbers cannot represent most decimal fractions exactly, so sums drift by fractions of a cent. Integers in cents are exact.

+What index does the nightly settlement query need, and why?

An index on status, provider, and capture time. The job selects captured payments per provider before a cutoff; the index makes that a range scan instead of a scan of every payment ever made.

+How does the batch job stay safe when it is rerun after a crash?

The batch row is unique per provider and day, so the rerun reuses it. The file is regenerated from payments assigned to that batch, and the provider deduplicates on the batch token.

+Why keep the authorization path synchronous?

The customer is waiting at the counter. A queue in the path adds latency under load, exactly when the line is longest. Only settlement and reporting are pushed to asynchronous jobs.

+What happens when the barista taps Capture twice?

The second request either repeats the first idempotency key and gets the stored response, or uses a new key and finds the payment already captured, so the conditional update changes nothing and the API returns a conflict. The customer is charged once.

+How should the system handle a tip larger than the allowed capture margin?

Validate the final amount against the margin before capture. Capture up to the allowed amount and charge the remainder separately, or request a new authorization.

Sources and further reading

NextOnline Chess Platform