The Forward Deployed

Anthropic Interview: The Design-Doc Review Round

A full guide to the Anthropic design-doc review round: how the hour runs, a reading method, two practice documents with planted flaws, ranked critiques, and the follow-ups interviewers drive.

By Reviewed

Part of the Anthropic system design question bank. The question is representative of the round. The practice documents and critiques are this site's own.

Problem statement

This round replaces the blank-page design with a critique. The interviewer shares a document, usually a few pages, sometimes close to a system the team runs. It has weaknesses on purpose, and some of them are absences. Your job is to find what matters, say why, propose a fix, and then defend your points as the interviewer pushes.

The round is reported in engineering-manager loops and, more often now, as the system-design slot for senior engineers. The inference server from the inference API question is a reported subject.

Clarifying questions

Ask these before you critique. Each one changes how you rank what you find.

  • What stage is this document at? A v1 for an internal tool tolerates shortcuts that a public launch does not.
  • Who are the users, and what do they need most? Latency-sensitive users rank latency flaws higher than cost flaws.
  • What constraints are fixed? If hardware, a vendor, or a deadline is pinned, critique inside it.
  • What scale must v1 handle, and what scale in a year? A design that is fine now and breaks at 10 times the load is a different finding from one that is broken today.
  • Is there an existing system this replaces? Migration risk may be the biggest risk in the document.

What makes the round hard

Three things trip people up.

First, everything looks like a problem. A deliberately flawed document has 20 things you could say. Saying all 20 at equal weight tells the interviewer you cannot rank risk.

Second, the worst flaws are often absences. The document says nothing about failure handling, or rollback, or who gets paged. You only see those if you read against a checklist.

Third, the round is a conversation. After your first points, the interviewer picks one and pushes: why does that matter, what would you do instead, what does your fix cost? A critique you cannot defend scores worse than one you never made.

So the driving tension is coverage versus judgment. You must read broadly enough to find the absences, and then lead with only the few findings that matter.

flowchart LR
  R[Read for intent]:::svc --> C[Read with a checklist]:::svc
  C --> L[Long list of findings]:::bad
  L --> K[Rank by impact x likelihood]:::svc
  K --> T[Top 3 with fixes]:::store
  T --> D[Defend under follow-ups]:::store
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef bad fill:#fbe9e4,stroke:#c4492d,color:#171717;
Key idea. The grade comes from ranking. Read broadly to find the absences, then lead with the few findings that would hurt most.

Key concepts

Severity is impact times likelihood

Rank each finding on two axes. How bad is it when it happens: data loss, an outage, a wrong answer, or a slow page? How likely is it: every day, at peak, or once a year? A likely data loss beats a rare slowdown. A flaw that caps the whole system beats a flaw in one component.

SeverityExampleHow to present it
BlockingData loss on a common failure, a hard throughput cap, no rollbackLead with it; ask for a change before launch
MajorA single point of failure, missing isolation between usersSecond tier; propose a fix and its cost
MinorNaming, an unclear diagram, a slightly slow pathMention briefly or skip
QuestionSomething that may be deliberateAsk before you criticize

Wrong versus missing

A wrong finding points at a sentence in the document and says it will fail. A missing finding points at a topic the document never covers. Both matter. Interviewers plant both, and the missing ones separate strong candidates, because they require a mental checklist.

Critique inside the constraints

If the document pins a constraint, such as a fixed batch function, no autoscaling, or a vendor, do not argue for removing it unless it makes the goal impossible. Redesigning the system from scratch shows you did not accept the problem. Improving the design within its limits shows you can work on a real team.

The anatomy of a good comment

A good review comment has three parts: the risk, the evidence in the document, and the change you want. Add the cost of the change when it is not obvious.

Key idea. Rank by impact times likelihood, look for absences as hard as for errors, respect pinned constraints, and write every finding as risk, evidence, change.

  1. How the hour runs

A typical shape, which you can propose if the interviewer does not set one:

MinutesWhat you do
0 to 5Ask the clarifying questions above
5 to 15Read: once for intent, once with the checklist
15 to 17Summarize the design in two sentences
17 to 30Present the top three findings, ranked, each with a fix
30 to 35List what is missing
35 to 55Follow-ups: the interviewer pushes on your points

The two-sentence summary matters more than it looks. It proves you understood the design, and it gives the interviewer a chance to correct a misreading before you build on it.

  1. A reading method

Read the document twice.

The first read is for intent. What problem does it solve, for whom, and what is the core mechanism? Do not take notes on flaws yet.

The second read is with a checklist. Mark each item as covered, weak, or absent.

AreaWhat to look for
Goals and non-goalsAre success measures stated? What is explicitly out of scope?
AlternativesWere other designs considered, and why were they rejected?
Scale and costAre there numbers? Do they add up?
Data flowCan you trace one request end to end?
Failure modesWhat happens when each component dies, slows down, or returns bad data?
ConsistencyCan two actors race? Is anything written twice?
IsolationCan one user or team hurt the others?
Rollout and rollbackHow does it ship, and how does it un-ship?
ObservabilityWhat is measured, what alerts, who is paged?
Security and privacyWho can read what? Are secrets and user data handled?
MigrationHow does traffic move from the old system?

Most planted absences sit in failure modes, rollout, and observability.

  1. Practice document 1: a batch inference gateway

Read it with a timer before you read the critique.

DESIGN: Batch inference gateway, v1

Goal: serve model M to internal teams through one HTTP endpoint.

Design:
- Clients POST a request. The gateway pushes it onto one FIFO list in Redis.
- A dispatcher wakes every 500 ms, pops up to 32 requests,
  and sends them as one batch to a free GPU host.
- Each GPU host runs one model replica.
- Results are written to Redis under the request id.
  Clients poll GET /result/<id> once per second.

Scaling: on-call adds GPU hosts by hand when the queue passes 1,000.

Rollout: new model versions go to all hosts at once, on Saturday.

Monitoring: a dashboard of GPU utilization per host.

Two-sentence summary: The gateway queues requests in one Redis list and a timer-driven dispatcher sends batches of up to 32 to free GPU hosts. Clients poll Redis for results, scaling is manual, and new versions deploy to every host at once.

flowchart LR
  C([Clients]):::user -->|POST| G[Gateway]:::svc --> Q[(Redis FIFO list)]:::store
  T((500 ms tick)):::bad --> D[Dispatcher]:::svc
  Q --> D -->|batch of up to 32| H[GPU hosts]:::store
  H --> RS[(Redis results)]:::store
  C -->|poll every 1 s| RS
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef bad fill:#fbe9e4,stroke:#c4492d,color:#171717;

  1. The ranked critique

4.1 Blocking: the tick caps throughput and adds latency

Risk. Dispatch happens only on the 500 ms tick, 32 requests at a time. That is at most 64 requests per second, however many GPU hosts exist. Every request also waits up to 500 ms before it reaches a GPU, even when the system is idle.

Evidence. "A dispatcher wakes every 500 ms, pops up to 32 requests."

Change. Flush a batch when it reaches the size limit or when the oldest request has waited a short timeout, whichever comes first, and dispatch whenever a GPU becomes free. The timeout bounds latency at low load. The size trigger removes the throughput cap at high load.

This leads because it caps the whole system. Adding hosts, which is the document's scaling plan, does nothing until this changes.

flowchart LR
  subgraph Before["Before: fixed tick"]
    direction LR
    T1((every 500 ms)):::bad --> P1[pop 32]:::svc
  end
  subgraph After["After: size or timeout"]
    direction LR
    Q2[(queue)]:::store --> F{32 queued OR<br/>oldest waited 50 ms?}:::svc
    F -->|yes| P2[dispatch to any free GPU]:::svc
  end
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef bad fill:#fbe9e4,stroke:#c4492d,color:#171717;

4.2 Major: one FIFO list lets one team starve the rest

Risk. A single batch job from one team fills the list, and every interactive caller behind it waits minutes.

Evidence. "The gateway pushes it onto one FIFO list."

Change. One queue per team, or per priority class, drained with weighted fair scheduling, plus a per-team quota at the gateway. Ask the author whether any callers are interactive. The answer sets how strict the priority must be.

4.3 Major: all-at-once rollout with no rollback

Risk. A bad model version reaches every host together. There is no stated way back, and Saturday means fewer people are around when it breaks.

Evidence. "New model versions go to all hosts at once, on Saturday."

Change. Canary on a few hosts, compare error rate, latency, and output checks against the old version, then roll forward in steps. Keep the old weights on disk so rollback is a restart. Gate traffic on a readiness check that confirms the right version loaded. Ship on a weekday.

4.4 What is missing

  • Failure handling. What happens to a batch when its GPU host dies mid-run? Nothing re-queues it, so those requests hang forever.
  • A single point of failure. Redis holds both the queue and the results. If it restarts without persistence, every queued request and unread result disappears.
  • Idempotency. A client that times out and retries creates a second request and a second GPU run.
  • Capacity math. No request rate, no latency target, no GPU count.
  • Useful monitoring. GPU utilization does not show user pain. Queue wait, p95 latency, error rate, and per-team usage do.
  • Result expiry. Nothing deletes results, so Redis memory grows until it fails.

4.5 Questions, not criticisms

Manual scaling and one-second polling may be deliberate for an internal v1. Ask. If traffic is modest and callers are batch jobs, polling is cheap to build and fine to run. Say you would accept it for v1, and name the point where it breaks: polling load grows with the number of waiting clients, and a one-second poll adds up to a second of latency. That answer shows judgment. Demanding server-sent events for a batch tool does not.

What separates answers: the first critique

WeakA flat list of everything

Twenty equal points, starting with naming and diagrams. No ranking, no fixes, nothing about what is absent.

GoodRanked findings with fixes

Leads with the tick cap, then starvation and rollout, each with evidence and a change, plus a list of missing topics.

StrongRanked, costed, and calibrated

Explains why the tick cap outranks everything (it caps the system and defeats the scaling plan), separates blocking from major, asks before criticizing deliberate v1 choices, and states the cost of each fix.

  1. Follow-ups the interviewer drives

5.1 "Why not just shorten the tick to 50 ms?"

Before reading on. It seems like an easy fix. What goes wrong?

A 50 ms tick raises the cap to 640 requests per second and cuts idle latency. But it still sends half-empty batches at low load, and at high load the queue can still hold more than 32 requests at the tick, so requests wait extra ticks. More important, the tick still ignores GPU availability. It may pop a batch when no GPU is free. The size-or-timeout trigger, combined with dispatch on GPU availability, adapts to both regimes. Say that a shorter tick is a fine stopgap for this week, and the trigger is the fix.

5.2 "The author says Redis is reliable enough. Push back or accept?"

Ask what "enough" means. If losing queued requests on a Redis restart is acceptable, because clients retry, then accept it and ask for client retries with idempotency keys. If it is not acceptable, the fix is modest: turn on append-only persistence and run a replica with failover. Show that you can accept a risk when the cost of fixing it exceeds the damage. Reviewers who fight every point lose trust.

5.3 "How would you know the fix worked?"

State the metrics before and after. Queue wait at p50 and p95. Batch fill ratio: requests per batch divided by 32. GPU busy time per host. Requests per second at saturation. The size-or-timeout change should raise the saturation throughput to roughly hosts × 32 / batch time, and cut idle-time latency to the timeout.

5.4 "What would you ask for in v2?"

Per-team quotas and priority queues. Autoscaling on queue wait. Server-sent events for interactive callers. A canary pipeline for model versions. Name them in priority order and say which one you would fund first. That is the manager-level signal the round looks for in EM loops.

What separates answers: follow-ups

WeakDefends every point equally

Treats each pushback as an attack, and argues for fixes whose cost exceeds their value.

GoodHolds the important points, concedes the small ones

Keeps the tick cap and rollback as must-fix, and accepts reasonable v1 shortcuts.

StrongConverts findings into a plan

Names the metrics that prove each fix, orders the v2 work by impact, and says what it would cost.

  1. Practice document 2: a usage metering pipeline

A second document, in a different domain, to practice on. Read it and write your top three before you read on.

DESIGN: Token usage metering, v1

Goal: bill customers for tokens used, daily.

Design:
- Each API server keeps an in-memory counter per API key.
- Every 60 seconds, each server writes its counters to the
  usage table with INSERT ... ON CONFLICT (api_key, minute)
  DO UPDATE SET tokens = <server's count>.
- At midnight UTC a job sums the usage table per key and
  sends invoices.

Scale: 200 API servers, 50,000 active keys.
Monitoring: the job emails the team if it fails.

6.1 The critique

Blocking: the upsert overwrites other servers' counts. Every server writes the same (key, minute) row, and SET tokens = <server's count> replaces the value. With 200 servers, the row ends up with whichever server wrote last, so usage is undercounted by up to 200 times. Change: key the row by (key, minute, server) or use tokens = tokens + excluded.tokens with an idempotent batch ID so a retried write does not double count.

Blocking: a server crash loses up to 60 seconds of usage. Counters live only in memory. Change: write usage events to a durable log as requests finish, and aggregate from the log. The log also gives an audit trail for billing disputes.

Major: no reconciliation. Nothing checks that billed usage matches what the API served. Change: compare daily totals with an independent source, such as request logs, and alert on differences above a threshold.

Missing: retries and idempotency for the flush, late data after midnight, time zones for customers billed on local days, and the monitoring of the counts themselves, not only job failure.

flowchart LR
  subgraph Doc["As written"]
    direction LR
    S1[Server 1: 40]:::svc --> U[(row key+minute<br/>SET tokens = count)]:::bad
    S2[Server 2: 55]:::svc --> U
  end
  subgraph Fix["Fixed"]
    direction LR
    E[usage events per request]:::svc --> L[(durable log)]:::store --> A[aggregate by key+minute]:::store
  end
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef bad fill:#fbe9e4,stroke:#c4492d,color:#171717;

The lesson carries over: the blocking finding is a correctness bug hiding in one line of SQL. Read every write path for races and overwrites.

6.2 Practice document 3: a model rollout pipeline

A third document, closer to infrastructure. Read it and write your ranked top three before reading on.

DESIGN: Model rollout pipeline, v1

Goal: ship new model versions to the serving fleet safely.

Design:
- A release manager marks a model version "approved" in the registry.
- A controller polls the registry every 5 minutes. When it sees a new
  approved version, it updates the fleet config to the new version.
- Each serving host watches the fleet config. On change, it stops serving,
  downloads the new weights from the registry bucket, loads them,
  and starts serving again.
- If more than 5% of requests fail for 10 minutes, the on-call engineer
  sets the fleet config back to the previous version.

Scale: 400 serving hosts, weights are 300 GB.
Monitoring: request error rate dashboard.

Two-sentence summary: Approval in the registry triggers a controller to flip one fleet-wide config, and every host independently reloads on the change. Rollback is manual, after 10 minutes of elevated errors.

flowchart LR
  RM[Release manager]:::user -->|approve| REG[(Registry)]:::store
  CTRL[Controller, 5 min poll]:::svc --> REG
  CTRL -->|flip version| CFG[(Fleet config)]:::bad
  CFG --> H1[Host 1: stop, download, load]:::svc
  CFG --> H2[Host 400: stop, download, load]:::svc
  H1 --> B[(Weights bucket)]:::store
  H2 --> B
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef bad fill:#fbe9e4,stroke:#c4492d,color:#171717;

6.3 The critique of document 3

Blocking: every host stops serving at once. One config flip makes all 400 hosts stop, download, and reload at the same moment. For the whole reload, the fleet serves nothing: an outage on every release. Change: roll in waves, such as 5% of hosts at a time, and only take a host out of rotation after its replacement capacity is ready.

Blocking: 400 hosts pull 300 GB from one bucket together. That is 120 TB at once from one source. If the bucket delivers 100 GB/s in aggregate, which is generous, the download alone takes 20 minutes, during which nothing serves. Change: waves reduce concurrent pulls, and peer-to-peer distribution removes the source bottleneck (see model weight distribution). Better still, download before stopping: a host fetches and verifies the new weights while still serving the old ones.

Major: rollback is slow and manual. Ten minutes of 5% errors on the whole fleet before a human acts, then another full reload to go back. Change: canary first, compare against the old version automatically on error rate, latency, and output quality checks, and stop the rollout without a human. Keep the old weights on local disk so rollback is a restart, not a download.

Missing: readiness checks before a host takes traffic; verification of downloaded weights by hash; evaluation gates before approval; what happens if the controller crashes mid-rollout; and monitoring beyond error rate, such as latency and quality metrics, which a bad model can degrade without raising errors.

The lesson: when two findings are both blocking, lead with the one users feel first. Here the fleet-wide stop is visible on every release; the bucket bottleneck makes it longer.

6.4 Handling disagreement

Before reading on. You say the rollout needs canaries. The author says canaries would slow releases from one day to three. How do you respond?

Engage with the cost instead of repeating the risk. Ask how often releases have caused incidents, and how long those incidents lasted. Then offer a smaller change that addresses most of the risk: an automated canary of 30 minutes on 5% of hosts adds under an hour, not two days. Propose that the canary runs automatically and only pages a human when a metric regresses.

If the author still disagrees, record the risk and the decision in the document. A reviewer's job is to make the tradeoff visible and the decision deliberate, not to win the argument. Interviewers, especially in manager loops, watch for exactly this: firm on the risk, flexible on the remedy, and clear about who decides.

What separates answers: disagreement

WeakRepeats the objection louder

Restates the risk and does not engage with the author's cost.

GoodProposes a cheaper remedy

Finds a smaller change that addresses most of the risk within the author's constraint.

StrongMakes the decision explicit

Quantifies both sides, offers options with their costs, and records the chosen tradeoff and its owner in the document.

  1. Variants

7.1 The engineering-manager loop

In EM loops the round replaces coding and sits beside execution and leadership rounds. Expect more weight on process: who reviews this, how the team would stage the work, what you would ask the author to change before approving, and how you would handle an author who disagrees. Frame findings as a review you would leave for a peer, and add a short plan: what blocks approval, what can follow in v2.

7.2 Pinned constraints

The document may state that batching is fixed, autoscaling is unavailable, and rate limiting is out of scope. Treat those as facts. Critique the dispatch and collation layer inside them. If one constraint truly makes the goal impossible, say so once, with the math, and move on.

7.3 Pressure delivery

Some interviewers give few hints and follow up only when you land on the point they want. Silence is not a signal to stop. Keep proposing concrete mechanisms and checking the checklist aloud until the interviewer engages.

7.4 Reviewing a document you partly disagree with

Sometimes the design is sound and you would simply have done it differently. Say so briefly and move on: "I would have used a queue here, but the author's approach works at this scale; not a blocker." Separating taste from risk is itself a strong signal. Spend the interview on findings that change outcomes.

7.5 A document with no clear flaw

If you find nothing blocking, say that clearly, then go deeper on the parts most likely to fail in production: failure modes, rollout, and observability. Ask what the team's last incident was, and check whether the design would have caught it. A reviewer who invents problems to fill time loses credibility.

  1. The transferable pattern

A design review is risk ranking under uncertainty. The same skill runs code review, incident review, and vendor evaluation: read for intent, read against a checklist, rank findings by impact and likelihood, write each as risk, evidence, change, and separate what blocks from what can wait. The checklist is the reusable asset. Build your own from the table in section 2 and extend it after every real incident you see.

Review: the 30-second answer

  • Summarize first. Two sentences prove you understood the design.
  • Rank. Lead with the finding that caps or breaks the whole system.
  • Hunt absences. Failure handling, rollout, and observability are where planted gaps live.
  • Respect constraints. Critique inside what the document pins.
  • Defend or concede. Hold blocking points, concede cheap v1 shortcuts, and name the metric that proves each fix.

Quiz

+Why does a flat list of correct findings score poorly?

The round measures judgment. A flat list shows you can spot issues but cannot tell which ones matter. Interviewers want to see the finding that would cause the most damage first, with the minor points short or skipped.

+In practice document 1, why does the dispatch tick outrank the single Redis instance?

The tick caps the whole system at 64 requests per second and defeats the scaling plan, every day, under normal load. The Redis risk is real but depends on a failure. Impact times likelihood puts the tick first.

+How should you handle a choice that looks weak but may be deliberate?

Ask before criticizing. If the author confirms it is a deliberate v1 shortcut, accept it, and name the point at which it breaks. That shows you can separate a real risk from a reasonable tradeoff.

+What made the usage-metering upsert a blocking bug?

Every server wrote the same row, and the update replaced the value instead of adding to it. The final count was whichever server wrote last. Billing would be wrong every minute, for every key.

+What are the three parts of a good review comment?

The risk, the evidence in the document, and the change you want. Add the cost of the change when it is not obvious.

+In practice document 3, why is the fleet-wide config flip blocking?

It makes every host stop, download, and reload at the same moment, so the fleet serves nothing for the duration of every release. That is an outage on each deploy.

+How should a reviewer respond when the author disagrees on cost?

Engage with the cost, offer a cheaper remedy that covers most of the risk, and if disagreement remains, record the tradeoff and the decision owner in the document.

Sources and further reading

NextPrompt Playground