Anthropic System Design Interview Questions
Reported Anthropic system design interview questions: LLM inference batching, model weight distribution, design-doc review, prompt playground, and chat.
By The Forward Deployed editorial teamReviewed
This bank covers the system-design pool reported for software, ML, and infrastructure roles at Anthropic. The forward-deployed loop has its own shape, covered in the Anthropic FDE interview guide. Use this page when your loop includes a system-design round, or when you are weighing an engineering role at Anthropic next to the FDE role. Each question below links to a full analysis and solution. Answer the question out loud before you open it.
The format matters as much as the question. The round is reported to run in a shared document, where you type your reasoning instead of drawing boxes. Interviewers pick one or two depth areas and skim the rest. A generic walk through requirements, entities, API, schema, and scale spends the hour on the wrong things. Find the depth area early and lead with it.
The questions
Design an LLM inference API that batches requests across GPUs
The most reported question in the pool. Clients send generation requests, streaming or not. Route them to model replicas and batch them to keep GPUs busy while tail latency stays bounded.
Know the modern serving baseline well enough to explain it. Continuous batching lets new requests join a running decode batch. Paged KV-cache memory lets requests of very different lengths share GPU memory. Prefill and decode get separate handling, because decode is the bottleneck. Then come the operational probes. GPUs take minutes to start, so what protects the front door while new capacity boots? A strong answer watches queue depth, tightens rate limits by tier as the queue grows, and sheds free-tier traffic first. Expect follow-ups on a GPU that fails mid-batch, idempotency keys for retries, and where safety checks sit in the pipeline.
A narrower variant fixes the GPU contract. You get a batch call you cannot change and design only the dispatch layer around it. For practice, assume each call takes up to 50 inputs and always returns in 200 ms, and traffic is 2,000 requests per second. A full batch serves 50 / 0.2 = 250 requests per second per GPU, so 8 GPUs is the floor. With half-full batches you need 16. The batcher decides which number you get. Flush when the queue reaches 50 or when a short timeout expires, whichever comes first. At low traffic, batches never fill, and the timeout becomes the main lever on latency.
Prepare with Cost and latency, which covers batching, caching, and latency budgets. Practice the GPU-count estimate until you can do it without notes.
Read the full analysis and solution.
Review a design document
A newer format. The interviewer shares a design document, sometimes close to the team's real work, gives you time to read, and asks for a critique. It appears in engineering-manager loops and, more and more, as the system-design slot for senior engineers. The inference server above is a reported subject.
Review it as a senior engineer would in a real comment thread. Rank the risks and lead with the one that would hurt most. Name what is missing, because the document is incomplete on purpose. Offer an alternative with a concrete tradeoff, and notice when the author already considered an angle. If the document pins a constraint, such as fixed batching or no autoscaling, critique inside it. A checklist speeds up the read: goals and non-goals, alternatives, scale and cost, failure modes, rollout and rollback, observability, security, migration. Codebase and learning drills train the same skill on code.
Read the full analysis and solution.
Design a prompt playground
Labeled system design, run as product design. The deployment diagram is given and fixed, and traffic is low. The interest is the product: prompts, versions, conversations, and sharing.
Sketch the schema fast. Make versions immutable with a parent pointer; that gives you history, branches, and undo without locks. Sharing takes most of the hour: public links, links that need a login, explicit access per account, and revocation. Say how long a permission can stay stale after revocation, and why that bound is acceptable. The other depth path is very large prompts. Keep the body in object storage behind a reference, upload it straight from the client, store diffs between versions with periodic snapshots, and render long text lazily. Close with export, which turns a saved version into a real API call. Prompting and structured output covers what a prompt version should freeze.
Read the full analysis and solution.
Distribute a large model checkpoint to hundreds of hosts
One source holds a checkpoint of hundreds of gigabytes. Every host in the cluster needs a copy, fast. Each link has a fixed bandwidth, and some hosts will fail during the rollout.
State the lower bound first. For practice, take a 400 GB checkpoint, 10 Gbps links (1.25 GB/s), and 200 hosts. Every host must download the whole file, so no design beats 400 / 1.25 = 320 seconds. A naive pull pushes 200 copies through the source's one link: 64,000 seconds, close to 18 hours. A binary tree, where each host forwards a full copy to two children, spends about 640 seconds per level across seven levels, close to 75 minutes. Split the file into chunks and forward each chunk the moment it lands, and every link works at once. That swarm approaches the 320-second bound.
Walk that progression out loud. The derivation earns more credit than naming BitTorrent in the first minute. Then the follow-ups: per-chunk hashes, slow or dead peers, and recovery without a central coordinator, where peers share which chunks they hold and pull the rarest missing pieces first. The last probe is readiness. Route traffic only to hosts that have loaded and verified the right model version. Deployment covers readiness gates and rollback.
Read the full analysis and solution.
Design a one-to-one chat system
One-to-one messages only, one device per user. The round covers delivery, presence, offline queueing, history, and the choice of message bus.
Put durability in a per-recipient inbox table, written before you publish. The live transport can then be fast and lossy, and a client that reconnects drains its inbox. Presence comes from heartbeats and an expiring key; writing a last-seen time on every heartbeat melts the database. The Kafka discussion is the most reported place to struggle. Know partitions, consumer groups and rebalancing, in-sync replicas and leader election, and log compaction, and argue when Redis pub/sub is enough. Promise ordering per conversation only.
Read the full analysis and solution.
Design a data platform for stakeholders with different sensitivity levels
Ingest data at scale, process it, and serve teams with different needs and different clearance. It needs no specialized ML knowledge.
Cover batch and streaming ingestion with replay and schema evolution. Keep transforms idempotent. Store data in a lakehouse layer with partitioning and cost tiers. Control access by row and column, tie it to a sensitivity label on each dataset, and keep an audit trail. Estimate storage from events per second, bytes per event, and retention. Bad data needs a blast-radius answer: how you detect it, quarantine it, and backfill. Security and compliance covers classification and access control.
Read the full analysis and solution.
Design an ML configuration system
A research-track option. Design, and sometimes partly implement, configuration for training runs: schemas, inheritance and overrides, validation, and reproducibility. Define reproducibility in concrete terms. Resolve the full config, hash it, pin the environment and data versions, and show that the same run relaunches weeks later on another machine. Then argue the tradeoff between short, flexible configs and strict typed validation.
Read the full analysis and solution.
What carries across the pool
Type your numbers and tradeoffs into the document as you go, because that document is what the interviewer reads. Pick the depth area within the first few minutes. Bring back-of-envelope arithmetic for GPUs, bandwidth, and storage; three of the seven questions turn on it. For the customer-facing version of these skills, go back to the Anthropic FDE interview guide.
Frequently asked questions
What system design questions does Anthropic ask?
Reported questions include an LLM inference API with GPU batching, a design-document review, a prompt playground, distributing a model checkpoint across a cluster, a one-to-one chat system, a data platform with sensitivity tiers, and, on the research track, an ML configuration system. The inference API is the most reported. Confirm the format with your recruiter.
What is the Anthropic inference API interview question?
You design an API that takes generation requests and batches them across GPUs, keeping the GPUs busy while tail latency stays bounded. Expect depth on continuous batching, KV-cache memory, prefill and decode, GPU-count estimates, and how you protect the service while new GPUs start. The page above works through the GPU-count arithmetic.
Does the Anthropic FDE loop use these questions?
The FDE loop centers on a customer-conversation simulation, a Claude deployment design round, and a values interview. The questions on this page are reported for the engineering system-design rounds. They train the same design skills. See the Anthropic FDE interview guide.
