The Forward Deployed

Anthropic Interview: Distribute Model Weights to a GPU Cluster

A full solution to the Anthropic model weight distribution question: the lower bound, naive pull, tree broadcast, pipelined chunks, the peer swarm, coordinator-free recovery, verification, and readiness-gated routing.

By Reviewed

Part of the Anthropic system design question bank. The question is representative of the round. The analysis and solution are this site's own.

Problem statement

The system moves one large file from one source to many destinations, verifies every copy, loads it onto GPUs, and tells the traffic router which hosts are ready. Model training and the serving engine itself are out of scope.

Clarifying questions

  • What units is the bandwidth in? Gigabits and gigabytes differ by a factor of 8, and that error spreads into every answer. Confirm before any math.
  • Are links full duplex? Can a host upload and download at full speed at the same time, or do both share one budget? It changes the tree math.
  • How big is the file, and how many hosts? For practice: 400 GB and 200 hosts, with a follow-up at 1,000 hosts.
  • What is the network shape? Flat, or racks with slower links between them?
  • How often do new versions ship? Hourly updates change the design compared with monthly ones.
  • What failure rate should we assume? For practice: 1% to 5% of hosts fail during one rollout.

What makes distribution hard

The source has one link. Hundreds of hosts each need the whole file. If every copy comes from the source, the source link carries hundreds of full copies and becomes the bottleneck. Everything else sits idle.

The other links are the untapped capacity. Every host that already holds data has an upload link doing nothing. The design question is how to put all those links to work at the same time, while hosts fail and some links are slow.

So the driving tension is coordination versus parallelism. More hosts sharing more pieces in parallel gets closer to the physical limit, but it needs more coordination about who has what and who sends to whom.

flowchart LR
  S[(Source<br/>one link)]:::store --> H1[Host 1]:::svc
  S --> H2[Host 2]:::svc
  S --> H3[Host ...]:::svc
  S --> H4[Host 200]:::svc
  H1 -. idle upload .-> X((unused)):::bad
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef bad fill:#fbe9e4,stroke:#c4492d,color:#171717;
Key idea. Distribution is limited by links. The source link is one; the hosts' upload links are hundreds. A good design keeps all of them busy.

Key concepts

The lower bound

Before any design, compute the best any design could do. Every host must receive every byte through its own download link. If the file is S bytes and a link carries B bytes per second, no host can finish before S / B. That is the target.

Store-and-forward versus pipelining

In store-and-forward, a host receives the whole file before it sends any of it on. In pipelining, a host forwards each piece as soon as it arrives. Pipelining turns a chain of full-file waits into a chain of small-piece waits.

Chunks and manifests

Split the file into chunks, such as 64 MB. A manifest lists every chunk with its hash, and the source signs it. Hosts can verify each chunk independently, fetch chunks from any peer, and resume after a crash by fetching only the chunks they miss.

Rarest first

In a swarm, each host chooses which chunk to fetch next. Fetching the chunk that the fewest peers hold spreads every chunk quickly and avoids a last chunk that only the source still has.

Key idea. Compute the bound first. Then chunk the file so hosts can pipeline, verify, and fetch from anyone.

  1. Requirements

Before reading on. List the requirements. Then name the property you must never violate, and the constraint that sets the design.

1.1 Functional requirements

  • Distribute a new checkpoint from one source to every host in the cluster.
  • Verify every byte on every host.
  • Survive host failures and slow links without operator action.
  • Load verified weights onto GPUs and report readiness per model version.
  • Send traffic only to hosts that are ready for the requested version.

1.2 Non-functional requirements

  • Speed. Finish close to the lower bound.
  • Correctness. No host ever serves corrupted or wrong-version weights.
  • Resilience. A 5% host failure rate slows the rollout by little and stops it never.
  • Observability. Operators can see progress per host and find stuck hosts in seconds.

1.3 The constraint versus the property

Correctness is the property. A host that serves corrupted weights produces wrong answers quietly, which is worse than a slow rollout. Link bandwidth is the constraint. Every design is measured against how close it comes to keeping all links busy.

Key idea. Never serve bad weights. Design to keep every link busy.

  1. Back-of-the-envelope estimation

For practice: a 400 GB checkpoint, 200 hosts, 10 Gbps full-duplex links on every host and on the source. 10 Gbps is 1.25 GB/s.

2.1 The lower bound

400 GB / 1.25 GB/s = 320 seconds, a little over five minutes. No design beats this.

2.2 Naive pull

Every host downloads from the source. The source must send 200 × 400 GB = 80 TB. At 1.25 GB/s that is 64,000 seconds, close to 18 hours.

2.3 Binary tree, store-and-forward

The source sends the full file to two hosts. Each of those sends it to two more, and so on. A host sends to two children at once, so each child gets half the upload link, and each level takes 2 × 320 = 640 seconds. Seven levels reach 2 + 4 + ... + 128 = 254 hosts, which covers 200. Total: 7 × 640 = 4,480 seconds, about 75 minutes.

2.4 Pipelined chunks

With 64 MB chunks, the file has 400 GB / 64 MB = 6,250 chunks. One chunk crosses one link in 64 MB / 1.25 GB/s ≈ 0.05 seconds. In a pipelined chain of depth d, the last host finishes at about 320 seconds plus d × 0.05 seconds. Even a depth of 200 adds only 10 seconds. The design gets within a few percent of the bound.

2.5 At 1,000 hosts

The bound does not change: 320 seconds. Naive pull grows to 90 hours. A binary tree needs 9 or 10 levels, about 100 minutes. The pipelined swarm stays near 320 seconds, because every new host adds upload capacity as well as demand. That is the scaling argument for peer-to-peer.

Key idea. 320 seconds is the bound, 18 hours is naive, 75 minutes is a tree, and a pipelined swarm gets within seconds of the bound at any cluster size.

  1. API design

The system has a control plane and a data plane.

3.1 Control plane

POST /v1/rollouts
  {model_version, manifest_ref, target_hosts: selector, max_concurrent_loads}
  -> {rollout_id}
GET  /v1/rollouts/:id
  -> {hosts_total, hosts_downloading, hosts_verified, hosts_ready,
      hosts_failed, bytes_total, bytes_done, stuck_hosts[]}
POST /v1/rollouts/:id/pause | resume | abort

3.2 Data plane between peers

GET  /chunks/:model_version/bitmap        -> bitmap of chunks this host holds
GET  /chunks/:model_version/:chunk_index  -> chunk bytes (range requests allowed)
POST /peers/gossip                        {host_id, model_version, bitmap_digest, peers_seen[]}

3.3 Readiness

GET /health/ready?model_version=v42
  -> 200 only if weights for v42 are fully verified, loaded on GPUs,
     and a smoke-test prompt returned a sane answer
Key idea. Operators talk to a rollout. Hosts talk to each other in chunks. The router talks to a readiness check that proves the right version is loaded.

  1. Data model

manifest (signed by the source)
  model_version, total_size, chunk_size,
  chunks: [{index, sha256, size}], signature

host state (local disk + memory)
  model_version, bitmap[6250], chunk files on NVMe,
  in_flight: {chunk_index -> peer, started_at},
  peer_stats: {peer_id -> throughput, failures, last_seen}

rollout (control plane DB)
  rollout_id, model_version, status, created_at
rollout_host
  rollout_id, host_id, state: pending|downloading|verified|loading|ready|failed,
  bytes_done, last_progress_at, error

The bitmap is small: 6,250 bits is under 1 KB, so hosts can share it often.

  1. High-level design

5.1 Every host pulls from the source

flowchart LR
  S[(Source)]:::store --> H1[Host 1]:::svc
  S --> H2[Host 2]:::svc
  S --> H3[Host 200]:::svc
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;

It is simple and it takes 18 hours, because the source link carries every copy.

5.2 Fix 1: hosts forward to each other in a tree

Each host that has the file sends it to two more.

flowchart TB
  S[(Source)]:::store --> A[Host A]:::new
  S --> B[Host B]:::new
  A --> C[Host C]:::new
  A --> D[Host D]:::new
  B --> E[Host E]:::new
  B --> F[Host F]:::new
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
  classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;

Now the source sends only two copies, and the number of copies doubles each level. But each host waits for the whole file before forwarding, and leaf hosts, about half the cluster, never upload at all. 75 minutes.

5.3 Fix 2: chunk the file and pipeline

Split the file into chunks and forward each chunk as it lands. A host starts sending chunk 1 to its children while it still receives chunk 2.

sequenceDiagram
  participant S as Source
  participant A as Host A
  participant C as Host C
  S->>A: chunk 1
  S->>A: chunk 2
  A->>C: chunk 1
  S->>A: chunk 3
  A->>C: chunk 2
  Note over S,C: every link busy at the same time

The tree's full-file wait per level shrinks to one chunk time. But the tree is fragile: if host A dies, its whole subtree stalls. And leaves still never upload.

5.4 Fix 3: a swarm

Drop the fixed tree. Each host fetches any missing chunk from any peer that has it, and serves any chunk it holds to any peer that asks. Hosts pick chunks rarest first and keep several downloads in flight from different peers.

flowchart LR
  S[(Source)]:::store --> H1[Host]:::svc
  S --> H2[Host]:::svc
  H1 <--> H2
  H1 <--> H3[Host]:::svc
  H2 <--> H4[Host]:::svc
  H3 <--> H4
  H3 <--> H5[Host]:::svc
  H4 <--> H5
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;

Every host uploads, including the ones that joined last, so upload capacity grows with the cluster. No single host failure stalls anyone, because each chunk has many sources.

5.5 Fix 4: verification, loading, and readiness

A chunk counts only after its hash matches the signed manifest. When every chunk verifies, the host loads weights onto its GPUs, runs a smoke-test prompt, and reports ready for that exact version. The router only sends traffic for a version to hosts that report ready for it.

5.6 The composed design

sequenceDiagram
  autonumber
  participant Op as Operator
  participant CP as Control plane
  participant H as Host
  participant P as Peers
  participant R as Router
  Op->>CP: start rollout v42
  CP->>H: manifest ref + seed peer list
  H->>H: verify manifest signature
  loop until all chunks held
    H->>P: exchange bitmaps (gossip)
    H->>P: fetch rarest missing chunks
    H->>H: verify each chunk hash
    H-->>P: serve held chunks to others
  end
  H->>H: load weights to GPUs, smoke test
  H->>CP: ready v42
  CP->>R: host ready for v42
  R->>H: route v42 traffic
Key idea. Start with the physics, remove the source bottleneck with forwarding, remove full-file waits with chunks, remove fragile trees with a swarm, and gate traffic on verified readiness.

  1. Deep dives

6.1 Coordinator or no coordinator

Before reading on. Your swarm uses a central tracker that assigns chunks and tracks progress. The interviewer removes it. How do hosts find peers, detect dead ones, and recover their outstanding chunks?

A central tracker is easy to explain. It knows every host's bitmap, assigns each host a list of peers and chunks, and reassigns work when a host goes quiet. It is also a single point of failure and a scaling hot spot.

Without it, hosts gossip. Every second or so, each host sends a small message to a few random peers: its bitmap digest and the peers it has heard from recently. Information spreads through the cluster in a number of rounds that grows with the logarithm of the host count. Each host builds a local view of who holds which chunks.

Failure detection comes from the same messages. If a peer misses several gossip rounds and stops answering chunk requests, neighbors mark it dead. Any chunk a host was fetching from the dead peer goes back into its missing set, and it fetches it from someone else. The source stays reachable as the fallback for any chunk nobody else holds.

What separates answers: coordination

WeakNeeds the coordinator to work

Cannot say how the swarm continues when the tracker dies.

GoodGossip for membership and bitmaps

Explains peer discovery and failure detection through gossip, with the source as a fallback.

StrongWeighs both and bounds convergence

Keeps an optional tracker for speed and observability, falls back to gossip when it fails, and explains why gossip spreads state in a logarithmic number of rounds.

6.2 Slow peers and the endgame

Before reading on. 199 hosts finish in 330 seconds. One host takes 20 minutes. Why, and what fixes it?

Two causes are common. The host fetched from a slow peer, such as one with a degraded network card, and waited. Or it was the last to need a few rare chunks, and only overloaded peers held them.

Fixes:

  • Track throughput per peer. Drop peers that deliver far below the median, and back off exponentially before retrying them.
  • Keep several requests in flight. A slow peer then delays one request, while the others continue.
  • Endgame mode. When a host has only a few chunks left, it requests each remaining chunk from several peers at once and cancels the duplicates when the first copy arrives. A little waste buys a predictable finish.

6.3 Racks and topology

Data centers often have fast links inside a rack and slower, shared links between racks. If peers are chosen at random, most chunk traffic crosses racks and overloads those links.

Make peer choice topology-aware. Prefer peers in the same rack. Let a small number of hosts per rack pull each chunk from other racks, then spread it inside the rack. Cross-rack traffic then carries roughly one copy per rack, not one copy per host.

flowchart LR
  subgraph R1["Rack 1"]
    A1[Host]:::svc <--> A2[Host]:::svc
    A2 <--> A3[Host]:::svc
  end
  subgraph R2["Rack 2"]
    B1[Host]:::svc <--> B2[Host]:::svc
    B2 <--> B3[Host]:::svc
  end
  A1 <-->|one cross-rack flow per chunk| B1
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;

6.4 Verification and bad data

The source signs the manifest. Every host checks the signature before it trusts any chunk hash. Every chunk is hashed as it arrives, and a mismatch means the chunk is discarded and fetched from another peer. A peer that sends several bad chunks is quarantined: other hosts stop fetching from it, and the control plane flags it for repair, because a bad disk or memory fault is the usual cause.

Verification also protects against the worst operational mistake: a host that loads a mix of two versions. Every chunk carries its model version in its path, and the manifest hash ties them together.

6.5 Readiness-gated routing

Before reading on. The rollout is half done. A request for v42 arrives. How does the router avoid a host that has the file but has not loaded it, or still serves v41?

Readiness is a per-version state, reported by the host after it has verified every chunk, loaded the weights onto its GPUs, and passed a smoke test. The router keeps a map from model version to ready hosts and routes only to that set. A host that is still loading appears in neither the v42 set nor, once it unloads v41, the v41 set.

Rollouts should keep capacity. Load the new version on a subset of hosts while the rest serve the old version, shift traffic as hosts become ready, then continue. Never unload the old version from more hosts than the remaining ones can cover.

6.6 Observability

Operators need one screen: hosts by state, bytes done, estimated finish, and the list of stuck hosts. A host is stuck when its bytes have not moved for a set time, such as 30 seconds. For each stuck host, show its peers and their throughput, which usually points at the cause in seconds. Distinguish a host that is slow from one that is waiting on a full disk or a failed GPU, because the fixes differ.

6.7 Chunk size

Before reading on. Why 64 MB? What goes wrong with 1 MB chunks, and with 4 GB chunks?

Chunk size trades pipeline delay against overhead.

Small chunks shorten the pipeline: a chunk crosses a hop quickly, so data reaches the far side of the swarm sooner. But every chunk has fixed costs: a request, a hash check, bitmap bookkeeping, and gossip bytes. With 1 MB chunks, a 400 GB file has 400,000 chunks. Bitmaps grow to 50 KB per host, gossip carries more data, and per-request overhead starts to eat bandwidth.

Large chunks cut overhead but lengthen the pipeline and waste more on failures. A 4 GB chunk takes 3.2 seconds per hop at 1.25 GB/s; across a depth of 20 hops, that adds about a minute. A peer that dies mid-chunk wastes up to 4 GB of transfer. And a file with only 100 chunks gives rarest-first too few pieces to spread.

In between, 16 to 256 MB is the usual range. At 64 MB, per-hop time is 0.05 seconds, the bitmap is under 1 KB, and a lost transfer wastes at most 64 MB. Say that the number comes from this tradeoff, and that you would tune it by measuring completion time on a test cluster.

6.8 Worked timeline for the swarm

Put numbers on the swarm to show it really approaches the bound.

  • t = 0 s. 200 hosts receive the signed manifest and a seed list of peers. All bitmaps are empty. The source starts sending different chunks to different hosts, rarest first, so no two hosts start with the same chunk.
  • t = 0 to 10 s. The source pushes about 12.5 GB, about 195 chunks, each to one host. Every chunk now exists in one place outside the source.
  • t = 10 to 30 s. Hosts exchange bitmaps and fetch from each other. The number of copies of early chunks doubles roughly every chunk time. Every host is now downloading at close to full speed, from several peers at once, and uploading chunks it holds.
  • Steady state. Each host downloads at about 1.25 GB/s and uploads at about the same, because every chunk it receives is wanted by others. The source keeps injecting new chunks at its full rate.
  • t ≈ 320 s. The source has sent every chunk at least once. Most hosts hold nearly the whole file.
  • t ≈ 320 to 340 s. Endgame: hosts fetch their last few chunks from several peers at once. The slowest hosts finish.

The total is within about 5% of the 320-second bound, and 5% of 200 hosts failing along the way changes little, because every chunk has many holders by then.

  1. Variants

7.1 Shared upload and download budget

If upload and download share one link budget, a host that uploads at full speed cannot download at full speed. The lower bound changes: the cluster's total link capacity must carry roughly one download per host plus the uploads that feed them. The swarm still wins, but finishing time roughly doubles compared with full duplex. State the new bound before the new design.

7.2 Frequent updates

If a new version ships every few hours and most weights change little, send only the chunks whose hashes changed. Hosts keep the previous version's chunks and fetch the difference. Prewarm by starting the transfer before the version is announced for traffic.

7.3 Terabyte checkpoints

At several terabytes, local disk space and loading time join bandwidth as limits. Stream chunks straight into GPU memory where the serving engine supports it, and shard loading across GPUs in parallel.

7.4 At ten times the cluster size

With 2,000 hosts, the bound stays 320 seconds, and the swarm stays near it, because each new host adds upload capacity. What changes is coordination. Gossip still converges in a logarithmic number of rounds, a few more than at 200 hosts. Cross-rack and cross-building links become the bottleneck, so topology-aware peer choice becomes essential. The control plane's rollout view shows percentages and a list of stuck hosts, since nobody reads 2,000 rows. And the router shifts traffic by percentage of ready hosts, in steps, instead of waiting for all of them.

  1. The transferable pattern

Weight distribution is broadcast with peer amplification: one source, many receivers, and receivers that become senders. The same shape serves container image distribution, dataset staging for training, game patches, and software updates across fleets. The method transfers: compute the bound, find the idle capacity, chunk to pipeline, and verify everything.

Review: the 30-second answer

  • Bound first. File size over link speed: 320 seconds here.
  • Naive pull is 18 hours; a tree is 75 minutes. The source link and full-file waits are the reasons.
  • Chunk, pipeline, swarm. Every host uploads, rarest chunks first, several peers at a time.
  • Survive without a coordinator. Gossip bitmaps and liveness; the source is the fallback.
  • Verify and gate. Signed manifest, per-chunk hashes, and routing only to hosts ready for that version.

Quiz

+Why compute the lower bound before designing?

It gives every design something to be measured against. Here, no host can finish before file size divided by link speed, 320 seconds. A design at 75 minutes is clearly far from optimal, and one at 330 seconds is clearly close.

+Why is a store-and-forward tree slow?

Each host must receive the whole file before sending it on, so each level adds a full-file transfer time, and shared upload links double it. Leaf hosts, half the cluster, never upload at all.

+How does chunking get close to the bound?

Hosts forward each chunk as soon as it arrives, so every link works at once. The pipeline adds only one chunk time per hop, a fraction of a second, on top of the single-file transfer time.

+Why fetch the rarest chunks first?

It spreads every chunk across many hosts quickly, so no chunk ends up held only by the source or a few overloaded peers near the end.

+What must be true before the router sends v42 traffic to a host?

Every chunk verified against the signed manifest, the weights loaded onto the GPUs, a smoke test passed, and the host reporting ready for v42 specifically.

+What goes wrong with very small chunks?

Per-chunk overhead grows: more requests, more hash checks, larger bitmaps, and more gossip traffic. At some point the overhead consumes a meaningful share of bandwidth.

+Why does the swarm stay near the bound as the cluster grows?

Every new host adds upload capacity along with its demand. Total transfer capacity grows with the host count, so the time per host stays close to file size divided by link speed.

Sources and further reading

NextOne-to-One Chat System