The Forward Deployed

Anthropic Interview: Design a One-to-One Chat System

A full solution to the Anthropic one-to-one chat question: capacity math, WebSocket gateways, user routing, presence, the inbox table, Kafka internals versus Redis pub/sub, sharding, ordering, and failures.

By Reviewed

Part of the Anthropic system design question bank. The question is representative of the round. The analysis and solution are this site's own.

Problem statement

Users open the app, see who is online, send messages, and see replies within half a second. A user who was offline gets everything they missed when they reconnect. Group chat, media uploads, and end-to-end encryption are out of scope unless the interviewer adds them.

Clarifying questions

  • How many users? For practice: 50 million daily active users.
  • How many messages? About 20 sent per user per day.
  • Delivery target? Under 500 ms end to end when both users are online.
  • Ordering? Messages in one conversation appear in send order for both users.
  • History? Kept for a year, loaded page by page.
  • One device per user? Yes, for now. Expect a follow-up that adds a second device.

What makes chat hard

A chat message crosses at least two servers. The sender's connection lands on one gateway, the recipient's on another, and one of them may not be connected at all. The system must find the recipient's gateway, deliver in order, and never lose a message, while keeping the delivery path fast.

Those goals pull against each other. The fast path, pushing through an in-memory pub/sub layer, can lose messages. The durable path, writing to disk and reading back, is slower. So the driving tension is latency versus durability. The standard answer uses both: a durable write that guarantees delivery, and a fast, lossy broadcast that provides speed.

flowchart LR
  A([Alice]):::user <-->|WebSocket| G1[Gateway 1]:::svc
  B([Bob]):::user <-->|WebSocket| G2[Gateway 2]:::svc
  G1 --> CS[Chat service]:::svc
  CS --> DB[(Messages + inbox)]:::store
  CS --> PS[[Pub/sub]]:::svc --> G2
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
Key idea. Guarantee delivery with a durable write, and deliver fast with a lossy broadcast. The broadcast may drop messages because the durable layer catches them.

Key concepts

Persistent connections

A WebSocket keeps one TCP connection open in both directions, so the server can push messages without the client asking. Long polling fakes this with repeated requests, which wastes requests and adds latency. For chat, WebSocket is the default.

Delivery guarantees

At-most-once delivers zero or one time; messages can be lost. At-least-once delivers one or more times; duplicates are possible. Exactly-once is what users want. Between machines it is built from at-least-once delivery plus deduplication at the receiver.

Pub/sub versus a log

Pub/sub, as in Redis, forwards a message to whoever is subscribed right now and keeps nothing. A log, as in Kafka, stores messages in order on disk; consumers read at their own pace and can replay. They solve different problems, and the interviewer will ask you to compare them.

Presence

Presence means showing who is online. It needs a way to notice that a connection died without a clean close, which is what heartbeats do.

Key idea. WebSockets for push, at-least-once plus deduplication for exactly-once, pub/sub for speed, a durable table for guarantees, heartbeats for presence.

  1. Requirements

Before reading on. List the requirements. Which property can never break, and which constraint drives the design?

1.1 Functional requirements

  • Send a text message to another user.
  • Receive messages in real time when online.
  • Receive missed messages on reconnect.
  • See whether a contact is online, and when they were last seen.
  • Load conversation history, newest first, page by page.

1.2 Non-functional requirements

  • No lost messages. A message the server acknowledged is delivered eventually.
  • No duplicates on screen. Each message shows once.
  • Ordered per conversation.
  • Fast. Under 500 ms end to end at p95 when both users are online.
  • Available. 99.9% or better.

1.3 The constraint versus the property

Durability of acknowledged messages is the property. Users forgive a slow message; they do not forgive a lost one. Connection count is the constraint. Millions of open sockets decide how many gateways you run and how you find a user's gateway.

  1. Back-of-the-envelope estimation

QuantityValueDerivation
Daily active users50 milliongiven
Messages per day1 billion50 M × 20
Average send rateabout 11,600 per second1 B / 86,400
Peak send rateabout 35,000 per second3 × average
Message size300 bytestext plus metadata
Storage per day300 GB1 B × 300 B
Storage per yearabout 110 TBbefore replication
Concurrent connections at peak5 million10% of daily users
Gatewaysat least 10about 500,000 idle sockets per host

Two numbers shape the design. Five million sockets mean many gateways, so you need a way to find which one holds a user. And 35,000 writes per second is well within a sharded database, so durable writes on the send path are affordable.

Key idea. Connection count sets the gateway fleet; write rate is modest enough to afford a durable write on every send.

  1. API design

Before reading on. The client sends a message and the network drops before the acknowledgment arrives. The client retries. How does the server avoid storing it twice?

The client generates an ID for each message before sending. The server stores it with a unique constraint on the sender and that ID. A retry with the same ID finds the existing message and returns the same acknowledgment.

3.1 WebSocket frames

client -> server
  {type: "send", client_msg_id, to_user_id, body}
  {type: "ack", message_id}                     // delivery ack for received messages
  {type: "ping"}

server -> client
  {type: "sent", client_msg_id, message_id, conversation_id, server_ts}
  {type: "message", message_id, conversation_id, from_user_id, body, server_ts}
  {type: "presence", user_id, online, last_seen}
  {type: "pong"}

3.2 HTTP endpoints

GET /v1/conversations?cursor=               -> recent conversations with last message
GET /v1/conversations/:id/messages?before=:message_id&limit=50
GET /v1/inbox                               -> undelivered messages (used on reconnect)

  1. Data model

conversations
  id (hash of the two user ids, sorted), user_a, user_b, last_message_id, updated_at

messages
  conversation_id, message_id, sender_id, client_msg_id, body, server_ts
  primary key (conversation_id, message_id)
  unique (sender_id, client_msg_id)

inbox
  user_id, message_id, conversation_id, created_at, expires_at
  primary key (user_id, message_id)

users
  id, last_seen_at

message_id increases within a conversation. A time-ordered ID with a sequence, or a counter per conversation, both work. Sorting by it gives send order.

The conversation ID is derived from the two user IDs, so both users compute the same ID without a lookup, and a second conversation between the same pair cannot appear.

  1. High-level design

5.1 One server holds all connections

flowchart LR
  A([Alice]):::user <--> S[Chat server<br/>in-memory user map]:::svc
  B([Bob]):::user <--> S
  S --> DB[(Messages)]:::store
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;

With one server, Alice's message finds Bob's connection in a local map. It cannot hold 5 million sockets, and it is a single point of failure.

5.2 Fix 1: many gateways, and a way to find a user

Spread connections across many gateways behind a layer-4 load balancer. Now Alice and Bob are on different gateways, and Alice's gateway does not know where Bob is.

The simplest fix: each gateway subscribes to a pub/sub channel per connected user, named after the user ID. To send to Bob, publish to Bob's channel. Whichever gateway holds Bob receives it.

flowchart LR
  A([Alice]):::user <--> G1[Gateway 1]:::svc
  B([Bob]):::user <--> G2[Gateway 2]:::svc
  G1 -->|publish user:bob| PS[[Redis pub/sub]]:::new
  PS -->|subscribed to user:bob| G2
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;

Without per-user channels, the sender's gateway would have to broadcast to every gateway, which scales badly as the fleet grows.

5.3 Fix 2: a durable inbox

Pub/sub loses messages when Bob is offline or his gateway restarts at the wrong moment. Add a durable write before the publish: store the message, and insert a row into Bob's inbox. Bob's client acknowledges each delivered message, and the server deletes the inbox row. On reconnect, Bob reads his inbox and gets everything he missed.

flowchart LR
  G1[Gateway 1]:::svc --> CS[Chat service]:::svc
  CS -->|1. write message + inbox row| DB[(Messages, inbox)]:::new
  CS -->|2. ack to sender| G1
  CS -->|3. publish| PS[[Pub/sub]]:::svc --> G2[Gateway 2]:::svc
  G2 -->|4. deliver| B([Bob]):::user
  B -->|5. ack| G2 -->|6. delete inbox row| DB
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;

5.4 Fix 3: presence

Gateways refresh a presence key per connected user with a time-to-live. When heartbeats stop, the key expires and the user shows offline. Presence changes publish to the user's contacts who are online.

5.5 The composed design

sequenceDiagram
  autonumber
  actor A as Alice
  participant G1 as Gateway 1
  participant C as Chat service
  participant D as Database
  participant P as Pub/sub
  participant G2 as Gateway 2
  actor B as Bob
  A->>G1: send {client_msg_id, to: Bob, body}
  G1->>C: send
  C->>D: insert message (dedup on client_msg_id)
  C->>D: insert inbox row for Bob
  C-->>G1: sent {message_id}
  G1-->>A: sent
  C->>P: publish user:bob
  P->>G2: message
  G2->>B: message
  B->>G2: ack message_id
  G2->>D: delete inbox row
Key idea. Many gateways, per-user channels to find the recipient, a durable inbox before the publish, and client acknowledgments to clear it.

  1. Deep dives

6.1 Kafka or Redis for the message bus

Before reading on. The interviewer asks why you did not use Kafka. What is your answer, and do you know how Kafka works well enough to defend it?

This is the most reported place candidates struggle. Know Kafka's internals first.

A Kafka topic splits into partitions. Each partition is an ordered, append-only log stored on disk. Each partition has replicas on several brokers. One replica is the leader and takes all reads and writes. The replicas that are caught up with it form the in-sync replica set. With acks=all, a write is acknowledged only after every in-sync replica has it. If the leader dies, the controller picks a new leader from the in-sync set, so no acknowledged message is lost.

Consumers form consumer groups. Each partition is read by exactly one consumer in a group, so the group splits the work. When a consumer joins or leaves, the group rebalances and partitions move between consumers, which pauses consumption briefly. Each consumer tracks its offset in each partition and can rewind to replay. Log compaction keeps only the latest record per key, which suits state such as the latest presence per user.

Now the comparison for this system.

QuestionRedis pub/subKafka
LatencyAbout a millisecondTens of milliseconds with acks=all
Stores messagesNoYes, replayable
A channel per userCheap: a channel is just a subscriber listNot feasible as topics; per-topic metadata does not scale to millions
DeliveryAt most onceAt least once
Routing to one gatewayBuilt in, by channelNeeds your own partition-to-gateway mapping

For one-to-one chat, Redis pub/sub wins. It routes directly to the recipient's gateway and adds almost no latency. It loses messages, but the inbox already guarantees delivery. Kafka's durability would duplicate what the inbox gives you, and routing users to gateways through a fixed set of partitions means building a mapping and living with rebalances.

Kafka earns its place when other systems need the message stream: search indexing, abuse detection, analytics. Then the chat service writes to Kafka as well, and those consumers read at their own pace. Gateways still deliver through Redis.

What separates answers: the message bus

WeakNames Kafka and stops

Picks Kafka for scale, then cannot explain partitions, replicas, or consumer groups.

GoodExplains both and picks one

Knows Kafka's log, leader and in-sync replicas, and consumer groups, and picks Redis for delivery because the inbox covers durability.

StrongUses each for its strength

Delivers through Redis, feeds other consumers through Kafka, and explains why per-user topics fail and what a rebalance costs.

6.2 Presence without melting the database

Before reading on. Each client heartbeats every 20 seconds. Where do you store "last seen," and how often do you write it?

Heartbeats every 20 seconds from 5 million connections are 250,000 per second. Writing last_seen_at to the database on each one would be 250,000 writes per second of low-value data.

Instead, the gateway refreshes a Redis key presence:{user_id} with a time-to-live of about twice the heartbeat, 40 seconds. If the key exists, the user is online. When heartbeats stop, the key expires on its own. Write last_seen_at to the database once, when the connection closes, and have gateways write it with a conditional update so a stale gateway cannot overwrite a newer value after a reconnect.

Heartbeats run at the application level, with a pong deadline of a few seconds. TCP keepalive detects dead connections only after minutes, which is too slow for a presence dot.

To tell contacts, publish presence changes to each contact's channel, but only to online contacts, and batch changes for users with many contacts.

6.3 Ordering

Before reading on. Alice sends two messages quickly. Could Bob see them in the wrong order?

He could, if they took different paths. Guarantee order per conversation, and only per conversation. The chat service assigns increasing message IDs within a conversation, and the client inserts messages by ID, not by arrival time. If a gap appears, such as message 7 arriving before 6, the client waits briefly, then fetches the missing range from the server.

Global order across all conversations is expensive and nobody can see it. Say so, and decline to build it.

6.4 Sharding

Shard messages by conversation ID, so one conversation's history lives on one shard and loads with one range query. Shard inbox by user ID, so a reconnecting user reads one shard. Different keys for different access patterns is normal; say why each key was chosen.

A heavy conversation, such as two bots chatting nonstop, stays on one shard. At one-to-one scale that is rarely a problem. If it becomes one, rate limit the sender.

6.5 Failures

  • Gateway crash. Its clients reconnect through the load balancer to another gateway, subscribe to their channels, and read their inboxes. Nothing is lost.
  • Pub/sub node failure. Live delivery pauses for users on that node's channels. The inbox still holds every message, and clients catch up on reconnect or on the next poll.
  • Database shard failover. Sends to that shard fail for seconds. The client retries with the same client message ID, so no duplicates.
  • Chat service crash between the write and the publish. The message is stored and in the inbox. Bob gets it on his next inbox read. Clients read their inbox on reconnect and every few minutes while connected, which bounds the delay.

What separates answers: reliability

WeakPub/sub as the only path

Delivers only through pub/sub, so an offline user or a restarted node loses messages.

GoodInbox plus acknowledgments

Writes an inbox row before publishing and deletes it on client acknowledgment.

StrongEvery failure walked through

Explains client message IDs for idempotent retries, per-conversation ordering with gap repair, and what happens at each crash point.

6.6 Finding the recipient's gateway at scale

Before reading on. Your Redis pub/sub cluster must route to 5 million connected users. How is the cluster laid out, and what happens when a user reconnects to a different gateway?

Each connected user has a channel named user:{id}, subscribed by the gateway that holds the user's socket. A Redis cluster shards channels by hash, so publishing to user:bob goes to one node, which forwards to Bob's gateway. With 5 million channels, each is just a subscription entry; the memory is small.

The gateways are the side to watch. Each gateway subscribes to hundreds of thousands of channels, one per local user. When a gateway restarts, all its users reconnect elsewhere and resubscribe at once: a burst of subscriptions. Spread reconnections with randomized backoff on the client, so a restart becomes a ramp, not a spike.

When Bob reconnects to a different gateway, the new gateway subscribes to user:bob, and the old gateway, if still alive, unsubscribes when the old socket closes. For a moment both may be subscribed; Bob may get a message twice, which the client ignores by message ID.

An alternative avoids per-user channels: a routing table in Redis maps each user to a gateway ID, and the chat service publishes to a channel per gateway. It uses fewer channels, but the table must be updated on every connect and disconnect, and a stale entry routes to the wrong gateway. Per-user channels are simpler; mention the alternative and why you did not choose it.

flowchart LR
  CS[Chat service]:::svc -->|publish user:bob| RC[(Redis cluster: shard by channel hash)]:::store
  RC -->|subscriber: gateway 7| G7[Gateway 7]:::svc --> B([Bob]):::user
  B -. reconnects .-> G3[Gateway 3]:::svc
  G3 -. subscribe user:bob .-> RC
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;

6.7 History, pagination, and retention

History loads newest first, 50 messages at a time. The query is a range scan on (conversation_id, message_id) in descending order, starting before the oldest loaded message ID. A cursor-based page never skips or repeats messages when new ones arrive, which offset-based pages do.

At 110 TB a year, keep recent months on fast storage and move older partitions to cheaper storage. Most conversations only read the last few pages, so a small hot tier serves nearly all reads. Retention rules, such as deleting after one year, are enforced by dropping old time partitions, which is far cheaper than deleting rows one by one.

The conversation list, "recent chats," needs its own index: per user, conversations ordered by last message time. Update it on every send for both participants. It is a small, write-heavy table; keep it in a store that handles frequent updates well.

What separates answers: routing and storage

WeakBroadcasts to all gateways

Has no way to find the recipient's gateway, and loads history with offsets.

GoodPer-user channels and cursor pages

Routes by per-user channels and pages history by message ID.

StrongPlans for churn and growth

Handles reconnect storms with jittered backoff, explains duplicate subscriptions during reconnects, tiers storage by age, and keeps a per-user recent-conversations index.

  1. Variants

7.1 A second device

Key inbox rows by user and device. Each device acknowledges independently. Cap devices per user, such as five, to bound fan-out. Read state is per user and syncs across devices through a small "read up to message X" event.

7.2 Group chat

Small groups can reuse the design: publish to each member's channel. Large groups need a different fan-out, publishing once per group and letting each gateway deliver to its local members. That is the Slack design.

7.3 End-to-end encryption

The server stores ciphertext and cannot read bodies. Delivery, inbox, and ordering are unchanged. Search moves to the client, and a second device needs key distribution between devices.

7.4 At ten times the users

At 500 million daily users, sends reach about 350,000 per second at peak and connections about 50 million. Gateways grow to about 100 hosts; Redis pub/sub grows to a larger cluster sharded by channel. The database grows to many shards by conversation ID, and the inbox table stays small because rows are deleted on acknowledgment. The new pressure is regional: users are global, so run full stacks per region, keep each user's inbox in their home region, and forward cross-region messages between regional chat services.

  1. The transferable pattern

Chat is durable write, then fast broadcast, with a per-recipient queue as the safety net. The same shape appears in notifications, live comments, and order-status updates: guarantee with storage, deliver with a fast path that may drop, and let receivers acknowledge and deduplicate.

Review: the 30-second answer

  • WebSocket gateways, per-user pub/sub channels. The sender's gateway reaches the recipient's in one publish.
  • Durable inbox before publish. Acknowledged messages survive every crash.
  • Client message IDs. Retries do not duplicate.
  • Presence in Redis with a time-to-live. Last seen written once, on disconnect.
  • Order per conversation only. Increasing IDs and gap repair on the client.

Quiz

+Why can the pub/sub layer be allowed to lose messages?

Because every message is written to the database and to the recipient's inbox before it is published. If the publish is lost, the recipient still gets the message on reconnect or the next inbox read.

+How does the system deliver exactly once from the user's point of view?

Delivery is at least once, from both pub/sub and the inbox. The client ignores any message ID it already has, and the server ignores a resent client message ID. Together they give exactly-once on screen.

+Why not write last_seen on every heartbeat?

At 5 million connections and a heartbeat every 20 seconds, that is 250,000 database writes per second for data nobody needs that fresh. An expiring Redis key shows online status, and one write on disconnect records last seen.

+What happens in Kafka when a partition leader dies?

The controller elects a new leader from the in-sync replicas. With acks=all, every acknowledged message is on all in-sync replicas, so none is lost.

+Why is a Kafka topic per user a bad idea?

Each topic and partition carries metadata the cluster must track. Millions of topics exceed what the cluster can manage. Grouping users onto fixed partitions works, but then you must map partitions to gateways and live with rebalances.

+Why use cursor-based pagination for history?

New messages arriving between page loads shift offsets, so offset pages skip or repeat messages. A cursor, the oldest message ID loaded, always continues from the same place.

+What happens when a gateway restarts with hundreds of thousands of users?

All its users reconnect to other gateways and resubscribe. Randomized client backoff spreads the reconnections into a ramp, and clients drop any duplicate messages by ID during the overlap.

Sources and further reading

NextData Platform with Sensitivity Tiers