Anthropic Interview: Design a One-to-One Chat System
A full solution to the Anthropic one-to-one chat question: capacity math, WebSocket gateways, user routing, presence, the inbox table, Kafka internals versus Redis pub/sub, sharding, ordering, and failures.
By The Forward Deployed editorial teamReviewed
Part of the Anthropic system design question bank. The question is representative of the round. The analysis and solution are this site's own.
Problem statement
Users open the app, see who is online, send messages, and see replies within half a second. A user who was offline gets everything they missed when they reconnect. Group chat, media uploads, and end-to-end encryption are out of scope unless the interviewer adds them.
Clarifying questions
- How many users? For practice: 50 million daily active users.
- How many messages? About 20 sent per user per day.
- Delivery target? Under 500 ms end to end when both users are online.
- Ordering? Messages in one conversation appear in send order for both users.
- History? Kept for a year, loaded page by page.
- One device per user? Yes, for now. Expect a follow-up that adds a second device.
What makes chat hard
A chat message crosses at least two servers. The sender's connection lands on one gateway, the recipient's on another, and one of them may not be connected at all. The system must find the recipient's gateway, deliver in order, and never lose a message, while keeping the delivery path fast.
Those goals pull against each other. The fast path, pushing through an in-memory pub/sub layer, can lose messages. The durable path, writing to disk and reading back, is slower. So the driving tension is latency versus durability. The standard answer uses both: a durable write that guarantees delivery, and a fast, lossy broadcast that provides speed.
flowchart LR A([Alice]):::user <-->|WebSocket| G1[Gateway 1]:::svc B([Bob]):::user <-->|WebSocket| G2[Gateway 2]:::svc G1 --> CS[Chat service]:::svc CS --> DB[(Messages + inbox)]:::store CS --> PS[[Pub/sub]]:::svc --> G2 classDef user fill:#e6efec,stroke:#315e55,color:#171717; classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
Key idea. Guarantee delivery with a durable write, and deliver fast with a lossy broadcast. The broadcast may drop messages because the durable layer catches them.
Key concepts
Persistent connections
A WebSocket keeps one TCP connection open in both directions, so the server can push messages without the client asking. Long polling fakes this with repeated requests, which wastes requests and adds latency. For chat, WebSocket is the default.
Delivery guarantees
At-most-once delivers zero or one time; messages can be lost. At-least-once delivers one or more times; duplicates are possible. Exactly-once is what users want. Between machines it is built from at-least-once delivery plus deduplication at the receiver.
Pub/sub versus a log
Pub/sub, as in Redis, forwards a message to whoever is subscribed right now and keeps nothing. A log, as in Kafka, stores messages in order on disk; consumers read at their own pace and can replay. They solve different problems, and the interviewer will ask you to compare them.
Presence
Presence means showing who is online. It needs a way to notice that a connection died without a clean close, which is what heartbeats do.
Key idea. WebSockets for push, at-least-once plus deduplication for exactly-once, pub/sub for speed, a durable table for guarantees, heartbeats for presence.
- Requirements
Before reading on. List the requirements. Which property can never break, and which constraint drives the design?
1.1 Functional requirements
- Send a text message to another user.
- Receive messages in real time when online.
- Receive missed messages on reconnect.
- See whether a contact is online, and when they were last seen.
- Load conversation history, newest first, page by page.
1.2 Non-functional requirements
- No lost messages. A message the server acknowledged is delivered eventually.
- No duplicates on screen. Each message shows once.
- Ordered per conversation.
- Fast. Under 500 ms end to end at p95 when both users are online.
- Available. 99.9% or better.
1.3 The constraint versus the property
Durability of acknowledged messages is the property. Users forgive a slow message; they do not forgive a lost one. Connection count is the constraint. Millions of open sockets decide how many gateways you run and how you find a user's gateway.
- Back-of-the-envelope estimation
| Quantity | Value | Derivation |
|---|---|---|
| Daily active users | 50 million | given |
| Messages per day | 1 billion | 50 M × 20 |
| Average send rate | about 11,600 per second | 1 B / 86,400 |
| Peak send rate | about 35,000 per second | 3 × average |
| Message size | 300 bytes | text plus metadata |
| Storage per day | 300 GB | 1 B × 300 B |
| Storage per year | about 110 TB | before replication |
| Concurrent connections at peak | 5 million | 10% of daily users |
| Gateways | at least 10 | about 500,000 idle sockets per host |
Two numbers shape the design. Five million sockets mean many gateways, so you need a way to find which one holds a user. And 35,000 writes per second is well within a sharded database, so durable writes on the send path are affordable.
Key idea. Connection count sets the gateway fleet; write rate is modest enough to afford a durable write on every send.
- API design
Before reading on. The client sends a message and the network drops before the acknowledgment arrives. The client retries. How does the server avoid storing it twice?
The client generates an ID for each message before sending. The server stores it with a unique constraint on the sender and that ID. A retry with the same ID finds the existing message and returns the same acknowledgment.
3.1 WebSocket frames
client -> server
{type: "send", client_msg_id, to_user_id, body}
{type: "ack", message_id} // delivery ack for received messages
{type: "ping"}
server -> client
{type: "sent", client_msg_id, message_id, conversation_id, server_ts}
{type: "message", message_id, conversation_id, from_user_id, body, server_ts}
{type: "presence", user_id, online, last_seen}
{type: "pong"}3.2 HTTP endpoints
GET /v1/conversations?cursor= -> recent conversations with last message GET /v1/conversations/:id/messages?before=:message_id&limit=50 GET /v1/inbox -> undelivered messages (used on reconnect)
- Data model
conversations id (hash of the two user ids, sorted), user_a, user_b, last_message_id, updated_at messages conversation_id, message_id, sender_id, client_msg_id, body, server_ts primary key (conversation_id, message_id) unique (sender_id, client_msg_id) inbox user_id, message_id, conversation_id, created_at, expires_at primary key (user_id, message_id) users id, last_seen_at
message_id increases within a conversation. A time-ordered ID with a sequence, or a counter per conversation, both work. Sorting by it gives send order.
The conversation ID is derived from the two user IDs, so both users compute the same ID without a lookup, and a second conversation between the same pair cannot appear.
- High-level design
5.1 One server holds all connections
flowchart LR A([Alice]):::user <--> S[Chat server<br/>in-memory user map]:::svc B([Bob]):::user <--> S S --> DB[(Messages)]:::store classDef user fill:#e6efec,stroke:#315e55,color:#171717; classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
With one server, Alice's message finds Bob's connection in a local map. It cannot hold 5 million sockets, and it is a single point of failure.
5.2 Fix 1: many gateways, and a way to find a user
Spread connections across many gateways behind a layer-4 load balancer. Now Alice and Bob are on different gateways, and Alice's gateway does not know where Bob is.
The simplest fix: each gateway subscribes to a pub/sub channel per connected user, named after the user ID. To send to Bob, publish to Bob's channel. Whichever gateway holds Bob receives it.
flowchart LR A([Alice]):::user <--> G1[Gateway 1]:::svc B([Bob]):::user <--> G2[Gateway 2]:::svc G1 -->|publish user:bob| PS[[Redis pub/sub]]:::new PS -->|subscribed to user:bob| G2 classDef user fill:#e6efec,stroke:#315e55,color:#171717; classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;
Without per-user channels, the sender's gateway would have to broadcast to every gateway, which scales badly as the fleet grows.
5.3 Fix 2: a durable inbox
Pub/sub loses messages when Bob is offline or his gateway restarts at the wrong moment. Add a durable write before the publish: store the message, and insert a row into Bob's inbox. Bob's client acknowledges each delivered message, and the server deletes the inbox row. On reconnect, Bob reads his inbox and gets everything he missed.
flowchart LR G1[Gateway 1]:::svc --> CS[Chat service]:::svc CS -->|1. write message + inbox row| DB[(Messages, inbox)]:::new CS -->|2. ack to sender| G1 CS -->|3. publish| PS[[Pub/sub]]:::svc --> G2[Gateway 2]:::svc G2 -->|4. deliver| B([Bob]):::user B -->|5. ack| G2 -->|6. delete inbox row| DB classDef user fill:#e6efec,stroke:#315e55,color:#171717; classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;
5.4 Fix 3: presence
Gateways refresh a presence key per connected user with a time-to-live. When heartbeats stop, the key expires and the user shows offline. Presence changes publish to the user's contacts who are online.
5.5 The composed design
sequenceDiagram
autonumber
actor A as Alice
participant G1 as Gateway 1
participant C as Chat service
participant D as Database
participant P as Pub/sub
participant G2 as Gateway 2
actor B as Bob
A->>G1: send {client_msg_id, to: Bob, body}
G1->>C: send
C->>D: insert message (dedup on client_msg_id)
C->>D: insert inbox row for Bob
C-->>G1: sent {message_id}
G1-->>A: sent
C->>P: publish user:bob
P->>G2: message
G2->>B: message
B->>G2: ack message_id
G2->>D: delete inbox rowKey idea. Many gateways, per-user channels to find the recipient, a durable inbox before the publish, and client acknowledgments to clear it.
- Deep dives
6.1 Kafka or Redis for the message bus
Before reading on. The interviewer asks why you did not use Kafka. What is your answer, and do you know how Kafka works well enough to defend it?
This is the most reported place candidates struggle. Know Kafka's internals first.
A Kafka topic splits into partitions. Each partition is an ordered, append-only log stored on disk. Each partition has replicas on several brokers. One replica is the leader and takes all reads and writes. The replicas that are caught up with it form the in-sync replica set. With acks=all, a write is acknowledged only after every in-sync replica has it. If the leader dies, the controller picks a new leader from the in-sync set, so no acknowledged message is lost.
Consumers form consumer groups. Each partition is read by exactly one consumer in a group, so the group splits the work. When a consumer joins or leaves, the group rebalances and partitions move between consumers, which pauses consumption briefly. Each consumer tracks its offset in each partition and can rewind to replay. Log compaction keeps only the latest record per key, which suits state such as the latest presence per user.
Now the comparison for this system.
| Question | Redis pub/sub | Kafka |
|---|---|---|
| Latency | About a millisecond | Tens of milliseconds with acks=all |
| Stores messages | No | Yes, replayable |
| A channel per user | Cheap: a channel is just a subscriber list | Not feasible as topics; per-topic metadata does not scale to millions |
| Delivery | At most once | At least once |
| Routing to one gateway | Built in, by channel | Needs your own partition-to-gateway mapping |
For one-to-one chat, Redis pub/sub wins. It routes directly to the recipient's gateway and adds almost no latency. It loses messages, but the inbox already guarantees delivery. Kafka's durability would duplicate what the inbox gives you, and routing users to gateways through a fixed set of partitions means building a mapping and living with rebalances.
Kafka earns its place when other systems need the message stream: search indexing, abuse detection, analytics. Then the chat service writes to Kafka as well, and those consumers read at their own pace. Gateways still deliver through Redis.
What separates answers: the message bus
WeakNames Kafka and stops
Picks Kafka for scale, then cannot explain partitions, replicas, or consumer groups.
GoodExplains both and picks one
Knows Kafka's log, leader and in-sync replicas, and consumer groups, and picks Redis for delivery because the inbox covers durability.
StrongUses each for its strength
Delivers through Redis, feeds other consumers through Kafka, and explains why per-user topics fail and what a rebalance costs.
6.2 Presence without melting the database
Before reading on. Each client heartbeats every 20 seconds. Where do you store "last seen," and how often do you write it?
Heartbeats every 20 seconds from 5 million connections are 250,000 per second. Writing last_seen_at to the database on each one would be 250,000 writes per second of low-value data.
Instead, the gateway refreshes a Redis key presence:{user_id} with a time-to-live of about twice the heartbeat, 40 seconds. If the key exists, the user is online. When heartbeats stop, the key expires on its own. Write last_seen_at to the database once, when the connection closes, and have gateways write it with a conditional update so a stale gateway cannot overwrite a newer value after a reconnect.
Heartbeats run at the application level, with a pong deadline of a few seconds. TCP keepalive detects dead connections only after minutes, which is too slow for a presence dot.
To tell contacts, publish presence changes to each contact's channel, but only to online contacts, and batch changes for users with many contacts.
6.3 Ordering
Before reading on. Alice sends two messages quickly. Could Bob see them in the wrong order?
He could, if they took different paths. Guarantee order per conversation, and only per conversation. The chat service assigns increasing message IDs within a conversation, and the client inserts messages by ID, not by arrival time. If a gap appears, such as message 7 arriving before 6, the client waits briefly, then fetches the missing range from the server.
Global order across all conversations is expensive and nobody can see it. Say so, and decline to build it.
6.4 Sharding
Shard messages by conversation ID, so one conversation's history lives on one shard and loads with one range query. Shard inbox by user ID, so a reconnecting user reads one shard. Different keys for different access patterns is normal; say why each key was chosen.
A heavy conversation, such as two bots chatting nonstop, stays on one shard. At one-to-one scale that is rarely a problem. If it becomes one, rate limit the sender.
6.5 Failures
- Gateway crash. Its clients reconnect through the load balancer to another gateway, subscribe to their channels, and read their inboxes. Nothing is lost.
- Pub/sub node failure. Live delivery pauses for users on that node's channels. The inbox still holds every message, and clients catch up on reconnect or on the next poll.
- Database shard failover. Sends to that shard fail for seconds. The client retries with the same client message ID, so no duplicates.
- Chat service crash between the write and the publish. The message is stored and in the inbox. Bob gets it on his next inbox read. Clients read their inbox on reconnect and every few minutes while connected, which bounds the delay.
What separates answers: reliability
WeakPub/sub as the only path
Delivers only through pub/sub, so an offline user or a restarted node loses messages.
GoodInbox plus acknowledgments
Writes an inbox row before publishing and deletes it on client acknowledgment.
StrongEvery failure walked through
Explains client message IDs for idempotent retries, per-conversation ordering with gap repair, and what happens at each crash point.
6.6 Finding the recipient's gateway at scale
Before reading on. Your Redis pub/sub cluster must route to 5 million connected users. How is the cluster laid out, and what happens when a user reconnects to a different gateway?
Each connected user has a channel named user:{id}, subscribed by the gateway that holds the user's socket. A Redis cluster shards channels by hash, so publishing to user:bob goes to one node, which forwards to Bob's gateway. With 5 million channels, each is just a subscription entry; the memory is small.
The gateways are the side to watch. Each gateway subscribes to hundreds of thousands of channels, one per local user. When a gateway restarts, all its users reconnect elsewhere and resubscribe at once: a burst of subscriptions. Spread reconnections with randomized backoff on the client, so a restart becomes a ramp, not a spike.
When Bob reconnects to a different gateway, the new gateway subscribes to user:bob, and the old gateway, if still alive, unsubscribes when the old socket closes. For a moment both may be subscribed; Bob may get a message twice, which the client ignores by message ID.
An alternative avoids per-user channels: a routing table in Redis maps each user to a gateway ID, and the chat service publishes to a channel per gateway. It uses fewer channels, but the table must be updated on every connect and disconnect, and a stale entry routes to the wrong gateway. Per-user channels are simpler; mention the alternative and why you did not choose it.
flowchart LR CS[Chat service]:::svc -->|publish user:bob| RC[(Redis cluster: shard by channel hash)]:::store RC -->|subscriber: gateway 7| G7[Gateway 7]:::svc --> B([Bob]):::user B -. reconnects .-> G3[Gateway 3]:::svc G3 -. subscribe user:bob .-> RC classDef user fill:#e6efec,stroke:#315e55,color:#171717; classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
6.7 History, pagination, and retention
History loads newest first, 50 messages at a time. The query is a range scan on (conversation_id, message_id) in descending order, starting before the oldest loaded message ID. A cursor-based page never skips or repeats messages when new ones arrive, which offset-based pages do.
At 110 TB a year, keep recent months on fast storage and move older partitions to cheaper storage. Most conversations only read the last few pages, so a small hot tier serves nearly all reads. Retention rules, such as deleting after one year, are enforced by dropping old time partitions, which is far cheaper than deleting rows one by one.
The conversation list, "recent chats," needs its own index: per user, conversations ordered by last message time. Update it on every send for both participants. It is a small, write-heavy table; keep it in a store that handles frequent updates well.
What separates answers: routing and storage
WeakBroadcasts to all gateways
Has no way to find the recipient's gateway, and loads history with offsets.
GoodPer-user channels and cursor pages
Routes by per-user channels and pages history by message ID.
StrongPlans for churn and growth
Handles reconnect storms with jittered backoff, explains duplicate subscriptions during reconnects, tiers storage by age, and keeps a per-user recent-conversations index.
- Variants
7.1 A second device
Key inbox rows by user and device. Each device acknowledges independently. Cap devices per user, such as five, to bound fan-out. Read state is per user and syncs across devices through a small "read up to message X" event.
7.2 Group chat
Small groups can reuse the design: publish to each member's channel. Large groups need a different fan-out, publishing once per group and letting each gateway deliver to its local members. That is the Slack design.
7.3 End-to-end encryption
The server stores ciphertext and cannot read bodies. Delivery, inbox, and ordering are unchanged. Search moves to the client, and a second device needs key distribution between devices.
7.4 At ten times the users
At 500 million daily users, sends reach about 350,000 per second at peak and connections about 50 million. Gateways grow to about 100 hosts; Redis pub/sub grows to a larger cluster sharded by channel. The database grows to many shards by conversation ID, and the inbox table stays small because rows are deleted on acknowledgment. The new pressure is regional: users are global, so run full stacks per region, keep each user's inbox in their home region, and forward cross-region messages between regional chat services.
- The transferable pattern
Chat is durable write, then fast broadcast, with a per-recipient queue as the safety net. The same shape appears in notifications, live comments, and order-status updates: guarantee with storage, deliver with a fast path that may drop, and let receivers acknowledge and deduplicate.
Review: the 30-second answer
- WebSocket gateways, per-user pub/sub channels. The sender's gateway reaches the recipient's in one publish.
- Durable inbox before publish. Acknowledged messages survive every crash.
- Client message IDs. Retries do not duplicate.
- Presence in Redis with a time-to-live. Last seen written once, on disconnect.
- Order per conversation only. Increasing IDs and gap repair on the client.
Quiz
+Why can the pub/sub layer be allowed to lose messages?
Because every message is written to the database and to the recipient's inbox before it is published. If the publish is lost, the recipient still gets the message on reconnect or the next inbox read.
+How does the system deliver exactly once from the user's point of view?
Delivery is at least once, from both pub/sub and the inbox. The client ignores any message ID it already has, and the server ignores a resent client message ID. Together they give exactly-once on screen.
+Why not write last_seen on every heartbeat?
At 5 million connections and a heartbeat every 20 seconds, that is 250,000 database writes per second for data nobody needs that fresh. An expiring Redis key shows online status, and one write on disconnect records last seen.
+What happens in Kafka when a partition leader dies?
The controller elects a new leader from the in-sync replicas. With acks=all, every acknowledged message is on all in-sync replicas, so none is lost.
+Why is a Kafka topic per user a bad idea?
Each topic and partition carries metadata the cluster must track. Millions of topics exceed what the cluster can manage. Grouping users onto fixed partitions works, but then you must map partitions to gateways and live with rebalances.
+Why use cursor-based pagination for history?
New messages arriving between page loads shift offsets, so offset pages skip or repeat messages. A cursor, the oldest message ID loaded, always continues from the same place.
+What happens when a gateway restarts with hundreds of thousands of users?
All its users reconnect to other gateways and resubscribe. Randomized client backoff spreads the reconnections into a ramp, and clients drop any duplicate messages by ID during the overlap.
Sources and further reading
- Apache Kafka documentation: design covers the log, replication, in-sync replicas, and log compaction.
- Redis pub/sub documents fire-and-forget delivery semantics.
- RFC 6455: The WebSocket Protocol defines the connection, framing, and ping and pong frames.
- The group-chat version is solved in design Slack.
