The Forward Deployed

OpenAI Interview: Design Slack

A full solution to the OpenAI Slack question: service boundaries, the data model, the durable-then-publish message path, adaptive fan-out for small and huge channels, Redis versus Kafka, multi-device sync, notifications, files, deletion, and workspace isolation.

By Reviewed

Part of the OpenAI system design question bank. The question is representative of the round. The analysis and solution are this site's own.

Problem statement

Users belong to workspaces, one per company. Inside a workspace they send direct messages and post in channels. Messages appear on every member's open devices within a fraction of a second. Offline devices catch up when they reconnect. Mentions trigger notifications. Search, apps, and calls are out of scope unless added.

Clarifying questions

  • Scale? For practice: 20 million daily users across 500,000 workspaces, 50 messages sent per user per day.
  • Channel sizes? Most under 50 members; some company-wide channels with 50,000.
  • Devices? Up to about five per user: desktop, laptop, phone, tablet, web.
  • Delivery target? Under 500 ms to online members.
  • History? Kept per the workspace's retention setting; loaded page by page.
  • Deletion? Senders can delete their messages; admins can too.
  • Isolation? Workspaces must never see each other's data; some large customers require data residency.

What makes team chat hard

Three problems combine.

Fan-out varies wildly. A direct message goes to one person. A message in a 50,000-member channel goes to 50,000 people, many online on several devices. One delivery strategy cannot serve both: per-recipient delivery is exact for small channels and explodes for huge ones.

Every user has several devices, each with its own view of what it has seen. A message must reach all of them, and reading it on one should clear the unread badge on the others.

And delivery must be fast and durable at the same time. The fast path through in-memory pub/sub can drop messages; the durable path through storage is slower.

So the driving tension is fan-out cost versus delivery latency, handled by switching fan-out strategy with channel size and separating durability from speed.

flowchart LR
  C([Clients: many devices]):::user <-->|WebSocket| GW[Gateways]:::svc
  GW --> MS[Message service]:::svc --> DB[(Messages)]:::store
  MS --> PS[[Pub/sub]]:::svc --> GW
  MS --> NS[Notification service]:::svc --> PUSH([Mobile push]):::user
  CH[Channel service: membership]:::svc --> MS
  classDef user fill:#e6efec,stroke:#315e55,color:#171717;
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
Key idea. Store the message once, durably, then deliver it fast with a fan-out strategy that depends on channel size.

Key concepts

Service boundaries

Separate services by data shape and load. Messages are huge in volume and append-heavy. Membership (who is in which channel) is small, read constantly, and cached. Gateways hold connections and no business logic. Notifications decide when to push to phones.

Push and pull delivery

Push sends each message to recipients as it happens. Pull lets clients fetch new messages when they look. Push is fast and costly per recipient; pull is cheap and needs the client to ask. Chat uses push for small audiences and a mix for huge ones.

Cursors

A device's cursor for a channel is the ID of the last message it has. On reconnect, it asks for messages after its cursor in each channel, which finds every gap.

Time-sortable IDs

Message IDs that sort by time, such as a timestamp plus a sequence, make a channel's history one ordered range and make cursors simple.

Key idea. Split services by data shape, deliver by push for small audiences and by signal-plus-pull for huge ones, and let cursors repair every gap.

  1. Requirements

Before reading on. List the requirements. What must never happen, and what drives the architecture?

1.1 Functional requirements

  • Send and receive direct messages and channel messages.
  • Deliver to all of a user's online devices; catch up offline devices.
  • Show unread counts and sync read state across devices.
  • Notify on mentions and direct messages, respecting preferences.
  • Share files in messages; delete messages.
  • Keep workspaces isolated.

1.2 Non-functional requirements

  • No lost messages after the server acknowledges them.
  • Order within a channel.
  • Latency under 500 ms to online recipients.
  • Scale to 50,000-member channels.
  • Isolation and residency per workspace.

1.3 The constraint versus the property

Durable, ordered delivery is the property. Fan-out cost is the constraint. It forces different strategies for different channel sizes.

  1. Back-of-the-envelope estimation

  • Messages: 20 M users × 50 = 1 billion a day; about 11,600 per second on average, about 35,000 at peak.
  • Storage: 1 B × 500 bytes = 500 GB a day, about 180 TB a year before replication.
  • Connections: 10 million online at peak, with 1.5 devices each, about 15 million sockets. At 500,000 per gateway host, about 30 gateway hosts, plus headroom.
  • Fan-out: most messages go to small channels; say the average delivery is 20 device connections. 35,000 × 20 = 700,000 deliveries per second at peak.
  • The big channel: one message in a 50,000-member channel, with 60% online on 1.5 devices each, is 45,000 socket writes. If that channel sees 5 messages per second during an announcement, that is 225,000 writes per second from one channel.
Key idea. The volume of messages is modest; the volume of deliveries is 20 times larger, and one busy large channel can rival everything else.

  1. API design

Before reading on. Should the client send messages over the WebSocket or over HTTP?

Either works. HTTP POST for sends is simpler to retry, authenticate, and load balance; the WebSocket is for receiving. Many chat systems send over HTTP and receive over the socket. Include a client message ID either way, for deduplication.

3.1 HTTP

POST /v1/channels/:id/messages   {client_msg_id, text, file_ids?, thread_ts?}
  -> {message_id, ts}
DELETE /v1/channels/:id/messages/:message_id
GET  /v1/channels/:id/messages?after=:message_id&limit=100    // catch-up
GET  /v1/channels/:id/messages?before=:message_id&limit=50    // history
POST /v1/channels/:id/read       {last_read_message_id}
POST /v1/files                   {name, size, mime} -> {file_id, upload_url}

3.2 WebSocket events

server -> client
  message.new       {channel_id, message}
  message.deleted   {channel_id, message_id}
  channel.activity  {channel_id, latest_message_id}     // for huge channels
  read.updated      {channel_id, last_read_message_id}  // from another device
  presence.changed  {user_id, status}
client -> server
  subscribe.focus   {channel_id}                        // user is viewing this channel
  ping

  1. Data model

workspaces (id, name, plan, region, kms_key_ref)
users      (workspace_id, id, name, ...)
channels   (workspace_id, id, name, type: public|private|dm, member_count, created_at)
members    (workspace_id, channel_id, user_id, role, last_read_message_id, notify_pref)
           two access paths: by channel (fan-out) and by user (sidebar)
messages   (workspace_id, channel_id, message_id, sender_id, client_msg_id,
            text, file_ids, thread_root_id, edited_at, deleted_at)
           primary key (channel_id, message_id); message_id sorts by time
           unique (channel_id, sender_id, client_msg_id)
devices    (user_id, device_id, push_token, platform, last_seen_at)
files      (workspace_id, id, uploader_id, storage_key, size, scan_status, deleted_at)

  1. High-level design

5.1 One chat server

All clients connect to one server that stores messages and forwards them to connected members. It works for a small office and fails at 15 million sockets and at the first crash.

5.2 Fix 1: gateways and a message service

Many stateless gateways hold sockets. A message service writes messages to a database sharded by channel ID, then publishes them. Gateways deliver to their local connections.

5.3 Fix 2: per-user fan-out for small channels

For direct messages and small channels, publish the message to one pub/sub topic per recipient user. Each gateway subscribes to topics for the users it holds. Delivery is exact: only the gateways with recipients receive anything.

flowchart LR
  MS[Message service]:::svc -->|publish user:alice| PS[[Pub/sub]]:::svc
  MS -->|publish user:bob| PS
  PS --> G1[Gateway 1: holds alice]:::new
  PS --> G2[Gateway 2: holds bob]:::new
  classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
  classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;

5.4 Fix 3: per-channel fan-out for large channels

For a 50,000-member channel, per-user publishing means 50,000 publishes for one message. Switch to one publish per channel. Each gateway that holds at least one member subscribes to the channel topic once and delivers to its local members.

5.5 Fix 4: cursors and catch-up

Every device keeps a cursor per channel. On reconnect, it asks for messages after its cursor. Pub/sub may drop messages; the catch-up read repairs any gap.

5.6 Fix 5: notifications

The notification service decides for each recipient: are they active on any device? If so, the in-app delivery is enough. If not, and the message is a mention or direct message, send a mobile push, respecting quiet hours and channel preferences.

5.7 The composed design

sequenceDiagram
  autonumber
  actor A as Sender
  participant M as Message service
  participant D as Messages DB
  participant C as Channel service
  participant P as Pub/sub
  participant G as Gateways
  participant N as Notifications
  A->>M: POST message (client_msg_id)
  M->>D: insert (dedupe on client_msg_id), assign message_id
  M-->>A: message_id
  M->>C: member count, small or large?
  alt small channel or DM
    M->>P: publish to each member's user topic
  else large channel
    M->>P: publish once to channel topic
  end
  P->>G: deliver
  G->>G: write to local member sockets (all devices)
  M->>N: mentions and DMs
  N->>N: recipient active anywhere? preferences?
  N-->>N: send mobile push if needed
Key idea. Durable write first, then publish per user or per channel by size, deliver to every device, repair gaps with cursors, and push to phones only when nobody is looking.

  1. Deep dives

6.1 Adaptive fan-out

Before reading on. Where would you set the switch from per-user to per-channel fan-out, and what happens to the gateway side when you switch?

Per-user publishing costs one publish per member; per-channel costs one publish per message, plus a subscription per gateway that holds any member. Below a few dozen members, per-user is cheap and precise. Above that, per-channel wins. Pick a threshold, such as 50 members, and move channels across it as they grow, publishing to both paths briefly during the switch so no message is missed.

For per-channel fan-out, gateways need to know which channels their users belong to. When a user connects, the gateway loads the user's large-channel memberships and subscribes to those channel topics. A gateway with 500,000 users may subscribe to many channels, but each subscription is one entry, not one per user.

For the very largest channels, pushing every message to every open socket is often wasteful: most members are not looking at that channel. Send the full message only to members who have the channel in focus, and a light channel.activity signal to the rest. Their clients fetch messages when the user opens the channel. Unread badges still update from the signal.

What separates answers: fan-out

WeakOne strategy for all

Publishes per recipient for every channel, and falls over on a 50,000-member announcement.

GoodPer-user for small, per-channel for large

Switches strategy by size, and explains gateway subscriptions to channel topics.

StrongSize- and attention-aware

Also sends full messages only to members viewing the channel and signals to the rest, migrates channels across the threshold without gaps, and quantifies the load of a busy large channel.

6.2 Redis or Kafka between services and gateways

Redis pub/sub is low latency and fire-and-forget: nothing is stored, and subscribers that are not connected miss messages. That suits the gateway delivery path, because durability comes from the messages database and cursors repair gaps.

Kafka is a durable, partitioned log with consumer offsets and replay, at higher latency and operating cost. It suits consumers that must process every message: search indexing, compliance archiving, analytics, and bots. The message service writes to both: Redis for live delivery, Kafka for everything downstream.

Be ready to explain the broadcast decision explicitly: publish to all gateways (simple, wasteful), or only to gateways with interested users (requires subscription tracking). Topic-per-user and topic-per-channel with gateway subscriptions is the latter.

6.3 Multiple devices and read state

Before reading on. A user reads a channel on their laptop. How does their phone's unread badge clear?

Each device has its own cursor, used for catch-up. Read state is per user, stored as last_read_message_id on the membership row. When the laptop reports a read, the service updates the row and publishes read.updated to the user's topic. Every device of that user receives it and clears the badge.

Do not create a copy of each message per device. It multiplies storage by the device count and makes read state inconsistent. Messages are stored once per channel; devices track position.

Presence is per user, derived from any active device: a heartbeat refreshes an expiring key, and a status change publishes to interested users. Write "last seen" only on disconnect, never per heartbeat.

6.4 Notifications

The decision per recipient, per message:

  1. Is it a direct message, a mention, or a channel set to notify on everything? If not, stop.
  2. Is the user active on any device right now? If yes, the in-app delivery is enough; no push.
  3. Is it quiet hours for the user? If yes, hold it.
  4. Send a push to the user's mobile devices, and collapse repeated pushes for the same channel into one.

Delay pushes by a few seconds for users who were active recently, so a message read on the desktop does not also buzz the phone.

6.5 Files and deletion

Files upload directly to object storage with a presigned URL, then pass a malware scan. The message references the file ID. Downloads go through short-lived signed URLs, issued only after checking the requester is a member of the channel.

Deleting a message sets deleted_at, and the service publishes message.deleted. Online clients remove it at once; offline clients learn about it during catch-up, which includes deletions after their cursor. A background job removes the text and any files once retention rules allow. Enterprise retention holds may require keeping deleted content for compliance; model that as a separate, access-controlled archive.

6.6 Workspace isolation and scale

Before reading on. How do you make sure a bug cannot show one company's messages to another?

Put the workspace ID in every key and every table. Derive it from the authenticated session, never from a request parameter. Check it in every service, and have storage enforce it as part of the key. Pub/sub topic names include the workspace ID.

For scale, shard messages by channel ID within a region. For residency and noisy-neighbor isolation, place each workspace in a home region; large enterprises can get dedicated shards and their own encryption keys. A hot channel, such as the company-wide one during an all-hands, concentrates on one shard: cache its latest messages, and rate limit posts in very large channels.

What separates answers: isolation and scale

WeakIsolation as an afterthought

Mentions workspaces but lets IDs come from the client and shares caches across tenants.

GoodWorkspace in every key

Derives workspace ID from the session and scopes storage and topics by it.

StrongIsolation plus placement

Adds home regions for residency, dedicated shards and keys for large customers, and a plan for hot channels.

6.7 Unread counts at scale

Before reading on. A user is in 400 channels. Their sidebar shows an unread count for each. How do you compute 400 counts quickly, and keep them current?

Counting messages newer than last_read_message_id per channel with 400 queries on every sidebar load is too slow. Two cheaper pieces work together:

  • The latest message ID per channel is kept in the channel row and in a cache. The sidebar compares it with the user's last_read_message_id per channel. If they differ, the channel is bold. That is one batched read of 400 small values.
  • Exact counts only where they show. Mention badges and direct messages show numbers; ordinary channels only need bold or not. Keep a per-user counter of unread mentions per channel, incremented by the message service when it records a mention, and reset on read.

Updates arrive as events. A new message updates the channel's latest ID, and gateways push channel.activity to members, so sidebars turn bold live. A read on another device pushes read.updated.

6.8 Walking a message through a huge channel

For practice: an announcement in a 50,000-member channel, with 30,000 members online on 45,000 sockets across 60 gateways.

  1. The message service writes the message once, to the channel's shard, and gets its ID.
  2. It publishes once to the channel topic. All 60 gateways are subscribed, because each holds some members.
  3. Each gateway checks which of its local members have the channel open. Say 2,000 in total. They get the full message.
  4. The other 43,000 sockets get a small channel.activity event. Their clients turn the channel bold.
  5. The notification service sees 12 users mentioned by name, and @channel. For large channels, @channel is restricted by workspace policy; it sends pushes only to the 12 mentioned users who are not active.
  6. Members who open the channel later fetch the message by cursor.

One database write, one publish, 60 gateway fan-outs, and about 45,000 socket writes, most of them tiny.

What separates answers: large-audience mechanics

WeakRecounts on every load

Counts unread messages per channel with a query each time.

GoodLatest ID per channel

Compares latest message IDs with read positions for bold state and keeps mention counters.

StrongWalks the big channel

Shows the exact work for one announcement: one write, one publish, per-gateway fan-out, full messages only for viewers, and restricted broadcast mentions.

  1. Variants

7.1 Threads

A thread is a set of replies to one root message. Store replies in the same channel partition with thread_root_id, and let clients load a thread by root. Notify thread participants, not the whole channel.

Index messages from the Kafka stream into a search engine partitioned by workspace, with permission filtering on channel membership at query time.

7.3 Shared channels between companies

A channel shared by two workspaces needs a home workspace for storage, and a membership model that spans both, with each company's retention applied to its own users' copies.

7.4 At ten times the users

At 200 million daily users, deliveries pass 7 million per second at peak. Gateways number in the hundreds, and the pub/sub layer becomes a sharded cluster per region. The biggest change is regional: workspaces are homed in a region, and cross-region shared channels forward messages between regional message services. Search and analytics consume Kafka streams per region.

  1. The transferable pattern

Team chat is durable storage plus adaptive fan-out. Store each event once, then choose push per recipient, push per group, or signal-and-pull based on audience size and attention. The same choice appears in social feeds (fan-out on write versus on read), notifications, and live dashboards.

Review: the 30-second answer

  • Split services. Gateways, messages, membership, notifications.
  • Write once, then publish. Durable message first; pub/sub for speed; cursors repair gaps.
  • Adaptive fan-out. Per-user topics for small channels, per-channel topics for large, signal-and-pull for the largest.
  • Devices track cursors; users track read state. One message copy; read events sync badges.
  • Workspace in every key. From the session, enforced in storage, with home regions and dedicated shards for large tenants.

Quiz

+Why switch from per-user to per-channel publishing for large channels?

Per-user publishing costs one publish per member per message, which is 50,000 publishes in a large channel. Per-channel publishing costs one publish, and each gateway delivers to its own local members.

+Why can the pub/sub layer lose messages without losing data?

Messages are written durably before they are published, and each device catches up by reading messages after its cursor when it reconnects or notices a gap.

+Why not keep a copy of each message per device?

It multiplies storage by the number of devices and makes read state hard to keep consistent. Storing each message once per channel and tracking a cursor per device is enough.

+When should the notification service send a mobile push?

When the message is a direct message, a mention, or in a channel set to notify, the user is not active on any device, and it is not the user's quiet hours.

+Why keep user data and message data in separate services?

They have different shapes and loads. Messages are huge and append-heavy; users and memberships are small and read constantly for fan-out. Separating them lets each scale, cache, and fail on its own.

+How does the sidebar show which of 400 channels are unread without 400 count queries?

It compares each channel's latest message ID with the user's last read message ID, fetched in one batch. Exact counts are kept only for mentions and direct messages.

+In a 50,000-member channel, who receives the full message immediately?

Only members who currently have the channel open. Other online members get a small activity event, and fetch the message by cursor when they open the channel.

Sources and further reading

NextNearby Places Search