
Durgesh Tiwari
Author
A WhatsApp-like chat application looks simple from the user's perspective:
User types a message → Sends it → Receiver gets it.
But a distributed messaging system must answer much harder questions:
What happens if the receiver is offline?
What if the sender loses the network after sending?
What if a chat server crashes?
How do we prevent duplicate messages?
How do we preserve message ordering?
How do multiple devices stay synchronized?
How do we maintain millions of persistent connections?
These challenges make WhatsApp System Design a common system design interview problem.
This article designs a scalable WhatsApp-like messaging platform. It does not attempt to reproduce WhatsApp's private internal architecture.

Keep the initial scope focused on core messaging:
Users can send one-to-one messages.
Users can send group messages.
Online users receive messages near real time.
Offline users receive missed messages after reconnecting.
Users can retrieve message history.
Delivery and read receipts are supported.
Multiple devices remain synchronized.
Secondary features may include presence, typing indicators, media sharing, push notifications, search, and end-to-end encryption.
Important system qualities are:
Low message-delivery latency
High availability
Durable message storage
Horizontal scalability
Fault tolerance
Per-conversation ordering
Multi-device consistency
Secure communication
The main design principle is:
Real-time delivery makes chat fast; durable storage and synchronization make chat reliable.
Assume:
100 million daily active users
50 messages per active user per dayTotal messages:
100M × 50
= 5 billion messages/dayAverage throughput:
5,000,000,000 / 86,400
≈ 58,000 messages/secondPeak traffic may be several times higher.
However, raw message throughput is not the only challenge. A messaging platform may also need to maintain millions of long-lived client connections simultaneously.
A simplified architecture is:
Mobile / Web Clients
|
v
Load Balancer
|
+-----------+-----------+
| | |
v v v
Chat Server Chat Server Chat Server
| | |
+-----------+-----------+
|
v
Message Service
/ \
v v
Message Store Broker / PubSub
|
v
Recipient Chat Server
|
v
ReceiverSupporting services may include:
Authentication Service
User Service
Group Service
Presence Service
Connection Registry
Media Service
Notification Service
Object Storage
Cache
Search
MonitoringThe architecture should evolve only when requirements justify each component.
A messaging application needs fast server-to-client delivery. Repeated HTTP polling works, but it becomes inefficient at large scale.
With polling, clients repeatedly ask:
Client → Any new message?
Server → No
Client → Any new message?
Server → No
Client → Any new message?
Server → YesSuppose 10 million connected users poll every five seconds:
10,000,000 / 5
= 2 million requests/secondMost requests may contain no useful information.
Long polling improves this by keeping the HTTP request open until data is available:
Client
|
| GET /messages
v
Server waits...
|
New message arrives
|
v
ResponseIt is better than frequent polling, but a large chat application naturally benefits from persistent bidirectional communication.
A WebSocket maintains a long-lived bidirectional connection:
Client <====================> Chat Server
Persistent
ConnectionThe server can immediately push events to the client without waiting for another HTTP request.
For a WhatsApp-like messaging system, WebSockets are therefore a natural choice for real-time communication.

Not every operation needs WebSockets.
Use HTTP/REST for operations such as:
Login
Profile management
Creating conversations
Fetching message history
Group management
Media upload setup
Use WebSockets for real-time events such as:
New messages
Delivery acknowledgements
Read receipts
Typing indicators
Presence changes
Client
|
+------ HTTP ------> API Services
|
+---- WebSocket ---> Chat ServersThis keeps normal request-response APIs separate from long-lived real-time communication.
Suppose Alice establishes a WebSocket connection and the load balancer sends her to:
Chat Server 7The system needs to know:
Alice → Chat Server 7
Bob → Chat Server 21Without this mapping, the system would not know which server currently owns Bob's live connection.
Maintain a distributed registry:
user_id → active connection(s)Example:
alice → server_7 / connection_82
bob → server_21 / connection_91A fast distributed key-value store is suitable because this information is ephemeral connection state, not durable business data.
If a chat server crashes, the client reconnects and the mapping can be recreated.
A user may be connected from several devices:
Alice
|
+--> Phone
+--> Laptop
+--> TabletThe registry may therefore store:
alice
|
+--> server_7 / phone
+--> server_12 / laptop
+--> server_18 / tabletMessages and read state may need to synchronize across all authorized devices.

Suppose Alice sends Bob:
"Hello Bob"
A basic flow is:
Alice
|
v
Alice's Chat Server
|
v
Message Service
|
v
Authenticate + Validate
|
v
Deduplicate
|
v
Assign Message ID + Sequence
|
v
Persist Message
|
v
Route Delivery Event
|
v
Bob's Chat Server
|
v
BobThe critical rule is:
Do not consider a message safely accepted until it reaches the system's defined durability point.

A WebSocket message may look like:
{
"type": "message.send",
"conversation_id": "conv_123",
"client_message_id": "alice_987",
"content": "Hello Bob"
}The server should verify:
Authentication
Conversation membership
Message size
Content type
Rate limits
Business rules
The backend then assigns its canonical message information.
A simplified message may contain:
message_id
conversation_id
sender_id
client_message_id
sequence_number
message_type
content
created_atExample:
message_id = msg_891
conversation_id = conv_123
sender_id = alice
client_message_id = alice_987
sequence_number = 5001
message_type = TEXT
content = "Hello Bob"message_id uniquely identifies the message.
sequence_number tells clients where the message belongs inside the conversation.
The main entities are users, conversations, conversation members, and messages.
user_id
name
profile_photo
created_atconversation_id
conversation_type
created_atConversation types:
DIRECT
GROUPconversation_id
user_id
role
joined_at
last_read_sequencemessage_id
conversation_id
sender_id
client_message_id
sequence_number
message_type
content
created_atAt smaller scale, a relational database may be completely sufficient.
The important question is not:
Which database is most scalable?
The important question is:
What are the application's access patterns?
One of the most common reads is:
Get the newest messages for conversation X.At very large scale, append-friendly distributed storage may become attractive because chat messages:
Are continuously appended
Are rarely modified
Are usually read by conversation
Need horizontal partitioning
Start with requirements and access patterns before choosing a database technology.
A natural partition key is:
conversation_idbecause most reads are conversation-oriented.
Shard 1
conv_1
conv_8
Shard 2
conv_2
conv_9
Shard 3
conv_3
conv_10This keeps messages from the same conversation logically close.
However, extremely large or highly active groups can create hot partitions, which may require special handling.
Suppose Alice sends:
M1 = "Hi"
M2 = "How are you?"Distributed networks can cause M2 to arrive before M1.
We do not need one global message order for every conversation in the entire platform.
Instead, preserve ordering within each conversation.
Example:
conversation_123
M1 → sequence 501
M2 → sequence 502
M3 → sequence 503The client can display messages in:
501
502
503even if network delivery temporarily arrives out of order.
This avoids expensive global coordination.
Suppose the server stores Alice's message successfully, but the ACK is lost.
Alice sees a timeout and retries.
Without deduplication, the system might store:
Hello
HelloInstead, the client sends a stable:
client_message_id = alice_987Both the original request and retry contain the same ID:
Alice → alice_987 → Server
Alice → alice_987 → ServerThe server recognizes that this logical message already exists and returns the existing result rather than inserting another copy.
A durable uniqueness constraint should enforce this behavior.

A common state model is:
SENDING
|
v
SENT
|
v
DELIVERED
|
v
READThe exact product semantics must be defined.
The messaging backend has durably accepted the message.
At least one intended recipient device has received it according to the product's delivery semantics.
The recipient has reported that the message or conversation has been read.
Unsafe:
Alice
|
v
Server
|
ACK "sent"
|
Server crashes before persistenceSafer:
Receive Message
|
v
Durably Accept Message
|
v
ACK SenderThe exact durability point may be a database commit, durable log, or another reliable acceptance mechanism.
The important part is that the system defines it explicitly.
Suppose Bob is offline when Alice sends a message.
Alice
|
v
Message Service
|
v
Message Store
|
Bob connected?
/ \
Yes No
| |
Realtime Keep durable message
delivery |
v
Optional push notificationWhen Bob reconnects, the client synchronizes from its last known position.
Example:
Last synchronized sequence = 840The server returns:
841
842
843
...Conceptually:
GET /conversations/{id}/messages?after_sequence=840This handles network disconnects, app restarts, device sleep, server failures, and missed real-time events.
The WebSocket is not the durable source of truth. The message store is.
WebSocket and message storage have different responsibilities in a messaging system.
Aspect | WebSocket | Message Store |
|---|---|---|
Primary Role | Real-time message delivery | Durable message persistence |
Connection | Requires an active connection | Does not depend on an active connection |
Data Nature | Temporary delivery channel | Persistent source of message history |
Offline User | Cannot deliver while disconnected | Keeps the message for later synchronization |
Failure Recovery | Client reconnects after connection loss | Missing messages are fetched from stored history |
Main Goal | Low-latency communication | Reliability and recovery |
A complete messaging system uses both:
┌── Real-time delivery ──> WebSocket ──> Receiver
Sender ──> Message Service
└── Durable storage ─────> Message StoreIf the receiver is offline or the WebSocket connection fails, the message remains in durable storage and can be synchronized after the client reconnects.
WebSocket makes message delivery fast; durable storage makes message delivery recoverable and reliable.

Suppose:
Alice → Chat Server 1
Bob → Chat Server 98Chat Server 1 needs a way to deliver the event to the infrastructure serving Bob.
A broker or Pub/Sub layer can decouple servers:
Chat Server 1
|
v
Broker / PubSub
|
v
Chat Server 98It can help with:
Cross-server routing
Buffering
Fan-out
Partitioning
Asynchronous consumers
Failure isolation
Replay in architectures that support it
No.
A smaller system may work with:
WebSocket Servers
+
Message Database
+
Connection RegistryIntroduce a broker only when it solves a real problem such as cross-server routing, high throughput, fan-out, asynchronous processing, buffering, or consumer isolation.
Adding Kafka automatically without explaining why is a common system design mistake.
Suppose the Message Service performs:
1. Insert message into database
2. Publish delivery eventIf the database insert succeeds but the process crashes before publishing:
Message exists
but
Realtime delivery event is missingPublishing first creates the opposite risk.
One solution is to write the message and event record in the same database transaction:
Database Transaction
|
+--> Insert Message
|
+--> Insert Outbox EventThen:
Outbox Publisher
|
v
Message BrokerIf the publisher crashes, unpublished outbox entries can be retried later.
This creates a reliable bridge between durable message state and asynchronous delivery.

Internal event processing commonly uses:
At-least-once deliveryThis means an event may occasionally be processed more than once, so consumers need idempotency.
Do not casually claim true end-to-end exactly-once delivery across:
Client
Network
Server
Database
Broker
ReceiverA more realistic design is:
At-least-once infrastructure + stable IDs + deduplication + idempotency = effectively-once user-visible behavior where possible.
Suppose Bob has read through:
sequence = 500Instead of storing one read row for every message, store:
bob.last_read_sequence = 500Messages at or below that sequence are considered read according to the conversation semantics.
If:
latest_sequence = 550
last_read_sequence = 500a simplified unread count is:
550 - 500 = 50Real systems may adjust this for deleted messages, system events, membership changes, and filtered message types.
Not every real-time event needs durable storage.
Typing is ephemeral:
Alice
|
v
Chat Server
|
v
Realtime Typing Event
|
v
BobIf one typing event is lost, expensive recovery is usually unnecessary.
Presence answers:
Is Alice online?
When was Alice last active?A simple approach:
Connection established
|
v
Presence = ONLINE
|
v
Periodic heartbeat
|
v
Refresh TTL
|
No heartbeat
|
v
TTL expires
|
v
Presence = OFFLINEPresence state can live in a fast ephemeral distributed store.
Do not broadcast Alice's presence to every user.
Publish it only to users who need it, such as contacts, active conversation participants, or subscribed clients.
Presence is also a privacy problem.
A user may configure:
Everyone
Contacts
NobodyThe Presence Service should enforce these visibility rules.
Suppose Alice sends:
"Meeting at 5?"to a group containing Alice, Bob, Charlie, and David.
Alice
|
v
Message Service
|
v
Store Message Once
|
v
Group Membership
|
+--> Bob
+--> Charlie
+--> DavidSmall groups are relatively straightforward. Very large groups create a fan-out problem.
Generate recipient delivery work when the message is sent:
Group Message
|
+--> User 1
+--> User 2
+--> User 3
...Advantages:
Fast recipient reads
Simple delivery preparation
Disadvantages:
Expensive writes
Potentially massive work for huge groups
Store the group message once:
Group Message
|
v
Shared Group History
|
+--> User reads
+--> User readsAdvantages:
Cheaper write path
Better for huge groups
Disadvantages:
More work during reads
Real-time active-user delivery still needs another mechanism
A practical system can combine both approaches.
For smaller groups, push aggressively to members.
For very large groups:
Store once
↓
Push to active subscribers
↓
Offline users fetch laterThe choice depends on group size, latency requirements, and infrastructure cost.

Suppose a group has:
1,000,000 membersand all activity maps to one conversation_id partition.
That conversation can become a storage or broker hot spot.
Possible approaches include:
Sub-partitioning extremely large groups
Separating message storage from delivery fan-out
Hierarchical fan-out
Batching recipient work
Prioritizing active members
Allowing offline users to synchronize later
Sub-partitioning can make strict ordering harder, so this is a real scalability trade-off.
Large files should not be stored directly inside the primary message database.
Examples include images, videos, audio, and documents.
A better architecture is:
Client
|
v
Media Service
|
v
Object StorageThe message stores only a reference:
{
"message_type": "VIDEO",
"media_id": "media_782"
}A scalable flow is:
Client
|
| Request upload
v
Media Service
|
| Signed upload instruction
v
Client
|
v
Object StorageAfter the upload succeeds, the client sends a chat message containing the media reference.
This prevents large files from flowing through chat servers.
Media can be distributed through a CDN:
Object Storage
|
v
CDN
/ | \
India US EuropeThis improves download latency and reduces origin bandwidth.
When Bob is offline:
Message Service
|
v
Notification Service
|
v
Push Provider
|
v
Bob's PhoneA push notification tells the device that new activity exists.
It is not the durable source of the message.
When Bob opens the application, the client synchronizes with the messaging backend.
Mobile connections frequently move between Wi-Fi, cellular networks, and offline states.
A reconnect flow may look like:
Reconnect
|
v
Authenticate
|
v
Restore WebSocket
|
v
Send sync position
|
v
Fetch missing messagesSuppose a regional network outage affects five million users.
When connectivity returns, they should not all reconnect simultaneously.
Use:
Exponential Backoff
+
JitterExample:
Client A → 1.2 sec
Client B → 2.7 sec
Client C → 4.1 secThis spreads recovery load across time.
Suppose:
Alice → Server 7and Server 7 crashes.
Alice reconnects:
Alice
|
v
Load Balancer
|
v
Server 15The registry updates to:
Alice → Server 15The client then synchronizes missing messages from durable storage.
Because durable messages were not stored only in Server 7's memory, message history survives the crash.
A load balancer may use sticky routing to keep reconnecting clients on the same server when possible.
However, system correctness should not depend on:
Alice must always connect to Server 7Chat servers should remain replaceable.
Connection ownership should be maintained through distributed state rather than assumed server affinity.
WebSocket servers should be balanced using more than raw connection count.
Important signals include:
Active connections
Memory consumption
CPU
Network bandwidth
Messages per second
Connection churn
For example:
Server A:
100,000 mostly idle connections
Server B:
30,000 highly active connectionsThese servers may have very different resource requirements.
Suppose one recipient connection receives:
10,000 eventsbut can transmit only:
1,000 events/secondOne slow client must not consume unlimited server memory.
Possible controls include:
Bounded per-connection buffers
Disconnecting extremely slow clients
Dropping ephemeral events
Preserving durable messages for later synchronization
A typing event may be dropped.
A durable chat message should not simply disappear.
Rate limiting can protect the system by dimensions such as:
user_id
device_id
IP address
conversation_idPossible limits include:
Messages per second
New conversations per minute
Media uploads per hour
Group invitations per day
A real messaging platform may also need:
Block lists
User reports
Spam detection
Suspicious URL detection
Media scanning
Abuse controls
Non-critical analysis can run asynchronously so it does not unnecessarily slow the main messaging path.
Before accepting:
message.sendthe system must verify:
Who is the sender?
↓
Is the sender a member of this conversation?
↓
Can the sender perform this action?The same applies to message-history APIs.
For example:
GET /conversations/123/messagesmust verify that the requester is allowed to access the conversation.
Knowing a conversation ID is never sufficient authorization.
With end-to-end encryption, message content is encrypted on the sender's authorized device and decrypted only by authorized recipient devices.
Alice Device
|
Encrypt
|
v
Ciphertext
|
v
Server
|
v
Ciphertext
|
v
Bob Device
|
DecryptThe messaging infrastructure transports encrypted content without requiring plaintext access.
E2EE adds complexity around:
Key management
Multi-device synchronization
Device replacement
Group membership changes
Backups
Search
Moderation
Message history recovery
The central trade-off is:
Privacy
vs
Server-side functionalityWithout E2EE limitations, search indexing may happen asynchronously:
Message Store
|
v
Indexing Event
|
v
Search IndexMessage delivery does not need to wait for indexing.
The message can therefore be delivered immediately while becoming searchable shortly afterward.
This is a reasonable use of eventual consistency.
With E2EE, search may instead require on-device indexing or another privacy-preserving design.
Messages may change after their original delivery, so clients need a way to synchronize those updates.
Suppose Alice changes:
Meet at 5to:
Meet at 6The message may contain:
message_id
content/version
edited_atConnected participants receive an update event, while offline devices receive the latest state during synchronization.
Instead of immediately removing every copy, the system may write a tombstone:
message_id = 123
deleted = trueClients can display:
This message was deletedBackground cleanup can later remove eligible content according to retention rules.
Different products may support:
Permanent history
Time-based retention
Ephemeral messages
Tenant-specific retention
Disappearing messages
Old data may be archived, compacted, deleted, or moved to cheaper storage.
Retention should be included in storage and compliance planning.
Good cache candidates include:
User profiles
Conversation metadata
Group membership
Presence
Connection mappings
Recent conversation information
But:
Cache
≠
Durable Message SourceLosing the cache should primarily affect performance, not destroy chat history.
A global messaging platform needs to keep connection infrastructure reasonably close to users.
For example:
India Users → India Region
Europe Users → Europe Region
US Users → US RegionThe difficult case is a conversation whose participants are in different regions.
One possible model assigns each conversation a home region:
conversation_123
→ ap-southThat region owns authoritative sequencing and write operations for the conversation.
Advantages:
Simpler ordering
Simpler write ownership
Disadvantage:
Remote participants may incur additional cross-region latency
Allowing several regions to accept writes for the same conversation can reduce local write latency.
But it introduces:
Concurrent write coordination
Ordering complexity
Consistency complexity
Start with the simpler ownership model and introduce active-active writes only if requirements justify the additional complexity.
Some deployments may require certain data to remain in specific geographic regions.
This can affect:
Message placement
Replication
Backups
Search
Analytics
Disaster recovery

Important metrics include:
Active WebSocket connections
Connection attempts per second
Reconnection rate
Messages sent per second
Message persistence latency
Sender ACK latency
Sender-to-recipient delivery latency
Broker lag
Failed deliveries
Retry rate
Database P95/P99 latency
Hot partitions
Presence-store load
Media upload failures
Push-notification failures
A useful metric is:
delivery_latency
=
recipient_delivery_time
-
sender_accept_timeMonitor percentiles such as:
P50
P95
P99rather than averages alone.
Suppose Alice sends message M1.
The Message Service persists it, and a delivery event is published.
Bob's Chat Server crashes before delivering the message.
Bob later reconnects and reports:
last_received_sequence = 900The server finds:
M1 sequence = 901
and returns it during synchronization.
This demonstrates why:
Real-Time Delivery
+
Durable Synchronizationmust work together.
Alice
|
| message.send
v
Chat Server A
|
v
Authentication / Validation
|
v
Message Service
|
+----> Deduplicate client_message_id
|
+----> Assign message_id / sequence
|
+----> Persist message
|
+----> Write outbox event
|
v
ACK Alice
Outbox Publisher
|
v
Broker / PubSub
|
v
Connection Registry
|
Bob online?
/ \
Yes No
| |
v v
Server B Optional Push Notification
|
v
Bob
|
v
Delivery ACKThe sender request path should not wait for every downstream delivery operation to finish.
Alice
|
v
Chat Server
|
v
Message Service
|
v
Verify Group Membership
|
v
Assign Conversation Sequence
|
v
Store Message Once
|
v
Group Delivery Pipeline
|
+----> Active Bob
|
+----> Active Charlie
|
+----> Active David
|
+----> Offline members synchronize laterFor very large groups, delivery work should be partitioned and batched rather than executed synchronously inside the sender's request path.
Putting the major components together gives us:
+----------------+
| Mobile / Web |
| Clients |
+-------+--------+
|
v
+----------------+
| Global LB / |
| Gateway |
+-------+--------+
|
+-------------+-------------+
| | |
v v v
Chat Server Chat Server Chat Server
| | |
+-------------+-------------+
|
v
+----------------+
| Message Service|
+-------+--------+
|
+-------------+-------------+
| | |
v v v
Message Store Outbox Group Service
|
v
Broker / PubSub
|
v
Connection Registry
|
v
Recipient Chat Server
|
v
ClientSupporting systems include authentication, presence, notifications, media storage, CDN, rate limiting, search, and monitoring.
System design interviews usually begin with the overall distributed architecture before moving into classes or implementation details.
Aspect | HLD | LLD |
|---|---|---|
Focus | Overall distributed architecture | Classes and internal components |
Examples | Chat servers, message store, Pub/Sub, presence | Message, Conversation, repositories, services |
Main Question | How does the complete platform work? | How are individual components implemented? |
Typical Interview Start | Usually first | When specifically requested |
A possible LLD model may contain:
User
Conversation
DirectConversation
GroupConversation
ConversationMember
Message
TextMessage
MediaMessage
MessageRepository
MessageService
PresenceService
DeliveryServiceIf an interviewer asks "Design WhatsApp", begin with HLD unless they explicitly request object-oriented or low-level design.
A strong design should evolve with scale instead of starting with every possible component.
Stage 1
Client
↓
Single Chat Server
↓
SQL Database
Stage 2
Clients
↓
WebSockets
↓
Chat Server
Stage 3
Load Balancer
↓
Multiple Chat Servers
↓
Connection Registry
↓
Distributed Message Store
Stage 4
Broker / PubSub
Presence Service
Group Service
Media Service
Push Notifications
Transactional Outbox
Stage 5
Multi-Region Routing
Conversation Partitioning
Hot-Group Handling
Regional Storage
Cross-Region Delivery
Advanced ObservabilityThe principle is simple:
Add architectural complexity when a requirement or bottleneck justifies it.
Use persistent WebSocket connections between clients and horizontally scalable chat servers.
Persist messages through a Message Service, maintain distributed connection mappings, route delivery events between servers, and synchronize missing messages from durable storage after reconnects.
Add presence, groups, media, notifications, idempotency, partitioning, and multi-region architecture as requirements expand.
Chat requires frequent bidirectional real-time communication.
WebSockets allow servers to push messages immediately over long-lived connections instead of requiring repeated polling.
Polling repeatedly sends requests even when there is no new data.
WebSockets maintain persistent bidirectional communication and are better suited to high-frequency real-time chat.
Partition message storage horizontally using conversation-oriented access patterns.
Use append-friendly storage, pagination, retention, and archival where required. Store large media objects separately.
Maintain ordering where it matters: normally inside a conversation.
Assign conversation-level sequence numbers rather than requiring one global ordering sequence for the entire platform.
Store messages durably and optionally send push notifications.
When the user reconnects, synchronize messages after the client's last known sequence or checkpoint.
The client sends a stable client_message_id, and retries reuse the same ID.
The server deduplicates repeated submissions and returns the previously created message.
Connected users reconnect to healthy servers, connection mappings are recreated, and missing messages are recovered from durable message storage.
Maintain ephemeral presence based on active connections and heartbeats.
Use TTL expiration so stale presence records automatically become offline.
Distribute connections across many chat servers.
Track ownership using a distributed connection registry and balance using connection count, CPU, memory, bandwidth, message rate, and connection churn.
Persist the group message, verify membership, and deliver events to active members.
Small groups can use aggressive fan-out, while huge groups may use hybrid delivery.
Do not synchronously create one million recipient operations inside the sender's request.
Store the message once, partition and batch delivery work, prioritize active subscribers, and let offline users synchronize later.
Track each user's read progress for the conversation.
Instead of writing one read row per message, maintain a last_read_sequence where the semantics allow it.
Treat them as ephemeral real-time events.
They generally do not need durable storage or expensive retries.
Upload large files through a media service to object storage.
Store only the media reference inside the message and distribute media through a CDN.
Not necessarily.
Use a broker when cross-server routing, asynchronous consumers, buffering, fan-out, or scale justify it.
Do not promise true end-to-end exactly-once delivery.
Use at-least-once processing with stable identifiers, deduplication, idempotency, and durable synchronization to provide effectively-once user-visible behavior where possible.
Maintain several active connections for the same user.
Deliver events to relevant sessions and synchronize message and read state between authorized devices.
Encryption and decryption move to authorized client devices while servers transport ciphertext.
This increases complexity around key management, groups, backups, search, moderation, device replacement, and multi-device synchronization.
Use a reliable database-to-event bridge such as a transactional outbox.
If real-time delivery still fails, the recipient can recover the persisted message during synchronization.
Use exponential backoff and jitter.
Also protect authentication and connection infrastructure with capacity controls and rate limiting.
Keep connection infrastructure close to users, avoid unnecessary synchronous dependencies, route efficiently, cache ephemeral connection metadata, optimize durable writes, and monitor P95/P99 delivery latency.
Avoid these common design mistakes:
Treating WebSockets as durable message storage.
ACKing before the message reaches the defined durability point.
Ignoring retries and duplicate messages.
Using global ordering when conversation-level ordering is enough.
Forgetting offline synchronization.
Storing large media files in the primary message database.
Keeping connection or presence state only inside one chat server.
Fan-outing huge groups synchronously in the sender path.
Adding Kafka or other infrastructure without a clear reason.
Ignoring authorization, multi-device behavior, and failure recovery.
A strong design should explain what happens when servers crash, users reconnect, events are duplicated, or recipients are offline.
A WhatsApp-like system is not only about delivering messages quickly.
The core architecture is:
Sender
↓
Chat Server
↓
Message Service
↓
Durable Message Store
↓
Delivery / PubSub Layer
↓
Recipient Chat Server
↓
ReceiverAround this flow, the system adds connection tracking, ordering, idempotency, offline synchronization, group fan-out, media handling, presence, security, and failure recovery.
The key design principle is:
Real-time delivery makes messaging fast, while durable storage and synchronization make it reliable.