
Durgesh Tiwari
Author
A collaborative editor allows multiple users to edit the same document at the same time while seeing one another’s changes in near real time.
Google Docs-like collaboration looks simple from the user’s point of view:
Alice types "Hello"
↓
Bob sees "Hello"
Bob deletes a word
↓
Alice sees the deletion
Charlie moves his cursor
↓
Alice and Bob see Charlie's cursorBehind this simple experience is a difficult distributed-systems problem.
What happens when Alice and Bob edit the same part of the document at the same time?
What if Bob loses internet connectivity, continues editing offline, and reconnects later?
What if clients receive concurrent operations in different orders?
What if a collaboration server crashes after receiving an edit but before the edit becomes durable?
The core system design problem is:
How can multiple users edit the same document concurrently while ensuring that all participants eventually converge to the same valid document state?
Solving this requires more than a database and WebSockets. A real-time collaborative editor also needs conflict resolution, operation ordering, durable storage, reconnection handling, and scalable real-time delivery.
A collaborative editor is a document system where multiple users can read and modify the same document concurrently.
Suppose the current document is:
Hello WorldAlice and Bob both start from this version.
Alice inserts:
Beautifulwhile Bob deletes:
Worldat almost the same time.
A possible final document is:
Hello BeautifulThe challenge is that Alice may create her operation without knowing about Bob’s deletion, while Bob may create his operation without seeing Alice’s insertion.
These are concurrent edits.
A real-time collaborative document editor therefore needs a strategy for:
Operation Ordering
Conflict Resolution
Convergence
Reconnection
DurabilityWebSockets can move edits between users quickly, but they do not decide how concurrent edits should be combined.
That requires collaboration semantics such as Operational Transformation (OT) or CRDTs, which we will discuss later.
The core system should support:
Create and open documents.
Edit the same document concurrently.
Propagate edits in near real time.
Resolve concurrent edits consistently.
Show collaborator presence and cursor positions.
Persist accepted document changes.
Reconnect and synchronize after disconnection.
Enforce document access permissions.
Features such as comments, suggestions, version history, search, export, notifications, and advanced offline editing can be added later.
The main focus is real-time collaborative editing, not building every feature of a complete office suite.
The most important requirements are:
Low latency: Remote edits should appear quickly.
Convergence: Clients should eventually reach the same document state.
Durability: Accepted edits should survive failures.
High availability: Collaboration should recover quickly from node failures.
Scalability: Support many active documents and concurrent connections.
Fault tolerance: Recover from network and server failures.
Efficient synchronization: Reconnecting clients should not replay the full history.
Security: Only authorized users should read or modify documents.
For real-time collaboration, users should normally see remote edits quickly enough that editing feels interactive without sacrificing correctness or durability.
The central correctness property of collaborative editing is convergence.
Suppose Alice, Bob, and Charlie are editing the same document. Because edits take time to travel through the network, they may temporarily see different document states.
However, after all valid operations have been exchanged and processed, their replicas should eventually converge:
Alice Document
=
Bob Document
=
Charlie DocumentThis assumes each replica has received the same relevant operations and applied the collaboration rules correctly.
The system does not require every client to have an identical state at every moment. Temporary differences are expected during concurrent editing.
For example:
Alice inserts X
Bob inserts Y
↓
Concurrent Operations
↓
Conflict-Resolution Rules
↓
Same Final DocumentThe important requirement is that concurrent operations are resolved consistently so replicas do not permanently diverge.
A system that delivers edits quickly but leaves users with different final document states is not a correct collaborative editor.
Before designing the architecture, estimate the expected collaboration workload.
Assume:
Daily Active Users = 50 million
Concurrent Connections = 5 million
Active Documents = 500,000
Average Edit Rate = 10 operations/sec
per active documentThe estimated edit-processing rate is:
500,000 × 10
=
5 million operations/secThis is only a rough capacity-planning estimate. Real traffic will be uneven.
Most documents may have only a few active editors, while a small number of hot documents may have hundreds of editors and thousands of viewers.
This creates two different scaling problems:
A large number of independent documents.
A small number of extremely active documents.
Independent documents can be distributed across collaboration shards using document_id. Hot documents require additional fan-out and scaling strategies, which we will discuss later.
We also need to estimate persistent connection capacity.
If the system has:
5,000,000 concurrent WebSocket connectionsand we assume one gateway can safely support:
50,000 concurrent connectionsthen the rough minimum is:
5,000,000 / 50,000
=
100 WebSocket gatewaysThis is only a planning assumption. Actual gateway capacity depends on memory per connection, message rate, CPU, TLS overhead, network bandwidth, and the WebSocket implementation.
Production capacity also needs headroom for traffic spikes, maintenance, and server failures.
Collaborative editor scale depends on both persistent connection count and document operation rate.
The core data model should represent documents, edits, collaboration sessions, presence, and permissions without becoming unnecessarily complex.
Important entities include:
User
Document
Operation
DocumentSnapshot
Session
Presence
PermissionA simplified document record may look like:
Document
--------
document_id
owner_id
title
current_version
snapshot_pointer
created_at
updated_atEach accepted edit is represented as an operation:
Operation
---------
operation_id
document_id
user_id
base_version
operation_type
position
content
client_sequence
created_atFor example:
{
"operation_id": "op_9281",
"document_id": "doc_123",
"user_id": "user_17",
"base_version": 104,
"operation_type": "INSERT",
"position": 12,
"content": "hello",
"client_sequence": 918
}The operation_id gives the edit a stable identity and helps with retries and deduplication.
The base_version identifies the document version the client used when creating the operation.
For example:
Client Base Version = 104
Server Version = 106In a server-ordered OT-style design, this tells the collaboration service that the client created the operation without seeing some newer accepted edits. The operation may therefore need to be transformed against those edits before it is applied.
For CRDT-based collaboration, the operation model can be different and may use stable element identifiers or causal metadata instead of relying mainly on absolute positions and a single base version.
The exact representation also becomes richer for structured documents containing paragraphs, tables, images, comments, and formatting.
At a high level:
Document State
+
Collaborative Operations
+
Version / Causal Information
=
Synchronized DocumentA collaborative editor has two main paths: document management and real-time collaboration.
CLIENT
/ \
/ \
v v
REST API WebSocket
| |
v v
API Gateway WebSocket Gateway
| |
v v
Document Collaboration
Service Router
| |
v v
Metadata DB Collaboration Engine
|
OT / CRDT Logic
|
v
Durable Operation Log
|
+--------+--------+
| |
v v
ACK / Snapshot
Fan-Out Service
|
v
Snapshot StorageThe document management path handles operations such as:
Create a document.
Fetch document metadata.
Manage sharing and permissions.
Fetch the initial snapshot.
The real-time collaboration path handles:
Document edits.
Operation acknowledgements.
Remote edits.
Reconnection and synchronization.
The main edit flow is:
Client Edit
↓
WebSocket Gateway
↓
Collaboration Engine
↓
Order / Transform / Merge
↓
Durable Operation Log
↓
ACK + Fan-Out
↓
CollaboratorsThe exact Order / Transform / Merge step depends on whether the system uses an OT-style model, a CRDT, or another collaboration algorithm.
Presence and cursor updates follow a separate lightweight path:
Cursor / Presence Update
↓
WebSocket Gateway
↓
Presence Service
↓
Ephemeral TTL State
↓
Fan-OutThis separation matters because the two types of data have different guarantees:
Document Operations
→ Durable + Correctness-Critical
Presence / Cursor Updates
→ Ephemeral + Loss-TolerantA lost document edit can change the document permanently. A lost intermediate cursor update usually does not matter because a newer cursor position will replace it.
This architecture keeps the correctness-critical editing path separate from lightweight real-time collaboration signals.

A collaborative editor uses both REST APIs and WebSockets because they handle different types of communication.
REST works well for request-response operations such as:
POST /documents
GET /documents/{documentId}
POST /documents/{documentId}/permissionsThese APIs handle document creation, metadata, sharing, permissions, and initial document loading.
Real-time collaboration requires continuous, bidirectional communication.
Clients send:
Document Edits
Cursor Updates
Presence HeartbeatsThe server sends:
Remote Edits
Operation ACKs
Cursor Updates
Presence Changes
Permission UpdatesA persistent WebSocket connection is a natural fit:
Client
⇄
WebSocket Gateway
⇄
Collaboration ServiceIn short, REST handles normal document operations, while WebSockets handle the active real-time collaboration session.
When Alice opens doc_123, the client first needs a recent baseline of the document.
Alice
|
| Open doc_123
v
Document Service
|
| Snapshot + Version
v
AliceFor example:
Snapshot = "Hello World"
Version = 105Alice then joins the collaboration channel and provides her known version:
document_id = doc_123
known_version = 105The collaboration service can then send any operations that occurred after version 105.
A long-lived document may contain millions of historical operations. Replaying the complete operation history every time someone opens the document would be slow and expensive.
Instead, the client loads:
Recent Snapshot
+
Operations After Snapshot
=
Current DocumentFor example:
Snapshot at version 50,000
+
Operations 50,001–50,120
=
Current version 50,120This keeps document loading, reconnection, and recovery efficient.
A collaborative editor should send changes, not the entire document after every keystroke.
For example, if Alice types one character in a 2 MB document, the client can send:
INSERT "a" at position 103Similarly, a deletion can be represented as:
DELETE 4 characters at position 208This keeps real-time messages small and makes network traffic depend mainly on the size of the edit rather than the size of the entire document.
Concurrent editing is the main challenge that separates a collaborative editor from a normal document application.
Suppose Alice and Bob both start with:
ABCDEThey concurrently insert characters at the same logical position:
Alice inserts X
Bob inserts YDepending on the collaboration rules, the final document could be:
ABXYCDEor:
ABYXCDEEither result can be valid. What matters is that every replica applies the same conflict-resolution rules and eventually reaches the same document state.
Concurrent edits may temporarily create different local states, but they must not cause replicas to permanently diverge.
Two common approaches for handling concurrent edits are Operational Transformation (OT) and Conflict-free Replicated Data Types (CRDTs).
Operational Transformation adjusts operations when concurrent edits have changed the document state they were originally created against.
Suppose positions are zero-based and the document is:
ABCDEAlice inserts X before B:
Insert X at position 1The document becomes:
AXBCDEAt the same time, Bob creates an operation from the original document:
Delete character at position 3In the original document, position 3 refers to D.
After Alice's insertion, D moves to position 4. If Bob's original operation were applied unchanged, it could delete the wrong character.
OT transforms Bob's operation:
Delete position 3
↓
Delete position 4Bob's intended edit can now be applied to the updated document.

In a server-ordered OT design, document versions help identify operations created from an older state.
Suppose:
Server Version = 100Alice sends:
Operation A
base_version = 100The server accepts the operation and advances the document:
Version = 101Bob then sends:
Operation B
base_version = 100Because the current version is 101, the server knows Bob created his operation without seeing Alice's accepted edit.
The simplified flow is:
Receive Operation B
↓
Check base_version
↓
Load relevant newer operations
↓
Transform B against them
↓
Apply transformed B
↓
Persist accepted operation
↓
Assign next versionThe exact transformation rules depend on the operation types and document model.
A server-coordinated OT design can maintain a canonical sequence for each document:
Document doc_123
v101 → Alice insert
v102 → Bob delete
v103 → Charlie insertThis ordered stream makes synchronization, transformation, replay, and recovery easier to reason about.
Ordering is needed per document, not globally:
doc_A → its own operation sequence
doc_B → its own operation sequence
doc_C → its own operation sequenceIndependent documents can therefore be processed and scaled independently.
Users should not wait for a server round trip after every keystroke.
Instead, the client applies an edit locally first:
Alice types "a"
↓
Render locally immediately
↓
Send operation asynchronously
↓
Server processes operation
↓
ACK / ReconcileThis is optimistic local editing. It keeps typing responsive even when network latency is noticeable.
The client must therefore track operations that have not yet been confirmed by the server.
A simplified OT-style client may maintain:
Server-Confirmed State
+
Outstanding Operation A1
+
Buffered OperationsFor example, Alice may send A1 and continue typing before its ACK arrives. Those later edits remain buffered while A1 is outstanding.
If Bob's remote operation arrives during this time, Alice's client must reconcile it with her outstanding local operation before updating the visible document.
Once A1 is acknowledged, the next buffered operation can be sent using the updated state.
This is why collaborative editing requires coordination logic on both the client and server.
A Conflict-free Replicated Data Type (CRDT) takes a different approach to concurrent editing.
In many sequence CRDTs, document elements have stable logical identities instead of being identified only by absolute character positions.
Conceptually:
A(id=1)
B(id=7)
C(id=12)Suppose Alice and Bob concurrently insert characters between A and B:
X(id=alice-44)
Y(id=bob-91)The CRDT defines deterministic rules for integrating these concurrent changes so replicas eventually converge after receiving the required updates.
The exact ordering and delivery requirements depend on the CRDT algorithm. Therefore, it is not accurate to say that every CRDT can simply process operations in any order without additional rules.
CRDTs are especially useful when offline editing, independent replicas, or multi-region collaboration are important.
The trade-off is additional metadata and more complex state management. Some CRDT designs also require careful compaction or garbage collection as document history grows.

OT and CRDTs both handle concurrent collaboration, but they use different approaches.
Dimension | OT | CRDT |
|---|---|---|
Core idea | Transform concurrent operations | Merge replicated updates deterministically |
Position model | Often positions plus document versions | Often stable logical identifiers |
Coordination | Commonly server-coordinated | Can support more independent replicas |
Offline editing | Supported, but synchronization can be complex | Often a strong fit |
State overhead | Can remain relatively compact | Usually requires additional metadata |
Rich documents | Transformation rules become complex | Data model and merge rules become complex |
Convergence | Requires correct transformation and integration rules | Provided by the CRDT's merge semantics |
Main challenge | Transformation correctness | Metadata and merge complexity |
OT is not inherently centralized, and CRDTs are not inherently serverless or decentralized. These are architecture choices rather than definitions of the algorithms.
For this design, we will use a server-ordered OT-style model because it gives each document a clear sequencing authority and canonical operation history.
If offline-first editing or more independent multi-region collaboration becomes a major requirement, a CRDT-based model can be considered.
Once the collaboration model is defined, an edit can move through the system.
Suppose Alice and Bob are editing the same document:
Alice Server Bob
| | |
| Insert "X" | |
|--------------------->| |
| | Validate |
| | Order / Transform |
| | Persist |
| | |
|<------- ACK ---------| |
| |---- Insert "X" ---->|Alice applies the edit locally before receiving the ACK, so typing remains responsive.
The server then:
Receive Operation
↓
Authenticate + Authorize
↓
Validate
↓
Order / Transform / Merge
↓
Durably Append
↓
Assign Canonical Version
↓
ACK + Fan-OutThe exact merge step depends on whether the system uses OT, CRDTs, or another collaboration model.
Consider what happens if the server broadcasts an operation before it becomes durable:
Receive Operation
↓
Broadcast
↓
Server Crashes
↓
Operation Was Never PersistedBob may have already seen the edit, but after recovery the durable document no longer contains it.
To avoid this inconsistency, an accepted edit should reach the system's required durable boundary before the client is told that it is safely saved.
A safer flow is:
Process Operation
↓
Durably Append
↓
Assign Canonical Version
↓
ACK Author
+
Fan-Out to CollaboratorsDo not declare an edit safely accepted until it reaches the system's defined durable boundary.
Accepted document edits should be stored in a durable operation log instead of rewriting the complete document after every change.
For example:
Document doc_123
v101 INSERT(...)
v102 DELETE(...)
v103 INSERT(...)
v104 FORMAT(...)The operation log supports:
Document reconstruction.
Reconnection and replay.
Failure recovery.
Version history.
OT transformation against recent operations.
Debugging and auditing.
However, replaying the complete operation history becomes expensive as a document grows. This is why the system also needs snapshots.
A snapshot stores the materialized document state at a specific version.
Snapshot at version 100,000
+
Operations 100,001 onward
=
Current DocumentWhen a document opens or a collaboration node recovers, it can load the latest snapshot and replay only the operations created after it.
Snapshot frequency creates a trade-off.
Creating snapshots too often increases storage writes and background processing. Creating them too rarely increases document-load and recovery time.
A snapshot policy can consider:
Operations since the previous snapshot.
Document size.
Time since the previous snapshot.
Document activity.
For example, an active document might create a new snapshot after a configured number of operations or after a bounded amount of time.
Old operations should only be compacted or removed when the system is sure they are no longer required for recovery, synchronization, version history, or collaboration semantics.

Different types of collaborative-editor data have different access patterns and durability requirements, so they do not need to use the same storage system.
Data | Storage Role |
|---|---|
Document metadata | Relational or distributed metadata store |
Permissions | Strongly consistent metadata store |
Operations | Durable append-oriented operation store |
Snapshots | Blob/object storage |
Images and attachments | Object storage |
Presence | In-memory TTL store |
Connection mappings | Distributed ephemeral state |
A document metadata record may contain:
document_id
owner_id
title
latest_version
latest_snapshot_version
latest_snapshot_uri
created_at
updated_atPermissions can be stored separately:
document_id
principal_id
roleCommon roles include:
OWNER
EDITOR
COMMENTER
VIEWERFor the operation log, an important access pattern is:
Get operations for document D
where version > V
ordered by versionA useful logical key is therefore:
(document_id, version)This supports efficient replay and synchronization of a specific document.
The important design decision is the access pattern and required guarantees, not the database brand.
Most collaboration work is scoped to a document, so document_id is a natural partitioning key.
Conceptually:
document_id
↓
Routing Layer
↓
Collaboration ShardDifferent documents can be distributed across different collaboration shards, allowing the system to scale horizontally.
For active editing, operations for the same document should normally reach the same logical owner or sequencer in a server-ordered design.
doc_123
↓
Logical Document OwnerThe owner can maintain:
Current Version
Recent Operations
Active Session State
Document Coordination StateThis makes document-scoped ordering and conflict handling easier.
The document does not need to remain on the same physical server forever. Ownership can move during failures or rebalancing, but the system should prevent multiple owners from independently acting as the authoritative sequencer at the same time.
WebSocket gateways manage long-lived client connections rather than document conflict-resolution logic.
Their main responsibilities include:
Authentication.
Persistent connections.
Document subscriptions.
Heartbeats.
Message delivery.
Backpressure.
The Collaboration Engine separately handles document-level logic such as ordering, transformation, merging, and accepted operation state.
Clients
↓
WebSocket Gateways
↓
Collaboration Router
↓
Collaboration EngineSeparating these responsibilities allows WebSocket capacity and document processing to scale independently.
Suppose we have:
Concurrent Connections = 5,000,000
Planning Assumption:
50,000 connections per gatewayThe rough minimum would be:
5,000,000 / 50,000
=
100 gatewaysThis is only a capacity-planning estimate. Production requires additional gateways for traffic spikes, maintenance, and failures.
Actual capacity depends on factors such as memory per connection, CPU usage, TLS overhead, message rate, network bandwidth, and the WebSocket implementation.
A connection registry may track mappings such as:
user/session → gateway
document → interested gatewaysThis lets the collaboration layer publish an accepted update to the gateways serving that document instead of managing every client connection itself.
Most documents may have only a few active collaborators, but a popular document can have:
Hundreds of Editors
Thousands of ViewersA single document owner should not have to directly push every accepted edit to every connected client.
Instead, separate document sequencing from connection fan-out:
Document Sequencer
|
v
Pub/Sub
/ | \
v v v
WS-1 WS-2 WS-3
| | |
Clients Clients ClientsThe document sequencer handles the correctness-critical write path and establishes the accepted operation sequence.
The pub/sub and WebSocket gateway layers handle distribution.
For example, a document may have:
20 Editors
100,000 ViewersEditors use the full collaboration path because they submit operations that may conflict with one another.
Read-only viewers only need the accepted document updates, so they can use a scalable broadcast path.
Editors
↓
Collaboration Engine
↓
Document Sequencer
↓
Pub/Sub
↓
Gateway Fleet
↓
ViewersThis reduces fan-out pressure on the document sequencer.
However, fan-out scaling does not remove the write bottleneck of an extremely active document. If thousands of users are concurrently editing the same document, its operations still need compatible ordering or merge semantics.
That hot-write problem is different from the hot-viewer fan-out problem and may require additional collaboration or document-model strategies.

Presence tells collaborators who is currently active in the document.
Unlike document edits, presence is temporary and does not need durable storage.
A presence record may contain:
document_id
user_id
session_id
last_heartbeat
statusClients periodically send heartbeats. If a heartbeat is not received before its TTL expires, the session can be treated as offline.
This makes presence naturally approximate. For example, if Alice closes her laptop without sending a disconnect message, she may appear online for a few seconds until the heartbeat expires.
Cursor and selection updates are also ephemeral and high-frequency.
For collaborative documents, cursor positions should ideally use transformable or relative anchors rather than relying only on absolute character offsets, which can become stale after concurrent edits.
Cursor traffic is:
High Frequency
Ephemeral
Loss-TolerantIf collaborators miss positions 41, 42, 43, and 44 but receive the latest position 45, document correctness is unaffected.
Clients can therefore throttle or coalesce cursor updates:
Local Cursor Movement
↓
Throttle / Coalesce
↓
Send Latest PositionThis gives us an important separation:
Document Edits
↓
Durable + Correctness-Critical
Presence / Cursor Updates
↓
Ephemeral + Loss-TolerantCursor movements should not be stored in the durable document operation log.
A client may temporarily lose its connection while other users continue editing the document.
Suppose Alice disconnects at:
Version = 200while the server advances to:
Version = 230When Alice reconnects, she sends:
last_seen_version = 200If the missing history is still available, the server can return:
Operations 201–230Alice applies these operations and catches up with the current document.
If the operation gap is too large or older history has already been compacted, the server can instead send:
Newer Snapshot
+
Operations After Snapshot
=
Current Server StateThis prevents clients from replaying very large operation histories.
Reconnection becomes harder when Alice continues editing while offline.
Suppose the server advances from version 200 to 230, while Alice creates local operations:
Remote:
201 ... 230
Local:
A1
A2
A3When she reconnects, the system must reconcile the remote changes with her pending local edits.
Conceptually:
Last Synchronized State
+
Remote Changes
+
Pending Local Changes
↓
OT / CRDT Reconciliation
↓
Synchronized DocumentThe exact process depends on the collaboration algorithm.
CRDTs are often attractive for offline-first systems because replicas can maintain local changes and merge them later. OT-based systems can also support offline editing, but usually require careful rebasing and synchronization against newer server operations.
Network failures can make operation delivery uncertain.
Suppose Alice sends an operation, but the connection fails before she receives its ACK. She cannot safely assume that the server rejected the operation.
Every client-generated operation should therefore have a stable ID, for example:
operation_id = session_42:sequence_918Alice can retry the same operation using the same ID.
If the server already accepted it:
Duplicate operation_id
↓
Do Not Apply Again
↓
Return Previous Canonical ResultA practical delivery model is:
At-Least-Once Delivery
+
Stable Operation IDs
+
Deduplication
+
Idempotent ProcessingThis avoids trying to guarantee exactly-once network delivery while still ensuring that one logical edit is not applied twice.
The client should keep pending operations until they reach the required acknowledgment state.
For example:
Pending Queue
A1
A2
A3After A1 is safely acknowledged:
ACK A1
↓
Remove A1If the connection fails before the ACK:
Connection Lost
↓
Keep A1 Pending
↓
Reconnect
↓
Retry A1 with Same operation_idServer-side deduplication prevents the retry from applying the edit twice.
For stronger autosave and offline guarantees, pending operations can also be stored in browser or device storage. This allows them to survive a tab, browser, or device-app restart.
In a server-ordered OT design, accepted operations receive a sequence or version within each document.
doc_1 → independent sequence
doc_2 → independent sequence
doc_3 → independent sequenceThe system does not need a global order across unrelated documents. This allows different documents to be processed and scaled independently.
For CRDT-based systems, ordering and delivery requirements depend on the chosen algorithm. We should not assume that every CRDT can safely handle arbitrary duplicates or out-of-order updates.
A slow client must not create an unlimited outbound message queue.
Useful backpressure policies include:
Bound per-connection buffers.
Coalesce or drop stale cursor and presence updates.
Never silently drop required document operations.
Disconnect extremely slow consumers when necessary.
Force resynchronization when a client falls too far behind.
For example:
Client Falls Far Behind
↓
Stop Large Replay
↓
Fetch Recent Snapshot
↓
Apply Operation Tail
↓
Resume CollaborationThis protects the server while preserving document correctness.
Sending a separate network message for every keystroke creates unnecessary protocol and processing overhead.
For example, typing:
systemcould produce six small insert operations.
Instead, the client can briefly batch compatible edits:
INSERT "system"The trade-off is:
Larger Batch
↓
Less Network + Processing Overhead
+
Slightly Higher Propagation LatencyThe batching window should remain small so remote edits still feel interactive.
Batching must also preserve operation order and editing semantics. Unrelated or conflicting operations should not be combined simply to reduce message count.
A production collaborative editor usually supports more than plain text.
Paragraphs
Headings
Lists
Tables
Images
Links
Formatting
CommentsInstead of representing the document as one long string, a rich-text editor can use a structured tree:
Document
|
+-- Paragraph
| |
| +-- Text
|
+-- Heading
|
+-- Table
|
+-- Row
|
+-- CellOperations can then target meaningful document structures instead of relying only on absolute character positions.
For example:
Insert text into paragraph P
Delete table row R
Move block B
Apply bold formatting to rangeThis makes rich-text collaboration more expressive, but conflict-resolution rules also become more complex.
Structured nodes should have stable identities.
For example:
paragraph_id = p_381
table_id = t_42Stable identities make it easier to reference content even when surrounding parts of the document change.
They are useful for:
Comments
Suggestions
Cursor Anchors
Selections
Structured EditsFor example, a comment should not depend only on:
characters 200–219If another user inserts text before that range, those absolute positions may no longer point to the intended content.
Instead, comments, cursors, and selections can use relative or transformable anchors that move with document edits.
Technical convergence does not automatically mean the result will match user expectations.
Suppose:
Alice → Deletes paragraph P
Bob → Formats text inside PBoth replicas may converge to the same final state, but the product still needs to decide what that state should mean.
Similar conflicts include:
Delete versus edit.
Delete versus formatting.
Concurrent block moves.
Concurrent structural changes.
Comments attached to deleted content.
These are product semantics, not only distributed-systems problems.
The product defines the expected behavior, and the OT or CRDT implementation must apply those rules consistently across replicas.
The durable operation log and snapshots can also support document version history.
A historical state can be reconstructed using:
Snapshot Before Target Version
+
Operations Up to Target Version
=
Historical Document StateHowever, users should not see a separate revision for every keystroke.
User-facing history can group low-level operations into meaningful revisions based on factors such as time, author, or editing session.
A collaborative editor continuously persists accepted edits, so a traditional Save button becomes less important.
The client should still distinguish between different states:
Local Unsent
↓
Sent but Unacknowledged
↓
Durably AcceptedA visible Saved indicator should appear only when the relevant edits have reached the system's defined durable acceptance state.
A collaborative document can support roles such as:
OWNER
EDITOR
COMMENTER
VIEWERAuthorization must always be enforced by the server, not only by the client UI.
Suppose Bob's permission changes while he is actively editing:
EDITOR
↓
VIEWERThe permission update should reach the active collaboration session.
Permission Changed
↓
Update Active Session
↓
Reject Future Unauthorized EditsEven if Bob's client has not yet received the permission update, the server must check the current authorization policy before accepting new edits.
Sharing operations can remain on the normal metadata API path:
POST /documents/{documentId}/permissionsFor example:
{
"principal_id": "user_42",
"role": "EDITOR"
}Keeping sharing and permission management separate from the high-frequency keystroke path keeps the collaboration architecture simpler.
Large files such as images should not travel through the normal real-time document-operation channel.
A better upload flow is:
Client
↓
Request Upload Authorization
↓
Upload to Object Storage
↓
Receive Attachment ID
↓
Submit Document OperationThe collaborative operation then contains a small reference:
INSERT_IMAGE image_928Other clients can resolve that reference and load the image from the appropriate storage or delivery layer.
This keeps real-time collaboration messages small and prevents large uploads from blocking normal document edits.
Search, analytics, notifications, and similar workloads should stay outside the latency-sensitive collaboration path.
After a document change becomes durable, it can be published for asynchronous processing:
Durable Document Change
↓
Event Stream
/ | \
v v v
Search Analytics NotificationsOther downstream consumers may include:
Activity feeds.
Audit pipelines.
Data processing systems.
These systems can usually be eventually consistent. A user should not have to wait for search indexing or analytics processing before an edit is accepted.
A system should avoid independently writing an accepted document change and its downstream event:
Write Document
↓
Crash
↓
Event Never PublishedThis can leave downstream systems permanently unaware of the change.
If the authoritative database and outbox support the same transactional boundary, the system can atomically commit both:
Document Update
+
Outbox Event
↓
Single Transaction
↓
Background Publisher
↓
Event StreamThe publisher may deliver an event more than once, so downstream consumers should still be designed for idempotent processing where needed.
If the durable operation log is already the source of truth, another option is to derive downstream events directly from that log or through an appropriate change-data-capture mechanism.
The exact pattern depends on the storage architecture. The important requirement is to avoid a gap where an authoritative document change becomes durable but its required downstream event is permanently lost.
Global collaboration creates a trade-off between low latency and coordination complexity.
A simple approach is to assign each document a home region:
doc_123
↓
Asia RegionUsers can connect to nearby regional WebSocket gateways, while authoritative ordering for doc_123 remains in its home region.
Alice → Asia Gateway ──┐
│
Bob → Europe Gateway ├──→ Home Region → Document Sequencer
│
Carol → US Gateway ────┘This keeps document ownership, ordering, and recovery easier to reason about.
The trade-off is latency. A collaborator far from the document's home region may need to send edits across regions before they become authoritative.
A CRDT-based design can support more independent regional editing, but the exact benefits and coordination requirements depend on the chosen CRDT algorithm and product requirements.
A document may need to move to another owner or region because of rebalancing, failures, or traffic changes.
Routing can use techniques such as:
Virtual Shards
Consistent Hashing
Routing MetadataFor an active document, migration must preserve the canonical write sequence and prevent two owners from independently accepting authoritative writes.
A simplified handoff is:
Old Owner
↓
Stop / Drain Writes
↓
Transfer State
↓
Advance Ownership Epoch
↓
Update Routing
↓
New Owner Accepts WritesAn ownership epoch or fencing token can prevent a stale owner from continuing to accept writes after ownership has moved.
A collaborative editor must recover from failures without silently losing edits that were already declared durable.
If a collaboration node crashes:
Collaboration Node Fails
↓
WebSocket Disconnects
↓
Client Reconnects
↓
New Owner Assigned
↓
Load Recent Snapshot
↓
Replay Durable Operation Tail
↓
Resume SynchronizationClients provide their last acknowledged state and safely retry pending operations using the same stable operation IDs.
Because accepted operations are already durable, the replacement node can reconstruct the document from the snapshot and operation log.
If document metadata or permission storage becomes unavailable, new document opens, permission checks, or sharing changes may fail.
Whether existing sessions can continue depends on the product's security policy.
For example:
Permission Store Unavailable
↓
Strict Revocation Policy
↓
Restrict / Stop New EditsA system may use cached authorization information in some cases, but it should not weaken required revocation guarantees simply to maintain availability.
If offline editing is supported, clients can preserve local work during a network failure:
Network Lost
↓
Persist Local Pending Edits
↓
Reconnect
↓
Fetch Remote Changes
↓
Merge / Rebase
↓
Submit Pending EditsThe exact reconciliation process depends on the OT or CRDT model.
If offline editing is not supported, the client can enter a reconnecting or read-only state until communication is restored.
If pending edits exist only in memory, a browser or application crash can lose them.
For stronger offline and autosave guarantees:
Pending Operations
↓
Local Persistent Storage
↓
Browser / App Restarts
↓
Restore Pending Operations
↓
Reconnect and SynchronizeStable operation IDs allow restored operations to be retried without applying the same logical edit twice.
Different parts of a collaborative editor need different consistency guarantees.
Data | Consistency Need |
|---|---|
Document operations | Defined ordering/merge semantics with convergence |
Accepted edit durability | Defined durable acceptance boundary |
Permissions | Stronger consistency based on authorization policy |
Presence | Eventual / best effort |
Cursor updates | Best effort |
Search | Eventual |
Analytics | Eventual |
Notifications | Eventual |
Document edits need correctness and convergence, while temporary information such as presence can tolerate stale or missing updates.
Using the strongest consistency model everywhere would add unnecessary latency and cost. Using weak consistency everywhere would risk correctness and security.
Every edit creates a trade-off between fast acknowledgment and durable storage.
Waiting for synchronous replication across distant regions before acknowledging every edit can improve failure tolerance but also increase user-visible latency.
At the other extreme:
Edit
↓
Stored Only in Memory
↓
ACK Sent
↓
Server Crashes
↓
Edit LostA practical design defines a durable acceptance boundary that is strong enough for the product's data-loss requirements while remaining close enough to keep editing responsive.
Client Edit
↓
Local / Regional Durable Boundary
↓
ACK
↓
Additional Replication
↓
Disaster-Recovery CopiesThe exact replication policy depends on the product's durability and latency SLOs.
Collaborative documents may contain sensitive information, so security must cover both stored data and real-time communication.
Important controls include:
Encryption in transit and at rest.
Short-lived authenticated connection credentials.
Server-side authorization.
Secure attachment access.
Tenant isolation.
Audit logging for sensitive actions.
Operation and payload validation.
Rate limiting.
A malicious or buggy client may send excessive operations, oversized payloads, or very frequent presence updates.
Rate limits can cover:
Operations / User
Cursor Events / Session
Payload Size
Document Size
Active Sessions
Connection AttemptsWebSocket traffic should receive the same authentication, authorization, validation, and abuse protection as REST APIs.
Observability should measure both infrastructure health and the quality of the collaboration experience.
Important metrics include:
Active WebSocket connections.
Active collaborative documents.
Operations per second.
Operation ACK latency.
Edit propagation latency.
Reconnect and resync rate.
Transformation or merge failures.
Duplicate-operation rate.
Operation-log lag.
Snapshot age and generation latency.
Slow-consumer disconnects.
These metrics help identify whether problems come from connection handling, collaboration processing, persistence, or fan-out.
An important user-facing metric is how long a local edit takes to appear for another collaborator.
Alice Local Edit
↓
Network
↓
Server Processing
↓
Durable Acceptance
↓
Fan-Out
↓
Network
↓
Bob RenderClient clocks may not be perfectly synchronized, so a single timestamp difference is not always reliable.
Production systems can combine distributed trace IDs, client telemetry, server-side stage timings, and synchronized clocks where available to locate latency.
Document startup should also be measured end to end:
Open Document
↓
Metadata Fetch
↓
Snapshot Download
↓
Operation Replay
↓
WebSocket Join
↓
Time to EditableUseful measurements include snapshot download time, number of operations replayed, synchronization time, and overall time to editable.
Snapshot and compaction policies directly affect this startup latency, so these metrics can also help tune snapshot frequency.
Not every collaborative feature has the same importance.
If presence infrastructure fails, users should ideally continue editing even if cursor avatars temporarily disappear. Similarly, analytics or notification failures should not block document edits.
A useful priority model is:
Correctness-Critical
--------------------
Document Operations
Durability
Authorization
Real-Time Experience
--------------------
Edit Propagation
Reconnect / Resync
Best-Effort
-----------
Presence
Cursor Updates
Analytics
NotificationsThis separation keeps failures in secondary features from unnecessarily taking down the core editing path.
The complete architecture combines document management, real-time collaboration, durable storage, presence, fan-out, and asynchronous processing.
CLIENTS
/ | \
/ | \
v v v
REST APIs WebSocket Attachments
| | |
v v v
API Gateway WS Gateway Object Storage
| |
| v
| Collaboration Router
| |
v v
Document Service Document Owner /
| Sequencer
| |
| OT / CRDT Engine
| |
| v
| Durable Operation Log
| |
| +-----+------+
| | |
v v v
Metadata DB Snapshot Pub/Sub
Service / | \
| v v v
v WS WS WS
Snapshot Store | | |
Clients...The correctness-critical edit path is:
Client Edit
↓
WebSocket Gateway
↓
Collaboration Router
↓
Document Owner / Sequencer
↓
Authenticate / Authorize
↓
Order / Transform / Merge
↓
Durable Operation Log
↓
ACK + Fan-OutPresence follows a separate, lightweight path:
Client
↓
WebSocket Gateway
↓
Presence Service
↓
TTL / Ephemeral State
↓
Pub/Sub
↓
Interested ClientsDownstream systems also stay outside the critical edit path:
Durable Operation Log
↓
Event / CDC Path
/ | \
v v v
Search Analytics NotificationsThis architecture keeps document correctness and durability on the critical path while presence, search, analytics, and notifications can scale and fail independently.

The low-level design should focus on collaboration state and behavior rather than large inheritance hierarchies.
Useful components include:
Document
Operation
DocumentSession
CollaborationManager
OperationTransformer
OperationLog
SnapshotManager
PresenceSession
PermissionServiceAn OT-style operation can contain:
Operation
---------
operationId
documentId
authorId
baseVersion
type
payload
clientSequenceHere, operationId provides stable identity for retries and deduplication, while baseVersion identifies the document version against which the operation was created.
A document session may maintain:
DocumentSession
---------------
documentId
currentVersion
activeSessions
recentOperations
ownerEpochThe ownerEpoch helps identify the current valid owner when document ownership changes.
The collaboration manager may expose operations such as:
joinDocument()
submitOperation()
transformOperation()
acknowledgeOperation()
broadcastOperation()
resyncClient()
leaveDocument()The important LLD concerns are stable operation identity, version handling, idempotency, document ownership, session state, and clear separation of collaboration responsibilities.
The client needs an explicit state model because it may connect, synchronize, lose connectivity, or lose edit permission.
A simplified state machine is:
DISCONNECTED
↓
CONNECTING
↓
SYNCING
↓
ACTIVE
↙ ↘
RECONNECTING
READ_ONLYDuring SYNCING, the client establishes a consistent baseline and reconciles it with any pending local operations before normal collaboration resumes.
A typical lifecycle is:
Authenticate
↓
Fetch Metadata + Permission
↓
Establish Collaboration Session
↓
Load Snapshot / Known State
↓
Replay Missing Operations
↓
Reconcile Pending Local Edits
↓
Join Active Document Session
↓
ACTIVE
↓
Optimistic Local Editing
↓
Submit Stable-ID Operations
↓
Order / Transform / Merge
↓
Durably Persist
↓
ACK + Fan-OutThe exact order of fetching the snapshot and establishing the WebSocket connection can vary by implementation. What matters is that the client reaches a synchronized state before treating the collaboration session as fully active.
If the connection fails:
ACTIVE
↓
Connection Lost
↓
RECONNECTING
↓
Send Last Known Version
↓
Replay Missing Operations
or
Snapshot + Tail Resync
↓
Reconcile Pending Operations
↓
ACTIVETypical WebSocket message types include:
JOIN_DOCUMENT
DOCUMENT_OPERATION
OPERATION_ACK
REMOTE_OPERATION
CURSOR_UPDATE
PRESENCE_UPDATE
RESYNC_REQUIRED
PERMISSION_CHANGED
PING
PONGThese messages should not all have the same delivery guarantees.
DOCUMENT_OPERATION
OPERATION_ACK
REMOTE_OPERATION
↓
Correctness-Critical
CURSOR_UPDATE
PRESENCE_UPDATE
↓
Ephemeral / CoalescibleDocument operations must not be silently lost, while stale cursor and presence updates can be dropped or replaced by newer updates.
A collaborative editor does not need every component from day one. The architecture can evolve as requirements become more complex.
Stage 1 — Single-User Editor
Start with basic document storage:
Client
↓
Document API
↓
DatabaseStage 2 — Real-Time Collaboration
Add WebSockets, document sessions, and real-time operation messages so multiple users can see each other's changes.
Stage 3 — Concurrent Editing
Add document versions, OT or CRDT semantics, conflict-resolution rules, and operation deduplication.
Stage 4 — Durable Collaboration
Add a durable operation log, snapshots, pending client operations, and reconnect synchronization.
Stage 5 — Large-Scale Collaboration
Add a WebSocket gateway fleet, document-based sharding, pub/sub fan-out, presence infrastructure, backpressure, and hot-document handling.
Stage 6 — Global Collaboration
Add regional gateways, document home regions, cross-region replication, ownership migration, and stronger offline support.
Stage 7 — Rich Document Editor
Add structured documents, comments, suggestions, attachments, version history, search, and scalable delivery for large read-only audiences.
This progression makes the architecture easier to explain because each new component solves a specific problem introduced by the previous stage.
A collaborative editor involves several design choices where improving one property may increase complexity, latency, or cost elsewhere.
Decision | Option A | Option B | Main Trade-Off |
|---|---|---|---|
Collaboration model | OT | CRDT | Transformation logic vs replicated metadata and merge complexity |
Local editing | Wait for server | Optimistic editing | Simpler state management vs responsive typing |
Durability | ACK before durable persistence | ACK after durable boundary | Lower latency vs lower data-loss risk |
Operation delivery | Fine-grained | Micro-batched | Faster propagation vs lower protocol overhead |
Snapshots | Frequent | Infrequent | Faster recovery vs more storage and background work |
Sequencing | Document owner | More autonomous replicas | Simpler ordering vs regional/offline flexibility |
Temporary state | Persist more state | Keep it ephemeral | Recovery guarantees vs storage and processing cost |
Multi-region | Home region | More independent regional collaboration | Simpler coordination vs lower regional edit latency |
There is no single configuration that fits every collaborative editor.
The right choices depend on latency requirements, offline behavior, document complexity, durability guarantees, collaboration model, and expected scale.
Several mistakes can make a collaborative-editor design incorrect or unnecessarily complex.
Treating WebSockets as the solution to concurrent-edit conflicts.
Sending the complete document after every small edit.
Waiting for a server ACK before showing local typing.
Describing OT as always centralized or CRDTs as always decentralized.
Ignoring retries, duplicate operations, and idempotency.
Replaying the complete operation history instead of using snapshots.
Persisting cursor and presence updates like document edits.
Requiring global operation ordering across unrelated documents.
Ignoring slow consumers, hot documents, reconnects, or permission changes during active sessions.
Naming technologies such as Kafka, Redis, or a specific database without explaining the problem each component solves.
A strong collaborative editor design starts with convergence and operation semantics, then explains durability, recovery, real-time delivery, and document-scoped scaling.
A collaborative editor allows multiple users to edit the same document concurrently, see changes in near real time, and resolve concurrent edits so replicas eventually converge to the same logical state.
The main challenge is concurrent editing. Multiple users may modify the same part of a document before seeing each other's changes, so the system needs clear conflict-resolution and convergence rules.
Collaborative editing requires continuous bidirectional communication. Clients send document operations, cursor updates, and presence events while receiving remote edits and acknowledgments. WebSockets avoid repeated polling and keep a persistent communication channel open.
Documents can be large while individual edits are usually small. Sending operations such as insert or delete reduces network bandwidth, serialization work, and processing overhead.
Operational Transformation, or OT, transforms an operation when concurrent edits have changed the document state it was originally created against. The goal is to preserve the intended effect while keeping replicas consistent.
A CRDT is a replicated data structure designed with merge semantics that allow replicas to converge after receiving the required updates. Sequence CRDTs commonly use stable logical identities to help represent concurrent text edits.
It depends on the requirements. OT works well with server-coordinated editing and canonical operation ordering. CRDTs are often attractive when offline editing or more independent replicas are important. Document structure, metadata overhead, and implementation complexity also affect the decision.
Use optimistic local editing. Apply the user's edit immediately on the client and send the operation asynchronously instead of waiting for a server round trip before rendering it.
The collaboration algorithm transforms, merges, or otherwise reconciles the concurrent operations. The client then updates its local state according to the accepted collaboration semantics.
Give every logical operation a stable operation_id. If the client retries an already accepted operation, the server detects the duplicate and returns the previous canonical result instead of applying it again.
In a server-ordered design, accepted operations can receive document-scoped sequence numbers or versions.
doc_123
v101 → Operation A
v102 → Operation B
v103 → Operation CThere is no need for global ordering across independent documents.
The client sends its last known server version. The server returns the missing operations or, if the gap is too large, a newer snapshot plus the remaining operation tail. Pending local edits are then reconciled before normal editing resumes.
Persist pending edits locally while disconnected. After reconnecting, fetch remote changes and reconcile them with the local operations using the chosen OT or CRDT semantics.
Without snapshots, opening or recovering a long-lived document could require replaying its complete operation history. A snapshot provides a recent baseline, so only newer operations need to be replayed.
Snapshot frequency can depend on operation count, document size, elapsed time, and activity. The goal is to balance snapshot-generation cost against document-load and recovery latency.
Cursor and selection state should remain ephemeral rather than being written to the durable document log. For structured documents, relative or transformable anchors are safer than assuming absolute character offsets remain valid after concurrent edits.
Only the latest useful cursor position normally matters. Missing intermediate positions does not affect document correctness, so cursor events can be throttled, coalesced, or replaced by newer updates.
Use a fleet of WebSocket gateways and distribute persistent connections across them. Track document subscriptions and use pub/sub to fan accepted document updates out to the gateways serving interested clients.
Separate edit sequencing from fan-out. Editors use the full collaboration path, while passive viewers can receive accepted updates through a scalable pub/sub and gateway broadcast path.
document_id is a natural partitioning dimension because collaboration state, operation ordering, snapshots, and history are primarily document-scoped.
Clients reconnect to a healthy node. A new document owner reconstructs the current state from the latest snapshot and durable operation tail. Clients safely retry unacknowledged operations using their stable operation IDs.
Use durable operations and periodic snapshots. Historical states can be reconstructed from a snapshot before the target version plus the operations up to that version. User-facing revisions can group many low-level operations together.
Propagate the permission change to the active collaboration session and enforce the latest authorization policy on the server. A connected client must not be able to continue editing simply because its UI has stale permissions.
A simple design assigns each document a home region for authoritative sequencing while users connect through nearby regional gateways. More independent regional editing can be considered when latency, offline behavior, and collaboration requirements justify the additional complexity.
Monitor operation ACK latency, edit propagation latency, active WebSocket connections, operations per second, reconnect and resync rates, transformation or merge failures, duplicate operations, slow consumers, operation-log lag, snapshot health, and document-load latency.
The final design separates correctness-critical document processing from connection management and best-effort features.
USERS
|
+------------+------------+
| |
v v
REST APIs WEBSOCKET GATEWAYS
| |
v v
Document Service Collaboration Router
| |
v v
Metadata DB Document Owner /
Sequencer
|
OT / CRDT Engine
|
v
Durable Operation Log
|
+---------------+---------------+
| |
v v
Snapshot Service Pub/Sub
| / | \
v v v v
Snapshot Storage WS WS WS
| | |
Clients...Presence and cursor updates follow a separate ephemeral path:
Client
↓
WebSocket Gateway
↓
Presence Service
↓
TTL / Ephemeral State
↓
Pub/Sub
↓
CollaboratorsA document can be reconstructed from:
Latest Snapshot
+
Operations After Snapshot
=
Current DocumentAfter a disconnection, the client synchronizes using:
Last Known Server State
+
Missing Remote Operations
+
Pending Local Operations
↓
Reconciliation
↓
Synchronized ClientThe core edit path remains:
Client Edit
↓
Optimistic Local Apply
↓
WebSocket
↓
Document Owner / Sequencer
↓
Authorize + Validate
↓
Order / Transform / Merge
↓
Durable Append
↓
ACK + Fan-OutA collaborative editor is more than a document database with WebSockets. WebSockets provide real-time communication, OT or CRDT semantics handle concurrent changes, durable operations and snapshots provide recovery, and document-scoped sequencing, sharding, and fan-out allow the system to scale.
Together, these components form the foundation of a Google Docs-like real-time collaborative editor.