
Durgesh Tiwari
Author
An Instagram-like application looks simple from the user's perspective:
User uploads photo/video
↓
Followers open their feeds
↓
Post appearsBut designing this experience at large scale introduces difficult problems:
Where should billions of photos and videos be stored?
How should media be resized and transcoded?
How do we generate personalized feeds efficiently?
What happens when a celebrity with millions of followers publishes a post?
How do we serve viral media without overloading storage?
How do we rank and paginate a continuously changing feed?
How do we handle failures, retries, privacy changes, and deleted posts?

This article designs a scalable Instagram-like photo and video sharing platform. It focuses on architecture and engineering trade-offs rather than claiming to describe Instagram's private internal implementation.
Before designing the architecture, define the core product scope and the system qualities that matter most.
Functional requirements describe what users should be able to do with the platform.
Upload photo and video posts
Follow and unfollow users
View posts from followed accounts
Browse a user's posts
Paginate through the home feed
Reliably access uploaded media
Secondary features can include likes, comments, recommendations, search, hashtags, notifications, Stories, and Reels.
Non-functional requirements define how reliably and efficiently the system should operate at scale.
Low feed latency
High availability
Durable post and media storage
Massive read scalability
High upload throughput
Horizontal scalability
Fast global media delivery
Eventual consistency where acceptable
The system is generally read-heavy. A user may create very few posts but consume hundreds of posts and media objects.
Scale estimation helps identify which parts of the architecture will become bottlenecks first.
Assume:
500 million daily active users
50 million posts/dayAverage post-write throughput:
50,000,000 / 86,400
≈ 579 posts/secondPeak traffic can be significantly higher.
Suppose each active user opens the feed 10 times per day:
500M × 10
= 5 billion feed requests/day
≈ 57,870 feed requests/second on averageEach feed request may return 20–30 posts.
Therefore, the biggest challenges are:
Feed reads
Fan-out
Media traffic
Ranking
Caching
Viral content
The high-level design separates user data, post metadata, media storage, and feed generation so each workload can scale independently.
Mobile / Web Clients
|
v
+-------------+
| API Gateway |
+------+------+
|
+---------------+---------------+
| | |
v v v
User Service Post Service Feed Service
| | |
v v v
Social Graph Post Store Feed Store
|
v
Media Service
|
v
Object Storage
|
v
CDNSupporting systems include:
Event Broker
Fan-Out Workers
Ranking Service
Cache
Notification Service
Media Processing Workers
Search
MonitoringIntroduce these components only when a requirement or bottleneck justifies them.
The data model is relatively small, while most complexity comes from scale, fan-out, media delivery, and feed generation.
These entities represent the minimum data needed for users, posts, relationships, and feed references.
User
user_id
username
profile_photo
bio
created_atPost
post_id
author_id
caption
media_id
media_type
created_at
visibility
statusFollow
follower_id
followee_id
created_at
FeedEntry
user_id
post_id
score
created_atA feed entry should normally contain a lightweight post reference instead of duplicating the complete post.
FeedEntry
↓
post_idThe core APIs cover post creation, social relationships, and feed retrieval.
Create Post
POST /v1/posts
{
"media_id": "media_982",
"caption": "Beautiful sunset"
}Follow User
POST /v1/users/{user_id}/followUnfollow User
DELETE /v1/users/{user_id}/followGet Home Feed
GET /v1/feed?cursor=abc&limit=20Get User Posts
GET /v1/users/{user_id}/posts?cursor=abcMedia handling should be separated from normal API traffic because photos and videos are much larger than ordinary metadata.
Direct upload prevents application servers from becoming unnecessary proxies for large binary files.
Bad approach:
Client
↓
API Gateway
↓
Post Service
↓
Large Media Transfer
↓
StorageBetter approach:
Client
|
| Request upload
v
Media Service
|
| Signed upload instruction
v
Client
|
| Direct upload
v
Object StorageAfter upload, the client creates the post using the returned media reference.

Post metadata and media bytes have very different storage characteristics, so separating them improves scalability and cost efficiency.
Post Database
↓
Metadata
Object Storage
↓
Photos / VideosExample:
Post DB
post_id = p123
author_id = u8
media_id = m771
caption = "Sunset"Object storage:
m771/original.jpgUploaded media usually needs resizing, transcoding, validation, or optimization before it is ready for efficient delivery.
Large original images should be converted into multiple variants so clients do not download more data than needed.
Original Upload
↓
Object Storage
↓
Processing Event
↓
Image Worker
┌──┼────────┐
↓ ↓ ↓
Small Medium LargeTypical variants include thumbnail, profile/grid, feed resolution, and high resolution.
Video processing is more expensive because multiple resolutions, bitrates, thumbnails, and streaming variants may be required.
Original Video
↓
Media Queue
↓
Transcoding Workers
┌────┼─────┐
↓ ↓ ↓
360p 720p 1080pProcessed variants are stored back in object storage.

Media processing may take seconds or minutes, so it should not block the user's upload request.
Upload
↓
Store Original
↓
Create Processing Job
↓
Return Quickly
↓ asynchronous
Processing Workers
↓
Resize / Transcode / ScanMedia can move through states such as:
UPLOADING
↓
PROCESSING
/ \
v v
READY FAILEDA CDN prevents popular images and videos from repeatedly hitting origin storage and reduces latency for global users.
Object Storage
|
v
CDN
/ | \
v v v
Region Region RegionBenefits include:
Lower media latency
Lower origin bandwidth
Better scalability
Better viral-content handling
Post creation should persist the post quickly while moving expensive feed and media work to asynchronous pipelines.
Creator
↓
Request Media Upload
↓
Media Service
↓
Direct Upload to Object Storage
↓
Media Processing Queue
↓
Resize / Transcode Workers
↓
Media Ready
↓
Create Post
↓
Post Database
↓
PostCreated Event
↓
Feed GenerationThe creator should not wait while thousands or millions of follower feeds are updated.
The social graph stores follow relationships and supports both feed generation and follower fan-out.
This direction answers which accounts a user follows.
Alice
├── Bob
├── Charlie
└── DavidThis direction answers which users should receive content from a creator.
Alice
↑
├── Bob
├── Charlie
├── David
└── EmilyLarge systems may maintain structures optimized for both directions because both access patterns are important.
Feed generation decides which posts should appear for a viewer and when the computational cost should be paid.
There are two classic strategies: fan-out on read and fan-out on write.
The feed is assembled when the viewer requests it.
Alice requests feed
↓
Find followees
↓
Fetch recent posts
↓
Merge
↓
Rank
↓
ReturnAdvantages
Cheap post creation
No massive feed writes
Relationship changes naturally reflected
Disadvantages
Higher read cost
Higher feed latency
More runtime merging
The creator's post is inserted into follower feed stores asynchronously after publication.
Bob's Post
↓
Find Followers
/ | \
v v v
Alice Charlie Emma
Feed Feed FeedThis makes later feed reads much faster.
Advantages
Fast reads
Feed mostly precomputed
Disadvantages
Write amplification
Additional storage
Poor fit for huge creators

Pure fan-out on write becomes expensive when one creator has tens or hundreds of millions of followers.
For ordinary creators, eager asynchronous fan-out gives fast feed reads at manageable cost.
Post
↓
Fan-Out on Write
↓
Followers' Feed StoresFor celebrities, storing once and merging during feed reads avoids enormous write amplification.
Celebrity Post
↓
Author Timeline
↓
No Massive Eager Fan-OutAt read time:
Precomputed Feed
+
Celebrity Posts
+
Recommendations
↓
Merge
↓
RankThe threshold can depend on follower count, posting frequency, queue backlog, worker capacity, and actual fan-out cost.

An event broker decouples post creation from feed generation and absorbs temporary bursts in asynchronous work.
Post Service
↓
PostCreated Event
↓
Event Broker
↓
Fan-Out Workers
↓
Feed StoresIt helps with buffering, retries, worker scaling, and failure isolation.
The feed store contains lightweight candidate references, while the author timeline stores posts created by one specific user.
Precomputing unlimited feed history is unnecessary because most users consume only relatively recent content.
A feed store may keep:
Recent N entriesor:
Recent X daysOlder content can be reconstructed when required.
An author timeline contains one creator's posts, while the home feed combines content from many sources for one viewer.
Bob
├── Post 100
├── Post 92
└── Post 81Keeping these concepts separate simplifies feed reasoning.
Ranking determines which candidate posts should appear first after a manageable candidate set has been generated.
Candidate generation narrows the enormous post corpus into a smaller set that is practical to rank.
Sources may include:
Followed accounts
Celebrity timelines
Recommendations
Trending content
Explore candidates
Millions of Posts
↓
Candidate Generation
↓
~1,000
↓
Ranking
↓
~100Ranking uses signals that estimate how useful or relevant each candidate is for the viewer.
Possible signals include freshness, affinity, engagement, content type, and relevance.
score =
f(
freshness,
affinity,
engagement,
relevance
)Large systems often use multiple ranking stages so expensive models run only on the most promising candidates.
Candidates
↓
Cheap Filters
↓
Lightweight Ranker
↓
Advanced Ranker
↓
Feed MixerThe final mixer applies product-level constraints after ranking.
It can enforce creator diversity, content diversity, freshness, recommendation limits, and ad spacing.

The feed read path combines precomputed and dynamically generated candidates before returning a fully hydrated response.
Viewer
↓
Feed Service
|
+── Precomputed Social Feed
+── High-Fan-Out Creator Posts
+── Recommendations
↓
Merge + Deduplicate
↓
Privacy / Block Filtering
↓
Ranking
↓
Feed Mixer
↓
Batch Hydration
↓
Media CDN URLs
↓
ResponseHydration converts lightweight post IDs into the complete data required by the client UI.
Avoid N+1 queries.
Instead use:
Batch post lookups
Batch user lookups
Batch counter lookups
Caches
Denormalized read models where justified
Cursor-based pagination is better suited than large offsets for feeds that continuously change.
GET /v1/feed?cursor=xyz&limit=20Ranked feeds can change between pages, so the system may preserve a feed session or snapshot context.
feed_session_idor:
snapshot_timeThis reduces duplicate and skipped posts during scrolling.
Caching is essential because feed and metadata traffic is much larger than write traffic.
Good cache candidates include profiles, post metadata, feed pages, popular posts, and engagement counts.
Viral content creates hot keys and extremely high repeated read traffic.
Protect the backend using CDN, metadata caches, hot-key replication, request coalescing, and read replicas where appropriate.
A cache stampede occurs when many requests miss the same popular key at the same time.
Mitigations include:
Jittered TTL
Request coalescing
Stale-while-revalidate
Background refresh
Engagement data can often tolerate more eventual consistency than privacy or ownership data.
A durable uniqueness constraint such as:
(user_id, post_id)can prevent duplicate likes.
Displayed aggregate counts can be updated asynchronously.
Feed cards should normally show only a small comment preview.
Full comment history should be fetched separately to avoid unnecessary hydration work.
Precomputed feeds can contain stale post references after follow, unfollow, privacy, or block changes. Therefore, the feed store should not be treated as the final source of authorization.
When Alice follows Bob, the system can either show only Bob's future posts or asynchronously backfill some recent posts into Alice's feed.
Alice Follows Bob
↓
Update Social Graph
↓
Optional Background Backfill
↓
Add Bob's Recent PostsBackfilling improves the initial feed experience without blocking the follow request.
When Alice unfollows Bob, the authoritative social-graph relationship should be updated immediately.
Existing Bob posts may still remain in Alice's precomputed feed, so they should be filtered during feed reads and removed later through asynchronous cleanup.
Privacy and block changes must take effect even when stale feed entries still exist.
Feed Entry
↓
Current Authorization Check
↓
Privacy / Follow / Block Rules
↓
Allowed? → Return Post
Denied? → Filter PostA precomputed feed entry is only a candidate reference—it is not proof that the viewer is still authorized to see the post.
Feed references should point to authoritative post data so edits and deletions do not require rewriting millions of complete feed copies.
Caption or metadata changes update the canonical Post record.
Future hydration naturally retrieves the latest version.
Use a deleted state or tombstone instead of synchronously removing the post from millions of feeds.
Post.status = DELETEDThe read path filters it immediately, while cleanup can happen later.
Different datasets have different access patterns, so storage choices should follow workload requirements rather than one universal database choice.
Partitioning by author_id helps author-timeline reads but may create hot creators.
Partitioning by post_id spreads writes more evenly but requires another index for creator timelines.
user_id is a natural feed partition key because most feed queries are viewer-centric.
Hot users may still require additional caching or replication.
Asynchronous fan-out must handle retries and partial failures without creating duplicates or silently losing work.
A retried event should not create the same post twice in one user's feed.
A uniqueness key such as:
(user_id, post_id)can enforce idempotent writes.
The outbox pattern reliably connects post persistence with event publication.
Database Transaction
|
+── Insert Post
|
+── Insert Outbox EventThen:
Outbox Publisher
↓
Event Broker
↓
Fan-Out WorkersThis prevents a stored post from silently missing its feed event.

Backpressure protects the fan-out pipeline when incoming events arrive faster than workers can process them.
Monitor important signals such as:
Queue depth
Oldest event age
Fan-out latency
Worker throughput
If the backlog keeps growing, the system can autoscale workers, throttle lower-priority work, or temporarily delay non-critical feed precomputation. A queue can absorb short traffic spikes, but it should not be treated as unlimited capacity.
Eagerly maintaining feeds for users who have been inactive for a long time can create unnecessary fan-out work and storage writes.
Active Users
↓
Eager Fan-Out
Inactive Users
↓
Pull / Rebuild on ReturnThis allows the system to spend resources on users who are more likely to consume the generated feed while rebuilding inactive-user feeds when needed.

These systems extend the basic feed architecture by adding additional candidate sources and discovery mechanisms.
Recommended posts become another candidate source before ranking.
Explore is more discovery-oriented than the normal home feed and may rely on embeddings, trending content, and recommendation retrieval.
Search should use a dedicated asynchronously updated index instead of scanning the primary post store.
PostCreated Event
↓
Search Indexer
↓
Search IndexNotifications and media-safety checks are usually asynchronous workflows so they do not unnecessarily block post creation, likes, comments, or other primary user actions.
Notifications inform users about events such as likes, comments, follows, and mentions without blocking the original action.
Like / Comment / Follow Event
↓
Notification Service
↓
Push / In-App NotificationIf notification delivery fails, the original like, comment, or follow should still succeed.
Uploaded photos and videos may require safety checks before they are publicly available.
Common checks include:
Malware scanning
Policy and abuse detection
Copyright checks
Content moderation
Media can remain in a PROCESSING or restricted state until mandatory checks complete.
Uploaded media may contain sensitive metadata such as GPS location, device information, or timestamps.
The media-processing pipeline should remove sensitive EXIF metadata when required before serving the processed media publicly.
Rate limiting protects the platform from spam, bots, abusive behavior, and traffic spikes while keeping services available for legitimate users.
Dimension | Purpose |
|---|---|
User ID | Limit excessive actions from a specific account |
Device ID | Detect repeated abuse from the same device |
IP Address | Control suspicious traffic from a network source |
Endpoint | Apply different limits to uploads, likes, comments, follows, and search |
For example, media uploads can have stricter limits than normal feed reads because they consume significantly more storage and processing resources.
Global deployment reduces latency by serving users from nearby regions while maintaining clear ownership for writes.
A simple model assigns each user a home region for post creation.
Events then propagate asynchronously to other regions.
Feed propagation, counters, recommendations, search, and analytics can often tolerate small delays.
Authentication, authorization, privacy, blocking, and ownership require stronger correctness.
Feed availability should not depend completely on every personalization component being healthy.
A fallback chain can be:
Personalized Ranked Feed
↓
Cached Personalized Feed
↓
Chronological Followed Feed
↓
Recent Cached ContentA less personalized feed is usually better than no feed.
Observability should measure both user-facing latency and the health of asynchronous pipelines.
Important metrics include:
Feed P50/P95/P99 latency
Cache hit ratio
Feed generation latency
Fan-out queue depth and age
Media upload success rate
Media processing latency
Ranking latency
Hydration latency
CDN hit ratio
Database latency
Hot partitions
Duplicate feed entries
Useful freshness metrics include:
upload_to_ready_latencyand:
post_to_follower_visible_latencyThe complete write path combines direct media upload, durable metadata storage, reliable event publication, and asynchronous fan-out.
Creator
↓
Request Media Upload
↓
Media Service
↓
Object Storage
↓
Media Processing
↓
Create Post
↓
Post DB + Outbox
↓
Outbox Publisher
↓
Event Broker
↓
Fan-Out Workers
↓
Followers' Feed StoresThe feed read path merges multiple candidate sources, validates visibility, ranks results, and hydrates the final response.
Viewer
↓
Feed API
↓
Feed Cache / Feed Store
|
+── Precomputed Social Candidates
+── Celebrity Candidates
+── Recommendation Candidates
↓
Merge + Deduplicate
↓
Privacy / Block Filter
↓
Ranking
↓
Feed Mixer
↓
Batch Hydration
↓
Media URLs via CDN
↓
ViewerThe final design separates metadata, media, and feed workloads so each layer can scale according to its own access patterns.
CLIENTS
|
v
API Gateway
|
+----------------+----------------+
| | |
v v v
User Service Post Service Feed Service
| | |
v v v
Social Graph Post Store Feed Store
|
v
Post + Outbox
|
v
Outbox Publisher
|
v
Event Broker
|
v
Fan-Out WorkersMedia follows a separate path:
Client
↓
Media Service
↓
Object Storage
↓
Processing Workers
↓
CDNHLD explains the overall system architecture and how major components interact, while LLD focuses on the internal design and implementation of individual components.
Aspect | HLD | LLD |
|---|---|---|
Focus | Overall distributed architecture | Internal component design |
Level | System-level | Class/component-level |
Examples | Feed Service, CDN, Event Broker, Storage | Post, Media, Repository, Ranking Strategy |
Main Question | How does the platform work and scale? | How is a component implemented? |
Interview Usage | Usually discussed first | Discussed when explicitly requested |
For a typical “Design Instagram” system design interview, start with HLD and move to LLD only when the interviewer asks for deeper implementation details.
A strong system design starts simple and introduces additional components only when scale or product requirements create a real need.
Stage | Architecture | Why It Is Added |
|---|---|---|
Early Product | App Server + SQL + Object Storage | Simple starting point |
Growing Traffic | Multiple servers + CDN + Cache | Read/media scalability |
Large Social Graph | Feed Service + Broker + Fan-Out Workers | Precomputed feeds |
Large Creators | Hybrid Fan-Out + Ranking | Celebrity problem |
Global Scale | Multi-Region + Hot-Key Mitigation + Advanced Observability | Worldwide scale |
These questions cover the decisions and trade-offs most likely to appear in an Instagram-style system design discussion.
Separate media blobs from metadata, use direct object-storage uploads and CDN delivery, and generate feeds through a hybrid fan-out architecture.
It is better suited for large immutable photos and videos than storing binary data directly in ordinary database rows.
It serves popular media close to users and protects origin storage from repeated traffic.
Use asynchronous workers to resize images and transcode videos after the original media is uploaded.
Aspect | Fan-Out on Write | Fan-Out on Read |
|---|---|---|
Work happens | Post time | Read time |
Reads | Fast | More expensive |
Writes | Expensive | Cheaper |
Best fit | Normal creators | Huge creators |
Store their posts in the author timeline and merge them during feed reads instead of eagerly writing to millions of feeds.
Store lightweight post references keyed by viewer and hydrate complete content during reads.
Generate candidates first, then rank a smaller set using freshness, affinity, engagement, and relevance signals.
Batch-load posts, profiles, and counters, with caches or denormalized read models where justified.
Use cursor-based pagination and optionally feed snapshots for ranked feeds.
Use CDN, distributed caches, hot-key mitigation, and cache-stampede protection.
Mark the authoritative post deleted, filter it during hydration, and clean stale references asynchronously.
Use idempotent fan-out workers and a logical uniqueness key such as (user_id, post_id).
Use a transactional outbox or another reliable database-to-event publishing mechanism.
Use replicated services, durable queues, caching, and fallback feeds when ranking or personalization is unavailable.
The central decision is whether to pay feed-generation cost during writes or during reads.
These mistakes usually make an Instagram design either unnecessarily expensive or unreliable.
Storing large media directly in the primary database
Routing every media upload through application servers
Using fan-out on write for every celebrity
Generating every feed completely from scratch
Ignoring CDN and caching
Copying complete post data into every feed entry
Using large offset pagination
Ignoring retries and idempotency
Trusting stale feed entries for authorization
Ignoring backpressure and graceful degradation
Avoid adding Kafka, Redis, Cassandra, Elasticsearch, or any other technology without explaining the requirement it solves.
An Instagram-like system combines separate media, post-distribution, and feed-read pipelines that work together to provide fast and scalable content delivery.
Media:
Client → Object Storage → Processing → CDN
Post Distribution:
Post → Durable Store → Event Pipeline → Feed Stores
Feed Read:
Candidates → Filter → Rank → Hydrate → ViewerThe key design principle is:
Precompute when fan-out is affordable, pull when fan-out is enormous, and combine both approaches when the system needs fast reads as well as massive creators.