
Durgesh Tiwari
Author
Designing a Food Delivery System like DoorDash or Swiggy is an interesting system design problem because it combines several different workloads:
Restaurant Discovery
+
Menu and Availability
+
Order Management
+
Payments
+
Driver Dispatch
+
Real-Time Tracking
+
ETA PredictionAt first, the workflow looks simple:
Customer
↓
Choose Restaurant
↓
Add Food
↓
Place Order
↓
Restaurant Prepares Food
↓
Driver Picks Up
↓
Food DeliveredBut each step introduces a different system design challenge:
What happens if payment succeeds but the Order Service crashes?
How do we prevent two drivers from accepting the same delivery?
How do we find nearby drivers without scanning every driver in the city?
How do we process hundreds of thousands of driver location updates?
How do we calculate ETA when preparation time, driver location, and traffic keep changing?
These problems make food delivery system design much more than a simple ordering application.
Scope: We are designing a conceptual DoorDash/Swiggy-like platform for system design interviews, not the private production architecture of either company.

The system connects three main actors: customers, restaurants, and delivery partners.
Customer
Search nearby restaurants.
Browse menus and add items to a cart.
Place and pay for an order.
View order status.
Track the delivery partner.
Restaurant
Receive and accept or reject orders.
Update preparation status.
Mark orders ready for pickup.
Delivery Partner
Go online or offline.
Send location updates.
Receive and accept delivery offers.
Pick up and deliver orders.
Secondary features can include cancellations, refunds, scheduled orders, ratings, promotions, tips, notifications, and batched deliveries.
The core journey is:
Discovery
↓
Order
↓
Payment
↓
Restaurant Preparation
↓
Driver Dispatch
↓
Pickup
↓
Tracking
↓
DeliveryDifferent parts of the platform have different scalability, latency, and consistency requirements.
Workload | Important Requirements |
|---|---|
Restaurant Discovery | Low latency, high availability, read scalability |
Orders | Correctness, durability, idempotency |
Payments | Correctness, auditability, reconciliation |
Driver Location | Low latency, high write throughput, freshness |
Dispatch | Fast matching, geospatial efficiency, availability |
Tracking | Low latency, scalable real-time delivery |
The overall platform should also support horizontal scaling, fault isolation, graceful degradation, security, and regional operation.
Key idea: Different workloads need different guarantees. Restaurant discovery can tolerate some stale data, while orders and payments require much stronger correctness.
Before choosing databases, queues, or caching strategies, we need a rough idea of the traffic the system may handle. These are hypothetical numbers used only for system design estimation.
Assume:
Metric | Estimate |
|---|---|
Daily active customers | 50 million |
Browse sessions per customer/day | 5 |
Orders per day | 10 million |
Delivery partners | 2 million |
Active drivers at peak | 500,000 |
Restaurant browsing
50M × 5
= 250M browse sessions/day
≈ 2,900 sessions/sec on averageOrder traffic
10M / 86,400
≈ 116 orders/sec on averageActual traffic will not be evenly distributed throughout the day. Lunch and dinner hours can create much higher peaks.
Now consider driver location updates. If 500,000 active drivers send their location every five seconds:
500,000 / 5
= 100,000 location updates/secThis reveals an important characteristic of a food delivery system:
Order creation needs strong transactional correctness, while driver location tracking needs very high write throughput and low latency.
That is why we should not design both workloads in the same way.
Before designing the services, it helps to identify the main business objects and understand what each one represents.
Customer
Restaurant
Menu
MenuItem
Cart
Order
OrderItem
Payment
Delivery
DeliveryPartner
DeliveryOffer
LocationUpdateAn Order represents the customer's food purchase:
Order
-----
order_id
customer_id
restaurant_id
delivery_address
subtotal
tax
delivery_fee
total
payment_status
order_status
created_atA Delivery represents the physical fulfillment of that order:
Delivery
--------
delivery_id
order_id
driver_id
pickup_location
dropoff_location
status
assigned_at
picked_up_at
delivered_atThe two should remain separate:
Entity | Responsibility |
|---|---|
Order | Food items, price, payment, and order lifecycle |
Delivery | Driver assignment, pickup, tracking, and delivery lifecycle |
For example, an order may remain valid even if the original driver cancels and another driver is assigned.
Order O123
|
+--> Delivery Attempt 1 → Driver A → Cancelled
|
+--> Delivery Attempt 2 → Driver B → DeliveredSeparating Order from Delivery makes driver reassignment, retries, cancellations, and advanced features such as batched deliveries much easier to model.
The first version of the architecture should separate discovery, transactional ordering, and real-time location processing.
CUSTOMER APP
|
v
API Gateway
|
+------------------+------------------+
| | |
v v v
Search Service Order Service Tracking Service
| | |
v v v
Search Index Order DB Real-Time Gateway
|
v
Restaurant / Menu Cache
Order Events
|
v
Event Broker
|
+--------------+--------------+
| | |
v v v
Payment Dispatch Notification
Service Service Service
|
v
Geo Index
|
v
Delivery PartnersDriver location follows a separate high-write path:
Driver App
↓
Location Gateway
↓
Event Stream
|
+--> Latest Location Store
+--> Geo Index
+--> Dispatch
+--> ETA
+--> TrackingThis separation is important because transactional order data and GPS updates have completely different access patterns.
When a customer opens a food delivery app, one of the first questions the system needs to answer is:
Which restaurants can deliver to this customer?A simple approach would be to check every restaurant and calculate its distance from the customer. That may work for a small application, but it becomes expensive when thousands or millions of restaurants are available.
A better approach is to use a geospatial index, which groups restaurants by geographic area.
Customer Location
↓
Spatial Index
↓
Current + Neighboring Cells
↓
Candidate Restaurants
↓
Serviceability Check
↓
Ranking
↓
Nearby RestaurantsInstead of searching the entire restaurant database, the system first looks at restaurants in the customer's current geographic cell and nearby cells. This greatly reduces the number of candidates that need further processing.
Common spatial indexing approaches include:
Approach | How It Works |
|---|---|
Geohash | Encodes locations into hierarchical geographic cells |
Quadtree | Recursively divides geographic areas into smaller regions |
Hierarchical Cells | Organizes geographic space at multiple levels |
Database Spatial Index | Uses built-in database support for spatial queries |
Why do we also search neighboring cells?
Consider two locations that are physically very close but happen to fall on opposite sides of a cell boundary:
+----------------+----------------+
| | |
| Customer ● | ● Restaurant |
| | |
+----------------+----------------+
Cell A Cell BIf we search only the customer's cell, we may miss that restaurant even though it is nearby.
So the typical approach is:
Current Cell
+
Neighboring Cells
↓
Candidate Restaurants
↓
Exact Distance + ServiceabilityGeospatial indexing narrows the search space. The final decision still depends on actual distance, delivery area, and restaurant serviceability.

Geographic distance helps find nearby restaurants, but serviceability determines whether a restaurant can actually deliver to the customer.
Aspect | Geographic Distance | Serviceability |
|---|---|---|
Purpose | Find nearby restaurants | Check whether delivery is actually possible |
Based On | Customer and restaurant coordinates | Delivery zones, road access, distance limits, and operational rules |
Example | Restaurant is 2 km away | Restaurant does not deliver across that zone |
Used For | Candidate generation | Final delivery eligibility |
Decision | “Is this restaurant nearby?” | “Can this restaurant deliver here?” |
The discovery flow becomes:
Nearby Restaurants
↓
Serviceability Check
↓
Eligible RestaurantsDistance finds nearby candidates; serviceability decides who can actually deliver.
Restaurant search needs to understand both what the customer wants and where the customer is located.
A customer may search for:
pizza
biryani
burger
Chinese food
restaurant nameUnlike a normal text search, a food delivery app should not return restaurants that cannot serve the customer's location. So the system combines text search, geographic filtering, and serviceability checks.
A restaurant search document may contain:
restaurant_id
restaurant_name
cuisines
menu_item_names
location
rating
delivery_statusThe search flow becomes:
Customer Query
↓
Search Index
↓
Text Candidates
↓
Location + Serviceability Filter
↓
Ranking
↓
Restaurant ResultsRanking can consider text relevance, distance, restaurant availability, delivery ETA, popularity, and other product signals.
This allows a query like "pizza" to return restaurants that are not only relevant, but also practical for the customer to order from.
Restaurant menus are read frequently but updated much less often, making them good candidates for caching.
Customer
↓
Menu Service
↓
Cache
/ \
HIT MISS
| |
v v
Menu Menu DBOn a cache hit, the Menu Service can return the menu immediately. On a cache miss, it reads the authoritative data from the Menu Database and can cache the result for future requests.
Slightly stale menu data may be acceptable while the customer is browsing. But when the customer places an order, the system should revalidate important information such as price and item availability against the authoritative source.
Stage | Data Strategy | Why |
|---|---|---|
Browsing | Cache-friendly | Fast responses and lower database load |
Checkout | Authoritative validation | Prevent stale prices or unavailable items from creating incorrect orders |
Browse for speed; validate again when the customer commits the order.
The price and availability shown while browsing may change before the customer reaches checkout, so the backend must validate them again before creating the order.
Suppose the customer adds:
Paneer Pizza = ₹400Later, the restaurant changes the price to ₹450, but the client still sends:
{
"item_id": "P1",
"price": 400
}The server should never trust the client-provided price as the final amount.
During checkout:
Cart Item IDs
↓
Menu / Pricing Service
↓
Current Price + Availability
↓
Recalculate Total
↓
Create OrderBefore creating the order, the backend should revalidate:
Menu items and quantities.
Current prices.
Item availability.
Restaurant availability.
Delivery eligibility.
Fees, discounts, and coupons.
If anything has changed, the customer can be shown the updated cart before confirming the order.
The client expresses what the customer wants to buy; the server determines what can be sold and the final payable amount.
The Cart Service stores the customer's temporary selections before an order is created.
Cart
----
cart_id
customer_id
restaurant_id
items[]
coupon
estimated_totalA food delivery app commonly keeps items from a single restaurant in one cart because preparation and fulfillment are restaurant-specific.
The cart is different from an Order:
Aspect | Cart | Order |
|---|---|---|
Purpose | Temporary shopping state | Confirmed business transaction |
Data | Selected items and estimated total | Validated items, final amount, payment and status |
Changes | Frequently modified | Controlled through order-state transitions |
Durability | Can use a low-latency key-value/document store | Should be durably persisted |
Price | May contain an estimate or snapshot | Uses server-validated price |
The important transition happens at checkout:
Cart
↓
Validate Price + Availability
↓
Calculate Final Amount
↓
Create Durable OrderThe cart represents purchase intent; the Order represents the committed transaction.
The Place Order API converts a validated cart into a durable order. Because clients may retry requests after timeouts or network failures, the API should support an idempotency key.
POST /v1/orders
Idempotency-Key: 4d8f...
{
"restaurant_id": "R123",
"items": [
{
"menu_item_id": "M101",
"quantity": 2
}
],
"delivery_address_id": "A91",
"payment_method_id": "PM42"
}The client sends what the customer wants to purchase, but it should not decide authoritative values such as:
Final Price
Tax
Delivery Fee
Restaurant Eligibility
Final Payable TotalThe backend validates these values before creating the order.
Order creation must be safe to retry because a successful request can still appear to fail from the client's point of view.
Suppose the server creates the order successfully, but the response is lost:
Customer
↓
Place Order
↓
Order O123 Created
↓
Response Lost
✕
CustomerThe customer retries the request.
Without idempotency:
Request 1 → Order O123
Request 2 → Order O124The same purchase has now created two orders.
Instead, associate the request with a unique key:
(customer_id, idempotency_key)The server stores the result of the first successful request:
First Request
↓
Create Order O123
↓
Store Idempotency Result
Retry with Same Key
↓
Return Order O123In production, it is also useful to associate the key with a request hash. If the same idempotency key is reused with a different payload, the server should reject it rather than treating it as the original request.
Idempotency does not prevent retries; it makes retries safe.
This becomes especially important when order creation triggers payment or other external side effects.
An order moves through several business stages, so a single boolean such as is_completed is not enough to represent its lifecycle.
A simplified order state machine can be:
CREATED
↓
PAYMENT_PENDING
↓
PAID
↓
RESTAURANT_CONFIRMED
↓
PREPARING
↓
READY_FOR_PICKUP
↓
PICKED_UP
↓
DELIVEREDFailure and cancellation paths can include:
PAYMENT_FAILED
RESTAURANT_REJECTED
CANCELLED
REFUNDEDThe system should validate every transition instead of allowing arbitrary status updates.
For example:
PICKED_UP → DELIVERED ✓
CREATED → DELIVERED ✕Explicit states help the system answer important questions such as whether payment has completed, whether the restaurant has accepted the order, whether a driver can be assigned, and whether cancellation is still allowed.
The state machine protects the order lifecycle by allowing only valid business transitions.

Payment has its own lifecycle because payment state and order state can change independently.
Order Service
↓
Payment Service
↓
Payment ProviderA simplified payment lifecycle can include:
INITIATED
↓
AUTHORIZED
↓
CAPTURED
FAILED
UNKNOWN
REFUNDEDThe most important failure case is a timeout:
An HTTP timeout does not mean the payment failed.

For example, the payment provider may successfully process the charge, but the response may be lost:
Payment Service → Payment Provider
↓
Payment Success
↓
Response Lost ✕In this situation, the payment should enter an UNKNOWN or pending-resolution state rather than immediately being marked as failed.
The system can resolve uncertain payments using:
A stable provider transaction reference.
Provider status APIs.
Payment webhooks.
Periodic reconciliation.
Blindly retrying an uncertain payment can create duplicate charges unless the payment operation is also idempotent at the provider level.
Order creation, payment, and restaurant confirmation happen across different systems, so they cannot safely be wrapped inside one local database transaction.
This design would be incorrect:
BEGIN TRANSACTION
Create Order
Call Payment Provider
Notify Restaurant
COMMITA database rollback cannot undo an external payment or a message already delivered to a restaurant.
Instead, the system can use a Saga-style workflow, where each step is persisted and failures are handled through retries or compensating actions.
One possible flow is:
Create Order
↓
Authorize Payment
↓
Restaurant Accepts
↓
Capture Payment
↓
Dispatch DriverIf the restaurant rejects the order:
Restaurant Rejects
↓
Void Authorization
or Refund if Captured
↓
Cancel OrderThe exact timing of payment authorization and capture depends on business and payment-provider requirements.
The goal is not one giant distributed transaction, but a recoverable workflow with durable state and compensating actions.
Once the order reaches the appropriate state, it is sent to the restaurant for confirmation.
Order Event
↓
Restaurant Order Service
↓
Restaurant AppThe restaurant can accept or reject the order. If accepted, it can also provide an estimated preparation time:
ACCEPT
estimated_preparation_time = 20 minThe preparation estimate later helps the Dispatch Service decide when to assign a driver and helps the ETA Service estimate delivery time.
Restaurant confirmation should require an explicit acknowledgement. If the restaurant device is offline or does not respond, the system should use defined timeout and retry rules rather than silently assuming the order was accepted.
Restaurant acceptance is an explicit state transition, not an assumption.
Drivers continuously send GPS updates, creating a high-volume workload that should be processed separately from transactional order data.
Driver GPS
↓
Location Gateway
↓
Event Stream
↓
Location Processor
|
+--> Latest Location Store
+--> Geo Index
+--> Historical StorageA location update may contain:
{
"driver_id": "D882",
"lat": 26.8467,
"lng": 80.9462,
"accuracy_meters": 7,
"heading": 120,
"timestamp": 1789630000
}Each storage path serves a different purpose:
Component | Purpose |
|---|---|
Latest Location Store | Stores the driver's most recent known location |
Geo Index | Finds drivers near a restaurant or geographic area |
Historical Storage | Keeps location history for analytics or operational needs |
GPS updates are frequent and mostly ephemeral, so every raw update should not be written to the transactional Order Database.

Dispatch needs current location, while analytics may need historical movement. These are different access patterns.
Workload | Example Question | Data Path |
|---|---|---|
Dispatch | Where is Driver D882 now? | Latest Location Store / Geo Index |
Analytics | Where did Driver D882 travel yesterday? | Historical Storage |
Dispatch should use the latest-location path rather than scanning historical GPS data.
Location alone is not enough. Dispatch also needs to know whether that location is recent and whether the driver can accept a delivery.
Driver Condition | Dispatch Behavior |
|---|---|
Fresh + Available | Include as candidate |
Fresh + Busy | Exclude |
Stale Location | Exclude or lower confidence |
Offline | Exclude |
Driver states can include:
OFFLINE
AVAILABLE
OFFERED
ASSIGNED
PICKING_UP
DELIVERINGOther eligibility rules may consider vehicle type, delivery capacity, and current workload.
Once the restaurant confirms the order, the Dispatch Service needs to find a suitable delivery partner without scanning every driver in the city.
A geospatial driver index narrows the search to nearby drivers:
Restaurant Location
↓
Spatial Cell
↓
Current + Neighboring Cells
↓
Nearby Drivers
↓
Fresh + Eligible Drivers
↓
Candidate DriversGeospatial search only creates the candidate set. The Dispatch Service can then rank those candidates using signals such as pickup ETA, restaurant preparation time, driver workload, and delivery constraints.
Geospatial search finds possible drivers; dispatch logic selects the most suitable candidate.
After candidate drivers are identified, the system sends delivery offers and safely assigns the order to one driver.
Order Ready for Dispatch
↓
Find Nearby Drivers
↓
Filter Eligible Drivers
↓
Rank Candidates
↓
Send Delivery Offer
↓
Driver Accepts
↓
Atomic Assignment
↓
Delivery AssignedThe critical correctness boundary is atomic assignment. Even if multiple drivers receive or accept an offer around the same time, only one driver should be able to claim the delivery.

Multiple drivers may receive the same delivery offer, but only one should be allowed to claim it.
Suppose:
D1 → ACCEPT
D2 → ACCEPTThe system should perform an atomic conditional update against the authoritative delivery state:
UNASSIGNED
↓
ASSIGNED(driver=D1)The update succeeds only if the delivery is still UNASSIGNED. Once D1 succeeds, D2's attempt fails.
Possible implementations include:
Conditional database updates.
Compare-and-set operations.
Row-level locking.
Another atomic conditional-write mechanism supported by the datastore.
Candidate discovery can be approximate, but final driver assignment must be authoritative and atomic.
The nearest driver is not always the best driver. Dispatch should consider how quickly and efficiently a driver can complete the pickup, not just straight-line distance.
For example:
Driver A
1 minute away
moving away from restaurant
Driver B
2 minutes away
moving toward restaurantA matching score can consider:
Pickup ETA.
Driver availability and workload.
Restaurant preparation ETA.
Traffic conditions.
Vehicle type and capacity.
Travel direction.
Acceptance probability.
Batching opportunities.
Distance is useful for candidate generation; the final assignment should consider the overall delivery outcome.
Food delivery dispatch must coordinate driver arrival with restaurant preparation time.
Suppose:
Restaurant Prep Time = 25 min
Driver Arrival Time = 4 minAssigning the driver immediately could leave them waiting for about 21 minutes, reducing driver utilization.
Ideally:
Driver Arrival Time
≈
Food Ready TimeThe Dispatch Service should therefore consider both remaining preparation time and driver pickup ETA when deciding when to send an offer.
This is an important difference from ride matching: in food delivery, the pickup itself may not be ready yet.
Delivery ETA includes more than road travel time because several stages contribute to when the order reaches the customer.
Conceptually:
Delivery ETA
=
Restaurant Confirmation Delay
+
Remaining Preparation Time
+
Driver-to-Restaurant Time
+
Pickup Delay
+
Restaurant-to-Customer Travel TimeThe ETA Service can combine changing signals:
Prep Estimate
+
Driver Location
+
Route + Traffic
+
Historical Data
↓
ETA Service
↓
Predicted Delivery TimeAs preparation, driver location, and traffic conditions change, the ETA should be recalculated throughout the delivery.
Statistical or ML models can improve prediction accuracy, but the architecture should not depend on a specific model.

After driver assignment, the customer needs live delivery progress and the driver's latest relevant location.
Driver GPS
↓
Location Gateway
↓
Location Service
↓
Tracking Service
↓
Real-Time Gateway
↓
Customer AppThe tracking screen may show states such as:
Driver approaching restaurant
↓
Order picked up
↓
Driver heading to customer
↓
Driver arrivingFor an actively open tracking screen, WebSockets or another streaming mechanism can provide low-latency updates.
Persistent connections are useful for real-time tracking, but the rest of the application can continue using normal request-response APIs where streaming is unnecessary.
A single driver location update may be useful to several independent systems.
Instead of synchronously calling every consumer from the location request, publish the update to an event stream:
Driver Location
↓
Event Stream
|
+--> Dispatch
+--> ETA
+--> Tracking
+--> AnalyticsThis keeps location ingestion fast while allowing downstream consumers to process updates independently.
Location tracking must balance freshness with battery, network, and backend cost.
Update frequency can adapt based on:
Movement speed.
Delivery state.
Distance from pickup or drop-off.
Battery level.
Network quality.
GPS readings can also be inaccurate or jump unexpectedly. The location processor can reduce bad observations using:
Accuracy metadata.
Speed sanity checks.
Map matching.
Temporal smoothing.
Poor GPS readings should not immediately trigger incorrect dispatch or ETA decisions.
GPS can indicate where a driver probably is, but it does not prove that the food was actually picked up or delivered.
Pickup should therefore use an explicit durable state transition:
ASSIGNED
↓
ARRIVED_AT_RESTAURANT
↓
PICKED_UPSimilarly, delivery can follow:
PICKED_UP
↓
ARRIVED
↓
DELIVEREDDepending on product requirements, verification may use a pickup code, OTP, photo, signature, or customer confirmation.
These transitions should be idempotent because mobile clients may retry requests after network failures.
Notifications are asynchronous side effects of business events and should not control the correctness of the order lifecycle.
Order State Changed
↓
Event Broker
↓
Notification Service
↓
Push / SMS / EmailTypical notifications include restaurant acceptance, driver assignment, pickup, driver arrival, and delivery completion.
If notification delivery fails, it can be retried independently.
A notification failure should never roll back an otherwise valid order transition.
Updating business state and publishing an event are two separate writes, which creates a failure window.
Consider:
1. Update Order → PICKED_UP
2. Publish PICKED_UP EventIf the database update succeeds but event publishing fails, downstream services may never learn about the new state.
A Transactional Outbox solves this by writing both records in the same database transaction:
Database Transaction
|
+--> Update Order → PICKED_UP
|
+--> Insert Outbox Event
↓
Outbox Publisher
↓
Event BrokerThe publisher can retry delivery independently if the broker is temporarily unavailable.
This removes the common database + message broker dual-write problem without requiring a distributed transaction.

Even with reliable event publishing, message brokers may deliver the same event more than once. Consumers should therefore be designed for at-least-once delivery.
At-Least-Once Delivery
+
Idempotent Consumers
↓
Effectively-Once Business EffectEach event can include an event_id or another stable business identifier. Consumers can use it to detect duplicates or make repeated processing harmless.
For example, receiving the same ORDER_PICKED_UP event twice should not create two notifications, duplicate analytics records, or repeat another business action.
Aim for idempotent business outcomes rather than claiming global exactly-once execution across databases, brokers, payment providers, and external devices.
Cancellation becomes more difficult as an order moves further through its lifecycle because more systems and physical actions are involved.
Order Stage | Cancellation Consideration |
|---|---|
Created | Relatively easy to cancel |
Restaurant Confirmed | Food preparation may have started |
Driver Assigned | Driver assignment must also be updated |
Picked Up | Physical delivery is already underway |
A cancellation request can trigger a coordinated workflow:
Cancellation Request
↓
Validate Order State
↓
Cancellation Workflow
|
+--> Payment Compensation
+--> Restaurant Update
+--> Driver Update
+--> NotificationRefunds should have their own durable state rather than being represented by a simple refunded = true flag.
Refund
------
refund_id
payment_id
amount
status
provider_reference
idempotency_keyA simplified refund lifecycle can be:
REQUESTED
↓
PROCESSING
↓
SUCCEEDED
FAILED
UNKNOWNIf the refund provider times out, the result may be UNKNOWN and should be resolved using the provider reference, status API, webhook, or reconciliation.
Do not mark a refund as successful until the payment provider confirms the outcome.
Exact inventory control is optional because restaurants do not always track every menu item as a fixed unit quantity.
If precise inventory is required, concurrent orders must not oversell the remaining stock.
For example, if only two portions remain, two customers should not both be able to reserve two portions.
available >= requested_quantity
↓
Atomic Reservation / DecrementThe inventory check and reservation should be atomic so concurrent checkout requests cannot consume the same stock.
In a system design interview, first clarify whether exact item-level inventory is required. If the restaurant only manages items as available/unavailable, a simpler availability model may be sufficient.
Scheduled orders must be stored durably so they survive application restarts and infrastructure failures.
Scheduled Order
↓
Durable Scheduler
↓
Calculate Start Time
↓
Order WorkflowThe workflow should not simply start at the requested delivery time. It may need to begin earlier based on restaurant preparation time and expected delivery time.
Avoid keeping an application thread or in-memory timer waiting for hours.
A delivery partner may carry multiple compatible orders when doing so improves overall delivery efficiency.
Order A ─┐
├──> Driver
Order B ─┘Batching must consider constraints such as:
Food freshness.
Pickup readiness.
Delivery SLA.
Route detour.
Driver capacity.
Customer ETA.
The dispatch problem now changes from:
Who is the nearest driver?to:
Which driver + order combination
produces the best overall outcome?For a system design interview, start with one order per dispatch and introduce batching as an advanced optimization.
Food delivery traffic is bursty, and a popular restaurant can create a hot key or hot partition during peak hours.
Restaurant R1
↓
Large Traffic Spike
↓
Potential Hot PartitionIf all workloads are partitioned only by restaurant_id, one popular restaurant can overload a single partition.
Possible mitigations include:
Cache read-heavy menu data.
Separate menu reads from order processing.
Queue asynchronous work where appropriate.
Partition orders independently when needed.
Apply per-restaurant capacity controls.
Queues can absorb short asynchronous bursts, but queue depth alone does not tell us whether customers are waiting too long. For latency-sensitive workflows such as dispatch, also monitor oldest message age or queue delay.
A healthy queue is not just one that stores work—it must process that work within the required SLA.
When a downstream service slows down, the system should prevent that slowdown from spreading into critical order workflows.
For example, a slow Notification Service should create a notification backlog rather than block order creation:
Order Service
↓
Event Broker
↓
Notification BacklogDifferent workloads have different priorities:
Workload | Priority Reason |
|---|---|
Orders / Payments | Business correctness and money movement |
Dispatch | Directly affects fulfillment time |
Active Tracking | Affects live delivery experience |
Notifications | Usually retryable asynchronous side effect |
Analytics / Promotions | Can generally tolerate more delay |
Backpressure, bounded queues, retries, and workload isolation help protect critical paths when downstream consumers become slow.
A food delivery system has workloads with very different access patterns, so using the same storage technology for everything is usually unnecessary.
Workload | Suitable Conceptual Storage |
|---|---|
Orders / Payments | Transactional relational database |
Restaurant Search | Search index |
Current Driver Location | Low-latency geo/location store |
Historical Location Events | Analytical or event storage |
Restaurant / Menu Cache | Distributed cache |
Images | Object storage + CDN |
Events | Durable broker or log |
Instead of starting with:
SQL or NoSQL?start with:
What are the access patterns, scale, latency, and consistency requirements of this workload?
The storage choice should follow those requirements.
Orders require durable state transitions and transactional correctness, making a relational database a reasonable starting point.
Typical access patterns include:
Create Order
Get Order by ID
Get Customer's Recent Orders
Update Order State
Get Restaurant's Active OrdersUseful indexes may include:
PRIMARY KEY (order_id)
INDEX (customer_id, created_at)
INDEX (restaurant_id, status)As traffic grows, the Order Database can be partitioned according to actual access patterns and regional boundaries rather than introducing sharding prematurely.
Start with a storage model that guarantees order correctness, then partition when scale and access patterns justify it.
Food delivery is naturally local.
An order in one city normally does not need an active driver from another country.
Therefore:
Global Routing Layer
|
+--> Region A
+--> Region B
+--> Region CInside a region:
City
↓
Delivery Zone
↓
Spatial CellsDispatch queries can stay geographically local, giving smaller geo indexes, lower latency, and better failure isolation.
Partition boundaries should not become hard walls.
Restaurants near a boundary may serve customers in neighboring zones, so relevant cross-zone queries must still be supported.
An order can have a natural home region based on its restaurant and service market.
order_id
↓
home_regionLatency-sensitive writes for the order, dispatch, and tracking can remain close to that region.
Global services can still handle shared concerns such as authentication, configuration, catalog replication, and analytics.
This avoids unnecessary global consensus for inherently local deliveries.
Good cache candidates include:
Restaurant details.
Menus.
Cuisine/category pages.
Popular searches.
Delivery-zone metadata.
Be more careful with:
Order state.
Payment state.
Driver assignment.
Current driver location.
These values change frequently or require stronger correctness.
If cached, they need explicit TTL, invalidation, and source-of-truth rules.
Suppose a popular restaurant's menu cache expires exactly during dinner traffic.
Thousands of Customers
↓
Same Menu
↓
Cache Miss
↓
Menu DatabaseThis can overload the database.
Common mitigations include:
Request coalescing.
Jittered TTL.
Background refresh.
Stale-while-revalidate.
Large food images should not be stored directly in normal relational rows.
Image
↓
Object Storage
↓
CDN
↓
CustomerThe restaurant or menu database stores image metadata and object references.
Object storage stores the actual image, while the CDN provides fast delivery.
Delivery fees may depend on:
Distance
Demand
Driver Supply
Restaurant
Time
Promotions
Weather
Market RulesThe Pricing Service can produce an authoritative quote.
quote_id
amount
expires_at
pricing_versionThe quote should expire because market conditions can change.
At checkout, validate that the quote is still usable.
Supply-demand calculations can also be aggregated by geographic zone rather than scanning all drivers and orders globally.
Discovery is optimized for speed, while checkout must prioritize correctness. Therefore, both paths can use different consistency guarantees.
Aspect | Discovery | Checkout |
|---|---|---|
Goal | Fast browsing | Correct order creation |
Data Source | Cache can be used | Authoritative source |
Consistency | Eventual consistency is acceptable | Stronger correctness required |
Stale Data | Slightly stale data is tolerable | Must revalidate before commit |
Examples | Restaurant status, menu, estimated ETA | Price, availability, serviceability, final amount |
A restaurant shown as open during discovery may become unavailable before checkout. The backend should therefore revalidate critical data before creating the order.
Optimize discovery for speed and checkout for correctness.
Distributed payment failures are unavoidable.
Suppose internal state says:
UNKNOWNbut the provider says:
CAPTUREDA reconciliation process compares internal and provider records.
Internal Payment Records
↕
Payment Provider Records
↓
Resolve MismatchThis acts as a safety net for uncertain money-moving operations.
Mobile connectivity is unreliable, so one missed heartbeat should not immediately trigger driver reassignment.
The system can consider:
Last location timestamp.
Heartbeat status.
Grace period.
Current delivery stage.
The response depends heavily on the delivery stage:
Situation | Possible Action |
|---|---|
Short disconnect | Wait for the grace period |
Disconnected before pickup | Reassign the delivery if needed |
Disconnected after pickup | Trigger operational intervention rather than simple reassignment |
After pickup, the driver physically possesses the food, so the system may require support, replacement, refund, or another operational workflow.
Physical-world failures cannot always be solved with automatic retries or reassignment.
Different failures require different responses.
Failure | Expected Behavior |
|---|---|
Payment timeout | Resolve provider status; do not assume failure |
Restaurant offline | Retry/timeout; require explicit acknowledgement |
Driver disconnects before pickup | Grace period, then possible reassignment |
Driver cancels before pickup | Reopen dispatch and recalculate ETA |
Driver fails after pickup | Operational exception workflow |
Notification fails | Retry notification; keep order state |
Tracking WebSocket fails | Fall back to polling/last known position |
Location is stale | Exclude driver or reduce confidence |
Search fails | Existing orders, payments, and deliveries continue |
Service crashes after payment | Recover from durable state/reconciliation |
Failure isolation prevents a discovery or notification outage from becoming a full-platform outage.
A food delivery platform handles sensitive data such as payments, customer addresses, restaurant accounts, driver identities, and live locations, so access must be tightly controlled.
Important security controls include:
Authentication and authorization.
TLS and encryption at rest where appropriate.
Least-privilege access and secret management.
Rate limiting and audit logging.
Location retention and privacy policies.
Authorization should follow the business relationship:
User | Allowed Access |
|---|---|
Customer | Access their authorized orders and active delivery tracking |
Restaurant | Manage its own menu and orders |
Driver | Access and update assigned deliveries |
Driver location requires additional privacy protection. A customer may need the driver's current location during an active delivery, but should not receive permanent access to the driver's location history.
Location data should be exposed only to authorized users, for the required purpose, and for the necessary duration.
Rate limiting protects high-traffic and abuse-sensitive APIs without blocking legitimate usage.
Endpoint | Why Limit It |
|---|---|
Restaurant Search | Prevent excessive queries and scraping |
Promo Validation | Reduce abuse and repeated validation attempts |
Order Creation / Payment | Prevent duplicate or abusive requests |
Driver Location Updates | Control excessive GPS update traffic |
Limits can be applied by user, device, driver, restaurant, IP, or API key, depending on the endpoint.
Driver location APIs need different thresholds because frequent updates are expected during active deliveries.
Rate limits should match the workload rather than using one global limit for every API.
Monitor each important workflow separately.
Area | Important Metrics |
|---|---|
Orders | Creation latency, success rate, duplicate prevention, cancellations |
Restaurants | Acceptance latency, rejection rate, preparation delay |
Dispatch | Assignment latency, candidate count, acceptance rate, unassigned age |
Location | Updates/sec, freshness, GPS accuracy, stream lag |
ETA | Pickup ETA error, delivery ETA error |
Delivery | Restaurant wait time, driver idle time, late deliveries |
Payments | Authorization/capture success, unknown states, refunds, reconciliation mismatches |
End-to-end metrics matter too:
Order Placed
↓
Food DeliveredA collection of healthy microservices is not enough if the customer's food is consistently late.
The checkout path revalidates critical business data before creating a durable order.
Customer
↓
Checkout API
↓
Validate Restaurant + Menu
↓
Validate Price + Delivery
↓
Order Service
↓
Create Durable Order
↓
Payment Workflow
↓
Restaurant Confirmation
↓
Order ConfirmedExternal side effects should use appropriate idempotency, retry, or reconciliation mechanisms.

Dispatch combines restaurant readiness, nearby-driver discovery, ranking, and safe assignment.
Restaurant Confirms Order
↓
Preparation ETA
↓
Dispatch Scheduler
↓
Geo Index
↓
Nearby Eligible Drivers
↓
Rank Candidates
↓
Delivery Offers
↓
Driver Accepts
↓
Atomic AssignmentIf no driver accepts:
Expand Search
↓
New Candidates
↓
Retry Offers
↓
Recalculate ETAIf assignment continues to fail, the workflow can escalate or cancel according to product policy.
The tracking architecture separates high-volume driver location ingestion from low-latency updates sent to customers.
Driver Phone
↓
Location Gateway
↓
Event Stream
|
+--> Geo Index
+--> Dispatch
+--> ETA
+--> Tracking
↓
Real-Time Gateway
↓
Customer AppThe Location Gateway handles frequent location writes, while the Tracking path delivers relevant updates to customers watching active orders.
HLD and LLD answer different questions.
Aspect | High-Level Design | Low-Level Design |
|---|---|---|
Focus | Distributed architecture | Objects and internal behavior |
Examples | Orders, Payments, Dispatch, Tracking, ETA | Order, Payment, Delivery, MenuItem |
Main Concerns | Scale, failures, storage, communication | Classes, states, methods, relationships |
Interview Goal | Explain how the platform works at scale | Explain implementation structure |
For HLD, the strongest deep dives are usually order correctness, driver dispatch, real-time tracking, and ETA.
Do not spend most of an HLD interview designing Java classes unless asked.
The main business relationships can be represented as:
CUSTOMER
|
| places
v
ORDER -------- ORDER_ITEM -------- MENU_ITEM
| |
| v
| RESTAURANT
|
+-------- PAYMENT
|
+-------- DELIVERY -------- DELIVERY_PARTNERImportant relationships include:
Customer 1 ─── N Order
Restaurant 1 ─── N MenuItem
Order 1 ─── N OrderItem
MenuItem 1 ─── N OrderItem
Order 1 ─── 1..N Payment Attempts
Order 1 ─── 0..N Delivery Attempts
DeliveryPartner 1 ─── N DeliveriesAllowing multiple payment or delivery attempts makes failure and retry handling more realistic.
A good design should evolve with scale instead of starting with dozens of services.
Stage | Architecture |
|---|---|
Single-City MVP | Monolith + relational DB + database spatial queries |
Read Scale | Search index + menu cache + CDN |
Dispatch Scale | Location Service + Geo Index + Dispatch Service |
Real-Time Tracking | Location stream + real-time gateway |
Reliable Workflows | Event broker + outbox + idempotent consumers |
Regional Scale | Regional order/dispatch ownership and local geo indexes |
Complexity should be introduced because a specific workload requires it.
The major trade-offs are:
Location freshness vs battery: frequent updates improve tracking but consume device resources.
Fast discovery vs correctness: cache discovery aggressively, but revalidate checkout.
Dispatch quality vs latency: evaluating more drivers may improve matching but takes more computation.
Immediate dispatch vs driver waiting: early dispatch reduces risk of delay but can waste driver capacity.
Synchronous vs asynchronous processing: keep correctness-critical operations explicit while moving independent side effects off the request path.
Regional isolation vs global coordination: local ownership reduces latency, while cross-region operations add complexity.
These questions cover the major concepts behind a DoorDash/Swiggy-like architecture.
Separate restaurant discovery, ordering, payments, restaurant fulfillment, driver location ingestion, dispatch, tracking, and ETA because each has different workload and consistency requirements.
Use a geospatial index to retrieve restaurants from the customer's current and neighboring spatial cells, then perform exact serviceability filtering and ranking.
A nearby restaurant may not be deliverable because of road topology, delivery zones, traffic, or operational restrictions.
Keep authoritative menu data in persistent storage and cache read-heavy menu responses. Revalidate price and availability during checkout.
Prices can change and clients can be manipulated. The server should retrieve current authoritative prices and recalculate the final amount.
Require an idempotency key and enforce uniqueness for the customer's logical order request.
Use an explicit state machine with validated transitions such as CREATED → PAID → PREPARING → PICKED_UP → DELIVERED.
Treat the result as uncertain until the provider status is resolved through a stable reference, webhook, status query, or reconciliation.
Order creation, payment, restaurant confirmation, and dispatch cross independent systems and cannot be safely handled by one database transaction.
Maintain a geospatial index of fresh, available driver locations and search the restaurant's current and neighboring cells.
Use an atomic conditional transition from UNASSIGNED to ASSIGNED(driver_id).
No. Candidate ranking can consider pickup ETA, preparation time, workload, traffic, direction, vehicle type, and batching opportunities.
Dispatching too early makes the driver wait. Ideally, driver arrival should approximately match food readiness.
Combine remaining preparation time, driver-to-restaurant ETA, pickup delay, restaurant-to-customer travel time, traffic, and historical signals.
Use a dedicated Location Gateway and event stream, then update a Latest Location Store and Geo Index asynchronously.
Dispatch needs fast current-position queries, while analytics needs historical trajectories. These have different access patterns.
Publish driver location updates through the tracking pipeline to a real-time gateway using WebSockets or another streaming mechanism.
Use timestamps and freshness thresholds. Exclude or lower the confidence of drivers whose location is too old.
It prevents the database and message broker from becoming inconsistent when business state changes but event publication fails.
Usually not globally. At-least-once delivery with idempotent consumers and business uniqueness constraints is a more practical model.
Validate the current order state and execute a state-dependent workflow that may update the restaurant, driver, payment, and notifications.
Use heartbeat/location freshness plus a grace period. Reassign before pickup when appropriate; after pickup, treat it as an operational exception.
Persist the scheduled time and use a durable scheduler to start preparation, payment, and dispatch workflows at the appropriate calculated time.
Treat batching as a constrained optimization problem involving pickup readiness, route overlap, detour, capacity, food freshness, and delivery SLA.
Assign orders to a natural regional market and keep active driver locations, dispatch, tracking, and latency-sensitive processing close to that region.
Common mistakes in a food delivery system design include:
Scanning drivers or choosing the nearest driver without proper geospatial indexing and matching logic.
Ignoring atomic assignment, allowing multiple drivers to accept the same delivery.
Storing every GPS update in the transactional Order Database.
Trusting client-provided prices, availability, or final totals.
Treating a payment timeout as a payment failure or retrying payments blindly.
Forgetting idempotency for orders, payments, and event consumers.
Ignoring restaurant preparation time, driver-location freshness, or pickup ETA during dispatch.
Making every workflow synchronous or ignoring database + broker dual-write failures.
Choosing technologies such as Kafka, Redis, or NoSQL before explaining the workload and problem they solve.
The final architecture separates read-heavy discovery, transactional ordering, payment processing, driver dispatch, and real-time location tracking so each workload can scale according to its own requirements.
CUSTOMER
|
v
API GATEWAY
|
+-------------------+-------------------+
| | |
v v v
DISCOVERY ORDER SERVICE TRACKING
| | |
v v v
SEARCH INDEX ORDER DB REAL-TIME GATEWAY
| |
v v
RESTAURANT / MENU ORDER EVENTS
CACHE |
v
EVENT BROKER
|
+------------------+------------------+
| | |
v v v
PAYMENT RESTAURANT DISPATCH
|
v
GEO INDEX
|
v
DELIVERY PARTNER
DRIVER LOCATION PIPELINE
DELIVERY PARTNER
|
v
LOCATION GATEWAY
|
v
EVENT STREAM
|
+------------+------------+
| | |
v v v
GEO INDEX ETA TRACKING
|
v
REAL-TIME GATEWAY
|
v
CUSTOMERThe architecture separates workloads with very different characteristics:
Workload | Primary Requirement |
|---|---|
Discovery | Fast, scalable reads |
Orders | Transactional correctness and durability |
Payments | Correctness, idempotency, and reconciliation |
Restaurant Fulfillment | Reliable state coordination |
Driver Location | High write throughput and freshness |
Dispatch | Fast geospatial matching and atomic assignment |
Tracking | Low-latency real-time updates |
Notifications / Analytics | Asynchronous processing |
The central principle is:
A food delivery platform is not simply an ordering system. It is a real-time marketplace coordinating customers, restaurants, and moving delivery partners while payment, preparation time, location, and ETA change independently.
Once these workloads are separated, the food delivery system architecture becomes easier to scale, reason about, and explain in a system design interview.