🔢
HLD Fundamentals

Back of the Envelope Estimation

DAU Se QPS Aur Storage Tak
💡 Scale estimation SHAADI KA KHAANA plan karne jaisa hai. 100 mehmaan hain ya 10,000 — isi ek number se decide hota hai ki ek cook chahiye ya poora catering setup. Number galat to poora plan galat.

Estimation ka ek hi flow hai: DAU se requests per day, usse average QPS (÷ 86400), aur peak QPS (usually 2-3x average). Storage ke liye: per-record size × records per day × retention days. Bandwidth ke liye: response size × QPS.

Numbers ROUND rakho — 86400 second ko 100,000 maan lena bilkul acceptable hai aur interviewer isi ki ummeed karta hai. Maqsad precision nahi, ORDER OF MAGNITUDE hai. 1000 QPS aur 100,000 QPS ke architecture bilkul alag hote hain, aur bas yahi farak matter karta hai.

// Example: 100M DAU, har user 10 requests/day
Requests/day = 100M × 10        = 1B
Average QPS  = 1B / 100,000     ≈ 10,000 QPS
Peak QPS     = 10,000 × 3       ≈ 30,000 QPS

// Storage: har record 1 KB, 5 saal retention
Storage/day  = 1B × 1 KB        = 1 TB/day
5 saal       = 1 TB × 365 × 5   ≈ 1.8 PB

// Ye numbers batate hain: single DB kaafi NAHI hai -> sharding chahiye
🔢
Scale estimation SHAADI KA KHAANA plan karne jaisa hai. 100 mehmaan hain ya 10,000 — isi ek number se decide hota hai ki ek cook chahiye ya poora catering setup. Number galat to poora plan galat.
1 / 2
⚡ Quick Recap
  • DAU → requests/day → average QPS → peak QPS (2-3x)
  • Round numbers use karo — order of magnitude matter karta hai, precision nahi
  • Har estimate ke baad uska architectural conclusion bolo
Is page mein (2 subtopics)

Kuch numbers ratne se estimation bahut tez ho jaati hai. 1 din ≈ 100,000 second (86,400 ko round karo). 1 million writes/day ≈ 12 writes/sec. Memory read ~100 nanosecond, SSD read ~100 microsecond, network round trip within datacenter ~500 microsecond, cross-continent ~150 millisecond.

Storage ke liye: 1 character = 1 byte, ek chhota JSON record ~1 KB, ek photo ~200 KB, ek minute HD video ~50 MB. In numbers se koi bhi estimate 30 second mein ban jaata hai.

// Latency ka order of magnitude — design decisions yahin se aate hain
Memory read           ~100 ns
SSD random read       ~100 μs      (1000x memory se slow)
Datacenter round trip ~500 μs
Cross-continent RTT   ~150 ms      (cache/CDN kyun chahiye, ye batata hai)

Ye EK sawaal poora architecture badal deta hai. Read-heavy (news feed, e-commerce catalog, 100:1 read:write) ka jawab cache, read replicas aur CDN hai. Write-heavy (logging, metrics, IoT, location updates) ka jawab queue, batching aur append-optimized storage (LSM tree) hai.

Estimation ke turant baad ye ratio bolo — "ye 100:1 read-heavy system hai, isliye main caching par focus karunga." Isse tumhara baaki design apne aap justified lagta hai.

💡Tip: Ratio bolna mat bhoolo. Bina iske tumhaare cache aur replica arbitrary lagenge; ratio ke saath wo evidence-based decision ban jaate hain.