Back to Blog

System Design Interview: Framework, Building Blocks, Example

Published December 28, 2025
Updated October 4, 2026Technical Tips6 min read

By

Last updated October 4, 2026

System Design Interview: Framework, Building Blocks, Example

A system design interview asks you to design something like a URL shortener, a news feed, or a chat service in 45 minutes, on a whiteboard or shared document. There is no single right answer. Interviewers are looking at whether you ask good questions, break a vague problem into parts, choose components for reasons, and reason honestly about trade-offs. This guide gives you a framework, the building blocks with their trade-offs, and a worked example.

It replaces our separate short pages on caching, load balancing, sharding, message queues, SQL versus NoSQL, microservices and API design, because the interview tests how they fit together.

A framework for the 45 minutes

  1. Clarify requirements (5 min). Ask what the system must do (functional) and how well (non-functional: scale, latency, availability, consistency). Write down what is out of scope.
  2. Estimate scale (3-5 min). Rough numbers: users, requests per second, storage per year. These decide whether you need sharding or a single database.
  3. Define the API and data model (5 min). The two or three core endpoints and the main entities.
  4. Draw the high-level design (10 min). Clients, load balancer, services, database, cache, queue. Walk through one request end to end.
  5. Deep dive (15 min). Pick the hardest part (often the data layer or the hot path) and go deep. The interviewer will often steer you.
  6. Trade-offs and failure (5 min). What breaks, what you would monitor, what you would change at 10x scale.

The building blocks, and when to use each

Load balancer

Spreads requests across several identical servers so no one server is the bottleneck, and removes failed servers from rotation. Round-robin is the simple default; least-connections helps when requests vary in cost. Make the app servers stateless so any of them can handle any request.

Cache

Keep hot data in fast memory to cut latency and database load. The common pattern is cache-aside: read the cache, on a miss read the database and fill the cache. The hard parts are invalidation (stale data), eviction (LRU is the usual choice), and stampedes when a popular key expires. Say what you would cache, for how long, and what happens when it is stale.

Database: SQL or NoSQL

Do not pick by fashion. Use a relational database when you need transactions, joins, and a well-defined schema, which covers most products at moderate scale. Consider a NoSQL store when access is simple key-based lookups at very high volume, the schema varies, or you need to scale writes horizontally. Be ready to say what you give up, for example multi-row transactions or ad-hoc queries.

Replication and sharding

Replication copies data to other nodes for availability and read scaling; the trade-off is replication lag, so a read from a replica can be slightly stale. Sharding splits data across nodes to scale writes and storage. Choose a shard key with even distribution and one that matches your main access pattern; a bad key creates hot shards. Re-sharding later is painful, which is why you estimate scale up front.

Message queue

A queue decouples producers from consumers and absorbs bursts. Use it for work that does not need to finish before you respond: sending email, resizing images, updating search indexes. Discuss delivery guarantees (at-least-once means consumers must be idempotent) and what happens to messages that keep failing (a dead-letter queue).

CDN

Serves static assets and cacheable responses from locations near users. Mention it any time the design serves images, video, or large static files.

Rate limiting

Protects services from abuse and overload. A token bucket allows short bursts while enforcing an average rate; a fixed window counter is simpler but lets bursts through at window boundaries. Store counters in a fast shared store so several servers agree.

Monolith or microservices

A monolith is simpler to build, test and deploy, and is the right default for a small team. Microservices help when teams need to deploy independently or parts of the system scale very differently, at the cost of network calls, distributed failures and operational overhead. In an interview, start simple and split a service only when you can name the reason.

Consistency and availability

When a network partition happens you must choose between rejecting requests (consistency) and serving possibly stale data (availability). Say which one the product needs: a bank balance favors consistency; a social feed can tolerate staleness.

Worked example: a URL shortener

Requirements. Create a short link for a long URL; redirect from short to long; links should not collide; reads far outnumber writes. Out of scope: analytics, user accounts.

Estimates. Suppose 100 million new links a month and a 100:1 read-to-write ratio. That is roughly 40 writes per second and about 4,000 reads per second on average, with peaks higher. Storage for a few hundred bytes per link is tens of gigabytes a month. These are assumptions; say so, and let the interviewer adjust them.

API. POST /links with the long URL returns a short code; GET /{code} returns a 301 or 302 redirect. Use 302 if you want to count clicks or change targets later; 301 lets browsers cache the redirect.

Short code. A 7-character base62 code gives 62^7, about 3.5 trillion combinations. Options: hash the URL and take a prefix (needs collision handling), or generate a unique counter ID and base62-encode it (no collisions, but you need a distributed ID generator or ranges handed to each server).

Data model. A table keyed by code with the long URL and creation time. It is a pure key-value lookup, so either a relational database with an index on code or a key-value store works.

Scaling. Put a cache in front of the lookup, since popular links dominate traffic. Reads can go to replicas. If a single database stops being enough for writes, shard by code.

Trade-offs to mention. Custom aliases need uniqueness checks; expired links need a cleanup job; abuse (spam links) needs rate limiting and URL scanning.

Common mistakes

  • Jumping to components before requirements. You end up designing the wrong thing in great detail.
  • Name-dropping technology. "Kafka" is not an answer; "a queue so the upload request does not wait for thumbnail generation" is.
  • Ignoring numbers. A back-of-envelope estimate shows whether the design needs sharding at all.
  • Never mentioning failure. Always say what happens when a node, cache, or database goes down.
  • Over-engineering. Start with the simplest design that meets the requirements, then add parts only when you can say why.

How to practice

Pick one classic design per session (URL shortener, rate limiter, news feed, chat, file storage). Time yourself for 40 minutes, speak your thinking aloud, and draw on paper. Then compare your design to a reference and write down one missed trade-off. Doing it out loud with feedback is the closest thing to the real interview; our system design mock interview guide shows how to run these with an AI interviewer, and this page works through the URL shortener in more depth.

Frequently asked questions

How do I start a system design interview?

Clarify the requirements and constraints first. Ask who the users are, what the main use cases are, what scale to expect, and which qualities matter most (latency, availability, consistency).

How much should I memorize?

Understand the building blocks and their trade-offs rather than memorizing designs. If you can explain why you would pick a cache, queue or shard here, you can handle unfamiliar prompts.

Do I need to draw diagrams?

Yes, simple ones. Boxes and arrows for clients, services, and data stores make your reasoning easy to follow and easy for the interviewer to critique.

Is system design asked for junior roles?

Less often. For junior roles you may get a smaller object or API design question. System design carries more weight at mid-level and senior levels.

Share:
#TechnicalTips#InterviewPrep#CareerGrowth