System Design Mock sample questions with answers

10 questions from the System Design Mock practice bank, spread across its domains. Pick your answer, then open the explanation to see why each option is right or wrong.

  1. Question 1Data and Storage

    A single PostgreSQL primary serves users, products and orders for a marketplace, and it is saturated on writes. The three areas are rarely joined in the same query. Before sharding any table, which approach splits the load by function?

    • A

      Federation: give users, products and orders their own separate databases

    • B

      Vertical scaling: move the primary to a much larger machine

    • C

      Read replicas: add four replicas and send all reads to them while writes stay put

    • D

      Sharding: split every table by user id across eight identical clusters of databases

    Show the answer and explanation

    Answer: A

    Federation (functional partitioning) splits databases by function, so each handles less read and write traffic, gets better cache locality and can be scaled independently. The cost is that joins across functions move into the application and cross-database transactions become harder. It is often a first step before sharding the hottest function.

    Why the other options are wrong

    • B. Bigger hardware delays the problem but does not split the write load, and it has a ceiling.

    • C. Replicas add read capacity only; the bottleneck here is writes on the single primary.

    • D. Sharding splits rows within a function and adds cross-shard complexity; federation splits by function first.

  2. Question 2Distributed Systems and Reliability

    An API allows each client 1,000 requests per minute. Ten gateway nodes sit behind a layer-4 load balancer that spreads connections round-robin, so each node enforces 100 requests per minute per client on its own. One client switches to a single long-lived HTTP/2 connection, so all of its requests reach one node. What happens?

    • A

      The client is still limited to about 1,000 per minute, because the load balancer keeps spreading the individual requests on its connection across nodes

    • B

      The client is capped near 100 per minute, a tenth of its quota; limits need a shared counter or one owning node per client

    • C

      The client can now send about 10,000 per minute, because each of the ten gateway nodes grants it its own 1,000-request allowance

    • D

      Nothing changes for the client, because rate limits apply per IP address and its address has not changed

    Show the answer and explanation

    Answer: B

    Fixed per-node shares (limit / N) under-admit when traffic concentrates on fewer nodes, and over-admit if the node count changes without updating the shares. Common fixes are a central counter (for example Redis with an atomic increment or a Lua token bucket), consistent-hash routing so one node owns each client's limiter, or local buckets that sync usage periodically and accept brief overshoot.

    Why the other options are wrong

    • A. A layer-4 balancer pins a connection to one node, so every request on it hits the same local limiter.

    • C. Each node enforces 100, not 1,000, so the client can't exceed its quota that way here.

    • D. The limit here is per client per node, so where its requests land changes what the client experiences.

  3. Question 3Caching and Messaging

    A shop caches category listing pages under keys like cat:7:page:3:sort:price. A category has hundreds of such variants, and when any product in it changes, all of them must stop being served. What is the most practical invalidation approach?

    • A

      Run KEYS cat:7:* on each change and then delete every key the command returns

    • B

      Keep a generation number per category in the keys and increment it on each change

    • C

      Call FLUSHALL whenever any product changes, so no stale page can survive

    • D

      Give the listing pages a very short TTL and rely on expiry to hide changes

    Show the answer and explanation

    Answer: B

    Versioned (generation) keys turn "delete hundreds of keys" into one atomic INCR cat:7:gen. Readers fetch the generation and build keys such as cat:7:g42:page:3:sort:price. Orphaned entries are reclaimed by TTL or eviction. The cost is one extra lookup per read, which is often cached locally for a second or two.

    Why the other options are wrong

    • A. KEYS scans the whole keyspace and blocks Redis while it runs; it is unsafe in production.

    • C. Flushing the whole cache for one change causes a miss storm across every unrelated key.

    • D. That bounds staleness but costs far more misses, and changes still show up late.

  4. Question 4Foundations and Estimation

    An API keeps instances at a 60% CPU target, and one instance handles 1,000 requests per second at 100%. It runs 17 instances for a steady 10,000 requests per second. A marketing push will raise traffic to 25,000 requests per second within 2 minutes at 18:00, and a new instance takes 5 minutes to become ready. What should the team do?

    • A

      Nothing: target-tracking autoscaling will add instances as CPU rises and can keep up with the spike

    • B

      Lower the CPU target to 40% at 18:00 so that autoscaling responds more aggressively to the spike

    • C

      Scale ahead: have about 42 instances ready before 18:00 (25,000 ÷ 600), then let autoscaling take over

    • D

      Double the fleet to 34 instances at 18:00, which gives twice today's capacity for the coming spike

    Show the answer and explanation

    Answer: C

    Reactive autoscaling can only add capacity as fast as instances start. 17 instances max out at 17,000 req/s, so a 2-minute ramp to 25,000 overloads them for several minutes. For known events, scale on a schedule: 25,000 ÷ (1,000 × 0.6) ≈ 41.7, so about 42 instances, ready before the ramp. Predictive scaling and warm pools serve the same purpose for recurring patterns.

    Why the other options are wrong

    • A. With a 5-minute start-up, capacity lags a 2-minute ramp, and 17 instances top out at 17,000 req/s.

    • B. A more sensitive target still waits 5 minutes for every new instance to become ready.

    • D. 34 × 600 = 20,400 req/s at target, short of 25,000, and they would not be ready until 18:05.

  5. Question 5APIs and End-to-End Designs

    A team adds a user_id label to an HTTP request counter in a Prometheus-style monitoring system. The service has 20 endpoints, 5 status codes and 1 million active users. How many time series can this single metric now create, and what is the problem?

    • A

      About 1 million series, one for each user, which a modern time-series database handles comfortably

    • B

      Up to 100 million series, which blows up memory and query cost; per-user detail belongs in logs or traces

    • C

      Exactly 100 series, because label values are stored on samples and don't create any new series at all

    • D

      Up to 100 million series, which is fine because each sample compresses down to only a couple of bytes

    Show the answer and explanation

    Answer: B

    Cardinality is the product of each label's distinct values. Unbounded labels (user IDs, request IDs, full URLs) are the most common way to take down a metrics system. Keep labels low-cardinality and put high-cardinality context in logs, traces or exemplars.

    Why the other options are wrong

    • A. Every label multiplies the others: each endpoint, status and user combination is its own series.

    • C. In Prometheus-style systems every distinct label set is a separate series.

    • D. Compression shrinks samples, but the cost is per series: index entries, head memory and query fan-out.

  6. Question 6Data and Storage

    A PostgreSQL primary has synchronous_standby_names = 'standby1' and synchronous_commit = on. It has no other standbys. The standby1 server crashes. What happens to new write transactions on the primary?

    • A

      They commit normally, because PostgreSQL automatically falls back to asynchronous mode after a short timeout

    • B

      They fail immediately with an error, so that applications know durability on a second server is unavailable

    • C

      Their commits hang, waiting for a standby to confirm, until one reconnects or an operator changes the setting

    • D

      They are committed on the primary and queued, then replayed to standby1 and acknowledged when it comes back

    Show the answer and explanation

    Answer: C

    With a single named synchronous standby, the standby is a single point of failure for writes: commits wait for it, without limit. Production setups therefore list several candidates, such as synchronous_standby_names = 'ANY 1 (s1, s2)', so any one of them can confirm, or use a failover manager that rewrites the setting. MySQL semi-sync takes the other side of the trade-off and silently degrades to asynchronous replication after a timeout, risking data loss to stay available.

    Why the other options are wrong

    • A. That is MySQL semi-sync behaviour (rpl_semi_sync_master_timeout). PostgreSQL has no automatic fallback.

    • B. PostgreSQL doesn't reject them. The commit is written locally and then waits for confirmation.

    • D. Clients don't get an acknowledgement until a synchronous standby confirms. Nothing is acknowledged in the meantime.

  7. Question 7Distributed Systems and Reliability

    A lease library computes expiresAt = wallClockNow() + 10s and later checks wallClockNow() > expiresAt to decide whether it still holds the lease. Why should it use a monotonic clock for this instead?

    • A

      Monotonic clocks are synchronised across machines, so other nodes can verify the lease

    • B

      Monotonic clocks have nanosecond resolution, and a 10-second lease needs that precision

    • C

      The wall clock can be stepped by NTP or an operator, so the measured interval can be wrong

    • D

      Reading the wall clock is a slow system call, and lease checks run on every request in the hot path

    Show the answer and explanation

    Answer: C

    Measuring a duration needs a clock that never jumps. Wall-clock time can move backwards or leap forward when it is corrected, so elapsed-time checks on it are unsafe. A monotonic clock is right for measuring intervals on one machine, but it still cannot compare times between machines.

    Why the other options are wrong

    • A. Monotonic clocks are local to one process or machine. Their absolute values mean nothing on another node.

    • B. Resolution is not the problem. A 10-second lease works fine with millisecond resolution.

    • D. Reading either clock is cheap. The concern is correctness of elapsed time, not speed.

  8. Question 8Caching and Messaging

    Every anonymous visitor to a news site's home page triggers the same rendering work on the app servers, even though the page changes about once a minute. What is the cheapest effective fix at the web tier?

    • A

      Let the reverse proxy cache the rendered page for about a minute

    • B

      Add app servers behind the load balancer until rendering keeps up at peak

    • C

      Cache each SQL query result in Redis and keep rendering the page per request

    • D

      Render on the client and return raw JSON

    Show the answer and explanation

    Answer: A

    Web server (reverse proxy) caching: proxies such as NGINX or Varnish can cache whole responses and serve repeated requests without touching the app servers. For pages that are identical for all anonymous users, a short TTL, or microcaching for a few seconds, removes most of the load. Choose the cache key carefully and bypass the cache for logged-in or personalised content.

    Why the other options are wrong

    • B. That scales the wasted work instead of removing it.

    • C. This helps the database, but every visitor still pays the full rendering cost.

    • D. A big rewrite that shifts work to devices; the identical page is still produced over and over.

  9. Question 9Foundations and Estimation

    A news site averages 2,000 requests per second. A breaking story can drive traffic to 10× within two minutes, but adding capacity through autoscaling takes about five minutes. A CDN serves 95% of requests from cache when the content is cacheable. Which statements are correct? (Choose two.)

    Choose 2.

    • A

      The spike arrives before scaling finishes, so the origin needs pre-provisioned headroom or load shedding

    • B

      With a 95% CDN hit ratio, origin load at the 20,000 requests-per-second peak is about 1,000

    • C

      Autoscaling on 15-minute average CPU will react in time to protect the origin

    • D

      Provisioning for 2× the average is enough, because clients retry failed requests

    • E

      The origin still sees all 20,000 requests per second, because CDNs only cache static images

    Show the answer and explanation

    Answer: A and B

    Size for the shape of the spike, not just its height: if the ramp is faster than scaling, you need standing headroom, edge caching or load shedding. Keeping a high CDN hit ratio, through short TTLs with stale-while-revalidate and request coalescing, turns a 20,000 rps spike into about 1,000 rps at the origin.

    Why the other options are wrong

    • C. A 15-minute average barely moves in two minutes, so scaling would start far too late.

    • D. Retries add load during the spike. 2× covers 4,000 of the 20,000 requests per second.

    • E. CDNs cache any cacheable response, including HTML pages with suitable headers.

  10. Question 10APIs and End-to-End Designs

    Field reporters upload videos of 2–5 GB from phones on unreliable cellular networks. Today one interrupted upload means starting again from zero. Which three choices together make uploads reliable and efficient? (Choose three.)

    Choose 3.

    • A

      Use a multipart (chunked) upload, with the API issuing presigned URLs for each part

    • B

      Have the client record completed parts and their ETags, re-send only failed parts, then call complete

    • C

      Add a storage lifecycle rule that aborts incomplete multipart uploads after a few days

    • D

      Keep one PUT but raise every timeout to two hours

    • E

      Base64-encode the video into a JSON body sent to the API servers

    • F

      Stream the file over a WebSocket to an app server that writes it to disk

    Show the answer and explanation

    Answer: A and B and C

    Resumable uploads split the file into parts that can be retried independently and uploaded in parallel, ideally straight to object storage via presigned part URLs (S3 multipart upload, GCS resumable uploads, or the tus protocol). The client tracks progress and resumes after a crash or network drop, and a lifecycle policy cleans up uploads that never finish.

    Why the other options are wrong

    • D. Long timeouts don't stop cellular drops; any drop still restarts all 5 GB.

    • E. That adds about 33% overhead, pushes gigabytes through the API tier and still isn't resumable.

    • F. One connection drop still loses progress, and the app tier now relays all the bulk data.

Practise all 487 SYS-DES questions

Start with the free 15-question diagnostic. It shows where to focus, and your results carry over if you sign up.

Go to SYS-DES