Queues, Streams and Delivery Semantics
Queues and logs decouple producers from consumers, but every guarantee has fine print. Reason about ordering, redelivery and crashes at every step, and assume at-least-once delivery.
Key points
- 1
Kafka orders records only within a partition. Same key → same partition (hash mod partition count), and a partition goes to one consumer per group, so partitions cap a group's parallelism.
- 2
Commit-then-process is at-most-once; process-then-commit is at-least-once. Exactly-once effects come from idempotent processing or committing the result and the offset atomically.
- 3
Kafka exactly-once (idempotent producer + transactions + read_committed) covers read-process-write inside Kafka only. Emails, HTTP calls and external database writes need their own deduplication.
- 4
The dual-write problem (database + broker) is solved by a transactional outbox relayed via CDC or polling, or by making the log the source of truth. Kafka does not support XA, so two-phase commit with it is not an option.
- 5
SQS redelivers a message when its visibility timeout expires before it is deleted. Size the timeout above the worst-case job, or extend it with a heartbeat. maxReceiveCount then sends repeat failures to the dead-letter queue.
- 6
A Kafka consumer is evicted when the time between poll() calls exceeds max.poll.interval.ms. Heartbeats run on a background thread, so a slow batch shows up as rebalances and reprocessing, not session timeouts.
Common traps
Adding partitions to a keyed topic remaps keys and breaks per-key ordering for events in flight.
A dead-letter queue unblocks a partition but breaks per-key ordering; park the whole key when order matters.
SQS FIFO deduplication only covers duplicate sends within 5 minutes; it does not stop redelivery after a visibility timeout.
Read the source
Test yourself on Queues, Streams and Delivery Semantics
Ten questions, with the answer and explanation after each one.