Notification System Lab

One event, queue-per-channel fan-out, at-least-once delivery, idempotency, and dead-lettering — break it the way production breaks, then fix it.

RabbitMQ UI

System map

live · nodes pulse when they log an event
producerPOST /notifynotify.push—notify.email—push-worker—email-worker—providersimulated sendRedis—push.dlq—
active now dedupe path off messages parked worker down
1

Fan-out: one event, two channels

Not started

The producer publishes one event, but writes one durable message per channel queue — notify.push and notify.email. Each channel has its own worker, so a slow or dead channel never blocks the others. The producer never talks to the workers; the broker decouples them.

Setup

No flag changes needed — works with any setting.

Run

Fire one normal notification for user 7.

What to look for

SENT push notif:<id> -> user:7 from push-worker and SENT email notif:<id> -> user:7 from email-worker — the same notification id on both.

Equivalent CLI
curl -s -X POST localhost:3100/api/notify -H 'Content-Type: application/json' -d '{"mode":"default"}'
docker compose logs push-worker email-worker | grep SENT
2

The fail moment: crash between send and ack

Not started

RabbitMQ only forgets a message once the consumer acks it — that's at-least-once delivery. The push worker has a crash hook: it sends the notification, then dies before acking. Docker restarts it (restart: unless-stopped) and the broker redelivers the un-acked message to the fresh worker.

Setup

IDEMPOTENCY_ENABLEDneeds off · now ? ✗
CLAIM_LEASE_ENABLEDneeds off · now ? ✗

Run

Fire a notification that makes push-worker crash after sending, before acking.

Current flags differ from this step's setup — results will differ.

What to look for

SENT push notif:<id>, then CRASH — worker dying after send, BEFORE ack, then after the restart the same notif id is SENT again.

Equivalent CLI
curl -s -X PUT localhost:3100/api/lab/flags -H 'Content-Type: application/json' -d '{"idempotency":false,"claimLease":false}'
curl -s -X POST localhost:3100/api/notify -H 'Content-Type: application/json' -d '{"mode":"crash"}'
docker compose logs push-worker | grep -E 'SENT|CRASH'
3

The fix: an idempotency record

Not started

Before sending, the worker claims a Redis record notification:<id>:<channel>:<user> with SET NX EX 86400 (only succeeds if the key doesn't exist, expires in 24h). The crash and redelivery still happen, but the redelivery finds the record already claimed and skips the send.

Setup

IDEMPOTENCY_ENABLEDneeds on · now ? ✗
CLAIM_LEASE_ENABLEDneeds off · now ? ✗

Run

Turn on IDEMPOTENCY_ENABLED, then trigger the same crash again.

Current flags differ from this step's setup — results will differ.

What to look for

SENT push notif:<id>, CRASH — worker dying after send, BEFORE ack, then DUPLICATE DROPPED instead of a second SENT. Watch the key appear in the Redis panel.

Equivalent CLI
curl -s -X PUT localhost:3100/api/lab/flags -H 'Content-Type: application/json' -d '{"idempotency":true,"claimLease":false}'
curl -s -X POST localhost:3100/api/notify -H 'Content-Type: application/json' -d '{"mode":"crash"}'
docker compose logs push-worker | grep -E 'SENT|CRASH|DUPLICATE'
4

A gap in the fix: crash before the send

Not started

The claim is written before the provider is ever called, and it's one permanent fact — it can't tell "I'm working on this" apart from "I finished this." If the worker dies in that exact gap, the claim is already sitting in Redis when the message is redelivered.

Setup

IDEMPOTENCY_ENABLEDneeds on · now ? ✗
CLAIM_LEASE_ENABLEDneeds off · now ? ✗

Run

Keep idempotency on (claim-lease off) and crash the worker right after it claims, before it sends.

Current flags differ from this step's setup — results will differ.

What to look for

CRASH — worker dying AFTER claim, BEFORE send, then DUPLICATE DROPPED — and no SENT line anywhere for this notif id.

Equivalent CLI
curl -s -X PUT localhost:3100/api/lab/flags -H 'Content-Type: application/json' -d '{"idempotency":true,"claimLease":false}'
curl -s -X POST localhost:3100/api/notify -H 'Content-Type: application/json' -d '{"mode":"crash-claim"}'
docker compose logs push-worker | tail -6
5

Closing the gap: a claim that knows the difference

Not started

Split the claim into two states: a short-lived pending lease (10s TTL) written before the send, upgraded to a long-lived done marker (24h) only after the send succeeds. On redelivery, done means a real duplicate (drop); pending means the previous attempt never finished (retry).

Setup

IDEMPOTENCY_ENABLEDneeds on · now ? ✗
CLAIM_LEASE_ENABLEDneeds on · now ? ✗

Run

Turn on CLAIM_LEASE_ENABLED and trigger the same crash-after-claim.

Current flags differ from this step's setup — results will differ.

What to look for

CRASH after claim, then CLAIM RECOVERED on redelivery, then SENT push notif:<id> — exactly once. The Redis key goes pending → done.

Equivalent CLI
curl -s -X PUT localhost:3100/api/lab/flags -H 'Content-Type: application/json' -d '{"idempotency":true,"claimLease":true}'
curl -s -X POST localhost:3100/api/notify -H 'Content-Type: application/json' -d '{"mode":"crash-claim"}'
docker compose logs push-worker | tail -6
6

Retries, backoff, and the dead-letter queue

Not started

User 999's device token is permanently invalid — the provider rejects it every time. The worker retries with exponential backoff (0.5s, then 1s), and after the 3rd failed attempt publishes the message to push.dlq and acks the original. A poison message gets parked for inspection, it doesn't block the queue or retry forever.

Setup

No flag changes needed — works with any setting.

Run

Send a notification to user 999.

What to look for

RETRY 1/3 (0.5s backoff), RETRY 2/3 (1s backoff), then DEAD-LETTER notif:<id> -> push.dlq. The push.dlq count goes up — inspect it in the DLQ panel.

Equivalent CLI
curl -s -X POST localhost:3100/api/notify -H 'Content-Type: application/json' -d '{"mode":"bad-token"}'
docker compose logs push-worker | grep -E 'RETRY|DEAD-LETTER'
docker exec rabbitmq rabbitmqctl list_queues name messages