What makes retry and dead-letter strategies dependable?
Dependability comes from explicit record authority, safe delivery semantics, bounded recovery, and reconciliation—not from the number of endpoints. For bounded retry, poison-message isolation, and operator recovery, the design must explain what happens after duplicates, delay, partial failure, and an ambiguous timeout. Retries need a stopping rule and an operator destination.
Should this use a request, webhook, queue, or batch?
Use a request when the caller needs an immediate decision, a webhook when a source announces change, a queue when work needs isolation and retry, and a batch or reconciliation job when completeness matters more than immediacy. Many durable integrations use more than one pattern.
What should be tested beyond the happy path?
Test invalid and missing data, stale versions, duplicate events, reordering, throttling, permission changes, timeout after remote commit, and replay. The route risk—retrying permanent failures forever or dropping them silently—needs a concrete test rather than a sentence in a brief. Infinite retries hide permanent defects, while immediate discard silently loses commerce work.
What evidence belongs at handoff?
Provide payload examples, mapping rules, state diagrams, failure categories, dashboards, alert ownership, replay instructions, and a reconciliation report. Seed each failure class and verify retry, stop, alert, dead-letter, and replay behavior.