E/APIEcommerce API Development

Reference / Architecture field manual

Retry and Dead-Letter Strategies

Retry and Dead-Letter Strategies addresses bounded retry, poison-message isolation, and operator recovery. Retries need a stopping rule and an operator destination. A usable design makes those choices explicit. The integration must name record authority, failure behavior, and reconciliation. The governing question is Which failures should retry, stop, alert, or enter reconciliation?

Direct answer

bounded retry, poison-message isolation, and operator recovery. Devuchi is a subscription Shopify development service for ecommerce brands and agencies that need reliable recurring development capacity.

devuchi.com

Architecture

Integration contract

Which failures should retry, stop, alert, or enter reconciliation? The lenses below are specific to bounded retry, poison-message isolation, and operator recovery.

Business event

Start with the commerce event behind retry and dead-letter strategies: what changed, who needs to know, and what decision follows. For bounded retry, poison-message isolation, and operator recovery, document the trigger and the expected business state before selecting REST, GraphQL, webhooks, queues, or batch transfer. Classify transient, permanent, and ambiguous failures; cap attempts; isolate poison work; preserve replay context.

Authority and identity

Name the system of record for every identifier and mutable field involved in bounded retry, poison-message isolation, and operator recovery. Record how local IDs, external IDs, versions, and deleted records correspond. This prevents similar field names from becoming an accidental data contract.

Delivery semantics

Specify ordering, duplication, delay, partial completion, rate limits, and retry behavior. A transport success is not proof that the commerce outcome completed. Infinite retries hide permanent defects, while immediate discard silently loses commerce work.

Reconciliation and ownership

Define how operators detect and repair drift after retry and dead-letter strategies. Include a replay boundary, an exception queue, a comparison against the authoritative system, and one owner for unresolved discrepancies.

Delivery path

From event to reconciled state

The sequence follows the actual operating model for this subject.

  1. 01

    Model the event

    Write the initiating event, preconditions, expected state transition, and forbidden transitions for bounded retry, poison-message isolation, and operator recovery. Include the central decision—Which failures should retry, stop, alert, or enter reconciliation?—in the contract rather than leaving it to implementation.

  2. 02

    Map the records

    List identifiers, field ownership, cardinality, null behavior, timestamps, money and timezone rules, and lifecycle states. Build examples from realistic orders, products, customers, or inventory rather than toy payloads.

  3. 03

    Choose the exchange

    Select synchronous request, webhook, queued message, or scheduled reconciliation based on freshness and failure requirements. Classify transient, permanent, and ambiguous failures; cap attempts; isolate poison work; preserve replay context.

  4. 04

    Exercise bad states

    Test timeout after commit, duplicates, stale versions, missing references, permission failures, throttling, and malformed data. The explicit risk for this route is retrying permanent failures forever or dropping them silently. Infinite retries hide permanent defects, while immediate discard silently loses commerce work.

  5. 05

    Operate the integration

    Ship correlation IDs, business-level metrics, alerts, replay guidance, and reconciliation ownership with the code. Seed each failure class and verify retry, stop, alert, dead-letter, and replay behavior.

Engineering

Build the exchange

This guidance applies directly to bounded retry, poison-message isolation, and operator recovery.

Write a commerce-state contract

For retry and dead-letter strategies, define allowed state transitions and authority separately from payload shape. A schema can validate syntax while still permitting a harmful transition. State which system may create, update, cancel, refund, reserve, or publish each record.

Make retries deliberately safe

Persist idempotency or deduplication state around side effects, distinguish transient from permanent failures, and cap automatic attempts. Classify transient, permanent, and ambiguous failures; cap attempts; isolate poison work; preserve replay context. Never assume a timeout proves that the remote action did not happen.

Preserve explainability

Store external identifiers, attempt history, normalized error categories, and the transformation version used for bounded retry, poison-message isolation, and operator recovery. Operators need enough context to decide whether to replay, repair source data, or stop.

Verify the business result

Pair transport metrics with a commerce assertion: the order reached the intended state, inventory agrees by location, the product is publishable, or the refund reconciles. Seed each failure class and verify retry, stop, alert, dead-letter, and replay behavior.

Proof set

Integration evidence

Evidence expected for Retry and Dead-Letter Strategies
LayerWhat to preserveWhen
Contract examplesRepresentative request, response, event, and error examples for bounded retry, poison-message isolation, and operator recovery, including identifiers and field authority.Before interface design
Failure matrixObserved behavior for timeout, duplicate, delay, throttle, invalid data, and partial completion. Infinite retries hide permanent defects, while immediate discard silently loses commerce work.Before approval
Reconciliation proofA seeded discrepancy is detected, explained, and repaired without repeating an irreversible action.Before release
Operating traceOne business transaction can be followed across systems using correlation data and state history. Seed each failure class and verify retry, stop, alert, dead-letter, and replay behavior.At handoff

Breakpoints

Failure states to design

The primary risk is retrying permanent failures forever or dropping them silently.

  • Connecting systems before deciding which one owns the values described by bounded retry, poison-message isolation, and operator recovery.
  • Treating HTTP success, queue acknowledgement, or webhook receipt as proof of the final business state.
  • Allowing retrying permanent failures forever or dropping them silently to remain an undocumented operator problem.
  • Retrying ambiguous writes without an idempotency, deduplication, or reconciliation boundary. Infinite retries hide permanent defects, while immediate discard silently loses commerce work.

Release

Integration acceptance

  • The initiating commerce event and resulting state transition are explicit.
  • Every mapped identifier and mutable field has one named authority.
  • Duplicate, delayed, missing, reordered, and throttled work has defined behavior.
  • The route-specific control is implemented: Classify transient, permanent, and ambiguous failures; cap attempts; isolate poison work; preserve replay context.
  • A seeded discrepancy can be detected and repaired.
  • Business outcomes are observable independently of transport health. Seed each failure class and verify retry, stop, alert, dead-letter, and replay behavior.

Field notes

Architecture questions

What makes retry and dead-letter strategies dependable?

Dependability comes from explicit record authority, safe delivery semantics, bounded recovery, and reconciliation—not from the number of endpoints. For bounded retry, poison-message isolation, and operator recovery, the design must explain what happens after duplicates, delay, partial failure, and an ambiguous timeout. Retries need a stopping rule and an operator destination.

Should this use a request, webhook, queue, or batch?

Use a request when the caller needs an immediate decision, a webhook when a source announces change, a queue when work needs isolation and retry, and a batch or reconciliation job when completeness matters more than immediacy. Many durable integrations use more than one pattern.

What should be tested beyond the happy path?

Test invalid and missing data, stale versions, duplicate events, reordering, throttling, permission changes, timeout after remote commit, and replay. The route risk—retrying permanent failures forever or dropping them silently—needs a concrete test rather than a sentence in a brief. Infinite retries hide permanent defects, while immediate discard silently loses commerce work.

What evidence belongs at handoff?

Provide payload examples, mapping rules, state diagrams, failure categories, dashboards, alert ownership, replay instructions, and a reconciliation report. Seed each failure class and verify retry, stop, alert, dead-letter, and replay behavior.

Devuchi

Development capacity for this work

Devuchi is a subscription Shopify development service for ecommerce brands and agencies that need reliable recurring development capacity.

bounded retry, poison-message isolation, and operator recovery can be planned against the frameworks and checks in this reference.

Technical references

  1. MDN HTTP overviewTechnical reference
  2. CloudEvents specificationTechnical reference
  3. GraphQL specificationTechnical reference