I want the retry test to answer a concrete question: after the sender tries again, how many records or actions exist at the destination?
A green HTTP response alone does not answer it. Neither does a log that says “duplicate ignored” if the original work never completed. This guide turns the error-handling part of an integration brief into concrete acceptance checks.
What must the receiver expect from Shopify?
Shopify documents duplicate delivery, unordered events and a finite retry window. I treat those as inputs to the acceptance plan, rather than assuming delivery happens once and in order.
For HTTPS delivery, Shopify specifies a one-second connection timeout and a five-second total request timeout. No response or an error triggers eight retries over four hours. An Admin API-configured subscription is deleted after eight consecutive failures. Source: verify deliveries.
The same documentation names X-Shopify-Webhook-Id for delivery deduplication. Separate subscriptions to the same topic can have different delivery IDs but share X-Shopify-Event-Id, which correlates the merchant action. I still define a business key for the operation I am accepting.
Shopify does not guarantee ordering within a topic, or across topics for the same resource. It recommends organising webhooks using the X-Shopify-Triggered-At header or the payload’s updated_at. Because delivery is not always guaranteed, it also recommends periodic API reconciliation. Source: about webhooks.
What did I actually run?
I ran four fault scenarios on 8 October 2026 with Node v22.22.2. The fixtures are synthetic. The HTTP requests, client timeout, stopped receiver and process restart are real local operations. No Shopify store, client data or external service was involved.
The harness starts one receiver on a free 127.0.0.1 port. Its saved snapshot contains delivery keys, current order state, local effects and a pending inbox. Replacing one snapshot keeps these together for this single-process test. A successful response follows the save.
Each fixture has an invented integer version; this is a test rule, not a Shopify field. A higher version can update state. A lower version is recorded as stale and cannot overwrite it. An effect means an entry in the local effect ledger, not a payment, email or remote API call.
I ran node run.mjs from the harness directory with Node v22.22.2. It is one script with no packages. The script and results are included in the owner’s review packet; this page does not offer a public download. It prints JSON and writes results.json with captured HTTP statuses, outcomes, saved-state counts and assertion results. A failed safe-receiver assertion exits unsuccessfully.
What did the fault matrix show?
All four safe-receiver scenarios passed their stated checks. The table below is pasted verbatim from table.md, generated by the runner from its captured JSON values. It records local responses and saved effects, not expected Shopify behaviour or a production reliability score.
Comparison table — scroll horizontally to see all columns
| Fault injected | Actual output | Acceptance check |
|---|---|---|
| Two overlapping deliveries | 200 applied; 200 duplicate; effects=1 | One saved effect |
| Version 2 before version 1 | 200 applied; 200 stale; final=paid v2; effects=1 | Newer state and one effect |
| Timeout after commit, then retry | CLIENT_TIMEOUT; 200 duplicate; effects=1 | Commit before timeout; one effect |
| Queue, stop and restart | 202 queued; ECONNREFUSED; pending=1 to 0; 200 duplicate; 200 applied; effects=1/1 | Inbox drained; one effect per order |
The overlapping requests are handled on one event loop, with no asynchronous gap between checking and saving the key. This is not a multi-worker race test; that remains a proposed project check below.
For the timeout row, the receiver reports the commit after saving. The runner records the notification time and the client’s timeout time on its own monotonic clock, and asserts that the notification came first. The client gives up after 300 ms; the receiver delays its response for 1,200 ms. The runner then retries the same key. “I received no answer” does not mean “nothing happened.”
For recovery, one fixture carries a test flag that makes the receiver queue it in the saved inbox instead of applying it. The receiver then stops; a different delivery receives a connection refusal. Restart drains the inbox. Redelivering the queued fixture creates no extra effect, and retrying the refused fixture creates its first effect. The final counts are one effect for each of those two orders.
I ran the same acceptance checks against four deliberately broken HTTP receivers. An unconditional receiver without idempotency applied the same delivery twice: 200 applied; 200 applied; effects=2. A receiver using arrival order ended at created v1 after paid v2. Both failed the checks. Two further receivers returned HTTP 500 or added a second effect on the stale path; those also failed. The JSON records each failed assertion. If a negative control unexpectedly passes, the runner exits unsuccessfully.
How do I turn this into project acceptance?
I write the expected destination state before the test, then inspect it after each fault. The count, recovery record and owner of unresolved work belong in the acceptance agreement.
Comparison table — scroll horizontally to see all columns
| Check to agree | Evidence I would ask for |
|---|---|
| Same operation retried with the same and different delivery keys | One business effect, tied to the agreed operation key |
| Two workers race for the same operation | One committed result; the losing worker can explain its outcome |
| Older update follows a newer update | Current state remains correct; ambiguous versions go to review or a source read |
| Process stops before and after a destination write | A retry either completes missing work or finds the committed result |
| Downtime exceeds the sender’s retry window | Reconciliation finds missed records; replay is safe and bounded |
| Retries are exhausted | A visible failed record, a named owner and a documented replay decision |
These are proposed checks for the real integration. The local harness did not execute the multi-worker, external-write or long-downtime rows. For direct API integrations, I would agree them with the systems and task in view.
Keep the IDs separate. The delivery key identifies an incoming delivery. The business key identifies the action, such as issuing an invoice once. A destination event ID or transaction ID has its own contract. One key does not automatically cover all three layers.
What does this harness leave unproved?
It proves the listed local outcomes for one receiver process and one file snapshot. It does not prove end-to-end exactly-once delivery, production durability or the behaviour of a remote destination.
I did not test power loss, filesystem flush guarantees, a process killed during a save, HMAC verification, malformed payloads, database contention or a remote API accepting a write before our local commit. The graceful restart tests persisted state and inbox recovery; it is not a crash-safety certification.
For an external effect, I would use a destination-supported operation key where available, or a recoverable outbox with a destination lookup and an explicit uncertainty path. Saving “seen” before doing the work can lose work. Doing the work before saving “seen” can duplicate it. That boundary needs its own test.
What is the next useful step?
I start with the operation, its key and its acceptance evidence. Then I choose the recovery mechanism that the destination can actually support.
The API integrations and automation service describes the scope and handover. The integration brief is the place to agree these checks. If the destination is a measurement platform, use order reconciliation to check which orders arrived, and keep subscription renewals separate from duplicate deliveries.