Retries and delivery
The retry schedule, what counts as acknowledged, how to inspect past deliveries, and how to replay one.
Delivery is at-least-once. We keep trying until your endpoint acknowledges, which means a duplicate is not an anomaly - it is the expected cost of never losing an event.
What counts as acknowledged
Any 2xx response. Anything else - a 4xx, a 5xx, a connection failure, or a response that
takes too long - is a failed attempt and will be retried.
Return 200 as soon as you have verified the signature and durably recorded the event. Do the
real work afterwards, out of band. An endpoint that fulfils an order inline and takes four seconds
will be retried while it is still working, and you will process the same event twice.
The retry schedule
| LIVE | TEST | |
|---|---|---|
| First 4 attempts | Every 3 minutes | Hourly |
| Thereafter | Hourly | Hourly |
| Total window | 72 hours from the first attempt | 10 hours |
The window is time-based, not a fixed attempt count. We stop scheduling once the next attempt would fall outside it - so a delivery is never abandoned mid-window.
If a delivery exhausts its budget, we notify you by email that an event could not be delivered.
That is a signal to check your endpoint and then reconcile - not
to wait for another attempt, because we will not schedule one. You can still force one yourself
with POST /webhooks/deliveries/{id}/retry, which works on abandoned deliveries too.
Warning
72 hours is generous but finite. An endpoint down over a long weekend can exhaust the window on real events. A periodic reconcile sweep over
GET /transactions/listwith a date window is the backstop that makes an outage recoverable rather than a permanent gap.
Deduplicating
Every payload carries eventId, stable across retries of the same event. Claim it before doing
any work:
Create the table in your own database. The primary key on event_id is what makes the claim
atomic:
CREATE TABLE processed_webhook_events (
event_id VARCHAR(64) PRIMARY KEY,
received_at TIMESTAMP NOT NULL
);
Then claim each event with a single insert. Examples below use PostgreSQL syntax:
-- The primary key on event_id makes the claim atomic and the handler idempotent.
INSERT INTO processed_webhook_events (event_id, received_at)
VALUES ($1, now())
ON CONFLICT (event_id) DO NOTHING;
Note
On MySQL or MariaDB use
INSERT IGNORE INTO processed_webhook_events (event_id, received_at) VALUES (?, now())instead. Either way, treat "0 rows inserted" as "already processed".
Node.js
async function claimEvent(eventId) {
const { rowCount } = await db.query(
`INSERT INTO processed_webhook_events (event_id, received_at)
VALUES ($1, now()) ON CONFLICT (event_id) DO NOTHING`,
[eventId],
);
return rowCount === 1; // false => already processed, acknowledge and stop
}
Python
def claim_event(event_id: str) -> bool:
with db.cursor() as cur:
cur.execute(
"""INSERT INTO processed_webhook_events (event_id, received_at)
VALUES (%s, now()) ON CONFLICT (event_id) DO NOTHING""",
(event_id,),
)
return cur.rowcount == 1 # False => already processed, acknowledge and stop
PHP
<?php
function claimEvent(PDO $db, string $eventId): bool
{
$stmt = $db->prepare(
'INSERT INTO processed_webhook_events (event_id, received_at)
VALUES (:id, now()) ON CONFLICT (event_id) DO NOTHING'
);
$stmt->execute(['id' => $eventId]);
return $stmt->rowCount() === 1; // false => already processed, acknowledge and stop
}
Java
boolean claimEvent(String eventId) {
// false => already processed, acknowledge and stop
return jdbcTemplate.update(
"""
INSERT INTO processed_webhook_events (event_id, received_at)
VALUES (?, now()) ON CONFLICT (event_id) DO NOTHING
""",
eventId) == 1;
}
C#
async Task<bool> ClaimEventAsync(string eventId)
{
// false => already processed, acknowledge and stop
var rows = await connection.ExecuteAsync(
@"INSERT INTO processed_webhook_events (event_id, received_at)
VALUES (@eventId, now()) ON CONFLICT (event_id) DO NOTHING",
new { eventId });
return rows == 1;
}
Tip
Dedupe on
eventId, not ontransactionId. One transaction legitimately produces several events -TRANSACTION_CREATED, thenTRANSACTION_SUCCEEDED, plusCHECKOUT_SESSION_PAIDfor a hosted checkout. Keying on the transaction would drop events you need.
Keep claimed ids for at least the length of the retry window, with a wide margin - 30 days is a reasonable retention.
Inspecting deliveries
| Endpoint | Purpose |
|---|---|
GET /webhooks/events |
Events generated for your Integrator |
GET /webhooks/deliveries/{id} |
One delivery attempt, with the response we got back |
POST /webhooks/deliveries/{id}/retry |
Force an immediate retry |
GET /integrators/{id}/usage/summary |
Request totals plus webhook delivery statistics |
GET /webhooks/deliveries/{id} records the status code and response body your endpoint returned,
which is usually enough to diagnose a failure without adding logging on your side.
Requires the WEBHOOKS_MANAGE scope.
Replaying after an outage
- Fix the endpoint and confirm it returns
200to a test delivery. - Force a retry on the affected deliveries with
POST /webhooks/deliveries/{id}/retry, rather than waiting for the hourly schedule. - Abandoned deliveries can be force-retried the same way. Then reconcile anyway: list transactions over the outage window and compare against your own records, so nothing slips through. See Reconciliation.
Designing an endpoint that stays up
- Acknowledge before processing. Verify, persist the raw event, return
200, then queue. - No synchronous third-party calls in the handler. A slow downstream service becomes your webhook timeout.
- Handle events you do not recognise. New event types get added. Acknowledge and ignore
anything unfamiliar rather than erroring - an unknown event that returns
500is retried for 72 hours. - Do not allowlist by IP. Egress addresses can change; the signature is the authentication mechanism.
- Alert on your own failure rate, not just on ours. A handler silently returning
500looks fine from your side until the reconciliation gap shows up.
Next
Reconciliation - the playbook for when webhooks and your books disagree.