Cloudflare Queues Retries: Ack, Delay and Dead Letters

A Cloudflare Queues retry is the whole batch, not the one message that broke. If your consumer throws, every message in that batch comes back, including the ones you had already processed. Three things change that: message.ack() on the messages you finished, message.retry({ delaySeconds }) on the ones you did not, and a max_retries limit with a dead letter queue behind it so a poison message stops circulating.

The default is easy to miss because the handler looks per-message: you loop over an array and deal with one body at a time. The platform sees one promise for the lot. Cloudflare’s own description of the model is blunt: “messages within a batch are treated as all or nothing when determining retries. If the last message in a batch fails to be processed, the entire batch will be retried.” What follows is the mechanics behind that sentence, taken from the Queues docs as of October 2026.

How do a producer and a consumer get wired up?

A queue is created once with the CLI and then referenced from the Worker config. The producer side is a binding; the consumer side is a consumers entry that points the queue at the Worker containing a queue() handler. Both can live in the same Worker.

npx wrangler queues create order-events
{
  "queues": {
    "producers": [
      { "queue": "order-events", "binding": "ORDER_EVENTS" }
    ],
    "consumers": [
      {
        "queue": "order-events",
        "max_batch_size": 10,
        "max_batch_timeout": 5,
        "max_retries": 5,
        "dead_letter_queue": "order-events-dlq"
      }
    ]
  }
}

The producer call is one line. send() takes a body that survives structured clone, up to 128 KB, and an optional delaySeconds of up to 86400 (24 hours). sendBatch() takes up to 100 messages or 256 KB in total, whichever you hit first.

export default {
  async fetch(request: Request, env: Env): Promise<Response> {
    const payload = await request.json();
    await env.ORDER_EVENTS.send(
      { topic: request.headers.get("x-shopify-topic"), payload },
      { contentType: "json" }
    );
    return new Response(null, { status: 202 });
  },
};

Returning 202 before the work is done is the point. A webhook sender wants a fast acknowledgement; the queue absorbs the slow part. If that webhook is Shopify’s, the duplicate-delivery behaviour described in our post on Shopify webhook retries and idempotency still applies, and now there is a second source of duplicates: the queue itself.

What actually happens when the consumer throws?

The consumer is invoked with a MessageBatch. Its messages array is delivered in best-effort order, and Cloudflare says plainly that it “does not guarantee that messages will be delivered to a consumer in the same order in which they are published”. Each Message carries id, timestamp, body and attempts, where attempts starts at 1 on first delivery.

If the handler resolves, every message in the batch is acknowledged implicitly and deleted. If it rejects, or runs past the 15 minute wall clock limit, the batch is retried according to the consumer’s settings.

Here is the case that bites. Ten messages arrive. You process nine, the tenth throws, the handler rejects. All ten are redelivered. The nine you already handled run again, with attempts now at 2, and whatever side effect they had (an email, an order update, a charge) happens twice.

Messages are not removed “until the consumer has successfully consumed the message”, which is the right guarantee for durability and the wrong one for a side effect that is not idempotent. Either make the work safe to repeat, keyed on message.id or on an id inside the body, or acknowledge as you go so a late failure cannot drag finished work back in.

How do you ack and retry one message at a time?

Call ack() on each message once its work is committed, and retry() on each one you want back. The first call wins: if you ack() a message and the batch later fails, that message is not redelivered, and batch.retryAll() does not override an individual ack() made earlier.

interface OrderEvent {
  topic: string;
  payload: { id: number };
}

export default {
  async queue(batch: MessageBatch<OrderEvent>, env: Env): Promise<void> {
    for (const message of batch.messages) {
      try {
        await handleOrderEvent(message.body, env);
        message.ack();
      } catch (error) {
        if (isPermanent(error)) {
          // Do not spend retries on a message that can never succeed.
          console.error("dropping", message.id, error);
          message.ack();
          continue;
        }
        // Back off: 30s, 60s, 120s... capped at an hour.
        const delaySeconds = Math.min(30 * 2 ** (message.attempts - 1), 3600);
        message.retry({ delaySeconds });
      }
    }
  },
};

Two decisions are hidden in that handler. The first is what to do with a message that will never succeed: a malformed body, a deleted record. Retrying it five times and then dead-lettering it is one option; acknowledging it and logging is another. Acknowledging is cheaper, but the dead letter queue then stops being a complete record of failures, so pick one and be consistent.

The second is the back-off. message.attempts is the only counter you get, and delaySeconds on retry() accepts anything up to 24 hours, so exponential back-off is a one-liner. Without a delay, nothing holds a retried message back except the consumer’s own capacity, which is the behaviour you least want when the thing failing is an upstream rate limit.

How long does a retry wait, and for how long is it kept?

Three settings control timing.

retry_delay on the consumer is the default wait for any message that fails without an explicit delay. delaySeconds on retry() or retryAll() overrides it for that call, and tops out at 86400 seconds.

{
  "queues": {
    "consumers": [
      { "queue": "order-events", "retry_delay": 300 }
    ]
  }
}

Retention is separate and is per queue, not per consumer. The default is 345600 seconds (four days), adjustable from 60 seconds to 1209600 (14 days) with npx wrangler queues update order-events --message-retention-period-secs 1209600. On the Workers Free plan it is fixed at 24 hours. A long back-off and a short retention are a bad pair: a message delayed past its retention window is gone.

max_retries caps the attempts. The default is 3 and the platform limit is 100. When a message runs out, it is deleted, or written to the dead letter queue if one is configured.

Where do failed messages end up?

A dead letter queue is an ordinary queue that the consumer config names. Set dead_letter_queue and wrangler creates it if it does not exist; the CLI form is wrangler queues consumer add order-events my-worker --dead-letter-queue=order-events-dlq.

It is not a place things are kept indefinitely. Cloudflare’s docs say messages in a DLQ with no active consumer persist for “four (4) days before being deleted”. If you want to inspect failures a week later, either raise the retention on the DLQ with the same queues update flag, or attach a consumer to it that writes each message somewhere durable. D1 is the easy choice; the trade-offs between it and KV are in our comparison of Cloudflare KV, D1 and Durable Objects.

A DLQ consumer is also where to put the alert. A message landing there means max_retries attempts all failed, which is a better signal than one error log line, and it fires once per message rather than once per attempt.

Which batch and concurrency settings matter for retries?

max_batch_size (default 10, max 100) and max_batch_timeout (default 5 seconds, max 60) decide when a batch is handed over: “whichever limit is reached first will trigger the delivery of a batch”. A larger batch means fewer invocations and a bigger blast radius per failure, because without per-message acks one bad message takes the whole batch round again. If you cannot ack per message, keep batches small.

Concurrency scales automatically, up to 250 concurrent invocations, based on backlog growth and the ratio of failed to successful invocations. Retrying with retry() or retryAll() does not count as a failed invocation for that calculation, so a consumer that retries cleanly keeps scaling, while one that throws gets throttled. Set max_concurrency to 1 when the thing you call is the bottleneck. The docs describe that setting as choosing “for your backlog to grow, trading off increased overall latency in order to avoid overwhelming an upstream system”. A consumer that talks to a rate-limited API such as Shopify Admin is exactly that case.

What I would configure on day one

Per-message ack() in every consumer, no exceptions, with a try around each message rather than the loop. Exponential delaySeconds from message.attempts. max_retries at 5 rather than 3, because three attempts spread over a few minutes can all land inside one upstream outage. A dead letter queue on every consumer, with a tiny consumer of its own that writes the body and attempts to D1 and sends one alert. max_concurrency: 1 only where the downstream cannot take more, and left unset otherwise.

What I would not do is treat the queue as the idempotency layer. It is not one. The queue promises your message will not be lost; it does not promise it arrives once. If your handler cannot run twice safely, no combination of these settings will save you, and the fix belongs in the handler.

We build this kind of pipeline for clients on Workers, usually to take webhook ingestion or order processing out of the request path. If you are weighing Queues against a cron trigger or a Durable Object for a job you have in mind, our serverless development work covers that call.

Need this built properly?

Whoooop Ltd has spent 15+ years building and maintaining web applications in TypeScript, React, Node.js and serverless — the same ground this post covers.

Get in touch