How does ExponentialBackoffRetry works with ServiceBus Trigger for Azure function?

Viewed 2947

I want to implement a very simple behavior in my Azure Function: if there is an exception during handling, I want to postpone the next retry for some time. As far as I know there is no direct possibility for that in the Service Bus e.g. (unless one creates a new message), but Service Bus Trigger has a possibility for ExponentialBackoffRetry.

I have not found any documentation on how that might work with regards to Service Bus Connection. I.e. what happens with the message after the execution of the function fails.

One possible way is to keep the message in functions infrastructure and keep renewing the lock for the duration I suppose. Some more practical questions on what I am wondering about:

  1. How long can I use backoff retry (e.g. if I want retry to up to 7 days e.g. will that work?)
  2. What happens when host is being reset/restarted/scaled, do I lose this backoff due to implementation details or it is still somehow maintained?
3 Answers

From the documentation:

Using retry support on top of trigger resilience

The function app retry policy is independent of any retries or resiliency that the trigger provides. The function retry policy will only layer on top of a trigger resilient retry. For example, if using Azure Service Bus, by default queues have a message delivery count of 10. The default delivery count means after 10 attempted deliveries of a queue message, Service Bus will dead-letter the message. You can define a retry policy for a function that has a Service Bus trigger, but the retries will layer on top of the Service Bus delivery attempts.

For instance, if you used the default Service Bus delivery count of 10, and defined a function retry policy of 5. The message would first dequeue, incrementing the service bus delivery account to 1. If every execution failed, after five attempts to trigger the same message, that message would be marked as abandoned. Service Bus would immediately requeue the message, it would trigger the function and increment the delivery count to 2. Finally, after 50 eventual attempts (10 service bus deliveries * five function retries per delivery), the message would be abandoned and trigger a dead-letter on service bus.

For the exponential retries, you likely need to keep the total backoff time + processing to less than what a function can hold on to the message or else the lock will expire, and even successful processing will result in an exception and retry.

The way Service Bus locks messages today, exponential backoff on top of Azure Service Bus is not a great idea. Once durable terminus is possible (unlimited lock time w/o the need to renew), this will make much more sense.

Update: Functions retry feature is being deprecated.

The retry options apply to a single service operation performed by the Service Bus SDK and are intended to allow the SDK work around short-term transient issues, like the occasional network interruption. Other than configuring the SDK clients, the Functions infrastructure is unaware of the retries and would simply see the SDK taking a longer time to perform the requested read/publish operation.

The Functions infrastructure will apply any execution time limits imposed by the runtime or may decide to take action to guard against an unresponsive service operation. (disclaimer: I can speak to the Service Bus SDK, but don't have deep insight into the Functions runtime)

The retries from the Service Bus extensions aren't applied to your Function code; on an error in your code you'll end up in an exception scenario and, depending on configuration and trigger/binding use, will either see your message abandoned or the lock held until timeout.

I'm not sure of your exact scenario, but it seems like you may want to consider deferring the message to be read explicitly at a later time or re-enqueuing the message with a schedule so that the Function can read again at a specific point in the future.

It doesn't because that feature is soon to be removed from pretty much all triggers. At least that how I read the brand new updated documentation:

IMPORTANT: The retry policy support in the runtime for triggers other than Timer and Event Hubs is being removed after this feature becomes generally available (GA). Preview retry policy support for all triggers other than Timer and Event Hubs will be removed in October 2022.

As it stands now, your only option is to implement the retry logic yourself. It's fairly easy to to a basic retry + sleep loop in your code and you can leverage something like Polly to make it more robust. Just be wary about timeout issues in your function.

Another approach is to use scheduled messages where you publish the failing message again on the queue by give it a datetime when it should appear + you need to add some kind of custom "retry count" header which you increase each time to publish it and manually dead letter it when it has failed a number of times.

Related