Azure Table/Blob/Queue random Timeout on linux system (k8s .net core 3 app)

Viewed 523

This is my scenario:

Microsoft.Azure.Storage.Blob 11.2.0
Microsoft.Azure.Storage.Queue 11.2.0
Micorosoft.Azure.Cosmos.Table 1.0.7

I've moved a lot of my code from Azure function to Google k8s and Google Cloud, running the Core .Net app, basically with the same library built in .net Standard 2.0 without any problems.

After a few days, I notice a different behavior in the Linux system. Few calls interacting with Azure service (blob, table, queue) get timeouts (subsystem appears to fail, i tried different retry-police with same result). In 10,000 calls I get 10 to 50 errors (or very long calls 180 seconds, before I changed the timeouts). This happens in all Azure services: table, blob and queue.

I tried different solutions to find out why:

  • I instantiate the client (blobClient, TableClient..etc) every call, or recycle the same client but without difference
  • I change all timeouts to handle this behavior. I work on ServerTimeout and MaximumExecutionTime and put a layer on top, with my retry mechanism, so I can minimize errors. Now I have "only" a few calls of 20 seconds (instead of 2/3 sec for example).
  • I tried all solutions with similar problems found on Stackoverflow :D ... but nothing works (for now)

Same dll code run on azure function without any problems.

So i came to the conclusion, there is something in the http client, used internally by the azure sdk, that depends on the operating system you are running your code on. I think after a few articles it may be the Keep-Alive header, so I try on my composition root:

ServicePointManager.SetTcpKeepAlive (true, 120000, 10000);

but nothing changes.

Any ideas or suggestions? ... maybe I'm on the wrong path, or i've missed something.

1 Answers

UPDATE

After reading the last article linked by @KrishnenduGhosh-MSFT in the last comment i tried to change this setting:

ServicePointManager.DefaultConnectionLimit = 100;

This was the turning point.

Since it used to happen randomly, I'm still not 100% sure if the problem is solved. But after 50k calls, I'm pretty optimistic. Obviously in production will have another behavior, but I already expect it :)

UPDATE 2 - AFTER PUBLISH IN PROD

In the end, it doesn't work :( I had written in the comments, but it seems fair to update here (more readable). I still have long calls (abbreviated with MaximumExecutionTime), but I don't see the light at the end of the tunnel. Now I'm thinking about moving some Azure storage to Google storage, but haven't completely given up.

Related