Worker Service mysteriously stops doing work

Viewed 1689

My good ladies and gentlemen, I recently had my first go at Worker services in .Net Core 3.1, and only the second go at Windows services in general (first one was made in .Net Framework and works fine to this day). If anyone could maybe shed some light at what I'm missing in the example that I will provide that would be great.

So, to keep it simple, my problem is this: My supposed long (forever) running Worker service unexpectedly stops doing work at an arbitrary time of day, but still is shown as "Running" in service manager (that's probably how Windows deals with services). It doesn't necessarily have to be every day, but it stops doing work every now and then until I manually stop it and then restart it in Service Manager.

I have also stumbled upon this question which seemed to deal with my problem, but even after completely wrapping all of my service's code blocks in try-catchs, even on top-level, I still get nothing registered in my Log table, or even in the file I set up to write in if my DB connection fails. Service seems to just stop calling ExecuteAsync() method.

Ok here's how my code's logically structured, I have excluded implementation and I'm just showing what happens until DoWork is called:

public class Worker : BackgroundService
{
    private readonly IConfiguration _configuration;
    public Worker(IConfiguration configuration)
    {
        _configuration = configuration;
    }

    public override Task StartAsync(CancellationToken cancellationToken)
    {                      
        return base.StartAsync(cancellationToken);
    }

    protected override async Task ExecuteAsync(CancellationToken stoppingToken)
    {
        try
        {
            while (true)
            {
                try //paranoid try-catch
                {
                    await DoWork();
                    await Task.Delay(TimeSpan.FromSeconds(45), stoppingToken);
                }
                catch (Exception e)
                {
                    await Log(e, customMessage: "Proccess failed at top level.");
                }
            }
        }
        catch (Exception e)
        {
            await Log(e, customMessage: "Proccess failed at topmost level.");
        }

    }

    private async Task DoWork()
    {
        try
        {
            
        }
        catch (Exception e)
        {
            await Log(e);
        }
    }

    public async Task Log(Exception e, string user = null, string emailID = null, string customMessage = null)
    {
        
    }
}

As you can see, I am not handling cancellation, as in the question I linked above. Now that I think about it maybe I should, and something is inadvertently sending cancellation? The reason I didn't is because I'm not sure what events exactly signal the cancellation. Only the manual stopping of service, or something else maybe? And if it is the cancellation that was sent that caused my service to stop doing work, shouldn't it also stop my service from running?

Btw I just tested cancellation on dummy service which implements my logic with while(true) and it catches the stopping exception, even though it's a bit awkward, as it catches it and logs it multiple times before stopping, so I presume it may not be the cancellation token that is causing my DoWork not to fire.

Publish settings (I'd installed it on Windows Server 2012 R2 Standard)

3 Answers

Ok guys I'd fixed it. See comment below.

I'd guessed that what was causing deadlock was probably too many concurrent calls from different threads to database over the same connection.
Not that that I knew that would be the cause (and I still don't know and can only guess why this happens so if someone can clarify why this happens and why don't the calls get queued please do), but as I tried to fix it that seemed like a good starting point.

What I did was just limit possible concurrent calls to 1:

  1. Instantiate SemaphoreSlim on a class level:
    private static SemaphoreSlim Semaphore = new SemaphoreSlim(1);
  2. Insert a SemaphoreSlim.WaitAsync before each of my DB calls and its respective SemaphoreSlim.Release in a finally block after the call:
try
{
    await Semaphore.WaitAsync();
    var id = await sqlCommand.ExecuteScalarAsync().ToString();
}
finally
{
    Semaphore.Release();
}

I thought this would decrease the performance but to my pleasant surprise I felt no noticeable difference.

Also, I was tempted to set Semaphore's initial count to more than 1 thread but I figured if deadlock happens for many threads, then it might happen for 2-10 threads. Does anyone perhaps know anything more about this number? Is it processor related, SQL related, or perhaps C# related?

Have you implemented a dispose method to close the database connection after finishing the DoWork method? I had a deadlock problem using worker service and realized the database connection wasn’t disposed. After implementing a dispose method, it works for me to solve the problem.

Old question, and I don't know how or even if Manus's issue eventually resolved, but in my experience, exceptions in threads that aren't the main thread caused this problem for us in a Windows Service. And we didn't see this when running the same code as a Windows Forms. Try/catch wrapped around the main processing line does not catch them. We had to add it in the threaded method.

Related