Storing each async result in its own array element

Viewed 788

Let's say I want to download 1000 recipes from a website. The websites accepts at most 10 concurrent connections. Each recipe should be stored in an array, at its corresponding index. (I don't want to send the array to the DownloadRecipe method.)

Technically, I've already solved the problem, but I would like to know if there is an even cleaner way to use async/await or something else to achieve it?

    static async Task MainAsync()
    {
        int recipeCount = 1000;
        int connectionCount = 10;
        string[] recipes = new string[recipeCount];
        Task<string>[] tasks = new Task<string>[connectionCount];
        int r = 0;

        while (r < recipeCount)
        {
            for (int t = 0; t < tasks.Length; t++)
            {
                tasks[t] = Task.Run(async () => recipes[r] = await DownloadRecipe(r));
                r++;
            }

            await Task.WhenAll(tasks);
        }
    }

    static async Task<string> DownloadRecipe(int index)
    {
        // ... await calls to download recipe
    }

Also, this solution it's not optimal, since it doesn't bother starting a new download until all the 10 running downloads are finished. Is there something we can improve there without bloating the code too much? A thread pool limited to 10 threads?

2 Answers

There are many many ways you could do this. One way is to use an ActionBlock which give you access to MaxDegreeOfParallelism fairly easily and will work well with async methods

static async Task MainAsync()
{
   var recipeCount = 1000;
   var connectionCount = 10;
   var recipes = new string[recipeCount];

   async Task Action(int i) => recipes[i] = await DownloadRecipe(i);
   
   var processor = new ActionBlock<int>(Action, new ExecutionDataflowBlockOptions()
   {
      MaxDegreeOfParallelism = connectionCount,
      SingleProducerConstrained = true
   });

   for (var i = 0; i < recipeCount; i++)
      await processor.SendAsync(i);

   processor.Complete();
   await processor.Completion;
}

static async Task<string> DownloadRecipe(int index)
{
   ...
}

Another way might be to use a SemaphoreSlim

var slim = new SemaphoreSlim(connectionCount, connectionCount);

var tasks = Enumerable
   .Range(0, recipeCount)
   .Select(Selector);
   
async Task<string> Selector(int i)
{
   await slim.WaitAsync()
   try
   {
      return await DownloadRecipe(i)
   }
   finally
   {
      slim.Release();
   }
}

var recipes = await Task.WhenAll(tasks);

Another set of approaches is to use Reactive Extensions (Rx)... Once again there are many ways to do this, this is just an awaitable approach (and likely could be better all things considered)

var results = await Enumerable
        .Range(0, recipeCount)
        .ToObservable()
        .Select(i => Observable.FromAsync(() => DownloadRecipe(i)))
        .Merge(connectionCount)
        .ToArray()
        .ToTask();

Alternative approach to have 10 "pools" which will load data "simultaneously".

You don't need to wrap IO operations with the separate thread. Using separate thread for IO operations is just a waste of resources.
Notice that thread which downloads data will do nothing, but just waiting for a response. This is where async-await approach come very handy - we can send multiple requests without waiting them to complete and without wasting threads.

static async Task MainAsync()
{
    var requests = Enumerable.Range(0, 1000).ToArray();
    var maxConnections = 10;
    var pools = requests
        .GroupBy(i => i % maxConnections)
        .Select(group => DownloadRecipesFor(group.ToArray()))
        .ToArray();

    await Task.WhenAll(pools);

    var recipes = pools.SelectMany(pool => pool.Result).ToArray();
}

static async Task<IEnumerable<string>> DownLoadRecipesFor(params int[] requests)
{
    var recipes = new List<string>();
    foreach (var request in requests)
    {
        var recipe = await DownloadRecipe(request);
        recipes.Add(recipe);
    }

    return recipes;
}

Because inside the pool (DownloadRecipesFor method) we download results one by one - we make sure that we have no more than 10 active requests all the time.

This is little bit more effective than originals, because we don't wait for 10 tasks to complete before starting next "bunch".
This is not ideal, because if last "pool" finishes early then others it aren't able to pickup next request to handle.

Final result will have corresponding indexes, because we will process "pools" and requests inside in same order as we created them.

Related