Boost Asio experimental channel poor performance

Viewed 110

I wrote the following code to analyze experimental channel performance in a single thread application. On i7-6700HQ@3.2GHz It takes around 1 second to complete which shows a throughput of around 3M item per second.

The problem might be due to the fact that because asio is in single threaded mode the producer has to signal the consumer part and that leads to immediate resumption of consumer coroutine on every call to async_send(), but i don't know how to test to make sure this is the case and how we can avoid it in real applications. reducing channel buffer size even to 0 has no effect on the throughput which might be for the same reason.

#include <boost/asio.hpp>
#include <boost/asio/experimental/awaitable_operators.hpp>
#include <boost/asio/experimental/channel.hpp>

namespace asio = boost::asio;
using namespace asio::experimental::awaitable_operators;
using channel_t = asio::experimental::channel< void(boost::system::error_code, uint64_t) >;

asio::awaitable< void >
producer(channel_t &ch)
{
    for (uint64_t i = 0; i < 3'000'000; i++)
        co_await ch.async_send(boost::system::error_code {}, i, asio::use_awaitable);

    ch.close();
}

asio::awaitable< void >
consumer(channel_t &ch)
{
    for (;;)
        co_await ch.async_receive(asio::use_awaitable);
}

asio::awaitable< void >
experiment()
{
    channel_t ch { co_await asio::this_coro::executor, 1000 };
    co_await (consumer(ch) && producer(ch));
}
int
main()
{
    asio::io_context ctx {};
    asio::co_spawn(ctx, experiment(), asio::detached);
    ctx.run();
}

2 Answers

You can save a little by providing hints about the threading:

  • provide concurrency hint unsafe (BOOST_ASIO_CONCURRENCY_HINT_UNSAFE)
  • optionally disabling all threading - this will in practice probably not matter, it's just possible as long as you don't need any services that employ internal threads)
  • avoiding type erasure on the executor; this means replacing any_io_executor with the concrete executor type that you employ

I wrote a side-by-side benchmark with reduced message-count (30k) so that Nonius can sample 100 runs and do statistical analysis on the results:

//#define TWEAKS

#ifdef TWEAKS
#define BOOST_ASIO_DISABLE_THREADS 1
#endif

#include <boost/asio.hpp>
#include <boost/asio/experimental/awaitable_operators.hpp>
#include <boost/asio/experimental/channel.hpp>

#include <iostream>

namespace asio = boost::asio;

using namespace asio::experimental::awaitable_operators;
using boost::system::error_code;

using context    = asio::io_context;
#ifdef TWEAKS
using executor_t = context::executor_type;
using channel_t  = asio::experimental::channel<executor_t, void(error_code, uint64_t)>;
#else
using executor_t = asio::any_io_executor;
using channel_t  = asio::experimental::channel<void(error_code, uint64_t)>;
#endif

asio::awaitable<void> producer(channel_t& ch) {
    for (uint64_t i = 0; i < 30'000; i++)
        co_await ch.async_send(error_code {}, i, asio::use_awaitable);

    ch.close();
}

asio::awaitable<void> consumer(channel_t& ch) {
    for (;;)
        co_await ch.async_receive(asio::use_awaitable);
}

asio::awaitable<void> experiment() {
    asio::any_io_executor ex = co_await asio::this_coro::executor;
    channel_t ch { *ex.target<executor_t>(), 1000 };
    co_await (consumer(ch) && producer(ch));
}

void foo() {
    try {
#ifdef TWEAKS
        asio::io_context ctx{BOOST_ASIO_CONCURRENCY_HINT_UNSAFE};
#else
        asio::io_context ctx{1};
#endif

        asio::co_spawn(ctx, experiment(), asio::detached);

        ctx.run();
    } catch (std::exception& e) {
        std::cerr << "Exception: " << e.what() << "\n";
    }
}

#include <nonius/benchmark.h++>
#define NONIUS_RUNNER
#include <nonius/main.h++>

NONIUS_BENCHMARK( //
    "foo",        //
    [](nonius::chronometer cm) { cm.measure([] { foo(); }); })

The results per 30k batch (including construction and teardown) are:

So ~25% speed increase, and also much reduced variance.

Combining the series in one graph:

enter image description here

Thoughts

These are just the Asio technical tweaks. I might be missing some still.

I suspect you should be able to get much better throughput with smart buffering. I'm assuming you need the Asio integration for other reasons, making this the right choice.

It turned out the consumer and producer sides are scheduled in the event loop on each send/receive operation that's why channel size has no effect on the throughout.
I've changed the code to the following and now it can send 90M per seconds. but this is what I was expected from the implementation.

 asio::awaitable< void >
 producer(channel_t &ch)
 {
     for (uint64_t i = 0; i < 90'000'000; i++)
     {
         if (!ch.try_send(boost::system::error_code {}, i))
             co_await ch.async_send(boost::system::error_code {}, i, asio::use_awaitable);
     }
 
     ch.close();
 }
 
 asio::awaitable< void >
 consumer(channel_t &ch)
 {
     for (;;)
     {
         if (!ch.try_receive([](auto, auto) {}))
             co_await ch.async_receive(asio::use_awaitable);
     }
 }

I think the reason that this is not the default behavior of channels is that because there is no way for awaitables in asio to return true in await_ready() call they always have to suspend and initiate an asynchronous operation.

Related