I've been trying to beat C#'s ConcurrentStack implementation in terms of performance with several different implementations including the Boost lockfree stack, and couldn't even come close.
With the test code below Boost comes in at ~4.6s, my own best performing code which is using DCAS (InterlockedCompareExchange128) on x64 at ~3.7s, C# destroys them both with a time of ~2.3s. My compiler is MSVC, full optimizations.
#include "boost/lockfree/stack.hpp" void TestPerformanceConcurrentStackBOOST() {
System::Diagnostics::Stopwatch sw = new System::Diagnostics::Stopwatch();
sw.Start();
int numThreads = 1;
boost::lockfree::stack<int> stack(128);
for (int t = 0; t < numThreads; t++) {
for (int i = 0; i < 1000000; i++) {
for (int q = 0; q < 100; q++) {
stack.push(q);
}
for (int j = 0; j < 100; j++) {
int res;
stack.pop(res);
}
}
}
sw.Stop();
Console::WriteLine((long)sw.ElapsedMilliseconds); }
The C# test code:
public static void TestPerformanceConcurrentStack() {
System.Diagnostics.Stopwatch sw = new System.Diagnostics.Stopwatch();
sw.Start();
int numThreads = 1;
System.Collections.Concurrent.ConcurrentStack<int> stack = new System.Collections.Concurrent.ConcurrentStack<int>();
System.Collections.Generic.List<System.Threading.ManualResetEvent> lstthreads = new System.Collections.Generic.List<System.Threading.ManualResetEvent>();
for (int t = 0; t < numThreads; t++) {
System.Threading.ManualResetEvent mre = new System.Threading.ManualResetEvent(false);
System.Threading.ThreadPool.QueueUserWorkItem((a) => {
for (int i = 0; i < 1000000; i++) {
for (int q = 0; q < 100; q++) {
stack.Push(q);
}
for (int j = 0; j < 100; j++) {
int res;
stack.TryPop(out (res));
}
}
mre.Set();
});
lstthreads.Add(mre);
}
for (int t = 0; t < numThreads; t++) {
lstthreads[t].WaitOne();
}
sw.Stop();
Console.WriteLine((long)sw.ElapsedMilliseconds);
}
I'm reasonable sure it's not just bc of C#'s garbage collection since I also tested a C++ version with custom GC, without luck. I've checked the C# code and it's pretty straightforward, not much special going on. Performance profiling showed the Boost version spends ~60% of the time in the pop method, similar to my DCAS implementation.
Does anyone have a C++ concurrent stack that's faster than C# or do you have any suggestions on what I can try next to make this faster?
Any help is much appreciated. Thanks!