AFAIU there's a common idea that often lock-free code has higher overhead than locking one.
Although, also it seems there's an idea that lock-free algorithms are more scalable under contention.
If there are 2 cores, and 2 threads contending over something like a std::queue (+ mutex) vs. a boost::lockfree::queue (MPMC lock-free queue), locks will likely outperform lock-free (AFAIU because of the general overhead of lock-free algorithms). If I have 50 threads instead, locking version will be really slow because of all the context switching, explicit lock convoy etc.
AFAIU in this case a lock-free version might tend to perform better than a locking one; is this roughly true?
However having 50 threads on a 2 core machine is not a good idea for performance anyway; so I guess a more relevant question would be the same thing on a machine with at least 50 cores.
2-thread versions would use 2 cores at a time and I expect the outcome would be roughly the same; what about 50 cores contending for the same lock? The impact of locking seems to be higher than that of a 2-thread version (only 1 thread can do the job out of 50, not out of 2; so utilization would be not great AFAIU). So do lock-free versions tend to outperform locking ones in this case as well (although it's still 1 out of 50)?
Roughly I'm trying to understand the concept of 'more scalable under contention'; I'm assuming that lock-free might start to outperform locking in oversubscription scenarios;
and I'm asking whether a lock-free version tends to outperform locks with an increase of HW thread count.
P.S.: I know that measuring is the most important; this is all just handwaving; but maybe there are some general ideas around this...