What is Causing Big Performance Differential Between Different Zen 1 Processors for a Simple Benchmark?

Viewed 117

I am investigating the performance differentials of a very simple benchmark across three different Zen 1 processors, and I am observing massive differences in instruction per cycle and L2 cache misses between the Threadripper and the Epyc processors.

Question

On Epyc processors, the instruction per cycle (IPC) is much lower when the working set size is bigger than the last level cache (LLC).

Why can the performance be so different between the same generation Epyc and Threadripper processors?

Benchmark

https://github.com/llvm/llvm-test-suite/tree/master/MultiSource/Benchmarks/Olden/treeadd

The treeadd benchmark in the Olden suite. Source available as part of the LLVM test suite.

The benchmark recursively creates a full binary tree of 2^N - 1 nodes where N is the input, and performs the traversal 100 times, performing read-only operations on each node. The benchmark is compiled using clang from an llvm trunk fork that is a few month old.

The benchmark is a single-threaded application running on a single core.

Systems Under Test

System 1

CPU: Ryzen Threadripper 1950x (https://en.wikichip.org/wiki/amd/ryzen_threadripper/1950x)
Memory: 32GB DDR4 2666
OS: Ubuntu 20.04, kernel 5.4.0

System 2

CPU: Epyc 7601 (https://en.wikichip.org/wiki/amd/epyc/7601)
Memory: 128 GB DDR4 2133
OS: Ubuntu 20.04, kernel 5.8.0

System 3

CPU: Epyc 7371 (https://en.wikichip.org/wiki/amd/epyc/7371)
Memory: 128 GB DDR4 2133
OS: Ubuntu 20.04, kernel 5.4.48

Performance Summary

When running the benchmark with N = 22, the following table is obtained. (Wall Time, IPC and Cache References/Misses):

Wall Time, IPC and Cache References/Misses

Two binaries are built, one on Threadripper (TR bin), the other on Epyc 7601 (Epyc bin).

Regardless of which binary is used, Epyc IPCs are much lower than those on Threadripper. The wall time is there for reference, but because the frequencies of the processors are different, they should probably not be directly compared. Using taskset to pin the benchmark on one core does not produce significantly different numbers.

By the looks of it, the IPC differential is caused by the dramatic differences in L2 misses (perf event l2_cache_req_stat.ls_rd_blk_c). A usual suspect in this situation is that the sizes of L2 per core are different. However, according to the documentation, all the processors have the same 512KB L2 cache per core. Furthermore, I ran lmbench and obtained the read latencies on the three systems, shown in this figure below.

Cache Read Latencies

The two Threadripper charts are identical for ease of comparison vertically and horizontally. Note that the "memory wall" hit at about the same array sizes, further indicating that these processors have identically sized caches per core. On the other hand, the access latency for different strides hints at the possibility that Threadripper may have an unusually effective prefetcher. However, turning off the prefetcher on Threadripper only doubles the L2 miss rate (seen in the first figure), and it is still no where near the L2 miss rates on the Epyc machines.

Lastly, the working set size is varied to test the effect of cache sizes. This last figure shows the IPC when the working set sizes are varied. The benchmark is altered to run 1000 iterations to make sure it runs long enough to get stable numbers.

IPC On the Three Machines with Different Working Set Size

Although the Epyc machines seem to have 8MB LLC (per CCX of 4 cores), the performance falls off a cliff when the working set size goes from 3MB to 6MB.

Different Ways of Running the Benchmark

Other than simply running the benchmark, I tired two other ways.

  • Using taskset to pin the benchmark to a specific core.
  • using numactl to pin the benchmark to a specific core, and use the membind flag to use the memory closest to the core.

Neither produced significantly different profiles.

Guesses

Evidences indicate that the sizes of the caches on the processors are the same. Hence the massive difference in L2 miss rate may be due to cache partition of L3. Maybe Epyc in some situations limit the per core L3 to 4MB?

Summary

Significant IPC differences are observed for the same benchmark on Threadripper and Epyc machines. The difference seems to be caused by the unusually high L2 misses on the Epyc processors. However, the caches have the same sizes per core.

What could have caused the massive differences in L2 misses? Or maybe there is something else inducing the IPC difference?

0 Answers
Related