Measure performance for small application

Viewed 210

I wrote a statusbar for Sway WM (https://swaywm.org), basically print every second some information (loadavg, cpu usage, memory usage, time, ...).

The project (https://gitlab.com/Yellowhat/statusbar) is purely to learn about CI/CD and how to measure performance (if possible in a noisy environment).

Initially the project was all written in python but some time ago I decided to write a functionally equivalent in rust (https://gitlab.com/Yellowhat/statusbar/-/tree/master/tools/rust) (they print out exactly the same information).

Looking at the number of instructions run (I refs using cachegrind, slimmed version) per second of execution:

  • python: ~800k
  • rust: ~35k

I could conclude that the rust version is much "lighter" than the python one.

As the statusbar will run continuosly while the system is running, I want to understand the impact that it has on my system. So I decided to monitor my system using the following python script while idling (do not touch the computer for 1 hour, 100% brightness):

from time import sleep

while True:
    stat = open("/proc/stat").readlines()[0].split()[1:8]
    power = float(open("/sys/class/power_supply/BAT0/power_now").read().strip()) / 10**6
    print(f"{stat},{power}")
    sleep(1)
  • /proc/stat returns how much CPU time (jiffies) is spent in:
    • user: normal processes executing in user mode
    • nice: niced processes executing in user mode
    • system: processes executing in kernel mode
    • idle: twiddling thumbs
    • iowait: waiting for I/O to complete
    • irq: servicing interrupts
    • softirq: servicing softirqs
  • /sys/class/power_supply/BAT0/power_now returns how much power (micro Watt) is taken from the battery

I have compared 4 cases:

  1. tty: boot the computer and run the script from tty
  2. sway: boot the computer, start sway (without a bar), open the terminal and run the script
  3. sway barpy: boot the computer, start sway (with the python version), open the terminal and run the script
  4. sway barrs: boot the computer, start sway (with the rust version), open the terminal and run the script

The following graph shows the SUM(user_t, nice_t, system_t, iowait_t, irq_t, softirq_t) - SUM(user_t0, nice_t0, system_t0, iowait_t0, irq_t0, softirq_t0), (how much time was spent by all CPUs not idling) - (the initial value at the beginning of the test), versus time for the 4 cases:

CPU Time

As you can see, tty is the lowest, followed by sway, but interestingly the rust version is much higher than the python one.

The following graph shows the power versus time for the 4 cases:

Power

As you can see, tty has the lowest consumption, followed by sway, very close by rust and python at last. This is the kind of result I would expect.

The first graph would suggest that, in this instance, rust is not lighter than python.

Are there more reliable way to measure "light" applications? Am I misunderstanding the /proc/stat content?

Thanks

Update 1

The following graph shows the user_t - user_t0, nice_t - nice_t0, system_t - system_t0 versus time for python and rust:

User, Nice, System

user are similar instead system and nice are much higher for rust.

Update 2

As suggested by Peter Cordes I have run:

perf stat -a -d -e cpu-cycles,cycles,cycles:u,instructions,instructions:u -r 10 <binary>

The following tables summarise the results:

Version Duration cpu-cycles cycles cycles:u instructions instructions:u
python 0 285,248,768 285,492,709 215,681,759 323,852,866 278,792,657
python 60 3,038,749,610 3,046,770,477 1,360,477,472 2,425,399,588 1,612,134,730
python 120 5,802,965,874 5,818,929,489 2,496,192,062 4,536,305,443 2,890,941,664
rust 0 1,165,223 1,165,869 304,188 1,319,168 442,565
rust 60 3,791,712,516 3,799,515,523 1,654,232,537 3,355,917,423 2,073,096,815
rust 120 7,878,549,570 7,897,534,278 3,341,698,282 6,665,643,064 4,144,148,632

Removing the initial value (duration = 0) and dividing by the duration:

Version Duration cpu-cycles cycles cycles:u instructions instructions:u
python 60 45,891,681 46,021,296 19,079,929 35,025,779 22,222,368
python 120 45,980,976 46,111,973 19,004,253 35,103,771 21,767,908
rust 60 63,175,788 63,305,828 27,565,472 55,909,971 34,544,238
rust 120 65,644,870 65,803,070 27,844,951 55,536,032 34,530,884

Again, seems that rust version is running more cycles/instructions than the python one.

Update 3

$ valgrind \
    --tool=cachegrind \
    --cachegrind-out-file=/dev/null \
    --trace-children=yes \
    --I1=32768,8,64 \
    --D1=32768,8,64 \
    --LL=8388608,16,64 \
    --cache-sim=yes \
    --branch-sim=yes \
    <binary>

Run for 120 seconds:

  • python version:
I   refs:      328,117,776
I1  misses:      5,268,249
LLi misses:         15,274
I1  miss rate:        1.61%
LLi miss rate:        0.00%

D   refs:      135,851,173  (95,116,936 rd   + 40,734,237 wr)
D1  misses:      4,717,331  ( 4,145,263 rd   +    572,068 wr)
LLd misses:        167,298  (    57,266 rd   +    110,032 wr)
D1  miss rate:         3.5% (       4.4%     +        1.4%  )
LLd miss rate:         0.1% (       0.1%     +        0.3%  )

LL refs:         9,985,580  ( 9,413,512 rd   +    572,068 wr)
LL misses:         182,572  (    72,540 rd   +    110,032 wr)
LL miss rate:          0.0% (       0.0%     +        0.3%  )

Branches:       60,115,083  (56,103,351 cond +  4,011,732 ind)
Mispredicts:     5,661,255  ( 4,245,148 cond +  1,416,107 ind)
Mispred rate:          9.4% (       7.6%     +       35.3%   )
  • rust version:
I   refs:      100,950,027
I1  misses:        351,835
LLi misses:          5,859
I1  miss rate:        0.35%
LLi miss rate:        0.01%

D   refs:       44,227,512  (24,903,985 rd   + 19,323,527 wr)
D1  misses:        670,307  (   341,521 rd   +    328,786 wr)
LLd misses:         34,962  (    20,657 rd   +     14,305 wr)
D1  miss rate:         1.5% (       1.4%     +        1.7%  )
LLd miss rate:         0.1% (       0.1%     +        0.1%  )

LL refs:         1,022,142  (   693,356 rd   +    328,786 wr)
LL misses:          40,821  (    26,516 rd   +     14,305 wr)
LL miss rate:          0.0% (       0.0%     +        0.1%  )

Branches:       19,868,173  (18,608,630 cond +  1,259,543 ind)
Mispredicts:     1,289,209  (   771,555 cond +    517,654 ind)
Mispred rate:          6.5% (       4.1%     +       41.1%   )
0 Answers
Related