I wrote a statusbar for Sway WM (https://swaywm.org), basically print every second some information (loadavg, cpu usage, memory usage, time, ...).
The project (https://gitlab.com/Yellowhat/statusbar) is purely to learn about CI/CD and how to measure performance (if possible in a noisy environment).
Initially the project was all written in python but some time ago I decided to write a functionally equivalent in rust (https://gitlab.com/Yellowhat/statusbar/-/tree/master/tools/rust) (they print out exactly the same information).
Looking at the number of instructions run (I refs using cachegrind, slimmed version) per second of execution:
python: ~800krust: ~35k
I could conclude that the rust version is much "lighter" than the python one.
As the statusbar will run continuosly while the system is running, I want to understand the impact that it has on my system. So I decided to monitor my system using the following python script while idling (do not touch the computer for 1 hour, 100% brightness):
from time import sleep
while True:
stat = open("/proc/stat").readlines()[0].split()[1:8]
power = float(open("/sys/class/power_supply/BAT0/power_now").read().strip()) / 10**6
print(f"{stat},{power}")
sleep(1)
/proc/statreturns how much CPU time (jiffies) is spent in:user: normal processes executing in user modenice: niced processes executing in user modesystem: processes executing in kernel modeidle: twiddling thumbsiowait: waiting for I/O to completeirq: servicing interruptssoftirq: servicing softirqs
/sys/class/power_supply/BAT0/power_nowreturns how much power (micro Watt) is taken from the battery
I have compared 4 cases:
- tty: boot the computer and run the script from
tty - sway: boot the computer, start
sway(without a bar), open the terminal and run the script - sway barpy: boot the computer, start
sway(with thepythonversion), open the terminal and run the script - sway barrs: boot the computer, start
sway(with therustversion), open the terminal and run the script
The following graph shows the SUM(user_t, nice_t, system_t, iowait_t, irq_t, softirq_t) - SUM(user_t0, nice_t0, system_t0, iowait_t0, irq_t0, softirq_t0), (how much time was spent by all CPUs not idling) - (the initial value at the beginning of the test), versus time for the 4 cases:
As you can see, tty is the lowest, followed by sway, but interestingly the rust version is much higher than the python one.
The following graph shows the power versus time for the 4 cases:
As you can see, tty has the lowest consumption, followed by sway, very close by rust and python at last. This is the kind of result I would expect.
The first graph would suggest that, in this instance, rust is not lighter than python.
Are there more reliable way to measure "light" applications?
Am I misunderstanding the /proc/stat content?
Thanks
Update 1
The following graph shows the user_t - user_t0, nice_t - nice_t0, system_t - system_t0 versus time for python and rust:
user are similar instead system and nice are much higher for rust.
Update 2
As suggested by Peter Cordes I have run:
perf stat -a -d -e cpu-cycles,cycles,cycles:u,instructions,instructions:u -r 10 <binary>
The following tables summarise the results:
| Version | Duration | cpu-cycles | cycles | cycles:u | instructions | instructions:u |
|---|---|---|---|---|---|---|
| python | 0 | 285,248,768 | 285,492,709 | 215,681,759 | 323,852,866 | 278,792,657 |
| python | 60 | 3,038,749,610 | 3,046,770,477 | 1,360,477,472 | 2,425,399,588 | 1,612,134,730 |
| python | 120 | 5,802,965,874 | 5,818,929,489 | 2,496,192,062 | 4,536,305,443 | 2,890,941,664 |
| rust | 0 | 1,165,223 | 1,165,869 | 304,188 | 1,319,168 | 442,565 |
| rust | 60 | 3,791,712,516 | 3,799,515,523 | 1,654,232,537 | 3,355,917,423 | 2,073,096,815 |
| rust | 120 | 7,878,549,570 | 7,897,534,278 | 3,341,698,282 | 6,665,643,064 | 4,144,148,632 |
Removing the initial value (duration = 0) and dividing by the duration:
| Version | Duration | cpu-cycles | cycles | cycles:u | instructions | instructions:u |
|---|---|---|---|---|---|---|
| python | 60 | 45,891,681 | 46,021,296 | 19,079,929 | 35,025,779 | 22,222,368 |
| python | 120 | 45,980,976 | 46,111,973 | 19,004,253 | 35,103,771 | 21,767,908 |
| rust | 60 | 63,175,788 | 63,305,828 | 27,565,472 | 55,909,971 | 34,544,238 |
| rust | 120 | 65,644,870 | 65,803,070 | 27,844,951 | 55,536,032 | 34,530,884 |
Again, seems that rust version is running more cycles/instructions than the python one.
Update 3
$ valgrind \
--tool=cachegrind \
--cachegrind-out-file=/dev/null \
--trace-children=yes \
--I1=32768,8,64 \
--D1=32768,8,64 \
--LL=8388608,16,64 \
--cache-sim=yes \
--branch-sim=yes \
<binary>
Run for 120 seconds:
pythonversion:
I refs: 328,117,776
I1 misses: 5,268,249
LLi misses: 15,274
I1 miss rate: 1.61%
LLi miss rate: 0.00%
D refs: 135,851,173 (95,116,936 rd + 40,734,237 wr)
D1 misses: 4,717,331 ( 4,145,263 rd + 572,068 wr)
LLd misses: 167,298 ( 57,266 rd + 110,032 wr)
D1 miss rate: 3.5% ( 4.4% + 1.4% )
LLd miss rate: 0.1% ( 0.1% + 0.3% )
LL refs: 9,985,580 ( 9,413,512 rd + 572,068 wr)
LL misses: 182,572 ( 72,540 rd + 110,032 wr)
LL miss rate: 0.0% ( 0.0% + 0.3% )
Branches: 60,115,083 (56,103,351 cond + 4,011,732 ind)
Mispredicts: 5,661,255 ( 4,245,148 cond + 1,416,107 ind)
Mispred rate: 9.4% ( 7.6% + 35.3% )
rustversion:
I refs: 100,950,027
I1 misses: 351,835
LLi misses: 5,859
I1 miss rate: 0.35%
LLi miss rate: 0.01%
D refs: 44,227,512 (24,903,985 rd + 19,323,527 wr)
D1 misses: 670,307 ( 341,521 rd + 328,786 wr)
LLd misses: 34,962 ( 20,657 rd + 14,305 wr)
D1 miss rate: 1.5% ( 1.4% + 1.7% )
LLd miss rate: 0.1% ( 0.1% + 0.1% )
LL refs: 1,022,142 ( 693,356 rd + 328,786 wr)
LL misses: 40,821 ( 26,516 rd + 14,305 wr)
LL miss rate: 0.0% ( 0.0% + 0.1% )
Branches: 19,868,173 (18,608,630 cond + 1,259,543 ind)
Mispredicts: 1,289,209 ( 771,555 cond + 517,654 ind)
Mispred rate: 6.5% ( 4.1% + 41.1% )


