The Swift compiler can optimize parts of a program by simply not executing them since these parts may not be producing results that are used elsewhere. This is quite useful almost always except when doing performance analysis. In the following code, the aim is to find the execution time of the sum() function.
import Foundation
func timeit( body: () -> () ) -> UInt64 {
let t1 = DispatchTime.now().uptimeNanoseconds
body()
let t2 = DispatchTime.now().uptimeNanoseconds
return t2-t1
}
func sum<Number>(_ x: ContiguousArray<Number> ) -> Number where Number: Numeric {
var s = Number.zero
for k in 0..<x.count {
s += x[k]
}
return s
}
let x = ContiguousArray<Int>(1...1_000_000)
var z = 0
for k in 0..<5 {
print("\(k)th run:")
let t = timeit {
z = sum(x)
}
print("\t run time: \(t)")
print("\t sum: \(z)")
}
This produces the following output:
0th run:
run time: 744485
sum: 500000500000
1th run:
run time: 724117
sum: 500000500000
2th run:
run time: 673185
sum: 500000500000
3th run:
run time: 710229
sum: 500000500000
4th run:
run time: 784819
sum: 500000500000
However, if I simply remove z to which the result was assigned as in
for k in 0..<5 {
print("\(k)th run:")
let t = timeit {
sum(x)
}
print("\t run time: \(t)")
print("\t sum: \(z)")
}
Then the following output is produced:
0th run:
run time: 7404
sum: 0
1th run:
run time: 112
sum: 0
2th run:
run time: 27
sum: 0
3th run:
run time: 26
sum: 0
4th run:
run time: 26
sum: 0
Where you can see that the run times are less than 120 nanoseconds except for the first run. These run times are dramatically different than the above case where run times are about 700,000 nanoseconds. Notice that the sum is zero because the result was never captured - here this is desired because we are not interested in the produced result, we just want to know about the execution times.
The “Whole Module” mode with -O optimization level was my building settings. When I choose the "Incremental" mode with -Onone optimization level, then there is no discrepancy of run times. However, we would like to know the run times of the optimized code.
Are we always forced to capture some results from a timed closure for being able to measure elapsed time? I am looking for a way that does not involve capturing results.