I have written an app to emulate a classic computer. Despite being on the App Store for a couple of years, I have regularly tried to reduce the demand on CPU cores through testing with Time Profiler in Instruments. When comparing the results between real devices with significantly different specifications the CPU % utilisation shows reverse trends.
The annotated Xcode screenshots show the contrasting device specifications and CPU usage contradiction. At the time of writing, Xcode 10.2.1 is used and both devices have iOS 12.2.1 installed. Compile optimisations are applied even when running in debug mode. The same trend can be seen between other devices. Time Profiler shows the same percentages as Xcode. Although, interestingly when using the File > Recording Options… > Record Waiting Threads, then the iPad Mini 2 device drops to ~22% and the iPhone XS Max drops to ~28%.
Implementation details:
The app has two concurrent process threads for two distinct tasks:
- CPU simulation thread - processing the emulated computer instructions
- CRT display simulation thread - processing the raw emulated video signals and turning them into vectors graphics
To avoid the expensive overhead of repeatedly creating the two processes when there is work for a task, despatch semaphores are used to control when the processes sleep. Compile optimisations are applied even when running in debug mode.
Stripped back example code:
This code below demonstrates some of the principles for the purpose of this post. On my test devices the CPU usage % difference is not as pronounced but still contradictory as both the iPad Mini 2 and iPhone XS Max devices report ~120%, where I should expect the the more modern iPhone device to be a significantly lower value.
When recording waiting threads again the values are lower but this time more in line with the generation of the device, iPad Mini 2 = ~48% vs iPhone XS Max = ~35%. Again, this still does match my expectations given the difference in their processors.
Every time this demo code is run the average results can deviate for no obvious reason by at least as much as 5%. Which makes me doubt the general accuracy of the CPU usage %.
final class ViewController: UIViewController {
let processorDispatchSemaphore = DispatchSemaphore(value: 0)
let videoDispatchSemaphore = DispatchSemaphore(value: 0)
fileprivate var stopEmulation = false
fileprivate var lastTime: CFTimeInterval = 0.0
fileprivate var accumulatedCycles = 0
final var pretendVideoData: [Int] = []
final var pretendDisplayData: [Int] = []
override func viewDidLoad() {
super.viewDidLoad()
let displayLink = CADisplayLink(target: self, selector: #selector(displayUpdate))
displayLink.add(to: .main, forMode: RunLoop.Mode.common)
let concurrentEmulationQueue = DispatchQueue.global(qos: .userInteractive)
// CPU simulation thread
concurrentEmulationQueue.async() {
repeat {
// pause until a display refresh
self.processorDispatchSemaphore.wait()
// calculate the number of simulated computer clock
// clock cycles that would have been executed in the
// same time
let currentTime = displayLink.timestamp
let delta: CFTimeInterval = currentTime - self.lastTime
self.lastTime = currentTime
// Z80A Microprocessor clocked at 3.25MHz = 3,250,000 per second
// 1 second / 3250000 = 0.000000307692308
var emulationCyclesRequired = Int((delta / 0.000000307692308).rounded())
// safeguard:
// Time delay every 1/60th (0.0166667) of a second
// 0.0166667 / 0.000000307692308 = 54167 cycles
// let's say that no more than 3 times that should
// be allowed = 54167 * 3 = 162501
if emulationCyclesRequired > 162501 {
// even on slow devices the thread only need
// cap cycles whilst the CADisplayLink takes
// time to kick - so after a less second the
// app need not apply this safeguard
emulationCyclesRequired = 162501
print("emulation cycles capped")
}
// do some simulated work
// **** fake process filling code ****
for cycle in 0...emulationCyclesRequired {
if cycle % 4 == 0 {
self.pretendVideoData.append(cycle &+ cycle)
}
self.accumulatedCycles = self.accumulatedCycles &+ 1
if self.accumulatedCycles > 40000 {
// unpause the CRT display simulation thread
self.videoDispatchSemaphore.signal()
self.pretendVideoData.removeAll(keepingCapacity: true)
}
}
// **** **** ****
// thread is allowed to finish when app goes to the
// background or a non-sumiulation screen.
} while !self.stopEmulation
}
let concurrentDisplayQueue = DispatchQueue.global(qos: .userInteractive)
// CRT display simulation thread
// (edit) see comment to Rob - concurrentEmulationQueue.async(flags: .barrier) {
concurrentDisplayQueue.async(flags: .barrier) {
repeat {
self.videoDispatchSemaphore.wait()
// do some simulated work
// **** fake process filling code ****
for index in 0...1000 {
self.pretendDisplayData.append(~index)
}
self.pretendDisplayData.removeAll(keepingCapacity: true)
// **** **** ****
// thread is allowed to finish when app goes to the
// background or a non-sumiulation screen.
} while !self.stopEmulation
}
}
@objc fileprivate func displayUpdate() {
// unpause the CPU simulation thread
processorDispatchSemaphore.signal()
}
}
Questions:
- Why might the CPU usage % be higher for devices with faster CPUs? Any reason to think the results are not accurate?
- How could I better interpret the figures or get better benchmarks between devices?
- Why does Record Waiting Threads result in lower CPU usage percentages (but still not significantly different and sometimes higher for the faster device)?




