I remember that back in the days, when developing code for the first iPhone, XCode had a profiling tool called shark. While shark IIRC could do regular profiling using time sampling the code, there was one feature that was incredibly helpful when i was performance optimizing the hot bottleneck of my code.
The tool would show the assembly instructions line by line, and for each line show how many cpu cycles it would require to execute as well as how long it would have to stall for the previous instructions.
The cycles for instructions are shown as X:Y, with X being the overal number of cpu cycles and Y how many cycles it would require until the next instruction can be executed (so long as it doesn't depend on the result of this instruction). The 'Stall' shows how long the execution of the next line is stalled because it depends on the result of a previous instruction.
This allowed me to restructure my already heavily optimized code to make it again two or three times faster by pipelining the instructions optimally and hiding all the latencies with instructions.
Is there still such a tool that can do that? It's clear that the cycle timings depend on the particular cpu that executes it, but I guess such a tool would allow to select the architecture to show the timings for, or be a specific tool for a certain architecture (in my case I'm mainly interested in optimizing for Intel Xeon SP 1 and 2)
//Edit: while it's clear that modern cpu's are quite complex (being able to execute instructions out of order or having multiple execution units that can run in parallel) such analysis is still possible though and there are such tables for the instruction latencies (and accumulated instruction latency) for various architectures: https://www.agner.org/optimize/instruction_tables.pdf
