I have seen this discussed as one of the potential pitfalls in microbenchmarking. If you specify that @Measurement (or @Warmup) will run for a fixed amount of time, it means that, when comparing different runs (e.g., different platforms, different versions of the VM, etc.), you will get less of an "apples to apples" comparison because the runs will not be doing the same work!
If one run executes faster, then it will go through more operations in a given time. This introduces confounding factors that may skew your results: more opportunities for dynamic optimization, differences in caching, different statistics per operation, etc.
On the other hand, if you specify a fixed number of operations, then each run is doing exactly the same quantity of work, which improves confidence in the results.
So, what is the reason that JMH utilizes this method of specifying a fixed amount of time rather than a fixed number of operations? Is there some important reason driving this design decision that I am missing? I have searched both within the JMH documentation and in online discussions, but I have not been able to find an answer to this question.