I have a single atomic variable that multiple threads are loading, they perform some local calculations on it, then call an atomically fetch_and on it. They check that they were able to make there change before another thread did, if not, repeat using the updated value returned from fetch_and
Works a lot faster than a locked version. But would be nice if I could encourage multiple threads to align and not load the atomic number until fetch_and completes without forcing it.
Is this possible? Thinking it might be using a memory fence or two?