TL;DR
If you have just one writer and one reader threads, all you need is to use atomics with a relaxed memory ordering or acquire/release.
Details
On x86 it will be translated to a normal add/mov instructions, so there will be no performance impact.
Here is a normal counter increment:
example::normal_inc:
add dword ptr [rip + example::normal_u32], 1
ret
Here is an atomic counter increment with relaxed ordering:
example::atomic_inc:
add dword ptr [rip + example::atomic_u32], 1
ret
There is no difference on x86, hence no performance impact. But is the code correct?
Relaxed loads/stores do not guarantee the ordering across threads, but guarantee the order on the same thread and the atomicity. What does it mean?
For one writer and one reader case it means that if thread W updates the counters, thread R eventually will see the change, and the value will be valid, as the atomicity is guaranteed. For example, if the counter is 0, and thread W increases it to 1, and 2, it's guaranteed thread R eventually will see 2, and it will never see 42 or some other random number.
What is not guaranteed is that this number will by aligned with other atomic or non-atomic variables. Say, if thread W adds element to a list and then increases the counter, thread R might see those events in a reverse order, i.e. first the counter get increased, and then a new element appears in the list.
What is still guaranteed, is the order of events from point of view of thread W. Having the same example with a list, it's guaranteed that for thread W the list element will appear before the counter get increased, as all of those changes are happening within the same thread, not across different threads.
As x86 has quite strong memory ordering, even aquire/release ordering on the atomics still uses normal add/mov operations. See memory ordering on Wikipedia.
Acquire/release semantics guarantee not only the atomicity, but also the ordering. Having the example with the list, thread W adds a list element, and then releases a counter. When thread R acquires the counter, it's guaranteed that the list element is there. On x86 there is no additional cost for this guarantee.
See also the examples above on Godbolt: https://godbolt.org/z/4EsY4j