Within a single thread
The CPU has to preserve the illusion of instructions executing one at a time, in program order, for a single thread. This is the cardinal rule of OoO exec. This means tracking what program order was, and making sure loads always see values consistent with that, and that values eventually written to cache are also consistent.
This is very much like C++'s "as-if" rule, just with different observables that need to be preserved. (C++ is very restrictive in what other threads are legally allowed to observe, unlike CPU ISAs, but neither compile-time nor run-time memory-reordering can be explained by reordering source lines1)
Loads by this core snoop the store buffer, forwarding data from it if the load is reloading a store that hasn't committed yet.
And for any individual memory location, making sure its modification order matches program order, i.e. not reordering stores to the same location. So the final value after the dust settles is the last one in program order. And even observation by other threads will see a consistent modification order for that location; that's why std::atomic is able to provide the guarantee that a modification order exists for every object separately, not having extra changes to A then B then back to A if program order stored B then A. ISO C++ can guarantee this because all real-world CPUs also guarantee it.
A system call like munmap is a special case, but otherwise new/delete (and malloc/free) aren't special as far as the CPU is concerned: putting a block on the free list and having other code allocate it is just another case of messing around with pointer-based data structures. As always, the CPU tracks any reordering its doing to make sure loads see correct values.
Reuse by another thread
You're not wrong to worry about this; correctness doesn't happen for free here based on CPU architecture alone; a buggy libc could get this wrong and allow exactly the problems you describe. @ixSci's answer quotes the relevant part of the C++ standard. (Compile-time ordering of memory access wrt. calls to new/delete is also necessary, but that always has to happen for any non-inline function call that the compiler doesn't know is "pure"; any function might read or write memory so it has to be in sync.)
If the memory is placed on a global free-list that could be reused by another thread, a thread-safe allocator will have used sufficient synchronization to create a C++ happens-before relationship between the code that previously used then deleted the memory, and the code in another thread that just allocated it.
So any old-thread stores into this memory block will already be visible to the thread that just allocated the memory. So they won't step on its stores. If the new thread passes a pointer to this memory to a 3rd thread, it had better use acq/rel or consume/release synchronization itself to make sure that 3rd thread sees its stores, not still stores from the first thread.
Unmapping entirely so access to that virtual address faults
If the free involves a munmap that uses a syscall instruction to run kernel code that changes page tables (to invalidate a mapping so loads/stores to it would fault), that itself will provide sufficient serialization. Existing CPUs don't rename the privilege level, so they don't do out-of-order exec into the kernel through a syscall instruction.
It's up to the OS to do sufficient memory-barriering around modifying page tables, although on x86-64 invlpg is already a serializing instruction. (In x86 terminology, that means draining the ROB and store buffer, so all previous instructions are fully done executing with their results written back to L1d cache (for stores).) So there's no possibility of it reordering with earlier loads / stores that depend on that TLB entry, even apart from the switch to kernel mode.
(Switching into kernel mode doesn't necessarily drain the store buffer, though; the physical address of those stores are known. The TLB checks were done as the store-address uops were executed. So changes to the page tables don't affect the process of committing them to memory.)
Footnote 1: memory reordering isn't source reordering
BTW, memory reordering doesn't work like reordering statements in the C++ source or instructions in the asm machine code; memory reordering is about what other threads can observe as loads read from cache and stores eventually commit to cache at the far end of the store buffer. Reordering the source to try to explain this break the code, violating the as-if rule, but memory-reordering can produce such effects while still having the thread's operations see correct values for its own stores, e.g. by store-forwarding. That's because real-world ISAs don't have sequentially consistent memory models; you need extra ordering to recover SC. Even an in-order CPU pipeline can reorder loads with a cache that can hit-under-miss, for example, and even strongly-ordered x86 allows StoreLoad reordering: its memory model is basically program-order plus a store buffer with store-forwarding.
(There was discussion in comments about compile-time reordering and source ordering; the question didn't have this misconception.)
The C++ as-if rule is the same idea that CPUs follow as they execute, just that the ISA's rules are what govern the requirements on external observables. No ISA has memory-ordering rules as weak as ISO C++, e.g. they all guarantee a coherent shared cache, and many CPU ISAs don't have UB. (Although some do, e.g. calling it "unpredictable" behaviour. Much more often just an unpredictable or undefined result in some register; user/supervisor privilege separation requires there be limits on what behaviour is possible so user-space can't run some unsupported instruction sequence and maybe take over or crash the whole machine.)
Fun fact: on strongly-ordered x86 specifically, store and load ordering need to be more closely tied together than most ISAs; Intel calls the combination of store buffer + load buffer the Memory Order Buffer, because it also has to detect cases where a load took a value early, before it was architecturally allowed to (LoadLoad ordering), but then it turns out this core lost access to the cache line. Or in case of mis-speculation about store-forwarding, e.g. dynamically predicting that a load would be reloading a store from an unknown address, but then it turns out the store was non-overlapping. In either case, the CPU rewinds the out-of-order back-end back to a consistent retirement state. (This is called a pipeline nuke; this specific cause is counted by the machine_clears.memory_ordering perf event.)