If your arrays are usually aligned by 16, it's probably best not to do even more checking to look for the odd-start case, just use your unaligned version unless it's a lot worse for some reason.
However, if they're usually aligned by 8 (but unknown whether they're aligned by 16), then you may be able to get away with only checking for alignment by 8 and branchlessly handling the maybe-unaligned first iteration for the aligned case, see below. (Otherwise just fall back to your fully unaligned case.)
If overlap isn't a problem (e.g. c[] = a[]+b[], or a memset-like store or whatever), a good technique is to always do a first vector with unaligned load/store, then advance to the first aligned vector (add rdi, 16 / and rdi, -16). If the input was aligned, this won't overlap. Otherwise, it partially overlaps and the store buffer + L1d cache handle it efficiently.
This keeps the cost minimal for the aligned case, and avoids the chance of branch mispredicts.
Rounding a pointer up/down to an alignment boundary is cheap, just an and, but you do have the code-size cost of peeling a whole copy of the loop body. So it's not totally free as far as startup overhead, but at least this kind of startup overhead can overlap with a cache miss from the data.
But note that a lot of SIMD functions have multiple pointer inputs that can be misaligned relative to each other. In that case, the standard advice is to align the output and keep using movups for inputs. Although if the front-end is the bottleneck, you might choose to reach an alignment boundary for the input so you can fold a memory source operand into an ALU instruction like xorps xmm0, [rdi] and use movups store.
But if anything other than the front-end, e.g. cache or memory throughput, are a bottleneck, then you more often want to align the destination. Intel's optimization manual has some advice about this. Some of the reason is that load throughput is typically 2x store throughput (until IceLake), so the load hardware can more readily absorb the extra work for split loads. Also, storing a full cache line with fewer stores can help reduce cases where a line gets evicted (written back) but then you store to it again and it has to get fetched + dirtied and eventually written-back again, instead of just fetched.