You initial code is already pretty efficient. The thing is the CPython interpreter make it slow.
Indeed, the interpreter use reference-counted variable-sized integers that are expensive to manage. Thus, -num allocates a new integer object as well as num & ... and num -= temp. This means 3 expensive allocations are done. The yield is also a quite expensive operation (it cause a low-level context-switch).
Such overhead can be mostly removed by just-in-time compilers (JIT). The PyPy general-purpose JIT-based interpreter for example is able to mostly remove the overhead of object allocations (thanks also to a fast garbage collector) though PyPy does not yet optimize the yield very well. Alternatively, Numba can be used here. Numba is a JIT compiler meant to optimize numerical codes that can be used in code executed by CPython. For example the following code is a bit faster:
import numba as nb
@nb.njit('(uint64,)')
def true_bits(num):
while num:
temp = num & -num
num -= temp
yield temp
That being said, it is limited to 64-bit integers (or smaller) like Numpy does. Cython could also help by compiling the code ahead of time using a basic compiler. This is similar to writing your own C module expect you do not need to write C code and Cython make this process much easier.
If you want to optimize the code further, then you certainly need to use such tools in the caller function so not to pay the expensive overhead of function calls from the CPython interpreter (which are at least 10 times slower than native ones).
If this is not possible (hopeless situation), you can use the following approach with Numba:
@nb.njit('(uint64,uint64[::1])')
def true_bits(num, buffer):
cur = 0
buffer.fill(0)
while num:
temp = num & -num
num -= temp
buffer[cur] = temp
cur += 1
return cur
buffer = np.empty(64, dtype=np.uint64)
written_items = true_bits(154781, buffer)
# Result stored in in buffer[:written_items]
The idea is to write the result in a pre-allocated buffer (since creating Numpy arrays is slow in this context). Then the function writes the value in the buffer when needed and returns the number of written items. You can get the actual items with buffer[:written_items] or you can iterate on the array but be aware that doing this is almost as expensive as the computation itself (again due to the CPython interpreter). Still, it is faster than the initial solution.