import torch
Q = torch.rand((16,128)).cuda()
K = torch.rand((128,10)).cuda()
QK = (Q @ K)[:1]
_QK = Q[:1] @ K
error = QK - _QK
print(error.sum().item())
The output is
-0.000331878662109375
Mathematically, QK and _QK should be strictly equal, why here they are different ?