Experimenting with Python code for the sentence_transformers package, I noticed that the generated embeddings are not perfectly equal while working with really small datasets or operating in small chunks.
The actual difference is in the order of 1e-08 to 1e-10 in mean and 1e-05 in absolute sum, and seems to disappear the larger the chunks.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer('all-MiniLM-L6-v2')
sentences = [
'This framework generates embeddings for each input sentence',
'Sentences are passed as a list of string.',
'The quick brown fox jumps over the lazy dog.'
]
e1 = model.encode(sentences)
e2 = model.encode(sentences[0])
print(np.mean(e1[0] - e2)) # 1e-08
print(np.sum(np.abs(e1[0] - e2)))
# The order is also relevant
e3 = model.encode([
'This framework generates embeddings for each input sentence',
'The quick brown fox jumps over the lazy dog.'
'Sentences are passed as a list of string.',
])
print(np.mean(e1[0] - e3[0])) # 1e-10
print(np.sum(np.abs(e1[0] - e3[0])))
Working directly with the batch_size parameter, using a bigger dataset (10k sentences), the difference disappears with larger batches (the standard is 32):
b1 = model.encode(sentences, batch_size = 10)
b2 = model.encode(sentences, batch_size = 50)
b3 = model.encode(sentences, batch_size = 100)
b4 = model.encode(sentences, batch_size = 1000)
np.sum(np.abs(b1[0] - b2[0])) # Difference 1e-05
np.sum(np.abs(b2[0] - b3[0])) # No difference
np.sum(np.abs(b3[0] - b4[0])) # No difference
I noticed similar results with Hugging Face models. Are sentence embeddings replicable with small datasets, or using small batches in the encoding? Is there any hidden dependence on the whole dataset or on the chunk size for the embeddings themselves?