Say you have a folder with hundreds or thousands of .csv or .txt files that presumably contain different information, but you want to make sure that joe041.txt doesn't actually contain the same data as joe526.txt by accident.
Rather than loading everything into one file – which could be troublesome if each file has thousands of lines –, I've taken to using a Python script to essentially read each file in the directory and calculate a checksum that you can then compare between your thousands of files.
Is there a more efficient way to do this?
Even using filecmp for this seems to less efficient since the module only has file vs file and dir vs dir comparisons but no file vs dir commands – which means that to use it you'd have to iterate through x² times (all files in dir against all other files in dir).
import os
import hashlib
outputfile = []
for x in(os.listdir("D:/Testing/New folder")):
with open("D:/Testing/New folder/%s" % x, "rb") as openfile:
text=openfile.read()
outputfile.append(x)
outputfile.append(",")
outputfile.append(hashlib.md5(text).hexdigest())
outputfile.append("\n")
print(outputfile)
with open("D:/Testing/New folder/output.csv","w") as openfile:
for x in outputfile:
openfile.write(x)