I want to find and extract lines, that contains one of 10,000+ keywords from huge text files.
Wingrep is a good tool, but its keyword length is limited and it is single core. So far I did a simple python script that goes through each line and each keyword string with the in operator. This works well, but is incredibly slow because the huge keywords number. The code looks so far like this:
data_list = [# data here] # extracted from TXT file \n character removed
keyword_list = [#keywords here] # extracted from TXT file \n character removed
result_list = [] # empty array for the results
for line in data_list:
for keyword in keyword_list:
if keyword in line:
result_list.append(line)
#save result_list to txt file
Multiple CPU core version is still too slow(6x faster for 5 more core). I have a GTX 1070 and Cuda cores could be perfect for parallelization of this. Maybe Numba could do this, but it is out of my knowledge level.
Question: Anyone can recommend a method or a code that could use the parallel power of a GPU for string matching like this?
Edit: Examples for data and keywords:
data_list = ("usa.newyork.greenstreet.12", "france.paris.blueroad.8", "italy.rome.pinksquare.2","greenland.narsaq.whiteroad.15")
keywordlist = ("green", "black", "blue")
I need both results from Country and road place.