I wanna clean a Persian text from stop-words. I already have stop-word data that is provided on the below link. It seems to me, if I have a pre-built tree on stop-words, I could save lots of time. I want to search each word of text in this pre-built tree, if the word is in the tree I delete it from the text, if not I hold it.
O(n * l) to O(n*log(l)).
If you have better suggestions than the pre-built tree search, I would be grateful to share it with me.