Removing the non-english data

Viewed 533

I have some non-english words/sentences in my data. I tokenized my text and tried using nltk.corpus.words.words() but its not really helpful as it also removes the brand names, company names, like NLTK etc. I need some solid solution for the purpose.

Here's what I tried:

def removeNonEnglishWordsFunct(x):
    words = set(nltk.corpus.words.words())
    filteredSentence = " ".join(w for w in nltk.wordpunct_tokenize(x) \
                                if w.lower() in words or not w.isalpha())
    return filteredSentence


string = "NLTK testing man Apple Confiz Burj Al Arab Copacabana Palace Wは比較的新しくてきれいなのですが Sheraton hotelは時々 NYらしい小さくて清潔感のない部屋"

res = removeNonEnglishWordsFunct(string)
Output: testing man Apple Al Palace

Expected output: NLTK testing man Apple Confiz Burj Al Arab Copacabana Palace Sheraton hotel
2 Answers

You are kind of asking for the impossible here, you want it to be 'smart'.

We can make guesses about the kind of thing you want but all we can do is get closer to something that may be right, we will never deal with all the edge cases.

For example. Lets assume any word starting with a capital letter is an acronym:

def wordIsRomanChars(w):
    return w[0].upper() and all([ord(c) <128 or (ord(c) >= 65313 and ord(c) <= 65339) or (ord(c) >= 65345 and ord(c) <= 65371) for c in w])


def removeNonEnglishWordsFunc2(x):
    words = set(nltk.corpus.words.words())
    filteredSentence = " ".join(w for w in nltk.wordpunct_tokenize(x) \
                                if w.lower() in words or not w.isalpha() or wordIsRomanChars(w))
    return filteredSentence


string = "NLTK testing man Apple Confiz Burj Al Arab Copacabana Palace Wは比較的新しくてきれいなのですが Sheraton hotelは時々 NYらしい小さくて清潔感のない部屋"

res = removeNonEnglishWordsFunc2(string)
print(res)
"Gives: NLTK testing man Apple Confiz Burj Al Arab Copacabana Palace Sheraton"    

This is a good start but it doesn't find the 'hotel' as that is attached to none-roman characters.

We can get round this by ignoring none-roman characters:

def takeCharsUntilNotRoman(w):
    result = []
    for c in w:
        if ord(c) <128 or (ord(c) >= 65313 and ord(c) <= 65339) or (ord(c) >= 65345 and ord(c) <= 65371):
            result.append(c)
        else:
            break
    # Assume a word needs to be at least 2 chars long
    if len(result) > 1:
        return ''.join(result)
    return ''


def removeNonEnglishWordsFunct(x):
    words = set(nltk.corpus.words.words())
    filteredSentence = (takeCharsUntilNotRoman(w) for w in nltk.wordpunct_tokenize(x) \
                                if w.lower() in words or not w.isalpha() or w[0].upper())

    return ' '.join([a for a in filteredSentence if a])

res = removeNonEnglishWordsFunct(string)
print(res)
"Gives: NLTK testing man Apple Confiz Burj Al Arab Copacabana Palace Sheraton hotel NY"

This is closer to your suggested output but it has included 'NY' in the output as that was pulled out of a mixed asian-roman string. The logic could be further tweaked but it is difficult to know without knowing exactly what you need. We include 'Al' as a valid string so why is 'NY' not included?

Some other questions you will want to ask yourself are: Would we want english words and acronyms in the middle of a mixed asian-roman string instead of just the words at the beginning ?

I do not know the answer to this and you will have to tweak the above to come up with an answer that suits your case.

Hypothetically, if there is a way to show the million records on screen , would you be ok to browse through all the records manually? This is impossible, and also not needed. The best way is to build a quality check program, like scan your columns in pyspark for string patterns or special characters through regex and contains operation.. Refer this answer for an example in string matching in pyspark - pyspark query and sql pyspark query Better your query, better the results

Related