You are kind of asking for the impossible here, you want it to be 'smart'.
We can make guesses about the kind of thing you want but all we can do is get closer to something that may be right, we will never deal with all the edge cases.
For example. Lets assume any word starting with a capital letter is an acronym:
def wordIsRomanChars(w):
return w[0].upper() and all([ord(c) <128 or (ord(c) >= 65313 and ord(c) <= 65339) or (ord(c) >= 65345 and ord(c) <= 65371) for c in w])
def removeNonEnglishWordsFunc2(x):
words = set(nltk.corpus.words.words())
filteredSentence = " ".join(w for w in nltk.wordpunct_tokenize(x) \
if w.lower() in words or not w.isalpha() or wordIsRomanChars(w))
return filteredSentence
string = "NLTK testing man Apple Confiz Burj Al Arab Copacabana Palace Wは比較的新しくてきれいなのですが Sheraton hotelは時々 NYらしい小さくて清潔感のない部屋"
res = removeNonEnglishWordsFunc2(string)
print(res)
"Gives: NLTK testing man Apple Confiz Burj Al Arab Copacabana Palace Sheraton"
This is a good start but it doesn't find the 'hotel' as that is attached to none-roman characters.
We can get round this by ignoring none-roman characters:
def takeCharsUntilNotRoman(w):
result = []
for c in w:
if ord(c) <128 or (ord(c) >= 65313 and ord(c) <= 65339) or (ord(c) >= 65345 and ord(c) <= 65371):
result.append(c)
else:
break
# Assume a word needs to be at least 2 chars long
if len(result) > 1:
return ''.join(result)
return ''
def removeNonEnglishWordsFunct(x):
words = set(nltk.corpus.words.words())
filteredSentence = (takeCharsUntilNotRoman(w) for w in nltk.wordpunct_tokenize(x) \
if w.lower() in words or not w.isalpha() or w[0].upper())
return ' '.join([a for a in filteredSentence if a])
res = removeNonEnglishWordsFunct(string)
print(res)
"Gives: NLTK testing man Apple Confiz Burj Al Arab Copacabana Palace Sheraton hotel NY"
This is closer to your suggested output but it has included 'NY' in the output as that was pulled out of a mixed asian-roman string. The logic could be further tweaked but it is difficult to know without knowing exactly what you need. We include 'Al' as a valid string so why is 'NY' not included?
Some other questions you will want to ask yourself are: Would we want english words and acronyms in the middle of a mixed asian-roman string instead of just the words at the beginning ?
I do not know the answer to this and you will have to tweak the above to come up with an answer that suits your case.