How to do a Python split() on languages (like Chinese) that don't use whitespace as word separator?

Viewed 19149

I want to split a sentence into a list of words.

For English and European languages this is easy, just use split()

>>> "This is a sentence.".split()
['This', 'is', 'a', 'sentence.']

But I also need to deal with sentences in languages such as Chinese that don't use whitespace as word separator.

>>> u"这是一个句子".split()
[u'\u8fd9\u662f\u4e00\u4e2a\u53e5\u5b50']

Obviously that doesn't work.

How do I split such a sentence into a list of words?

UPDATE:

So far the answers seem to suggest that this requires natural language processing techniques and that the word boundaries in Chinese are ambiguous. I'm not sure I understand why. The word boundaries in Chinese seem very definite to me. Each Chinese word/character has a corresponding unicode and is displayed on screen as an separate word/character.

So where does the ambiguity come from. As you can see in my Python console output Python has no problem telling that my example sentence is made up of 5 characters:

这 - u8fd9
是 - u662f
一 - u4e00
个 - u4e2a
句 - u53e5
子 - u5b50

So obviously Python has no problem telling the word/character boundaries. I just need those words/characters in a list.

9 Answers

Best tokenizer tool for Chinese is pynlpir.

import pynlpir
pynlpir.open()
mystring = "你汉语说的很好!"
tokenized_string = pynlpir.segment(mystring, pos_tagging=False)

>>> tokenized_string
['你', '汉语', '说', '的', '很', '好', '!']

Be aware of the fact that pynlpir has a notorious but easy fixable problem with licensing, on which you can find plenty of solutions on the internet. You simply need to replace the NLPIR.user file in your NLPIR folder downloading a valide licence from this repository and restart your environment.

if str longer than 30 then take 27 chars and add '...' at the end
otherwise return str

str='中文2018-2020年一区6、8、10、12号楼_「工程建设文档102332号」'
result = len(list(str)) >= 30 and ''.join(list(str)[:27]) + '...' or str
Related