I'm dealing with Persian language poems (The language alphabet is almost the same as Arabic). For each line of the poem in my file, I want to parse its words and hold them in another list as monolithic words. The problem is that some words are separated by space, which I can easily handle with split(), but very few are separated by half-space or - \u200c.
For example, this is a string in the Persian language:
s = "سنگیتری"
The first word is "سنگی" and the second word is "تری". I wanted to separate each of them, but my problem is that I don't know how, and if I use s.split(), I get ['سنگی\u200cتری'] which is one word and also has \u200c, which should not be. (The two words in s are separated by \u200c instead of space and this is where the problem arises).
I should also repeat that I need words that are separated by space to be parsed too. So if it was s = "سنگی تری" (this time separated by space), I also need to handle it and parse it into "سنگی" and "تری". As I said, the latter is achievable by split() method.