I'd like to make MySQL full text search work with Japanese and Chinese text, as well as any other language. The problem is that these languages and probably others do not normally have white space between words. Search is not useful when you must type the same sentence as is in the text.
I can not just put a space between every character because English must work too. I would like to solve this problem with PHP or MySQL.
Can I configure MySQL to recognize characters which should be their own indexing units? Is there a PHP module that can recognize these characters so I could just throw spaces around them for the index?
Update
A partial solution:
$string_with_spaces =
preg_replace( "/[".json_decode('"\u4e00"')."-".json_decode('"\uface"')."]/",
" $0 ", $string_without_spaces );
This makes a character class out of at least some of the characters I need to treat specially. I should probably mention, it is acceptable to munge the indexed text.
Does anyone know all the ranges of characters I'd need to insert spaces around?
Also, there must be a better, portable way to represent those characters in PHP? Source code in Literal Unicode is not ideal; I will not recognize all the characters; they may not render on all the machines I have to use.