How to split Chinese char array by a specific delimiter `……`(non-single chars) in standard c or c++?

Viewed 83

When I tried to split a char array in standard c, the problem is that it cannot show full char (only return \n发\n办法 in this example) when input is chinese char with …… just like, 印发……办法 .However, it is okay if the input is abc……def or 印发...办法. Why and how to solve this problem? Any solution in c or c++ is helpful!

#pragma warning(disable:4996)
#include <string.h>
#include <stdio.h>
void split(char* str)
{
    char* token;
    const char delim[] = "……";
    token = strtok(str, delim); //it is a c++ method
    while (token != NULL)
    {
        printf("%s\n", token);
        token = strtok(NULL, delim);
    }
}

int main()
{
    char ipt1[] = "印发……办法";
    split(ipt1);
}
1 Answers

UTF-8 or other multibyte encoding represent ideographs, or ideograms, as sequence of multiple bytes. A single Chinese ideograph consists of multiple chars in UTF-8, for a single ideograph.

strtok doesn't know anything about multibyte ideographs. It recognizes delimiters as single chars. The second parameter to strtok is a character string, and every individual char value in it gets recognized as a delimiter.

The character …, encoded in UTF-8 is three chars:

E2 80 A6

Any one of those individual chars will be recognized by strtok as a valid delimiter for the string to be tokenized. These values will occur as part of other Chinese ideographs, resulting in strtok making mincemeat of the string that gets passed in for tokenization. strtok does not work with multibyte encodings.

If you need to implement this kind of tokenization using basic functions from the C library then the closest match would be strstr, which works in a completely different way. You'll need to reimplement this tokenization algorithm based on strstr.

Related