Is the POSIX `[:print:]` character class the same as Unicode `\P{C}`?

Viewed 117

Ruby's \p{print} character class in regex does not seem to match JavaScript's \P{C} class. For example, Ruby handles U+00AD (SOFT HYPHEN) as printable:

0x00ad.chr(Encoding::UTF_8).match?(/\p{print}/) # => true

whereas Node.js handles the character as not printable

/\P{C}/u.test(String.fromCharCode(0x00ad)) // false

I believe Ruby's Onigmo regex engine follows POSIX, and a website like this tells that POSIX [:print:] corresponds to Unicode \P{C}. A website like this tells that indeed U+00AD belongs to 'Other, Format' (Cf) category, a subset of 'Other', (C) category. But they do not match.

Is either side mistaken, or are POSIX [:print:] character class and Unicode \P{C} actually different things? If they are different, what are the idea behind the respective classes? And, is there a POSIX character class counterpart to Unicode \P{C}, or a Unicode counterpart to POSIX [:print:]?

0 Answers
Related