Why is Unicode Character 'DOUBLE LOW-9 QUOTATION MARK' (U+201E) not a quotation mark?

Viewed 495

Even though I stumbled over this because a PHP regex I wrote failed to match it in the way I expected, I'm not sure if this is the right place to ask. After all, the definition in PHP (and probably other Unicode-aware regex engines) seems to match the official categorization (cf. e.g. https://www.fileformat.info/info/unicode/char/201e/index.htm) and it is this official categorization I am unhappy with.

According to this, the DOUBLE LOW-9 QUOTATION MARK is categorized as Ps (therefore matched by /\p{Ps}/) and, in spite of its very name, not as Pi (initial quotation mark), for which is used in German. It didn't even make it into the less specific 'Punctuation, Initial quote (may behave like Ps or Pe depending on usage)' category. What could be the reason for this (mis)categorization? In what languages is it actually used as a Ps (i.e., similar to "(" or "[" or "{")?

But most importantly: What is a suitable regex that covers all kinds of quotation marks across all languages without enumerating too many individual codepoints?

1 Answers

The general categories Pi (Initial_Punctuation) and Pf (Final_Punctuation) are not used exclusively for quotation marks, just like Ps (Open_Punctuation) and Pe (Close_Punctuation) are not used exclusively for characters that aren’t quotation marks. Rather, Pi and Pf are used for pairs of characters where either one can be opening or closing depending on usage, whereas Ps characters are always opening and Pe characters are always closing (ignoring rare or specialised cases). Which of these general categories a character belongs to is based on these considerations and has nothing to do with whether it is a quotation mark, a bracket or something else.

U+201E DOUBLE LOW-9 QUOTATION MARK is categorised as Ps because there is no established orthography in the world where it can be used as a closing mark. It is always opening in practice. In contrast, U+201C LEFT DOUBLE QUOTATION MARK is categorised as Pi because it can be both an opening and a closing quote depending on which specific style of quotes you chose.

Unicode has a dedicated property for identifying quotation marks appropriately named Quotation_Mark. This property is defined independently from the general category values previously discussed.

Related