How to normalize fancy-looking unicode string in C#?

Viewed 1369

I receive from a REST API a text with this kind of style, for example

  • ?

  • ?

  • нσω тσ яємσνє тнιѕ ƒσηт ƒяσм α ѕтяιηg?

But this is not italic or bold or underlined since the type it's string. This kind of text make it failed my Regex ^[a-zA-Z0-9._]*$

I would like to normalize this string received in a standard one in order to make my Regex still valid.

1 Answers

You can use Unicode Compatibility normalization forms, which use Unicode's own (lossy) character mappings to transform letter-like characters (among other things) to their simplified equivalents.

In python, for instance:

>>> from unicodedata import normalize
>>> normalize('NFKD','       ')
'How to remove this font from a string'

# EDIT: This one wouldn't work
>>> normalize('NFKD','нσω тσ яємσνє тнιѕ ƒσηт ƒяσм α ѕтяιηg?')
'нσω тσ яємσνє тнιѕ ƒσηт ƒяσм α ѕтяιηg?'

Interactive example here.

EDIT: Note that this only applies to stylistic forms (superscripts, blackletter, fill-width, etc.), so your third example, which uses non-latin characters, can't be decomposed to ASCII.

EDIT2: I didn't realize your question was specific to C#, here's the documentation for String.Normalize, which does just that:

string s1 = "       "
string s2 = s1.Normalize(NormalizationForm.FormKD)
Related