I'm getting weather alerts from a weather service. Although the HTTP response claims to be UTF-8, clearly it contains some text like this:
Suðaustan 13-20 m/s og snjókoma með lélegu skyggni og versnandi akstursskilyrðum.
...that should look like this:
Suðaustan 13-20 m/s og snjókoma með lélegu skyggni og versnandi akstursskilyrðum.
...but has already been improperly decoded before it first reached me, being re-encoded as UTF-8 after being decoded improperly. Most of us have probably seen this kind of "mojibake" garbage before, and visually at least, it often has a lot of common characteristics — such as lots of à characters, ¢ signs and the like.
I'm using this code to fix it up right now:
// Check for UTF-8 wrongly decoded as Latin-1
if (/[\x80-\xC5]/.test(result)) {
const bytes = Buffer.from(result, 'latin1');
const altText = bytes.toString('utf8');
if (altText.length < result.length)
result = altText;
}
...and that's doing the job for now, but it's not a very sophisticated test.
Anyone know of a better method?