indexOf() gone crazy with Eastern languages

Viewed 134

I have a string of Arabic characters:

var txt="یہ ایک جملہ ہے۔";

And I want to find the position of a certain character (e.g. ج) in this string.

alert (txt.indexOf("ج"));

I tried using txt.indexOf() function but something really strange happens: if I specify the strings (both: the base string and the search string) in real-time (e.g. through an inputbox or form textbox) then it works as intended. However, when I specify the characters of the base string as a hard coded JavaScript line, then all hell breaks loose.

The characters appear as some bizarre ASCII values when I alert() them (appearing as Ößùīñè etc) and the indexOf result always returns -1 (not found). Initially I thought it's an issue with the js file encoding to ensure it supports the extended character set. Turns out the encoding is UTF-8 and the characters appear perfectly fine in the editor when I close and then reopen the file. The problem is only when processing them with JavaScript.

I am using notepad++ as the code editing software.

var txt="یہ ایک جملہ ہے۔";
console.log(txt.indexOf("ج"));

Any help would be greatly appreciated.

1 Answers

Notepad++ has 5 encodings, and as I mentioned in my comment, I have already recognised using PowerShell, that the default UTF-8 is not correct for all hungarian characters, and UTF-8-BOM is the right one.
I could reproduce the problem in Chrome, but not in Firefox, with the below (on desktop, I didn't check on other devices):
Save alert("یہ ایک جملہ ہے۔") to 5 files encoded differently and named after the encoding, and then save

<!DOCTYPE html>
<html>
<head>
</head>
<body>
<script src = "ANSI.js"></script>
<script src = "UCS2 BE BOM.js"></script>
<script src = "UCS2 LE BOM.js"></script>
<script src = "UTF-8.js"></script>
<script src = "UTF-8-BOM.js"></script>
</body>
</html>

to an html file and open to see that UTF-8-BOM shows the characters correctly, the UCS-2 types also, while the ANSI file can not show the characters even in Notepad++.
I would recommend you to set the default at Settings -> Preferences -> New document.
Regarding what happens when the input is from the site: I think as the browser creates the elements it is controlling their behaviour, and it encodes the characters so that it can reproduce them, so there is no encoding "conflict" like if it is coming from a file and is already encoded.

Related