I am working on translating a hashing algorithm from C# to Java, and it requires using the byte array of the string. The problem is when working with characters like 'ł' & 'ą', Java converts these letters into 2 characters and thus giving me 4 bytes instead of 2 that I am expecting.
I tried using string.codePointAt() instead of string.charAt(), but it keeps on processing those letters as 2 characters instead of 1. I thought Java uses 16bit Unicode same as C# & VB but why does it require 4 bytes for this letters when C# & VB were able to convert these as 2 bytes.
C# and VB reads the bytes of 'ł' as: [66, 1] (code below)
bytes = Encoding.Unicode.GetBytes("ł");
Console.WriteLine(string.Join(",", bytes));
Java reads the bytes of 'ł' as: [-59, 0, 26, 32] (code below)
String str = "ł";
byte[] B = str.getBytes(Charset.forName("UTF-16LE"));
System.out.println(Arrays.toString(B));
I even tried using StandardCharsets too, but still same issue.
Is there a way for Java to process these letters as a single UTF-16 character instead of separating it into 2 characters?
PS: I cannot also refactor the algorithm since it is already in use, and it just had to be done in our new Java too.
PPS: I tried normalizing the string but there are still differences, character "æ" is read with [-26,0] when C# outputs [230,0] for the character