How to force Java to use only 2 bytes per character for Unicode characters (eg. 'ł')?

Viewed 321

I am working on translating a hashing algorithm from C# to Java, and it requires using the byte array of the string. The problem is when working with characters like 'ł' & 'ą', Java converts these letters into 2 characters and thus giving me 4 bytes instead of 2 that I am expecting.

I tried using string.codePointAt() instead of string.charAt(), but it keeps on processing those letters as 2 characters instead of 1. I thought Java uses 16bit Unicode same as C# & VB but why does it require 4 bytes for this letters when C# & VB were able to convert these as 2 bytes.

C# and VB reads the bytes of 'ł' as: [66, 1] (code below)

 bytes = Encoding.Unicode.GetBytes("ł");
 Console.WriteLine(string.Join(",", bytes));

Java reads the bytes of 'ł' as: [-59, 0, 26, 32] (code below)

String str = "ł";
byte[] B = str.getBytes(Charset.forName("UTF-16LE"));
System.out.println(Arrays.toString(B));

I even tried using StandardCharsets too, but still same issue.

Is there a way for Java to process these letters as a single UTF-16 character instead of separating it into 2 characters?

PS: I cannot also refactor the algorithm since it is already in use, and it just had to be done in our new Java too.

PPS: I tried normalizing the string but there are still differences, character "æ" is read with [-26,0] when C# outputs [230,0] for the character

2 Answers

I have found the issue, as Ralf Kleberhoff have guessed, I wasn't using the proper file encoding that the Java compiler expects. My file was using UTF-16LE encoding so I just passed -encoding "UTF-16" when compiling the file.

javac -encoding "UTF-16" HashBrowns.java

Also, as Samuel Hunter had suggested, I converted the values to positive to make sure that I get the exact same values as I get with C# & VB6.

private int[] convertSignedBytesToUnsignedint(byte[] b)
{
    int[] intArr = new int[b.length];
    for (int i = 0; i < b.length; i++) {
        intArr[i] = b[i] & 0xff;
    }
    return intArr;
}

I am not sure on which code is more optimized, but I just wanted to post this here so I can share what worked with my situation.

While Java's internal character encoding is UTF-16BE, String#codePointAt(int) and String#getBytes() (with no supplied arguments) both use the default character encoding, which depends on the Java implementation and the platform it's on. You have the right idea to use String.getBytes(Charset.forName("UTF-16LE")), but I recommend you use String.getBytes(StandardCharsets.UTF_16LE) instead.

The second issue with C# returning [230,0] while Java returns [-26, 0]: technically, they are the same, bit-wise. However, C#'s byte array holds unsigned bytes, while Java's array holds signed bytes. Even though both Java and C# gives the same byte pattern, if you really want to express a positive value, you could store them in a short array instead:

        String str = "æ";

        byte[] byteArray = "æ".getBytes(StandardCharsets.UTF_16LE);
        short[] newByteArray = new short[byteArray.length];

        for (int i = 0; i < byteArray.length; i++) {
            byte c = byteArray[i];
            newByteArray[i] = (c >= 0) ? c : (short)(c + 256);
        }

        System.out.println(Arrays.toString(byteArray));
        // => [-26, 0]
        System.out.println(Arrays.toString(newByteArray));
        // => [230, 0]

FWIW, replacing æ with ł gives me [66, 1] both for the byte array and short array.

Although the code "converts" the array into unsigned "bytes", I would advise against doing this if you can, because the byte array gives the same pattern as C#, and promises the same number size.

Related