Why does it appear that this string is stored inline by value in an explicit layout class or struct?

Viewed 192

I have been doing some extremely unsafe and slightly useless messing with the System.Runtime.CompilerServices.Unsafe MSIL package that allows you to do a lot of things with pointers you can't in C#. I created an extension method that returns a ref byte, with that byte being the start of the Method Table pointer at the start of the object, which allows you to use any object in a fixed statement, taking a byte pointer to the start of the object:

public static unsafe ref byte GetPinnableReference(this object obj)
{
    return ref *(byte*)*(void**)Unsafe.AsPointer(ref obj);
}

I then decided to test it, using this code:

[StructLayout(LayoutKind.Explicit, Pack = 0)]
public class Foo
{
    [FieldOffset(0)]
    public string Name = "THIS IS A STRING";
}

[StructLayout(LayoutKind.Explicit, Pack = 0)]
public struct Bar
{
    [FieldOffset(0)]
    public string Name;
}

And then in the method

        var foo = new Foo();
        //var foo = new Bar { Name = "THIS IS A STRING" };

        fixed (byte* objPtr = foo)
        {
            char* stringPtr = (char*)(objPtr + (foo is Foo ?  : 12));

            for (var i = 0; i < foo.Name.Length; i++)
            {
                Console.Write(*(stringPtr + i /* Char offset */));
            }

            Console.WriteLine();
        }

        Console.ReadKey();

The really weird thing about this is that this successfully prints "THIS IS A STRING"? The code works like this:

  1. Get a byte pointer, objPtr, to the very start of the object
  2. Add 16 to get to the actual data
  3. Add another 16 to get past the string header to the string's actual data
  4. Add 4 to skip the first 4 bytes of the string, which are the int _stringLength (exposed to us as Length property)
  5. Interpret the result as a char pointer

EDIT: Important point - when switching foo to type Bar, I only add 12 rather than 36 bytes on (36 = 16 + 16 + 4). Why does it only have 8 bytes of header in the struct rather than 32 in the class? It would make sense that the struct has a smaller header (no syncblk i believe), but then why doesn't the string still have a 16 byte head? I would expect the offset to be 8 + 16 + 4 (28) rather than just 8 + 4 (12) However, this assumption makes a big flaw. It assumes the string is stored inline inside the class/struct. However, strings are reference types and only a reference to them is stored inside the object from my knowledge. Particularly, I thought reference types can only be put on the heap - and as this struct is a local variable I thought it was on the stack. If it wasn't, the code would surely look something more like this to get the stringPtr

byte** stringRefptr = objPtr + 16;
char* stringPtr = (char*)(*stringRefPtr + 20);

where you take the string reference as a byte** and then use it to get to the chars. And this still wouldn't make sense if the string internally was a char[] (I'm not sure if it is)

So why does this work, and print the string, even though it mistakenly assumes string is stored inline, when string is a reference type?

NOTE: Requires .NET Core 2.0+ with System.Runtime.CompilerServices.Unsafe nuGet package, and C# 7.3+.

1 Answers

Because strings are indeed stored inline. The problem with your assumption is that strings are not normal objects but handled as a special case by the CLR (probably for performance reasons).

And as for the objects, since the string is the only member this would naturally be the most efficient way to allocate the memory. Try adding more members after your string member and your code would break.

Here’s a few references in how strings are stored in the CLR

https://mattwarren.org/2016/05/31/Strings-and-the-CLR-a-Special-Relationship/

https://codeblog.jonskeet.uk/2011/04/05/of-memory-and-strings/

Edit: I didn’t check, but I believe your reasoning behind the offsets is off. 36 = 24 (size of object) + 8 (string header?) + 4 (size of int) while for the struct the24 bytes becomes 0 as it has no header.

Related