Why DLL function calls are not compiled to relative CALL instructions

Viewed 657

When I call a function DLLFunction(int), which is defined in a DLL. Visual Studio 2013 on my Intel X86 PC compiles it to the following instruction

CALL [__imp__DLLFunction@4] // Call absolute indirect address FF 15 00 90 40 00 CALL [00409000h] // original absolute CALL instruction FF 15 00 90 39 01 CALL [01399000h] // After address fixup by OS loader // __imp__DLLFunction@4 is the IAT entry address for DLLFunction, where there stores the address for DLLFunction().

IAT's RVA (Relative Visual Address) in the caller image is 0x9000, where there stores import functions' addresses. RVA Import Function Address 0x9000 0x60fd1014 // DLLFunction 0x9004 0x60fdxxxx // someOtherDLLFunction0 0x9008 0x60fdxxxx // someOtherDLLFunction1 ...

Why compiler does not generate relative CALL instruction?

If using relative CALL instruction, the loader does not need to fix up the addresses for all these CALL instructions like this.

2 Answers

CALL [__imp__DLLFunction@4] is not calling the usual stub that steers the control to the imported function through an indirect jump, it is calling the imported function directly through the pointer in the IAT.

This happens when the external function is annotated with __declspec(dllimport) (and possibly in any way that makes the compilers aware of the programmer intent).

Without it, the compiler generates a relative (near) call and the linker add the stub.

:401005 E806000000     call 401010h              ;Relative near call to the stub
... The stub ...
:401010 FF25F4B04000   jmp DWORD PTR [0040b0f4]  ;Indirect abs jump

With the intent clear, the code above transforms to

:401005 FF15F4B04000   call DWORD PTR [0040b0f4]

that is using an absolute indirect call.
This spares a jump but requires an additional fix-up at load time, a relative indirect call would be effectively better but, unfortunately, it doesn't exist.

x86-64 code can use RIP-relative addressing to mitigate the fix-up problems.

update: I didn't read the question carefully: the indirect call target is still inside the executable, and is itself an indirect jmp, according to the OP.

The answer below is discussing making the call rel32 go directly into the DLL.


That would require modifying the machine code at every call instruction to put in the right offset during dynamic loading of a DLL. (You don't know what address the DLL will be loaded at, so you don't know the rel32 distance between the executable and the DLL at link time of the executable.)

Using a table of function pointers puts all the relocation stuff in one place where it can be written efficiently while dynamic linking.

It also couldn't reach far enough in 64-bit code if a DLL was loaded more than 2GB away from code that needed to call it.


IIRC, Windows does support the scheme you describe for DLLs that you link normally (with the linker at compile time, not with a runtime dll import). Every DLL has a "preferred" load address, and code that calls it optimistically uses call rel32 instructions so every call site needs fixups if that address isn't available. These fixups happen while the process is being loaded. With ASLR enabled, DLLs won't load at the same address every time.

Once the process is already running, its code pages will be read-only, so if these fixups are needed, it's a problem. This is probably why dynamic DLL imports don't use this mechanism. (The implementation could use VirtualProtect to make code pages writeable for these fixups, but it wouldn't be thread-safe to have two different threads importing DLLs at the same time. One thread might make a page read-only after it was done, but while another thread was still writing fixups to that page, resulting in a fault.)

Also, cross-modifying code isn't in general safe. Other threads could be running instructions in the same function where you were applying a fixup. You could atomically store the new rel32 using an xchg or something. This might be safe.

BTW, on Linux, even "normal" libraries like libc are called through a level of indirection like this (function-pointer from the Global Offset Table). See Sorry state of dynamic libraries on Linux.

It's partly a tradeoff between overhead of runtime dynamic linking (load times) vs. performance after it's loaded.

Related