I am looking for some help understanding the parsing of the decompressed Linux image in parse_elf() in arch/x86/boot/compressed/misc.c. Specifically, I don't understand what memory regions the ELF segments are getting copied to and from. Below is some annotated code showing my (mis)understanding.
for (i = 0; i < ehdr.e_phnum; i++) { // For each segment...
phdr = &phdrs[i];
switch (phdr->p_type) { // Ignore all segments that
case PT_LOAD: // aren't labeled as loadable
#ifdef CONFIG_RELOCATABLE
dest = output; // Set `dest` to be equal to the base of the kernel
// after decompression and KASLR considerations
// Next, add to `dest` the difference between the physical address
// of the segment and where the kernel was told to be loaded by the
// kernel configuration file. It seems to me that this difference
// is equal to `phdr->p_offset`.
dest += (phdr->p_paddr - LOAD_PHYSICAL_ADDR);
#else
// If we aren't considering relocations then just use the physical
// address of the segment as the destination.
dest = (void *)(phdr->p_paddr);
#endif
// Copy, to the destination determined above, from the beginning
// of the decompressed kernel plus the offset to the segment until
// the end of the segment is reached.
memcpy(dest,
output + phdr->p_offset,
phdr->p_filesz);
break;
default: /* Ignore other PT_* */ break;
}
}
What is confusing is that in the case of relocation, memcpy's first and second argument seem like they will be the same and there is no point in calling parse_elf. Perhaps I am misunderstanding what LOAD_PHYSICAL_ADDR or phdr->p_paddr is, or the steps taken after the kernel is decompressed in place.
The non-relocation case makes more sense as we just need to copy from the decompressed kernel to a "hard-coded" address.