Your question is precise: I just want to know why fahr is going to be printing 0.0 in the whole table.
As you already understand, you have undefined behavior because you pass a float, which is actually converted to a double, to printf as an argument for which printf expects an int. Anything can happen, you get 0 and 0.0 as output for all lines, you could have gotten pretty much anything else or a crash...
To try and explain your observations, you must look into what is actually happening on your very system for this code. Such an analysis requires in depth knowledge on your system, ABI, compiler, compiler options, etc.
I modified your code and compiled it with Godbolt's Compiler Explorer and here are my observations for 2 configurations:
gcc version 9.3 for intel 64-bits, with optimisations disabled.
The code for the erroneous printf in 64-bits is this:
cvtss2sd xmm1, DWORD PTR [rbp-8]
cvtss2sd xmm0, DWORD PTR [rbp-4]
mov edi, OFFSET FLAT:.LC6
mov eax, 2
call printf
The code for the a modified argument (int)celcius which is what printf expects:
cvtss2sd xmm0, DWORD PTR [rbp-8]
movss xmm1, DWORD PTR [rbp-4]
cvttss2si eax, xmm1
mov esi, eax
mov edi, OFFSET FLAT:.LC6
mov eax, 1
call printf
In the 64-bit version, the erroneous code passes celcius and fahr as double values in floating point registers %xmm0 and %xmm1 respectively and passes the value 2 in %eax, whereas the correct code would pass fahr as a double in %xmm0 and celcius converted to an int in register %esi, and the value 1 in %eax.
The value in %eax, more precisely the contents of %al is the number of vector registers used to pass arguments. To implement the vararg api in printf, the compiler generates a prolog that uses this value to save the register arguments to the stack:
myprintf:
push rbp
mov rbp, rsp
sub rsp, 104
mov QWORD PTR [rbp-216], rdi
mov QWORD PTR [rbp-168], rsi
mov QWORD PTR [rbp-160], rdx
mov QWORD PTR [rbp-152], rcx
mov QWORD PTR [rbp-144], r8
mov QWORD PTR [rbp-136], r9
test al, al
je .L12
movaps XMMWORD PTR [rbp-128], xmm0
movaps XMMWORD PTR [rbp-112], xmm1
movaps XMMWORD PTR [rbp-96], xmm2
movaps XMMWORD PTR [rbp-80], xmm3
movaps XMMWORD PTR [rbp-64], xmm4
movaps XMMWORD PTR [rbp-48], xmm5
movaps XMMWORD PTR [rbp-32], xmm6
movaps XMMWORD PTR [rbp-16], xmm7
.L12:
So printf will read from [rbp-216] the int value expected for the %8d format and from [rbp-128] the double value for fahr.
The int value will be whatever %rsi happened to contain when printf was called, 0 from your observations. The double value should be whatever was passed in xmm0, so you would expect to see the value of celcius, and that's indeed what I observe on my system. Since you observe something quite different, there is a good chance your system does not use this 64-bit ABI.
gcc version 9.3 for intel 32-bits, with optimisations disabled.
In 32 bits, the arguments are all passed on the stack. When passing 2 float values, we have:
fld DWORD PTR [ebp-12]
fld DWORD PTR [ebp-16]
sub esp, 12
lea esp, [esp-8]
fstp QWORD PTR [esp]
lea esp, [esp-8]
fstp QWORD PTR [esp]
push OFFSET FLAT:.LC6
call printf
add esp, 32
and when passing an int and a float:
fld DWORD PTR [ebp-16]
movss xmm0, DWORD PTR [ebp-12]
cvttss2si eax, xmm0
lea esp, [esp-8]
fstp QWORD PTR [esp]
push eax
push OFFSET FLAT:.LC6
call printf
add esp, 16
So printf expects the int at [ebp+12] and the double at [ebp+16] but instead celcius was pushed as a doouble at [ebp+12] and fahr at [ebp+20].
The int read from [ebp+12] are in fact the 4 low order bytes of the celcius value. Since the celcius values are small integers, the 32 low order bits of their 64-bit floating point representation are all zeroes, hence the integer read is 0. Conversely, the double value read for fahr is misaligned: the first 4 bytes are the high 32-bits of the double value of celcius and the last 4 bytes are the low 32 bits of the double value fahr, which are 0 because fahr is also a small integral value. Hence the exponent part of this double value has all bits zero, so it is either the value 0.0 or an extremely small denormal value that converts to 0.0 with the %5.1f conversion format. Indeed I get the same output as you do in 32-bit mode.
You can experiment with a different format such as %g for fahr and check if my prediction for a very small value is correct.
Of course this forensic study is only relevant to a specific architecture and by no means condoned by the C Standard.