Introduction
Hello, as floating point arithmetic is very complex to define entirely, I will define some sub-areas that will help you to explore the other sub-areas of floating point arithmetic.
Definitions
Rounding Error
To crush an infinite number of real numbers into a finite number of bits, an approximate representation is required.
Even though there are an infinite number of integers, in most programs the result of integer calculations can be stored in 32 bits.
In contrast, for a fixed number of bits, most calculations with real numbers result in quantities that cannot be represented exactly with that many bits.
For this reason, the result of a floating-point calculation must often be rounded to fit back into its finite representation. This rounding error is the characteristic feature of floating point calculations.
Relative Error and Ulps
As rounding errors are inherent in floating point calculations, it is necessary to have a way to calculate this error.
For example, consider the floating point format with b = 10 and p = 3 used in this section.
If the result of a floating point calculation is 3.12 × 10-2 and the answer is .0314 in an infinite accuracy calculation, it is clear that it is in error by 2 units in the last digit.
Analogously, if the real number .0314159 is represented as 3.14 × 10-2, then it is in error by .159 units in the last digit.
In general, if the floating point number d.d...d × e is used to represent z, then it is in error by d.d...d - (z/e)p-1 units in the last place.
The term ulps is used as an abbreviation for "units in the last place".
If the result of a calculation is the floating point number nearest to the correct result, it can still be in error by up to 0.5 ulp.
Another way to measure the difference between a floating point number and the real number it approximates is the relative error, which is simply the difference between the two numbers divided by the real number. For example, the relative error of approximating 3.14159 by 3.14 × 100 is .00159/3.14159 .0005.