A Few Ponits About Floating Point Data 03-01-2014, 03:42 AM
#1
This is going to be my first informal post about something in programming for which I think is highly neglected despite it's importance. As proven through history, the severity of such neglect has proven disastrous as well. The concept is related to Floating Point Error, a term that many programmers have probably heard about, but few actually understand.
Firstly, there are typically/usually two standard types used for floating point datatypes in a few languages, single and double floating point datatypes.
VB.NET:
C#/C/C++:
.NET also has a datatype decimal/Decimal, and I'll explain how this is different later and why it should be used in some cases over a floating point value type.
Have you ever, whether a result of lack of experience, or by accident, written something like this in C?
And received a numeric value (as an integer), and didn't understand what was going on?
This is because all floating point values have an integer representation. The difference between 2 integral representations of a floating point value indicate the units in last place (ULP), which is a measurement of floating point separation. Some example code I've written is as follows:
Which would output:
The significance of such floating point error can be emphasized when you try to match these values.
Which would still output:
What's going on? The output is wrong! No... You're just not seeing the full precision.
printf() for instance will round for you. See this following code:
Which will output:
This ULP of 1 means these numbers are within the range such that they are as close as they can possibly get without being the same in float space.
For this same reason though, as demonstrated, this is exactly why programmers who compare floating point values with the overloaded == operator for instance, do not know what they are doing. You may get lucky some of the time, but there will be the odd case where floating point error will catch you off guard, and this is or can be very tricky to debug after the fact.
Next, what to use as an epsilon? In C/C++ for instance, within the float.h header, you've got such defines as DBL_EPSILON and FLT_EPSILON, which are calculated to be the best values to use for an epsilon based on the precision of float and double. For instance:
The exact representation of the full bits of the floating pint value depends on the machine, but the most common standard used for representing a floating point value is the IEEE-754 standard. Depending on whether it's a single floating point or a double floating point representation may determine whether the value is 4 bytes or 8 bytes, but the concept remains the same. For a 4 byte floating point datatype, 1 bit for the sign (+/-), an exponent of 8 bits for magnitude, and a mantissa which specify the actual digits of the number, and is a length of 23 bits. This leaves us with a (23 + 1 + 8 =) 32 bit value, which is 4 bytes on most modern machines. For a 64 bit value, the exponent becomes 11 bits, the sign stays 1 bit, which leaves 52 significant bits for the mantissa field.
In addition to the IEEE-754 standard, there also exists an analogous 96-bit extended precision format which falls under the IEEE-854 standard. I haven't gotten into this one much yet so I won't be explaining that here.
What most programmers should be doing with floating point evaluation is by doing a comparison including an epsilon, which makes up for more of a range based equality check, rather than a precise or an exact one. Why? Because floating point datatypes were never intended to be used with an exact precision, and they are unreliable where such precision is required.
0.6 for instance is not 0.6000... 0.6 is actually 0.599999999999999977795539507496869191527366638183...
A common question I see is "why and when do we use floating point datatypes then if they are so inaccurate?" good question, but when you understand that they are for approximations up to a certain pre-determined precision, you should be able to figure things out from there. A common rule of thumb is to not use floating point when dealing with currency calculations (money), and in VB.NET/C# for instance, you should be using the decimal/Decimal alternative. In other languages, if you want precision, you may have to understand complex floating point math and manipulate that to gain more precision. Such libraries exist out there which attempt to do this, but if you want to try this on your own, you'll need to understand the representation for starters.
Here's a fairly in depth link you can read if you wish: http://docs.oracle.com/cd/E19957-01/806-...dberg.html
And some demo code I wrote to show you the parts of the floating point value by IEEE 754 standard:
For this example, the output is:
If you're using decimal over double or float, you'll get more precision in .NET code because there are more bits to store a higher precision. decimal is 16 bytes in .NET whereas float is 4 bytes, and double is 8 bytes. Thus, although decimal is still really a floating point datatype, it'll yield much much better precision than it's relatives, float and double, due to more significant bits being available to store the digits and similarly, double has a better precision than float.
This is also why there are 2 pre-defined epsilons for each float and double in the float.h header for C/C++ programmers.
Firstly, there are typically/usually two standard types used for floating point datatypes in a few languages, single and double floating point datatypes.
VB.NET:
Code:
Single
DoubleC#/C/C++:
Code:
float
double.NET also has a datatype decimal/Decimal, and I'll explain how this is different later and why it should be used in some cases over a floating point value type.
Have you ever, whether a result of lack of experience, or by accident, written something like this in C?
Code:
float amount = 2.7;
printf("%d\n", amount);And received a numeric value (as an integer), and didn't understand what was going on?
This is because all floating point values have an integer representation. The difference between 2 integral representations of a floating point value indicate the units in last place (ULP), which is a measurement of floating point separation. Some example code I've written is as follows:
Code:
#include <stdio.h>
#include <stdlib.h>
int main()
{
float pi; __asm("fldpi" : "=t"(pi));
int i1 = *(int *)π
printf("%f = ", pi);
printf("0x%X (%d)\n", i1, i1);
float pi_approx = 3.14;
int i2 = *(int *)&pi_approx;
printf("%f = ", pi_approx);
printf("0x%X (%d)\n", i2, i2);
printf("\nULP = %d\n", abs(i1 - i2));
}Which would output:
Code:
3.141593 = 0x40490FDB (1078530011)
3.140000 = 0x4048F5C3 (1078523331)
ULP = 6680The significance of such floating point error can be emphasized when you try to match these values.
Code:
float pi_approx = 3.141593;Which would still output:
Code:
3.141593 = 0x40490FDB (1078530011)
3.141593 = 0x40490FDC (1078530012)
ULP = 1What's going on? The output is wrong! No... You're just not seeing the full precision.
printf() for instance will round for you. See this following code:Code:
#include <stdio.h>
#include <stdlib.h>
int main()
{
float pi; __asm("fldpi" : "=t"(pi));
float pi_approx = 3.141593;
printf(" pi: %.20f\npi_approx: %.20f\n", pi, pi_approx);
}Which will output:
Code:
pi: 3.14159274101257324219
pi_approx: 3.14159297943115234375This ULP of 1 means these numbers are within the range such that they are as close as they can possibly get without being the same in float space.
For this same reason though, as demonstrated, this is exactly why programmers who compare floating point values with the overloaded == operator for instance, do not know what they are doing. You may get lucky some of the time, but there will be the odd case where floating point error will catch you off guard, and this is or can be very tricky to debug after the fact.
Code:
if (abs(floatvalue1 - floatvalue2) <= epsilon) ... // GOOD
if (floatvalue1 == floatvalue2) ... // BADNext, what to use as an epsilon? In C/C++ for instance, within the float.h header, you've got such defines as DBL_EPSILON and FLT_EPSILON, which are calculated to be the best values to use for an epsilon based on the precision of float and double. For instance:
Code:
if (abs(val1 - val2) <= DBL_EPSILON) ...
if (abs(val1 - val2) <= FLT_EPSILON) ...The exact representation of the full bits of the floating pint value depends on the machine, but the most common standard used for representing a floating point value is the IEEE-754 standard. Depending on whether it's a single floating point or a double floating point representation may determine whether the value is 4 bytes or 8 bytes, but the concept remains the same. For a 4 byte floating point datatype, 1 bit for the sign (+/-), an exponent of 8 bits for magnitude, and a mantissa which specify the actual digits of the number, and is a length of 23 bits. This leaves us with a (23 + 1 + 8 =) 32 bit value, which is 4 bytes on most modern machines. For a 64 bit value, the exponent becomes 11 bits, the sign stays 1 bit, which leaves 52 significant bits for the mantissa field.
In addition to the IEEE-754 standard, there also exists an analogous 96-bit extended precision format which falls under the IEEE-854 standard. I haven't gotten into this one much yet so I won't be explaining that here.
What most programmers should be doing with floating point evaluation is by doing a comparison including an epsilon, which makes up for more of a range based equality check, rather than a precise or an exact one. Why? Because floating point datatypes were never intended to be used with an exact precision, and they are unreliable where such precision is required.
0.6 for instance is not 0.6000... 0.6 is actually 0.599999999999999977795539507496869191527366638183...
A common question I see is "why and when do we use floating point datatypes then if they are so inaccurate?" good question, but when you understand that they are for approximations up to a certain pre-determined precision, you should be able to figure things out from there. A common rule of thumb is to not use floating point when dealing with currency calculations (money), and in VB.NET/C# for instance, you should be using the decimal/Decimal alternative. In other languages, if you want precision, you may have to understand complex floating point math and manipulate that to gain more precision. Such libraries exist out there which attempt to do this, but if you want to try this on your own, you'll need to understand the representation for starters.
Here's a fairly in depth link you can read if you wish: http://docs.oracle.com/cd/E19957-01/806-...dberg.html
And some demo code I wrote to show you the parts of the floating point value by IEEE 754 standard:
Code:
#include <stdio.h>
#include <stdlib.h>
#include <limits.h>
int main()
{
const int width = 16;
double f; __asm("fldpi" : "=t"(f));
printf("%*s: %f\n", width, "Value", f);
int bits = sizeof(f) * CHAR_BIT;
int i = *(int *)&f;
printf("%*s: 0x%X (%d)\n", width, "Integer", i, i);
int sign_field = i >> (bits - 1);
printf("%*s: %i\n", width, "Sign", sign_field);
int significant_bit_width = bits - CHAR_BIT;
int exp_field = (i >> (significant_bit_width - 1)) & 0xFF;
printf("%*s: %i\n", width, "Exponent Field", exp_field);
printf("%*s: %i\n", width, "Exponent Value", exp_field == 0 ? -126 : exp_field - 127);
int mantissa_field = i & ((1 << (significant_bit_width - 1)) - 1);
printf("%*s: 0x%X (%d)\n\n", width, "Mantissa Field", mantissa_field, mantissa_field);
}For this example, the output is:
Code:
Value: 3.141593
Integer: 0x40490FDB (1078530011)
Sign: 0
Exponent Field: 128
Exponent Value: 1
Mantissa Field: 0x490FDB (4788187)If you're using decimal over double or float, you'll get more precision in .NET code because there are more bits to store a higher precision. decimal is 16 bytes in .NET whereas float is 4 bytes, and double is 8 bytes. Thus, although decimal is still really a floating point datatype, it'll yield much much better precision than it's relatives, float and double, due to more significant bits being available to store the digits and similarly, double has a better precision than float.
This is also why there are 2 pre-defined epsilons for each float and double in the float.h header for C/C++ programmers.



![[+]](https://sinister.li/images/modern/collapse_collapsed.png)