float single-precision floating-point numberIn the machine, it occupies 4 bytes, described using 32-bit binary.

double precision floating-point numberIn the machine it occupies 8 bytes, described by 64-bit binary.

Floating-point numbers are represented in exponential form inside the machine, decomposed into:sign, mantissa, exponent sign, exponentFour parts.

The sign bit occupies 1 binary digit, indicating whether the number is positive or negative.

The exponent sign occupies 1 binary bit, indicating whether the exponent is positive or negative.

The mantissa represents the significant digits of a floating-point number, 0.xxxxxxx, but does not store the leading 0 and the decimal point.

The exponent stores its significant digits.

The number of bits for the exponent and the number of bits for the mantissa are determined by the computer system.

Maybe the sign of the number plus the mantissa occupies 24 bits, and the sign of the exponent plus the exponent occupies 8 bits --float。

The sign plus mantissa occupies 48 bits, and the exponent sign plus exponent occupies 16 bits —double。

Once you know the placeholders for these four parts, estimate the size range in binary, then convert to decimal—that gives you the value range you want to know.

For programmers, the difference between double and float is that double has higher precision, with 16 significant digits, while float has 7 digits of precision. But double consumes twice as much memory as float, and double's computation speed is much slower than float. In C language, the names of math functions for double and float are different, so don't write them incorrectly. When single precision is sufficient, don't use double precision (to save memory and speed up computation).

Typebit countsignificant figuresNumerical range
float32 6-7 -3.4*10(-38)~3.4*10(38)
double 64 15-16 -1.7*10(-308)~1.7*10(308)
long double 128 18-19 -1.2*10(-4932)~1.2*10(4932)

In simple terms, Float is single precision, occupying 4 bytes in memory, with 7 significant digits (because of positive and negative signs, it's not 8 digits); on my computer and the VC++6.0 platform, the default display is 6 significant digits; double is double precision, occupying 8 bytes, with 16 significant digits, but on my computer and the VC++6.0 platform, the default display is likewise 6 significant digits.

Example: In C and C++, the following assignment statement:

float a=0.1; 

Compiler error:warning C4305: 'initializing' : truncation from 'const double ' to 'float '

Reason:In C/C++ (not sure if it’s only like this in VC++), for the 0.1 on the right side of the equals sign in the above statement, we think it’s a float, but the compiler treats it as a double (because decimal literals default to double), so it reports this warning. Usually change it to...0.1fThen it's fine.

My usual practice is to frequently use double, and I dislike using float.

In C and C#, single-precision type is used for floating-point data.floatand double-precision typedoubleto store,floatData occupies 32 bits,doubleData occupies 64 bits, we are declaring a variable.float f= 2.25fWhen it comes to allocation, how is memory allocated? If it were allocated randomly, wouldn't the world be in chaos? In fact, whether float or double, both follow IEEE standards in storage: float follows IEEE R32.24, and double follows R64.53.

Whether single-precision or double-precision, in storage they are divided into three parts:

  • Sign bit: 0 represents positive, 1 represents negative.
  • Exponent bits: used to store the exponent data in scientific notation, and uses biased representation.
  • Mantissa part: Mantissa part.

Source address: https://my.oschina.net/zd370982/blog/724265