• 1、(-1)^sDenotes the sign bit. When s=0, V is positive; when s=1, V is negative.
  • 2. M represents the significand, which is greater than or equal to 1 and less than 2.
  • 3. 2^E represents the exponent field.

For example, the decimal 5.0, written in binary as 101.0, is equivalent to 1.01×2^2. Then, according to the format of V above, we can derive s=0, M=1.01, E=2.

IEEE 754 specifies that for a 32-bit floating-point number, the highest 1 bit is the sign bit s, the next 8 bits are the exponent E, and the remaining 23 bits are the significand M.

IEEE 754 also has some special regulations for the significand M and the exponent E.

As mentioned earlier, 1 ≤ M < 2, that is, M can be written in the form 1.xxxxxx, where xxxxxx represents the fractional part. IEEE 754 stipulates that when M is stored inside the computer, the first digit of this number is always 1 by default, so it can be discarded, and only the following xxxxxx part is stored. For example, when saving 1.01, only 01 is stored, and when reading, the first digit 1 is added back. The purpose of this is to save one bit of significand. Taking a 32-bit floating-point number as an example, only 23 bits are left for M; after discarding the first 1, it is equivalent to storing 24 bits of significand.

As for the exponent E, the situation is more complicated.

First, E is an unsigned integer (unsigned int). This means that if E is 8 bits, its value range is 0~255; if E is 11 bits, its value range is 0~2047. However, we know that E in scientific notation can be negative, so IEEE 754 stipulates that the true value of E must be obtained by subtracting a bias from the stored value. For an 8-bit E, this bias is 127; for an 11-bit E, this bias is 1023.

For example, the E of 2^10 is 10, so when saving as a 32-bit floating-point number, it must be saved as 10+127=137, i.e., 10001001.

Then, the exponent E can be further divided into three cases:

  • (1) E is neither all 0s nor all 1s. In this case, the floating-point number is represented according to the rules above, i.e., the computed value of exponent E minus 127 (or 1023) yields the true value, and then the leading 1 is prepended to the significand M.
  • (2) E is all 0s. In this case, the exponent E of the floating-point number equals 1-127 (or 1-1023), and the significand M is no longer prepended with the leading 1, but restored to a decimal of 0.xxxxxx. This is done to represent ±0, as well as very small numbers close to 0.
  • (3) E is all 1s. In this case, if the significand M is all 0s, it means ±infinity (the sign depends on the sign bit s); if the significand M is not all 0s, it means the number is not a number (NaN).
  • Note:Decimal fractions stored in memory are extremely likely to lose precision. For example, 0.3, this guy, cannot by itself be converted into a finite binary representation.

    Original address: https://www.cnblogs.com/SimpleISP/p/5280362.html