How are decimals represented? -- Floating point numbers and IEEE 754

In this lecture you will understand: how decimals are represented in a world with only 0s and 1s, and why floating point numbers have precision errors.


Everyday analogy: scientific notation

You have certainly seen scientific notation in physics class:

Speed of light = 3.0 x 108m/s, electron mass = 9.1 x 10-31 kg。

This method can represent very large or very small numbers with very few digits.

Computers use exactly the same idea to represent decimals--Binary Scientific Notation。

Any binary decimal can be written as:

1.xxxxxxxxxx x 2yyyyy

This form only needs to store three things:

  • symbol: positive or negative sign?
  • mantissa(xxxxxxxxxx): significant digits part
  • Exponent(yyyyy): position of the decimal point

The IEEE 754 standard packs these three things into 32 bits in a fixed format.


The three-segment structure of IEEE 754 single precision

32 bits (single precision float) are precisely divided into three segments:

Sign Bit 0 1 bit
Exponent 00000000 8 bits
Mantissa bits Mantissa (Fraction) 00000000000000000000000 23 bits
Sign bit (1 bit)
Exponent bits (8 bits)
Mantissa bits (23 bits)

First part: sign bit (1 bit)

The simplest segment.

0represents positive numbers,1represents negative numbers.

Second segment: exponent bits (8 bits)

The exponent determines the position of the decimal point (this is where the name "floating point" comes from--the decimal point can "float").

But the exponent is not stored directly; instead, "actual exponent + 127" is stored.

This 127 is calledOffset (Bias)。

Why add an offset?

Because 8 bits can only represent 0~255 (unsigned), but the actual exponent needs to be negative (to represent numbers less than 1).

After adding the offset 127:

  • Stored value 127 corresponds to actual exponent 0 (127 - 127 = 0)
  • Storage value 0 corresponds to actual exponent -127
  • Storage value 255 corresponds to actual exponent +128

Third segment: mantissa bits (23 bits)

The mantissa stores the "xxxxx" part of "1.xxxxx", omitting the leading 1.

Because in binary scientific notation, the normalized form is always 1.xxx, so the leading 1 does not need to be stored.

This cleverly saves an extra 1 bit of precision.

The key to understanding the three-segment structure: floating point number = (-1)^symbol× 1.mantissa × 2(Exponent-127)This is the entire essence of IEEE 754.


Interactive demo: 32-bit floating point number dissector

Enter a floating point number and observe its three-segment structure in IEEE 754 single precision format.

The segmented bar will highlight which bits each of the three segments corresponds to.

IEEE 754 single precision dissector

Input floating point number:

Try entering 0.1 and observe the mantissa part.

The binary mantissa of 0.1 is an infinitely repeating decimal--23 bits are completely insufficient to store it.

This leads to a famous phenomenon in computing.


Why 0.1 + 0.2 does not equal 0.3

If you run this in Python or JavaScript0.1 + 0.2 == 0.3, the result isFalse。

This is not a bug, but an inevitable consequence of floating-point representation.

The reason is very simple:0.1 cannot be exactly represented in binary。

Just like in decimal, 1/3 = 0.33333... is an infinitely repeating decimal.

In binary, 0.1 (decimal) = 0.000110011001100110011... (binary), which is also an infinitely repeating fraction.

IEEE 754 only has 23 bits of mantissa to store this infinitely repeating decimal.

The part beyond 23 bits can only be truncated (rounded), which produces rounding error.

The value actually stored for 0.1 in float32 is approximately 0.100000001490116119384765625.

The value actually stored for 0.2 in float32 is approximately 0.20000000298023223876953125.

After adding the two and truncating, the result is not exactly 0.3.

Amount calculations must avoid floating-point numbers

In scenarios involving money, you should use integers (in units of "cents") or a dedicated Decimal type.

Python'sDecimalJava'sBigDecimalC#'sdecimalare all designed for this.

If you store money with floating point numbers, by the time you discover that the accumulated error of 0.01 has become 0.010000000000000002, the seeds of the problem were sown long ago.


Range of floating-point numbers and special values

IEEE 754 reserves several special combinations for the exponent and mantissa:

CategorySign bitExponent (storage)mantissaMeaning
zero0 or 10 (all zeros)All zeros+0 or -0
Subnormal number0 or 10 (all zeros)Not all zerosValues very close to 0
Normalized number0 or 11 ~ 254AnyNormal floating-point numbers
Infinity0 or 1255 (all 1s)All zeros+Infinity or -Infinity
NaNAny255 (all 1s)Not all zerosNon-numeric values (e.g., 0/0)

Approximate range of single precision float (32 bits):

  • Maximum positive value: approximately 3.4 x 1038
  • Smallest positive value (normalized): approximately 1.2 x 10-38
  • Smallest positive value (denormalized): approximately 1.4 x 10-45
  • Effective precision: approximately 7 significant decimal digits

Double precision double (64 bits) has a larger range and higher precision:

  • 11-bit exponent (bias 1023)
  • 52-bit mantissa
  • Effective precision: approximately 15~16 significant decimal digits

Three cases of floating-point precision loss

Case 1: decimal to binary conversion error

Such as the earlier 0.1 example.

Any decimal that is "clean" in decimal may be infinitely repeating in binary.

Only fractions whose denominator is a power of 2 (such as 0.5, 0.25, 0.125) can be represented exactly in binary.

Case 2: Adding a small number to a large number

The precision of floating point numbers is not uniformly distributed--the larger the value, the larger the smallest difference that can be distinguished.

For example, 16777216.0 + 1.0 = 16777216.0 (in float32, the result remains unchanged).

Because a 32-bit floating point number can only represent about 7 significant digits exactly, and 16777216 itself already takes up 8 digits.

The change produced by adding 1 falls into the precision range that the mantissa cannot represent, and is directly "swallowed".

Case 3: Subtracting two nearly equal numbers

Subtracting two very close floating point numbers can lead toCatastrophic loss of significant digits。

For example, 1.2345678 - 1.2345677 = 0.0000001, but the actual computation may yield 0.000000099999..., a result with a large error.

Knowing the precision limits of floating point numbers is not to make you "afraid" to use them, but to let you choose the right tool for the right scenario. Scientific computing, graphics rendering, machine learning--floating point numbers are completely sufficient in these fields. But finance, accounting, precise counting--stay away from floating point numbers.


Code demonstration: the secrets of floating point

Example

# Complete Demo of the Floating-Point Precision Issue (Example Demo)
import struct

def float_to_ieee754_binary(num):
    Convert a floating-point number to IEEE 754 single-precision three-segment representation
    # Use struct to pack the floating-point number into 4 bytes (big-endian)
    packed = struct.pack('>f', num)
    # Convert to a 32-bit binary string
    bits = ''.join(f'{b:08b}' for b in packed)

    sign = bits[0]
    exponent = bits[1:9]
    mantissa = bits[9:32]

    exp_stored = int(exponent, 2)
    exp_actual = exp_stored - 127

    # Calculate the decimal value of the mantissa
    frac = 0
    for i, bit in enumerate(mantissa):
        if bit == '1':
            frac += 2 ** -(i + 1)

    value = (-1 if sign == '1' else 1) * (1 + frac) * (2 ** exp_actual)

    return {
        'bits': bits,
        'sign': sign,
        'exponent': exponent,
        'mantissa': mantissa,
        'exp_stored': exp_stored,
        'exp_actual': exp_actual,
        'fraction': frac,
        'reconstructed': value
    }

# =============================================
# Demo 1: Comparing the Representation of 0.5 and 0.1
# =============================================
print("=" * 60)
print("Demo 1: Why 0.5 is exact but 0.1 is not")
print("=" * 60)

for num in [0.5, 0.1]:
    r = float_to_ieee754_binary(num)
    print(f"\nFloating-point: {num}")
    print(fSign bit: {r['sign']} ({'positive' if r['sign'] == '0' else 'negative'}))
    print(f"  Exponentbits: {r['exponent']} (Storage={r['exp_stored']}, real际={r['exp_actual']})")
    print(fMantissa bits: {r['mantissa']})
    print(fActual stored value: {r['reconstructed']})

# =============================================
# Demo 2: The classic 0.1 + 0.2 != 0.3
# =============================================
print("\n" + "=" * 60)
print(Demo 2: 0.1 + 0.2 != 0.3)
print("=" * 60)

# Display with higher precision
import decimal
ctx = decimal.getcontext()
ctx.prec = 50

a = decimal.Decimal(0.1)
b = decimal.Decimal(0.2)
c = decimal.Decimal(0.3)

print(f"Actual value of 0.1: {a}")
print(f"Actual value of 0.2: {b}")
print(f"Actual value of 0.3: {c}")
print(f"0.1 + 0.2  = {a + b}")
print(f"0.1 + 0.2 == 0.3 ? {0.1 + 0.2 == 0.3}")

# =============================================
# Demo 3: Large numbers 'swallow' small numbers
# =============================================
print("\n" + "=" * 60)
print("Demo 3: Large number plus small number is swallowed")
print("=" * 60)

big = 16777216.0
small = 1.0
result = big + small
print(f"{big} + {small} = {result}")
print(f"Result equals original number? {result == big}")

# Show that this number reaches the precision limit of float32
print(f"\nReason: {big} In float32 Mediumrequires {len(format(struct.unpack('>I', struct.pack('>f', big))[0], '032b')[:23])} bitsmantissa")
print(fThe minimum interval between two adjacent float32 values is approximately 2^24 / 2^23 = 2.0)

# =============================================
# Demo 4: Correct case conversion techniques
# =============================================
print("\n" + "=" * 60)
print("Demo 4: EXAMPLE example - Correct handling of amounts")
print("=" * 60)

from decimal import Decimal

# Wrong approach: using floating-point numbers
float_total = 0.1 + 0.2
print(fError (floating point): 0.1 + 0.2 = {float_total})

# Correct approach: use Decimal
decimal_total = Decimal('0.1') + Decimal('0.2')
print(f"correctconfirm(Decimal): Decimal('0.1') + Decimal('0.2') = {decimal_total}")

# Or use integers (in cents)
cents_total = 10 + 20  # 0.1 yuan = 10 fen, 0.2 yuan = 20 fen
print(f"correctconfirm(IntegerDivide): 10Divide + 20Divide = {cents_total}Divide = {cents_total/100}元")

# =============================================
# Demo 5: Special values of floating-point numbers
# =============================================
print("\n" + "=" * 60)
print("Demo 5: Special values of floating-point numbers")
print("=" * 60)

import math

# Infinity
print(f"float('inf')    = {float('inf')}")
print(f"1.0 / 0.0       = {1.0 / 0.0}")
print(f"float('inf') > 1e1000000 = {float('inf') > 1e1000000}")

# NaN
print(f"float('nan')     = {float('nan')}")
print(f"0.0 / 0.0        = {0.0 / 0.0}")
print(f"float('nan') == float('nan') ? {float('nan') == float('nan')}")
print(f" (NaN is not equal to any value, including itself!)")

# Check for NaN
print(f"math.isnan(float('nan')) = {math.isnan(float('nan'))}")

Practical advice: when to use what

ScenariosRecommended typeReason
Scientific computing, graphics renderingfloat or doubleSufficient precision, good hardware acceleration support
Financial amount calculationDecimal / BigDecimalExact decimal arithmetic, no rounding errors
Counters, indexesint / longExact integers, no precision issues
Comparing floating-point numbersComparing with toleranceDo not use == directly; use abs(a-b) < epsilon
Currency storage/transferIntegers (in cents)Avoid floating-point errors during serialization
other extensions