How are characters represented -- ASCII and Unicode

In this lecture you will understand: how text is represented in a computer that can only store numbers, and why opening a file sometimes produces garbled text.


Everyday analogy: room numbering

Imagine a building with 128 rooms.

The front desk does not manage rooms by names like "Presidential Suite" or "Garden Room".

They assign each room a number: 101, 102, 103...

Look up number 101 and you know: oh, that's a standard king room.

Computers handle text in exactly the same way:

Assign a numeric code to each character, then store and process this number.

This "numbering scheme" is character encoding.


ASCII: The origin of all encodings

ASCII(American Standard Code for Information Interchange,United StatesInformation交换StandardCode)诞生in 1963 year。

It uses 7 bits to represent one character, for a total of 2^7 = 128 characters.

Contains:

  • Control characters (0~31): invisible characters such as carriage return, line feed, and backspace
  • Printable characters (32~126): space, punctuation, digits 0-9, uppercase letters A-Z, lowercase letters a-z
  • DEL delete character (127)

ASCII complete table (0~127)

Key anchor points in ASCII

Character categoryASCII Rangefeature
'0' ~ '9'48 ~ 57 (0x30 ~ 0x39)The code for digit 0 is 48, not 0.
'A' ~ 'Z'65 ~ 90 (0x41 ~ 0x5A)Uppercase letters are arranged consecutively
'a' ~ 'z'97 ~ 122 (0x61 ~ 0x7A)Lowercase letters are arranged consecutively
Newline '\n'10 (0x0A)LF, Unix/Mac line break
Carriage return '\r'13 (0x0D)CR, the line break used by old Macs

The clever relationship between uppercase and lowercase

Note the binary representation of uppercase 'A' (65) and lowercase 'a' (97):

'A' = 01000001
'a' = 01100001
Only bit 5 differs (value = 32, bit numbering from right to left, starting from 0)

This is not a coincidence.

At the binary level, changing bit 5 from 0 to 1 directly accomplishes the uppercase-to-lowercase conversion.

This clever design allowed early teletype machines to switch between uppercase and lowercase modes with just a simple circuit.

Bit 5 is counted using the convention of numbering bits from right to left starting at 0 (i.e., the LSB is bit 0), which is the most common bit-numbering convention in computers:

Uppercase and lowercase letters differ by 32, so at the ASCII level,'a' - 'A' = 32,'A' | 0x20 = 'a'(Bitwise operation for case switching). This all stems from the clever layout of the ASCII encoding table.


ASCII's limitation: one byte is not enough

ASCII has only 128 characters, which is sufficient for English.

But when computers spread to Europe, problems arose:

French has é, German has ü, Spanish has ñ—these accented letters simply do not exist in ASCII.

So each country started "patching"—using the 128~255 range that ASCII did not use (with the 8th bit set to 1):

  • Western Europe: ISO-8859-1 (Latin-1), which added Western European characters such as e and u
  • Eastern Europe: ISO-8859-2, added Slavic letters
  • Greece: ISO-8859-7, added Greek letters
  • China: GB2312/GBK, uses two bytes to represent one Chinese character
  • Japan: Shift-JIS
  • Korea: EUC-KR

This became a disaster—The same binary data, when interpreted under different encodings, yields completely different text。

Open a Japanese document and see garbled text? That is the classic symptom of encoding mismatch.


Unicode: giving every character in the world a unique number

Unicode's goal is simple yet ambitious:

Assign a unique numeric code (Code Point) to every character in the world.。

Unicode is not 16-bit—its code point range is U+0000 to U+10FFFF, which can hold more than 1 million characters.

About 150,000 characters have been assigned so far, covering nearly all human writing systems, as well as emoji.

Unicode code point examples

CharacterUnicode code pointMeaningAssigned block
AU+0041Latin uppercase letter ABasic Latin alphabet (ASCII-compatible)
youU+4F60CJK Chinese character「你」CJK Unified Ideographs
OKU+597DCJK Chinese character「好」CJK Unified Ideographs
helloU+68D2Arabic “Hello”Arabic
exampleU+E0001Language tagTags block

Note: Unicode is just a "numbering scheme"—it tells you "character X has the number U+XXXX".

It does not specify "how the number U+XXXX is stored in a file".

Storing to files is the job of "encoding schemes" such as UTF-8 and UTF-16.

Many people confuse Unicode with UTF-8. Remember this distinction: Unicode is a dictionary (defining which character corresponds to which number), while UTF-8 is a writing rule (specifying how a number is written as a byte sequence).


UTF-8: the most popular Unicode encoding scheme

UTF-8's design is extremely clever. It is aVariable-length encoding:

  • ASCII characters (U+0000 ~ U+007F): stored in 1 byte
  • Most European and Arabic characters: use 2 bytes
  • Chinese, Japanese, and Korean: use 3 bytes
  • Supplementary characters (including most emoji): use 4 bytes

UTF-8 encoding rules

U+0000~U+007F 0xxxxxxx 1 byte, 7-bit data
U+0080~U+07FF 110xxxxx 10xxxxxx 2 bytes, 11-bit data
U+0800~U+FFFF 1110xxxx 10xxxxxx 10xxxxxx 3 bytes, 16-bit data
U+10000~U+10FFFF 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx 4 bytes, 21-bit data

Of the first byteleading bit patternIt tells the decoder: how many bytes this character occupies.

  • Starting with 0: single-byte character
  • Starting with 110: two-byte character
  • Starts with 1110: three-byte character
  • Starts with 11110: four-byte character
  • Subsequent bytes all start with 10

UTF-8’s biggest advantage:Fully compatible with ASCII。

Any valid ASCII text is also valid UTF-8 text.

This property allows UTF-8 to smoothly replace the old ASCII system without breaking any existing data.


Interactive demo: real-time encoding lookup

Enter any text below to view each character's Unicode code point, UTF-8 encoded bytes, and space occupied.

Quick fill: Chinese-English mix Uppercase A Lowercase a Numbers All English All Chinese Case comparison

Operation suggestions:

  1. Enter English-only text and observe that each character takes only 1 byte.
  2. Enter Chinese text and observe that each Chinese character takes 3 bytes.
  3. InputHello EXAMPLENote whether the ASCII codes of the letters are arranged consecutively.
  4. Compare the binary representations of uppercase and lowercase letters, and find the corresponding letter pairs that differ by only 1 bit (such as A and a).

Encoding garbling problem: principles and prevention

The essence of garbled text can be summarized in one sentence:Using the wrong dictionary to interpret data。

Suppose a file contains three bytes stored in UTF-8 encoding:E4 BD A0

  • Decode using UTF-8: these 3 bytes form one character → "you"
  • Decode with GBK: E4 BD is treated as a GBK two-byte character → might be some unrelated Chinese character
  • Decoding with Latin-1: E4, BD, A0 are treated as three independent Western European characters → garbled text like "a + 1/2 + nbsp"

The same binary data, decoded with three different dictionaries, yields three different texts.

How to avoid garbled text?

  1. Use UTF-8 uniformlyModern Web, APIs, and databases almost all default to UTF-8
  2. Explicitly declare encoding: Write in HTML<meta charset="UTF-8">, write in the HTTP response headerContent-Type: text/html; charset=utf-8
  3. editor settingsEditors such as VS Code and Sublime all support setting file encoding to UTF-8

BOM (Byte Order Mark) issue

Although UTF-8 does not need to worry about byte order (because it is byte-based), some Windows programs add a BOM at the beginning of UTF-8 files:EF BB BF。

These three bytes are used to mark that "this file is UTF-8 encoded".

But on Unix/Linux systems, the BOM can cause various problems:

  • The first line of a Shell script#!/bin/bashThe extra three invisible bytes at the front cause the script to fail to execute
  • If a PHP file has a BOM, it may trigger a "headers already sent" error
  • JSON parsers may reject JSON data that starts with a BOM

Recommendation: use the "no BOM" format for UTF-8 files (UTF-8 without BOM), which is also the default setting of most modern editors.


Code demonstration: encoding and decoding

Example

# Character encoding complete demonstration (example demo)
import sys

# =============================================
# Demo 1: Basic usage of ord() and chr()
# =============================================
print("=" * 60)
print("Demo 1: Conversion between characters and code points")
print("=" * 60)

# ord() gets the Unicode code point of a character
text = "Hello EXAMPLE"
for ch in text:
    cp = ord(ch)
    print(f'{ch}' → code point U+{cp:04X} (decimal {cp}))

print()

# chr() gets character based on code point
for cp in [65, 97, 48, 0x4F60, 0x597D]:
    ch = chr(cp)
    print(f" Code point U+{cp:04X} → character '{ch}'")

# =============================================
# Demo 2: encode as byte sequence (encode)
# =============================================
print("\n" + "=" * 60)
print(Demo 2: String → UTF-8 Byte Sequence)
print("=" * 60)

texts = ["A", "Hello", Hello, "EXAMPLE Tutorial"]

for text in texts:
    utf8_bytes = text.encode('utf-8')
    hex_str = utf8_bytes.hex(' ').upper()
    print(f"\nText: '{text}'")
    print(f" Character count: {len(text)}")
    print(fUTF-8 byte count: {len(utf8_bytes)})
    print(f" UTF-8 hexadecimal: {hex_str}")

    # Byte-by-byte analysis
    print(f" Byte Decomposition:")
    for i, b in enumerate(utf8_bytes):
        bin_str = format(b, '08b')
        # Determine Byte Type
        if b < 0x80:
            typ = "ASCII single byte"
        elif b >= 0xC0 and b < 0xE0:
            typ = "Double-Byte Leading Byte"
        elif b >= 0xE0 and b < 0xF0:
            typ = "Three-Byte Leading Byte"
        elif b >= 0x80 and b < 0xC0:
            typ = "Continuation bytes"
        else:
            typ = "Other"
        print(fByte {i}: 0x{b:02X} = {bin_str} ({typ}))

# =============================================
# Demo 3: GBK encoding (a non-Unicode encoding commonly used in China)
# =============================================
print("\n" + "=" * 60)
print(Demo 3: GBK Encoding (Compared to UTF-8))
print("=" * 60)

text = "Hello EXAMPLE"
utf8_bytes = text.encode('utf-8')
gbk_bytes = text.encode('gbk')

print(fText: '{text}')
print(f"  UTF-8 ({len(utf8_bytes)} byte): {utf8_bytes.hex(' ').upper()}")
print(f"  GBK   ({len(gbk_bytes)} byte): {gbk_bytes.hex(' ').upper()}")

# In GBK, each Chinese character occupies 2 bytes
print(f"\nNote: In GBK, Chinese characters occupy 2 bytes; in UTF-8, Chinese characters occupy 3 bytes.")
print(f" For Chinese text, GBK is more compact than UTF-8")
print(fBut GBK does not support all Unicode characters (such as emoji).)

# =============================================
# Demo 4: decode error scenario
# =============================================
print("\n" + "=" * 60)
print("Demo 4: decoding error scenario (cause of garbled text)")
print("=" * 60)

# Encode Chinese with UTF-8
original = Hello
utf8_data = original.encode('utf-8')
print(f"Original text: '{original}'")
print(fUTF-8 encoding: {utf8_data.hex(' ').upper()})

# Error 1: decoding UTF-8 data with GBK
try:
    decoded_wrong = utf8_data.decode('gbk')
    print(fDecoded with GBK: '{decoded_wrong}' (garbled!))
except Exception as e:
    print(fFailed to decode with GBK: {e})

# Error 2: Decode with Latin-1
decoded_latin = utf8_data.decode('latin-1')
print(fDecoded with Latin-1: '{decoded_latin}' (looks like garbled Western European characters))

# =============================================
# Demo 5: Create a reasonable decoder
# =============================================
print("\n" + "=" * 60)
print(Demo 5: Encoding detection and safe decoding)
print("=" * 60)

def safe_decode(data, encodings=['utf-8', 'gbk', 'latin-1']):
    """Try decoding with multiple encodings, return the first successful result"""
    for enc in encodings:
        try:
            return data.decode(enc), enc
        except (UnicodeDecodeError, LookupError):
            continue
    return data.decode('latin-1', errors='replace'), 'latin-1 (fallback)'

# Test
test_cases = [
    "Hello EXAMPLE".encode('utf-8'),
    "Hello World".encode('utf-8'),
    "Computer".encode('gbk'),
]

for data in test_cases:
    result, encoding = safe_decode(data)
    print(fData ({len(data)} bytes) decoded with {encoding}: '{result}')

print("\n" + "=" * 60)
print(Summary: Strings are Python's internal representation; encoding is the format for storage/transmission.)
print(Always do encoding conversion at the I/O boundary, and always use the str type internally.)
print("=" * 60)

Current status and future of character encoding

As of today,UTF-8 has become the de facto standard on the Internet.。

Over 98% of web pages use UTF-8 encoding.

The JSON specification requires UTF-8.

Almost all modern programming languages use UTF-8 by default to process text.

Unicode is still expanding.

New versions of Unicode add new characters each year, including: rare scripts, historical scripts, and new emoji.

Encoding has shifted from the question of "whether it can be represented" to "how to store and transmit efficiently".

If you start a new project today, choosing UTF-8 won't go wrong. It is an encoding scheme that is backward compatible with ASCII, cross-platform, and supports all languages. The only thing to note is—make sure your entire stack (frontend, backend, database) all uses UTF-8, and don't mix.

other extensions