How are characters represented -- ASCII and Unicode
In this lecture you will understand: how text is represented in a computer that can only store numbers, and why opening a file sometimes produces garbled text.
Everyday analogy: room numbering
Imagine a building with 128 rooms.
The front desk does not manage rooms by names like "Presidential Suite" or "Garden Room".
They assign each room a number: 101, 102, 103...
Look up number 101 and you know: oh, that's a standard king room.
Computers handle text in exactly the same way:
Assign a numeric code to each character, then store and process this number.
This "numbering scheme" is character encoding.
ASCII: The origin of all encodings
ASCII(American Standard Code for Information Interchange,United StatesInformation交换StandardCode)诞生in 1963 year。
It uses 7 bits to represent one character, for a total of 2^7 = 128 characters.
Contains:
- Control characters (0~31): invisible characters such as carriage return, line feed, and backspace
- Printable characters (32~126): space, punctuation, digits 0-9, uppercase letters A-Z, lowercase letters a-z
- DEL delete character (127)
ASCII complete table (0~127)

Key anchor points in ASCII
| Character category | ASCII Range | feature |
|---|---|---|
| '0' ~ '9' | 48 ~ 57 (0x30 ~ 0x39) | The code for digit 0 is 48, not 0. |
| 'A' ~ 'Z' | 65 ~ 90 (0x41 ~ 0x5A) | Uppercase letters are arranged consecutively |
| 'a' ~ 'z' | 97 ~ 122 (0x61 ~ 0x7A) | Lowercase letters are arranged consecutively |
| Newline '\n' | 10 (0x0A) | LF, Unix/Mac line break |
| Carriage return '\r' | 13 (0x0D) | CR, the line break used by old Macs |
The clever relationship between uppercase and lowercase
Note the binary representation of uppercase 'A' (65) and lowercase 'a' (97):
'a' = 01100001
Only bit 5 differs (value = 32, bit numbering from right to left, starting from 0)
This is not a coincidence.
At the binary level, changing bit 5 from 0 to 1 directly accomplishes the uppercase-to-lowercase conversion.
This clever design allowed early teletype machines to switch between uppercase and lowercase modes with just a simple circuit.
Bit 5 is counted using the convention of numbering bits from right to left starting at 0 (i.e., the LSB is bit 0), which is the most common bit-numbering convention in computers:

Uppercase and lowercase letters differ by 32, so at the ASCII level,
'a' - 'A' = 32,'A' | 0x20 = 'a'(Bitwise operation for case switching). This all stems from the clever layout of the ASCII encoding table.
ASCII's limitation: one byte is not enough
ASCII has only 128 characters, which is sufficient for English.
But when computers spread to Europe, problems arose:
French has é, German has ü, Spanish has ñ—these accented letters simply do not exist in ASCII.
So each country started "patching"—using the 128~255 range that ASCII did not use (with the 8th bit set to 1):
- Western Europe: ISO-8859-1 (Latin-1), which added Western European characters such as e and u
- Eastern Europe: ISO-8859-2, added Slavic letters
- Greece: ISO-8859-7, added Greek letters
- China: GB2312/GBK, uses two bytes to represent one Chinese character
- Japan: Shift-JIS
- Korea: EUC-KR
This became a disaster—The same binary data, when interpreted under different encodings, yields completely different text。
Open a Japanese document and see garbled text? That is the classic symptom of encoding mismatch.
Unicode: giving every character in the world a unique number
Unicode's goal is simple yet ambitious:
Assign a unique numeric code (Code Point) to every character in the world.。
Unicode is not 16-bit—its code point range is U+0000 to U+10FFFF, which can hold more than 1 million characters.
About 150,000 characters have been assigned so far, covering nearly all human writing systems, as well as emoji.
Unicode code point examples
| Character | Unicode code point | Meaning | Assigned block |
|---|---|---|---|
| A | U+0041 | Latin uppercase letter A | Basic Latin alphabet (ASCII-compatible) |
| you | U+4F60 | CJK Chinese character「你」 | CJK Unified Ideographs |
| OK | U+597D | CJK Chinese character「好」 | CJK Unified Ideographs |
| hello | U+68D2 | Arabic “Hello” | Arabic |
| example | U+E0001 | Language tag | Tags block |
Note: Unicode is just a "numbering scheme"—it tells you "character X has the number U+XXXX".
It does not specify "how the number U+XXXX is stored in a file".
Storing to files is the job of "encoding schemes" such as UTF-8 and UTF-16.
Many people confuse Unicode with UTF-8. Remember this distinction: Unicode is a dictionary (defining which character corresponds to which number), while UTF-8 is a writing rule (specifying how a number is written as a byte sequence).
UTF-8: the most popular Unicode encoding scheme
UTF-8's design is extremely clever. It is aVariable-length encoding:
- ASCII characters (U+0000 ~ U+007F): stored in 1 byte
- Most European and Arabic characters: use 2 bytes
- Chinese, Japanese, and Korean: use 3 bytes
- Supplementary characters (including most emoji): use 4 bytes
UTF-8 encoding rules
Of the first byteleading bit patternIt tells the decoder: how many bytes this character occupies.
- Starting with 0: single-byte character
- Starting with 110: two-byte character
- Starts with 1110: three-byte character
- Starts with 11110: four-byte character
- Subsequent bytes all start with 10
UTF-8’s biggest advantage:Fully compatible with ASCII。
Any valid ASCII text is also valid UTF-8 text.
This property allows UTF-8 to smoothly replace the old ASCII system without breaking any existing data.
Interactive demo: real-time encoding lookup
Enter any text below to view each character's Unicode code point, UTF-8 encoded bytes, and space occupied.
Operation suggestions:
- Enter English-only text and observe that each character takes only 1 byte.
- Enter Chinese text and observe that each Chinese character takes 3 bytes.
- Input
Hello EXAMPLENote whether the ASCII codes of the letters are arranged consecutively. - Compare the binary representations of uppercase and lowercase letters, and find the corresponding letter pairs that differ by only 1 bit (such as A and a).
Encoding garbling problem: principles and prevention
The essence of garbled text can be summarized in one sentence:Using the wrong dictionary to interpret data。
Suppose a file contains three bytes stored in UTF-8 encoding:E4 BD A0
- Decode using UTF-8: these 3 bytes form one character → "you"
- Decode with GBK: E4 BD is treated as a GBK two-byte character → might be some unrelated Chinese character
- Decoding with Latin-1: E4, BD, A0 are treated as three independent Western European characters → garbled text like "a + 1/2 + nbsp"
The same binary data, decoded with three different dictionaries, yields three different texts.
How to avoid garbled text?
- Use UTF-8 uniformlyModern Web, APIs, and databases almost all default to UTF-8
- Explicitly declare encoding: Write in HTML
<meta charset="UTF-8">, write in the HTTP response headerContent-Type: text/html; charset=utf-8 - editor settingsEditors such as VS Code and Sublime all support setting file encoding to UTF-8
BOM (Byte Order Mark) issue
Although UTF-8 does not need to worry about byte order (because it is byte-based), some Windows programs add a BOM at the beginning of UTF-8 files:EF BB BF。
These three bytes are used to mark that "this file is UTF-8 encoded".
But on Unix/Linux systems, the BOM can cause various problems:
- The first line of a Shell script
#!/bin/bashThe extra three invisible bytes at the front cause the script to fail to execute - If a PHP file has a BOM, it may trigger a "headers already sent" error
- JSON parsers may reject JSON data that starts with a BOM
Recommendation: use the "no BOM" format for UTF-8 files (UTF-8 without BOM), which is also the default setting of most modern editors.
Code demonstration: encoding and decoding
Example
import sys
# =============================================
# Demo 1: Basic usage of ord() and chr()
# =============================================
print("=" * 60)
print("Demo 1: Conversion between characters and code points")
print("=" * 60)
# ord() gets the Unicode code point of a character
text = "Hello EXAMPLE"
for ch in text:
cp = ord(ch)
print(f'{ch}' → code point U+{cp:04X} (decimal {cp}))
print()
# chr() gets character based on code point
for cp in [65, 97, 48, 0x4F60, 0x597D]:
ch = chr(cp)
print(f" Code point U+{cp:04X} → character '{ch}'")
# =============================================
# Demo 2: encode as byte sequence (encode)
# =============================================
print("\n" + "=" * 60)
print(Demo 2: String → UTF-8 Byte Sequence)
print("=" * 60)
texts = ["A", "Hello", Hello, "EXAMPLE Tutorial"]
for text in texts:
utf8_bytes = text.encode('utf-8')
hex_str = utf8_bytes.hex(' ').upper()
print(f"\nText: '{text}'")
print(f" Character count: {len(text)}")
print(fUTF-8 byte count: {len(utf8_bytes)})
print(f" UTF-8 hexadecimal: {hex_str}")
# Byte-by-byte analysis
print(f" Byte Decomposition:")
for i, b in enumerate(utf8_bytes):
bin_str = format(b, '08b')
# Determine Byte Type
if b < 0x80:
typ = "ASCII single byte"
elif b >= 0xC0 and b < 0xE0:
typ = "Double-Byte Leading Byte"
elif b >= 0xE0 and b < 0xF0:
typ = "Three-Byte Leading Byte"
elif b >= 0x80 and b < 0xC0:
typ = "Continuation bytes"
else:
typ = "Other"
print(fByte {i}: 0x{b:02X} = {bin_str} ({typ}))
# =============================================
# Demo 3: GBK encoding (a non-Unicode encoding commonly used in China)
# =============================================
print("\n" + "=" * 60)
print(Demo 3: GBK Encoding (Compared to UTF-8))
print("=" * 60)
text = "Hello EXAMPLE"
utf8_bytes = text.encode('utf-8')
gbk_bytes = text.encode('gbk')
print(fText: '{text}')
print(f" UTF-8 ({len(utf8_bytes)} byte): {utf8_bytes.hex(' ').upper()}")
print(f" GBK ({len(gbk_bytes)} byte): {gbk_bytes.hex(' ').upper()}")
# In GBK, each Chinese character occupies 2 bytes
print(f"\nNote: In GBK, Chinese characters occupy 2 bytes; in UTF-8, Chinese characters occupy 3 bytes.")
print(f" For Chinese text, GBK is more compact than UTF-8")
print(fBut GBK does not support all Unicode characters (such as emoji).)
# =============================================
# Demo 4: decode error scenario
# =============================================
print("\n" + "=" * 60)
print("Demo 4: decoding error scenario (cause of garbled text)")
print("=" * 60)
# Encode Chinese with UTF-8
original = Hello
utf8_data = original.encode('utf-8')
print(f"Original text: '{original}'")
print(fUTF-8 encoding: {utf8_data.hex(' ').upper()})
# Error 1: decoding UTF-8 data with GBK
try:
decoded_wrong = utf8_data.decode('gbk')
print(fDecoded with GBK: '{decoded_wrong}' (garbled!))
except Exception as e:
print(fFailed to decode with GBK: {e})
# Error 2: Decode with Latin-1
decoded_latin = utf8_data.decode('latin-1')
print(fDecoded with Latin-1: '{decoded_latin}' (looks like garbled Western European characters))
# =============================================
# Demo 5: Create a reasonable decoder
# =============================================
print("\n" + "=" * 60)
print(Demo 5: Encoding detection and safe decoding)
print("=" * 60)
def safe_decode(data, encodings=['utf-8', 'gbk', 'latin-1']):
"""Try decoding with multiple encodings, return the first successful result"""
for enc in encodings:
try:
return data.decode(enc), enc
except (UnicodeDecodeError, LookupError):
continue
return data.decode('latin-1', errors='replace'), 'latin-1 (fallback)'
# Test
test_cases = [
"Hello EXAMPLE".encode('utf-8'),
"Hello World".encode('utf-8'),
"Computer".encode('gbk'),
]
for data in test_cases:
result, encoding = safe_decode(data)
print(fData ({len(data)} bytes) decoded with {encoding}: '{result}')
print("\n" + "=" * 60)
print(Summary: Strings are Python's internal representation; encoding is the format for storage/transmission.)
print(Always do encoding conversion at the I/O boundary, and always use the str type internally.)
print("=" * 60)
Current status and future of character encoding
As of today,UTF-8 has become the de facto standard on the Internet.。
Over 98% of web pages use UTF-8 encoding.
The JSON specification requires UTF-8.
Almost all modern programming languages use UTF-8 by default to process text.
Unicode is still expanding.
New versions of Unicode add new characters each year, including: rare scripts, historical scripts, and new emoji.
Encoding has shifted from the question of "whether it can be represented" to "how to store and transmit efficiently".
other extensionsIf you start a new project today, choosing UTF-8 won't go wrong. It is an encoding scheme that is backward compatible with ASCII, cross-platform, and supports all languages. The only thing to note is—make sure your entire stack (frontend, backend, database) all uses UTF-8, and don't mix.