HTML Character Set
To display an HTML page correctly, the browser must know which character set (character encoding) to use.
HTML Character Sets
In HTML, what is the correct character encoding?
The default character encoding in HTML5 is UTF-8.
This has not always been the case. The character encoding for the early web was ASCII.
Later, from HTML 2.0 to HTML 4.01, ISO-8859-1 was recognized as the standard.
With the arrival of XML and HTML5, UTF-8 finally came, solving a large number of character encoding problems.
Below is a brief overview of character encoding standards.
In the beginning: ASCII
Computer information (numbers, text, pictures) is stored electronically in binary 1s and 0s (01000101).
To standardize the storage of alphanumeric characters, ASCII (full name: American Standard Code for Information Interchange) was created. It defines a unique binary 7-bit number for each stored character, supporting digits 0-9, upper/lowercase English letters (a-z, A-Z), and some special characters such as ! $ + - ( ) @ < > .
Since ASCII uses one byte (7 bits for the character, 1 bit for transmission parity control), it can only represent 128 different characters. Of these characters, 32 are reserved for other control purposes.
The biggest disadvantage of ASCII is that it excludes non-English letters.
ASCII is still widely used today, especially in large computer systems.
For more information about ASCII, seeComplete ASCII Reference Manual。
In Windows: ANSI
ANSI (also called Windows-1252) is the default character set in Windows 95 and earlier Windows systems.
ANSI is an extension of ASCII, adding international characters. It uses a full byte (8 bits) to represent 256 different characters.
Since ANSI became the default character set in Windows, all browsers support ANSI.
For more information about ANSI, seeComplete ANSI Reference Manual。
In HTML 4: ISO-8859-1
Because most countries use characters outside ASCII, the default character encoding was changed to ISO-8859-1 in the HTML 2.0 standard.
ISO-8859-1 is an extension of ASCII, adding international characters. Like ANSI, it uses a full byte (8 bits) to represent 256 different characters.
![]() |
When a browser detects ISO-8859-1 in a web page, it usually defaults to ANSI, because ANSI is essentially equivalent to ISO-8859-1 except that ANSI has 32 extra characters. |
|---|
If an HTML 4 page uses a character set other than ISO-8859-1, it needs to be specified in the <meta> tag, as follows:
Example
![]() |
The default character set in HTML5 is UTF-8. |
|---|
For more information about ISO-8859-1, seeComplete ISO-8859-1 Reference Manual。
In HTML5: Unicode (UTF-8)
Since the character sets listed above are limited and incompatible in multilingual environments, the Unicode Consortium developed the Unicode Standard.
The Unicode Standard covers (almost) all characters, punctuation marks, and symbols.
Unicode enables the processing, storage, and transport of text, independent of platform and language.
The default character encoding in HTML5 is UTF-8.
For more information about Unicode (UTF-8), seeComplete Unicode Reference Manual。
Other extensions
