Linux iconv command

Linux 命令大全Linux Command Encyclopedia


iconv is a command-line tool in the Linux system, used to convert the content of text files between different character encodings. It can handle various common character encoding formats, such as UTF-8, GB2312, ISO-8859, etc., solving the problem of text encoding incompatibility between different systems.


Why character encoding conversion is needed

Character encoding issues often lead to the following situations:

  • Text files copied from Windows to Linux appear garbled
  • Web content displays abnormally in different browsers
  • Source code files developed cross-platform encounter encoding issues
  • Compatibility issues when handling internationalized multilingual text

The iconv command was designed precisely to solve these problems.


Basic Syntax

iconv [选项] -f 原编码 -t 目标编码 [输入文件]

Common Option Parameters

Option Description
-f Specify the original file's character encoding (from)
-t Specify the target character encoding to convert to (to)
-o Specify the output file (default output to standard output)
-l List all supported encoding formats
-c Silently ignore unconvertible characters (error by default)
--verbose Display detailed information during conversion

Supported Encoding Formats

To view all encoding formats supported by the system, run:

iconv -l

Common encoding formats include:

  • UTF-8
  • GB2312
  • GBK
  • GB18030
  • BIG5
  • ISO-8859-1 (Latin-1)
  • ASCII
  • EUC-JP (Japanese)
  • SHIFT_JIS (Japanese)

Practical Application Examples

Example 1: Basic Encoding Conversion

Convert a GB2312-encoded file to UTF-8:

iconv -f GB2312 -t UTF-8 input.txt -o output.txt

Example 2: Handling Standard Input/Output

Convert text through a pipe:

cat gb2312_file.txt | iconv -f GB2312 -t UTF-8

Example 3: Ignoring Unconvertible Characters

iconv -f GBK -t UTF-8//IGNORE input.txt -o output.txt

Example 4: Batch Converting All Files in a Directory

Example

for file in *.txt; do
    iconv -f GB2312 -t UTF-8 "$file" -o "utf8_${file}"
done

Common Problem Solutions

Problem 1: Encoding Recognition Error

If you don't know the original encoding of a file, you can first try common encodings:

Example

# Try GB2312
iconv -f GB2312 -t UTF-8 input.txt

# If it fails, try GBK
iconv -f GBK -t UTF-8 input.txt

Problem 2: Garbled Characters After Conversion

The encoding may have been specified incorrectly, try:

Example

# Use //TRANSLIT to handle special characters
iconv -f GBK -t UTF-8//TRANSLIT input.txt

# Or use //IGNORE to ignore unconvertible characters
iconv -f GBK -t UTF-8//IGNORE input.txt

Problem 3: Insufficient Memory When Converting Large Files

For large files, you can use split processing:

Example

split -l 10000 bigfile.txt part_
for part in part_*; do
    iconv -f GB2312 -t UTF-8 "$part" -o "utf8_${part}"
done
cat utf8_part_* > bigfile_utf8.txt

Best Practice Recommendations

  1. Backup the original file: Backup before conversion to prevent data loss
  2. Test the conversion: First test the conversion effect with a small file
  3. Standardize encoding: Use a unified encoding standard in projects (UTF-8 recommended)
  4. Check the result: After conversion, usefilecommand to check file encoding
  5. Automation: Write commonly used conversion commands into scripts for easy reuse

Combining with Other Tools

Combining with find command for batch conversion

find . -name "*.txt" -exec bash -c 'iconv -f GB2312 -t UTF-8 "{}" > "{}.utf8"' ;

Combining with vim to check encoding

vim -c "set fileencoding" filename.txt

Using file command to detect encoding

file -i filename.txt

Advanced Tips

Converting filename encoding

Example

# Convert GBK-encoded filenames to UTF-8
convmv -f GBK -t UTF-8 --notest *.txt

Handling HTML/XML files

Example

# Preserve the encoding declaration in the file
iconv -f GB2312 -t UTF-8 input.html |
sed 's/charset=gb2312/charset=utf-8/i' > output.html

Creating an encoding conversion alias

In~/.bashrcAdd in:

Example

alias gb2utf8='iconv -f GB2312 -t UTF-8'
alias big52utf8='iconv -f BIG5 -t UTF-8'

Then executesource ~/.bashrcto activate the alias.


Summary

iconv is a powerful tool for handling text encoding issues in Linux systems. Through this article, you should be able to:

  1. Understand the basic concepts of character encoding conversion
  2. Master the basic syntax and common options of the iconv command
  3. Solve encoding conversion problems in daily work
  4. Apply advanced techniques to handle complex scenarios

Remember to always backup before processing important files, and verify results after conversion. UTF-8, as a universal encoding standard, is the recommended choice for most modern applications.


Linux 命令大全Linux Command Encyclopedia

Other Extensions