Linux split command
splitsplit is a file splitting tool built into Linux that can split a file by line count, size, or a specified quantity.
splitThe command is used to split a large file into several smaller files, making it easy to transfer, store, or process in parallel.
The split files can be reassembled using thecatcommand to merge them back into the original file.
When log files are too large to open with an editor, when a large file needs to be uploaded in chunks, or when data needs to be processed in parallel,splitsplit is the preferred solution.
Split files are named by default withxaa、xab、xacand other alphabetic sequences.
splitsplit does not delete or modify the original file, so the split operation is safe. However, when splitting into many small files, make sure there is sufficient disk space.
Command syntax
splitThe basic syntax of split is as follows:
split [选项] [输入文件] [输出文件前缀]
If no input file is specified, data is read from standard input.
If no output file prefix is specified, the default prefix isxused.
splitThe common options of split are listed below:
| Option | Description | Example |
|---|---|---|
-b, --bytes=大小 | Split the file by the specified number of bytes | split -b 100M large.log |
-l, --lines=行数 | Split the file by the specified number of lines | split -l 1000 data.txt |
-n, --number=数量 | Split into the specified number of small files | split -n 5 data.txt |
-a, --suffix-length=N | Specify the suffix length (default is 2) | split -a 3 -l 100 data.txt |
-d, --numeric-suffixes | Use numeric suffixes instead of alphabetic suffixes | split -d -l 100 data.txt |
-C, --line-bytes=大小 | Split by size while keeping each line complete | split -C 10M log.txt |
--verbose | Display detailed information about the split process | split --verbose -b 10M file.bin |
Supported size units include:K(KB)、M(MB)、G(GB)、T(TB) and other suffixes.
Difference between -b and -C:
| Option | Behavior | Applicable scenario |
|---|---|---|
-b | Strictly cuts by byte count and may truncate in the middle of a line | Binary files or scenarios where line breaks do not matter |
-C | Tries to cut at line boundaries when reaching the size limit, ensuring each line is complete | Text logs, CSV, and other scenarios where lines must remain intact |
When splitting text files, prefer
-Cover-bto avoid a line being cut off and split between two files.
Detailed usage
Split by line count
Splitting by line count is the most intuitive way and is suitable for processing log files or CSV data.
# 生成一个包含 5000 行的测试文件 $ seq 1 5000 > data.txt # 每 1000 行拆分为一个小文件 $ split -l 1000 data.txt part_ # 查看生成的文件 $ ls -lh part_* -rw-r--r-- 1 example example 3.9K May 19 14:30 part_aa -rw-r--r-- 1 example example 3.9K May 19 14:30 part_ab -rw-r--r-- 1 example example 3.9K May 19 14:30 part_ac -rw-r--r-- 1 example example 3.9K May 19 14:30 part_ad -rw-r--r-- 1 example example 3.9K May 19 14:30 part_ae
After running, 5 files are generated, each with exactly 1000 lines, and the filenames usepart_as the prefix.
Split by file size
Splitting by size is suitable for cutting binary files or limiting the maximum size of a single file.
# 生成一个 10MB 的测试文件 $ dd if=/dev/urandom of=test.bin bs=1M count=10 # 按每 2MB 拆分 $ split -b 2M test.bin chunk_ # 查看拆分结果 $ ls -lh chunk_* -rw-r--r-- 1 example example 2.0M May 19 14:32 chunk_aa -rw-r--r-- 1 example example 2.0M May 19 14:32 chunk_ab -rw-r--r-- 1 example example 2.0M May 19 14:32 chunk_ac -rw-r--r-- 1 example example 2.0M May 19 14:32 chunk_ad -rw-r--r-- 1 example example 2.0M May 19 14:32 chunk_ae
Specify the number of splits
Use-nto divide the file evenly into the specified number of small files.
# 将文件均分为 3 份 $ split -n 3 data.txt equal_ # 查看各文件行数 $ wc -l equal_* 1667 equal_aa 1667 equal_ab 1666 equal_ac 5000 total
The file is distributed as evenly as possible; a file with a total of 5000 lines is divided into about 1667 lines per file.
When the total number of lines is not evenly divisible, the earlier files get a few extra lines and the later files get a few fewer.
Use numeric suffixes
The default alphabetic suffix (aa、ab...) is not intuitive enough; you can use-dto switch to numeric suffixes.
# 使用数字后缀,并指定后缀长度为 3 $ split -d -a 3 -l 1000 data.txt example_ # 生成的文件名 $ ls example_* example_000 example_001 example_002 example_003 example_004
Here, a suffix length of 3 means up to 1000 files are supported (from 000 to 999). When splitting into many small files, you can increase the value of-athe option.
Split text while preserving line integrity
Use-Cto split by size while ensuring that no line is truncated.
# 生成一个包含不同长度行的测试日志 $ for i in $(seq 1 100); do echo "Line $i: EXAMPLE testing data $(head -c $((RANDOM % 50 + 10)) /dev/urandom | base64)"; done > log.txt # 按 1KB 拆分,同时保持行完整性 $ split -C 1K log.txt log_ # 检查每个文件的最后一行是否完整(都以换行符结尾) $ for f in log_*; do echo "$f: $(tail -c 1 $f | xxd | grep -c '0a')"; done
-CThis ensures that the last line of each split file is complete, and lines are never cut in half.
Merge split files
The split files can be merged losslessly to restore the original file:
# 合并所有拆分文件,还原为原始文件 $ cat part_* > restored.txt # 验证合并后的内容与原始文件一致 $ diff data.txt restored.txt && echo "文件一致,合并成功" 文件一致,合并成功
When merging,catcat concatenates them in the alphabetical order of shell wildcard expansion, which is exactly thesplitorder in which split generates the files.
FAQ
The split files still occupy the same total disk space as the original file (actually, there is some additional filesystem metadata overhead), so make sure there is enough free disk space.
When using numeric suffixes, if the number of split files exceeds the maximum representable by the suffix length (e.g.,
-a 2up to 100 files),splitsplit will exit with an error. Increase the value of the-aparameter to fix it.
If you need to split by specific content (such as by a delimiter),splitsplit does not support it; use thecsplitcsplit command instead.
When reading data from standard input,splitsplit cannot estimate the total size, so the-noption for specifying the number of splits does not support standard input mode.
Related commands
| Command | Description |
|---|---|
cat | Merge files or output file contents |
csplit | Split files by content (regular expression) |
wc | Count the lines, words, and bytes in a file |
dd | Convert and copy files, and extract data by block size |
head | Output the beginning of a file |
tail | Output the end of a file |
Linux Command Reference