Linux awk command
awk is a language for processing text files and a powerful text analysis tool.
awk reads text files line by line and provides programming-language-like features, such as:
- Variable definition and calculation
- Conditional judgments and loops
- String processing and formatted output
These features make AWK very efficient at processing structured text (such as CSV, log files).
It is called awk because it takes the first characters of the family names of its three founders: Alfred Aho, Peter Weinberger, and Brian Kernighan.
Syntax
awk options 'pattern {action}' file
Option parameter description:
options: are some options used to controlawkbehavior.pattern: is a pattern used to match input data. If omitted, thenawkit will operate on all lines.{action}: is the action to execute on lines that match the pattern. If omitted, the default action is to print the entire line.
Description of options parameters:
-F <分隔符>or--field-separator=<分隔符>: Specifies the delimiter for input fields, default is space. Use this option to specify a field separator different from the default.-v <变量名>=<值>: Setawkinternal variable values. This option can be used to pass external values toawkvariables in the script.-f <脚本文件>: Specify a file containingawkscript. This allows you to write larger scripts in a fileawkscript, and then use the-foption to load it.-Vor--version: Displayawkversion information.-hor--help: Displayawkhelp information, including options and usage examples.
The following are some common awk command usages:
Print the whole line:
awk '{print}' filePrint a specific column:
awk '{print $1, $2}' file
Specify columns with a delimiter:
awk -F',' '{print $1, $2}' file
Print the number of lines:
awk '{print NR, $0}' file
Print lines that meet a line number condition:
awk '/pattern/ {print NR, $0}' file
Calculate the sum of a column:
awk '{sum += $1} END {print sum}' file
Print the maximum value:
awk 'max < $1 {max = $1} END {print max}' fileFormatted output:
awk '{printf "%-10s %-10s\n", $1, $2}' file
Basic Usage
The content of log.txt is as follows:
2 this is a test 3 Do you like awk This's a test 10 There are orange,apple,mongo
Usage 1: Basic syntax
awk '{ [pattern] action }' filenames
Note:The awk script part must be wrapped using single quotes' 'to wrap.
Example 1: Output the 1st and 4th columns of each line
awk '{print $1, $4}' log.txt
Output:
2 a 3 like This's 10 orange,apple,mongo
Example 2: Formatted output
printf can customize the output format (similar to printf in C language):
awk '{printf "%-8s %-10s\n", $1, $4}' log.txt
Output:
2 a 3 like This's 10 orange,apple,mongo
Usage 2: Specify delimiter -F
-FThe option is used to specify the delimiter between columns, equivalent to the built-in variableFS (Field Separator)。
Example 1: Use comma,as the delimiter
awk -F, '{print $1, $2}' log.txt
Output:
2 this is a test 3 Do you like awk This's a test 10 There are orange apple
Example 2: Use the built-in variable FS to set the delimiter
awk 'BEGIN { FS="," } {print $1, $2}' log.txt
The output result is the same:
2 this is a test 3 Do you like awk This's a test 10 There are orange apple
Example 3: Use multiple delimiters (space or comma)
[ ,]Means that both spaces and commas are used as delimiters.
awk -F '[ ,]' '{print $1, $2, $5}' log.txt
Output:
2 this test 3 Do awk This's a 10 There apple
Usage 3: Use -v to set variables
-vThe option is used to pass external variables before executing the AWK script, commonly used for dynamic assignment.
Syntax format:
awk -v 变量名=值 '{动作}' 文件名
Example 1: Define a numeric variable and use it in arithmetic
awk -va=1 '{print $1, $1+a}' log.txt
Output:
2 3 3 4 This's 1 10 11
Explanation:
- -v a=1 means defining variable a=1;
- $1+a means the first column plus the variable a.
When $1 is not a number (such as This's), the result is empty or output as is.
Example 2: Define multiple variables at the same time
awk -va=1 -vb=s '{print $1, $1+a, $1b}' log.txt
Output:
2 3 2s 3 4 3s This's 1 This'ss 10 11 10s
Explanation:
- -vb=s defines a string variable b;
- $1b means concatenating the first column and the variable b (e.g., 2s, 10s).
Usage 4: Use -f to call an external AWK script
When the AWK script is long, you can write the logic into a separate file and load it via the -f parameter.
Syntax format:
awk -f 脚本文件 文件名
Suppose we have a script file cal.awk with the following content:
{ print $1, $1 + 10 }
Execute the command:
awk -f cal.awk log.txt
Output:
2 12 3 13 This's 10 10 20
Explanation:
- -f cal.awk means executing the external script cal.awk;
- AWK will run the actions defined in the script for each line of log.txt.
Operators
| Operator | Description |
|---|---|
| = += -= *= /= %= ^= **= | Assignment |
| ?: | C conditional expression |
| || | Logical OR |
| && | Logical AND |
| ~ and !~ | Match regular expression and do not match regular expression |
| < <= > >= != == | Relational operators |
| Space | Concatenation |
| + - | Addition, subtraction |
| * / % | Multiplication, division, and remainder |
| + - ! | Unary plus, minus, and logical negation |
| ^ *** | Exponentiation |
| ++ -- | Increment or decrement, as prefix or suffix |
| $ | Field reference |
| in | Array member |
Filter lines where the first column is greater than 2
$ awk '$1>2' log.txt #命令 #输出 3 Do you like awk This's a test 10 There are orange,apple,mongo
Filter lines where the first column equals 2
$ awk '$1==2 {print $1,$3}' log.txt #命令
#输出
2 is
Filter lines where the first column is greater than 2 and the second column equals 'Are'
$ awk '$1>2 && $2=="Are" {print $1,$2,$3}' log.txt #命令
#输出
3 Are you
Built-in Variables
| Variable | Description |
|---|---|
| $n | The nth field of the current record, fields separated by FS |
| $0 | The complete input record |
| ARGC | The number of command-line arguments |
| ARGIND | The position of the current file in the command line (starting from 0) |
| ARGV | Array containing command-line arguments |
| CONVFMT | Numeric conversion format (default %.6g) ENVIRON associative array of environment variables |
| ERRNO | Description of the last system error |
| FIELDWIDTHS | Field width list (separated by spaces) |
| FILENAME | Current file name |
| FNR | Line number counted separately for each file |
| FS | Field separator (default is any space) |
| IGNORECASE | If true, case-insensitive matching is performed |
| NF | The number of fields in a record |
| NR | The number of records already read, i.e., the line number, starting from 1 |
| OFMT | Output format for numbers (default %.6g) |
| OFS | Output field separator, default value is the same as the input field separator. |
| ORS | Output record separator (default is a newline) |
| RLENGTH | The length of the string matched by the match function |
| RS | Record separator (default is a newline) |
| RSTART | The first position of the string matched by the match function |
| SUBSEP | Array subscript separator (default is /034) |
$ awk 'BEGIN{printf "%4s %4s %4s %4s %4s %4s %4s %4s %4s\n","FILENAME","ARGC","FNR","FS","NF","NR","OFS","ORS","RS";printf "---------------------------------------------\n"} {printf "%4s %4s %4s %4s %4s %4s %4s %4s %4s\n",FILENAME,ARGC,FNR,FS,NF,NR,OFS,ORS,RS}' log.txt
FILENAME ARGC FNR FS NF NR OFS ORS RS
---------------------------------------------
log.txt 2 1 5 1
log.txt 2 2 5 2
log.txt 2 3 3 3
log.txt 2 4 4 4
$ awk -F\' 'BEGIN{printf "%4s %4s %4s %4s %4s %4s %4s %4s %4s\n","FILENAME","ARGC","FNR","FS","NF","NR","OFS","ORS","RS";printf "---------------------------------------------\n"} {printf "%4s %4s %4s %4s %4s %4s %4s %4s %4s\n",FILENAME,ARGC,FNR,FS,NF,NR,OFS,ORS,RS}' log.txt
FILENAME ARGC FNR FS NF NR OFS ORS RS
---------------------------------------------
log.txt 2 1 ' 1 1
log.txt 2 2 ' 1 2
log.txt 2 3 ' 2 3
log.txt 2 4 ' 1 4
# 输出顺序号 NR, 匹配文本行号
$ awk '{print NR,FNR,$1,$2,$3}' log.txt
---------------------------------------------
1 1 2 this is
2 2 3 Are you
3 3 This's a test
4 4 10 There are
# 指定输出分割符
$ awk '{print $1,$2,$5}' OFS=" $ " log.txt
---------------------------------------------
2 $ this $ test
3 $ Are $ awk
This's $ a $
10 $ There $
Using regular expressions, string matching
# 输出第二列包含 "th",并打印第二列与第四列
$ awk '$2 ~ /th/ {print $2,$4}' log.txt
---------------------------------------------
this a
~ indicates the start of a pattern. // contains the pattern.
# 输出包含 "re" 的行 $ awk '/re/ ' log.txt --------------------------------------------- 3 Do you like awk 10 There are orange,apple,mongo
Ignore case
$ awk 'BEGIN{IGNORECASE=1} /this/' log.txt
---------------------------------------------
2 this is a test
This's a test
Pattern negation
$ awk '$2 !~ /th/ {print $2,$4}' log.txt
---------------------------------------------
Are like
a
There orange,apple,mongo
$ awk '!/th/ {print $2,$4}' log.txt
---------------------------------------------
Are like
a
There orange,apple,mongo
awk script
Regarding awk scripts, we need to pay attention to two keywords: BEGIN and END.
- BEGIN{ This is where statements to be executed before processing are placed }
- END { This is where statements to be executed after processing all lines are placed }
- { This is where statements to be executed when processing each line are placed }
Suppose there is such a file (student grade sheet):
$ cat score.txt Marry 2143 78 84 77 Jack 2321 66 78 45 Tom 2122 48 77 71 Mike 2537 87 97 95 Bob 2415 40 57 62
Our awk script is as follows:
$ cat cal.awk
#!/bin/awk -f
#运行前
BEGIN {
math = 0
english = 0
computer = 0
printf "NAME NO. MATH ENGLISH COMPUTER TOTAL\n"
printf "---------------------------------------------\n"
}
#运行中
{
math+=$3
english+=$4
computer+=$5
printf "%-6s %-6s %4d %8d %8d %8d\n", $1, $2, $3,$4,$5, $3+$4+$5
}
#运行后
END {
printf "---------------------------------------------\n"
printf " TOTAL:%10d %8d %8d \n", math, english, computer
printf "AVERAGE:%10.2f %8.2f %8.2f\n", math/NR, english/NR, computer/NR
}
Let's take a look at the execution result:
$ awk -f cal.awk score.txt NAME NO. MATH ENGLISH COMPUTER TOTAL --------------------------------------------- Marry 2143 78 84 77 239 Jack 2321 66 78 45 189 Tom 2122 48 77 71 196 Mike 2537 87 97 95 279 Bob 2415 40 57 62 159 --------------------------------------------- TOTAL: 319 393 350 AVERAGE: 63.80 78.60 70.00
Some more examples
The AWK hello world program is:
BEGIN { print "Hello, world!" }
Calculate file size
$ ls -l *.txt | awk '{sum+=$5} END {print sum}'
--------------------------------------------------
666581
Find lines longer than 80 characters in a file:
awk 'length>80' log.txt
Print the 9x9 multiplication table
seq 9 | sed 'H;g' | awk -v RS='' '{for(i=1;i<=NF;i++)printf("%dx%d=%d%s", i, NR, i*NR, i==NR?"\n":"\t")}'
Other ExtensionsMore content:
Linux Command Reference