Linux awk command

Linux 命令大全Linux Command Reference

awk is a language for processing text files and a powerful text analysis tool.

awk reads text files line by line and provides programming-language-like features, such as:

  • Variable definition and calculation
  • Conditional judgments and loops
  • String processing and formatted output

These features make AWK very efficient at processing structured text (such as CSV, log files).

It is called awk because it takes the first characters of the family names of its three founders: Alfred Aho, Peter Weinberger, and Brian Kernighan.

Syntax

awk options 'pattern {action}' file

Option parameter description:

  • options: are some options used to controlawkbehavior.
  • pattern: is a pattern used to match input data. If omitted, thenawkit will operate on all lines.
  • {action}: is the action to execute on lines that match the pattern. If omitted, the default action is to print the entire line.

Description of options parameters:

  • -F <分隔符>or--field-separator=<分隔符>: Specifies the delimiter for input fields, default is space. Use this option to specify a field separator different from the default.

  • -v <变量名>=<值>: Setawkinternal variable values. This option can be used to pass external values toawkvariables in the script.

  • -f <脚本文件>: Specify a file containingawkscript. This allows you to write larger scripts in a fileawkscript, and then use the-foption to load it.

  • -Vor--version: Displayawkversion information.

  • -hor--help: Displayawkhelp information, including options and usage examples.

The following are some common awk command usages:

Print the whole line:

awk '{print}' file

Print a specific column:

awk '{print $1, $2}' file

Specify columns with a delimiter:

awk -F',' '{print $1, $2}' file

Print the number of lines:

awk '{print NR, $0}' file

Print lines that meet a line number condition:

awk '/pattern/ {print NR, $0}' file

Calculate the sum of a column:

awk '{sum += $1} END {print sum}' file

Print the maximum value:

awk 'max < $1 {max = $1} END {print max}' file

Formatted output:

awk '{printf "%-10s %-10s\n", $1, $2}' file

Basic Usage

The content of log.txt is as follows:

2 this is a test
3 Do you like awk
This's a test
10 There are orange,apple,mongo

Usage 1: Basic syntax

awk '{ [pattern] action }' filenames

Note:The awk script part must be wrapped using single quotes' 'to wrap.

Example 1: Output the 1st and 4th columns of each line

awk '{print $1, $4}' log.txt

Output:

2 a
3 like
This's
10 orange,apple,mongo

Example 2: Formatted output

printf can customize the output format (similar to printf in C language):

awk '{printf "%-8s %-10s\n", $1, $4}' log.txt

Output:

2        a
3        like
This's
10       orange,apple,mongo

Usage 2: Specify delimiter -F

-FThe option is used to specify the delimiter between columns, equivalent to the built-in variableFS (Field Separator)。

Example 1: Use comma,as the delimiter

awk -F, '{print $1, $2}' log.txt

Output:

2 this is a test
3 Do you like awk
This's a test
10 There are orange apple

Example 2: Use the built-in variable FS to set the delimiter

awk 'BEGIN { FS="," } {print $1, $2}' log.txt

The output result is the same:

2 this is a test
3 Do you like awk
This's a test
10 There are orange apple

Example 3: Use multiple delimiters (space or comma)

[ ,]Means that both spaces and commas are used as delimiters.

awk -F '[ ,]' '{print $1, $2, $5}' log.txt

Output:

2 this test
3 Do awk
This's a
10 There apple

Usage 3: Use -v to set variables

-vThe option is used to pass external variables before executing the AWK script, commonly used for dynamic assignment.

Syntax format:

awk -v 变量名=值 '{动作}' 文件名

Example 1: Define a numeric variable and use it in arithmetic

awk -va=1 '{print $1, $1+a}' log.txt

Output:

2 3
3 4
This's 1
10 11

Explanation:

  • -v a=1 means defining variable a=1;
  • $1+a means the first column plus the variable a.

When $1 is not a number (such as This's), the result is empty or output as is.

Example 2: Define multiple variables at the same time

awk -va=1 -vb=s '{print $1, $1+a, $1b}' log.txt

Output:

2 3 2s
3 4 3s
This's 1 This'ss
10 11 10s

Explanation:

  • -vb=s defines a string variable b;
  • $1b means concatenating the first column and the variable b (e.g., 2s, 10s).

Usage 4: Use -f to call an external AWK script

When the AWK script is long, you can write the logic into a separate file and load it via the -f parameter.

Syntax format:

awk -f 脚本文件 文件名

Suppose we have a script file cal.awk with the following content:

{ print $1, $1 + 10 }

Execute the command:

awk -f cal.awk log.txt

Output:

2 12
3 13
This's 10
10 20

Explanation:

  • -f cal.awk means executing the external script cal.awk;
  • AWK will run the actions defined in the script for each line of log.txt.

Operators

Operator Description
= += -= *= /= %= ^= **= Assignment
?: C conditional expression
|| Logical OR
&& Logical AND
~ and !~ Match regular expression and do not match regular expression
< <= > >= != == Relational operators
Space Concatenation
+ - Addition, subtraction
* / % Multiplication, division, and remainder
+ - ! Unary plus, minus, and logical negation
^ *** Exponentiation
++ -- Increment or decrement, as prefix or suffix
$ Field reference
in Array member

Filter lines where the first column is greater than 2

$ awk '$1>2' log.txt    #命令
#输出
3 Do you like awk
This's a test
10 There are orange,apple,mongo

Filter lines where the first column equals 2

$ awk '$1==2 {print $1,$3}' log.txt    #命令
#输出
2 is

Filter lines where the first column is greater than 2 and the second column equals 'Are'

$ awk '$1>2 && $2=="Are" {print $1,$2,$3}' log.txt    #命令
#输出
3 Are you

Built-in Variables

Variable Description
$n The nth field of the current record, fields separated by FS
$0 The complete input record
ARGC The number of command-line arguments
ARGIND The position of the current file in the command line (starting from 0)
ARGV Array containing command-line arguments
CONVFMT Numeric conversion format (default %.6g) ENVIRON associative array of environment variables
ERRNO Description of the last system error
FIELDWIDTHS Field width list (separated by spaces)
FILENAME Current file name
FNR Line number counted separately for each file
FS Field separator (default is any space)
IGNORECASE If true, case-insensitive matching is performed
NF The number of fields in a record
NR The number of records already read, i.e., the line number, starting from 1
OFMT Output format for numbers (default %.6g)
OFS Output field separator, default value is the same as the input field separator.
ORS Output record separator (default is a newline)
RLENGTH The length of the string matched by the match function
RS Record separator (default is a newline)
RSTART The first position of the string matched by the match function
SUBSEP Array subscript separator (default is /034)
$ awk 'BEGIN{printf "%4s %4s %4s %4s %4s %4s %4s %4s %4s\n","FILENAME","ARGC","FNR","FS","NF","NR","OFS","ORS","RS";printf "---------------------------------------------\n"} {printf "%4s %4s %4s %4s %4s %4s %4s %4s %4s\n",FILENAME,ARGC,FNR,FS,NF,NR,OFS,ORS,RS}'  log.txt
FILENAME ARGC  FNR   FS   NF   NR  OFS  ORS   RS
---------------------------------------------
log.txt    2    1         5    1
log.txt    2    2         5    2
log.txt    2    3         3    3
log.txt    2    4         4    4
$ awk -F\' 'BEGIN{printf "%4s %4s %4s %4s %4s %4s %4s %4s %4s\n","FILENAME","ARGC","FNR","FS","NF","NR","OFS","ORS","RS";printf "---------------------------------------------\n"} {printf "%4s %4s %4s %4s %4s %4s %4s %4s %4s\n",FILENAME,ARGC,FNR,FS,NF,NR,OFS,ORS,RS}'  log.txt
FILENAME ARGC  FNR   FS   NF   NR  OFS  ORS   RS
---------------------------------------------
log.txt    2    1    '    1    1
log.txt    2    2    '    1    2
log.txt    2    3    '    2    3
log.txt    2    4    '    1    4
# 输出顺序号 NR, 匹配文本行号
$ awk '{print NR,FNR,$1,$2,$3}' log.txt
---------------------------------------------
1 1 2 this is
2 2 3 Are you
3 3 This's a test
4 4 10 There are
# 指定输出分割符
$  awk '{print $1,$2,$5}' OFS=" $ "  log.txt
---------------------------------------------
2 $ this $ test
3 $ Are $ awk
This's $ a $
10 $ There $

Using regular expressions, string matching

# 输出第二列包含 "th",并打印第二列与第四列
$ awk '$2 ~ /th/ {print $2,$4}' log.txt
---------------------------------------------
this a

~ indicates the start of a pattern. // contains the pattern.

# 输出包含 "re" 的行
$ awk '/re/ ' log.txt
---------------------------------------------
3 Do you like awk
10 There are orange,apple,mongo

Ignore case

$ awk 'BEGIN{IGNORECASE=1} /this/' log.txt
---------------------------------------------
2 this is a test
This's a test

Pattern negation

$ awk '$2 !~ /th/ {print $2,$4}' log.txt
---------------------------------------------
Are like
a
There orange,apple,mongo
$ awk '!/th/ {print $2,$4}' log.txt
---------------------------------------------
Are like
a
There orange,apple,mongo

awk script

Regarding awk scripts, we need to pay attention to two keywords: BEGIN and END.

  • BEGIN{ This is where statements to be executed before processing are placed }
  • END { This is where statements to be executed after processing all lines are placed }
  • { This is where statements to be executed when processing each line are placed }

Suppose there is such a file (student grade sheet):

$ cat score.txt
Marry   2143 78 84 77
Jack    2321 66 78 45
Tom     2122 48 77 71
Mike    2537 87 97 95
Bob     2415 40 57 62

Our awk script is as follows:

$ cat cal.awk
#!/bin/awk -f
#运行前
BEGIN {
    math = 0
    english = 0
    computer = 0
 
    printf "NAME    NO.   MATH  ENGLISH  COMPUTER   TOTAL\n"
    printf "---------------------------------------------\n"
}
#运行中
{
    math+=$3
    english+=$4
    computer+=$5
    printf "%-6s %-6s %4d %8d %8d %8d\n", $1, $2, $3,$4,$5, $3+$4+$5
}
#运行后
END {
    printf "---------------------------------------------\n"
    printf "  TOTAL:%10d %8d %8d \n", math, english, computer
    printf "AVERAGE:%10.2f %8.2f %8.2f\n", math/NR, english/NR, computer/NR
}

Let's take a look at the execution result:

$ awk -f cal.awk score.txt
NAME    NO.   MATH  ENGLISH  COMPUTER   TOTAL
---------------------------------------------
Marry  2143     78       84       77      239
Jack   2321     66       78       45      189
Tom    2122     48       77       71      196
Mike   2537     87       97       95      279
Bob    2415     40       57       62      159
---------------------------------------------
  TOTAL:       319      393      350
AVERAGE:     63.80    78.60    70.00

Some more examples

The AWK hello world program is:

BEGIN { print "Hello, world!" }

Calculate file size

$ ls -l *.txt | awk '{sum+=$5} END {print sum}'
--------------------------------------------------
666581

Find lines longer than 80 characters in a file:

awk 'length>80' log.txt

Print the 9x9 multiplication table

seq 9 | sed 'H;g' | awk -v RS='' '{for(i=1;i<=NF;i++)printf("%dx%d=%d%s", i, NR, i*NR, i==NR?"\n":"\t")}'

More content:

Linux 命令大全Linux Command Reference

Other Extensions