Linux bpftrace Command

Linux 命令大全Linux Command Reference


bpftrace is an advanced tracing tool based on eBPF (extended Berkeley Packet Filter). It allows developers to dynamically observe and analyze the runtime state of Linux systems without modifying kernel code.

eBPF is a revolutionary technology in the Linux kernel. It provides a secure virtual machine environment that can run user-defined code in the kernel. bpftrace is built on top of eBPF, providing a simpler, higher-level abstraction layer.


Core Advantages of bpftrace

Real-time System Observation

  • No need to restart the system or applications
  • Extremely low performance overhead
  • Can observe both kernel and userspace programs

Flexible Probing Capabilities

  • Supports multiple probe point types: function entry/exit, timers, hardware events, etc.
  • Can trace system calls, network events, disk I/O, etc.

Simple Scripting Language

  • AWK-like syntax with a gentle learning curve
  • Rich set of built-in functions and variables
  • Supports conditional filtering and aggregate statistics

bpftrace Installation and Configuration

Installation Methods

Example

# Ubuntu/Debian
sudo apt install bpftrace

# CentOS/RHEL
sudo yum install bpftrace

# Build from source
git clone https://github.com/iovisor/bpftrace.git
mkdir bpftrace/build && cd bpftrace/build
cmake ..
make
sudo make install

Verify Installation

sudo bpftrace -e 'BEGIN { printf("Hello, bpftrace!n"); }'

bpftrace Basic Syntax

A bpftrace program consists of probe points and associated actions. The basic structure is as follows:

probe /filter/ {
    action
}

Probe Point Types

Probe Point Type Description Example
kprobe Kernel function entry kprobe:vfs_read
kretprobe Kernel function return kretprobe:vfs_read
uprobe Userspace function entry uprobe:/bin/bash:readline
tracepoint Kernel static tracepoint tracepoint:syscalls:sys_enter_open
interval Timer-triggered interval:s:5
software Software events software:faults:major

Common Built-in Variables

  • pid: Current process ID
  • tid: Current thread ID
  • comm: Current process name
  • nsecs: Nanosecond timestamp
  • arg0-argN: Function arguments
  • retval: Function return value

bpftrace Practical Examples

1. Tracing System Calls

Example

# Count the number of open system calls
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_open { @[comm] = count(); }'

2. Analyzing Function Execution Time

Example

# Measure the execution time of vfs_read
sudo bpftrace -e '
kprobe:vfs_read { @start[tid] = nsecs; }
kretprobe:vfs_read /@start[tid]/ {
    @times = hist(nsecs - @start[tid]);
    delete(@start[tid]);
}'

3. Monitoring Process File Access

Example

# Trace files opened by a specific process
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_openat /pid == 1234/ { printf("%s -> %sn", comm, str(args->filename)); }'

4. Counting TCP Connections

Example

# Count TCP connections by process
sudo bpftrace -e 'kprobe:tcp_connect { @[comm] = count(); }'

bpftrace Advanced Features

1. Map Functionality

bpftrace provides a variety of built-in map types for data aggregation:

Example

# Count histogram
@hist = hist(nsecs);

# Calculate average
@avg = avg(nsecs);

# Count unique values
@unique = count();

2. Conditional Filtering

Example

# Only trace read calls of a specific process
tracepoint:syscalls:sys_enter_read /pid == 1234/ {
    printf("PID %d reading %d bytesn", pid, args->count);
}

3. Multiple Probe Combination

Example

# Trace the entire process from socket creation to connection
kprobe:sock_alloc {
    @socket[tid] = 1;
}

kprobe:tcp_connect /@socket[tid]/ {
    printf("socket %d connecting to %s:%dn", args->sock->__sk_common.skc_dport,
           ntop(args->sock->__sk_common.skc_daddr),
           args->sock->__sk_common.skc_dport);
    delete(@socket[tid]);
}

bpftrace Best Practices

  1. Limit trace scope: Use PID or command name filtering to reduce system overhead
  2. Avoid excessive printing: Too many printf calls can affect performance
  3. Use aggregation: Use count(), sum(), and other aggregate functions whenever possible
  4. Clean up resources: Long-running scripts should periodically clean up map data
  5. Security considerations: bpftrace requires root privileges; be cautious when running unknown scripts

bpftrace Comparison with Other Tools

Tool Advantages Disadvantages
bpftrace Flexible, high-performance, easy to use Requires root privileges
strace Simple, no compilation required High performance overhead
perf Comprehensive features, low overhead Steep learning curve
SystemTap Powerful Requires compilation, complex configuration

Recommended Learning Resources

  1. bpftrace Official Documentation
  2. bpftrace One-Liner Tutorial
  3. bpftrace Examples Repository

bpftrace is a powerful tool for system performance analysis and troubleshooting. By practicing these examples and mastering its core concepts, you will be able to gain a deeper understanding of Linux system runtime behavior and optimize it.


Linux 命令大全Linux Command Reference

Other Extensions