Linux bpftrace Command
bpftrace is an advanced tracing tool based on eBPF (extended Berkeley Packet Filter). It allows developers to dynamically observe and analyze the runtime state of Linux systems without modifying kernel code.
eBPF is a revolutionary technology in the Linux kernel. It provides a secure virtual machine environment that can run user-defined code in the kernel. bpftrace is built on top of eBPF, providing a simpler, higher-level abstraction layer.
Core Advantages of bpftrace
Real-time System Observation
- No need to restart the system or applications
- Extremely low performance overhead
- Can observe both kernel and userspace programs
Flexible Probing Capabilities
- Supports multiple probe point types: function entry/exit, timers, hardware events, etc.
- Can trace system calls, network events, disk I/O, etc.
Simple Scripting Language
- AWK-like syntax with a gentle learning curve
- Rich set of built-in functions and variables
- Supports conditional filtering and aggregate statistics
bpftrace Installation and Configuration
Installation Methods
Example
sudo apt install bpftrace
# CentOS/RHEL
sudo yum install bpftrace
# Build from source
git clone https://github.com/iovisor/bpftrace.git
mkdir bpftrace/build && cd bpftrace/build
cmake ..
make
sudo make install
Verify Installation
sudo bpftrace -e 'BEGIN { printf("Hello, bpftrace!n"); }'
bpftrace Basic Syntax
A bpftrace program consists of probe points and associated actions. The basic structure is as follows:
probe /filter/ {
action
}
Probe Point Types
| Probe Point Type | Description | Example |
|---|---|---|
kprobe |
Kernel function entry | kprobe:vfs_read |
kretprobe |
Kernel function return | kretprobe:vfs_read |
uprobe |
Userspace function entry | uprobe:/bin/bash:readline |
tracepoint |
Kernel static tracepoint | tracepoint:syscalls:sys_enter_open |
interval |
Timer-triggered | interval:s:5 |
software |
Software events | software:faults:major |
Common Built-in Variables
pid: Current process IDtid: Current thread IDcomm: Current process namensecs: Nanosecond timestamparg0-argN: Function argumentsretval: Function return value
bpftrace Practical Examples
1. Tracing System Calls
Example
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_open { @[comm] = count(); }'
2. Analyzing Function Execution Time
Example
sudo bpftrace -e '
kprobe:vfs_read { @start[tid] = nsecs; }
kretprobe:vfs_read /@start[tid]/ {
@times = hist(nsecs - @start[tid]);
delete(@start[tid]);
}'
3. Monitoring Process File Access
Example
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_openat /pid == 1234/ { printf("%s -> %sn", comm, str(args->filename)); }'
4. Counting TCP Connections
Example
sudo bpftrace -e 'kprobe:tcp_connect { @[comm] = count(); }'
bpftrace Advanced Features
1. Map Functionality
bpftrace provides a variety of built-in map types for data aggregation:
Example
@hist = hist(nsecs);
# Calculate average
@avg = avg(nsecs);
# Count unique values
@unique = count();
2. Conditional Filtering
Example
tracepoint:syscalls:sys_enter_read /pid == 1234/ {
printf("PID %d reading %d bytesn", pid, args->count);
}
3. Multiple Probe Combination
Example
kprobe:sock_alloc {
@socket[tid] = 1;
}
kprobe:tcp_connect /@socket[tid]/ {
printf("socket %d connecting to %s:%dn", args->sock->__sk_common.skc_dport,
ntop(args->sock->__sk_common.skc_daddr),
args->sock->__sk_common.skc_dport);
delete(@socket[tid]);
}
bpftrace Best Practices
- Limit trace scope: Use PID or command name filtering to reduce system overhead
- Avoid excessive printing: Too many printf calls can affect performance
- Use aggregation: Use count(), sum(), and other aggregate functions whenever possible
- Clean up resources: Long-running scripts should periodically clean up map data
- Security considerations: bpftrace requires root privileges; be cautious when running unknown scripts
bpftrace Comparison with Other Tools
| Tool | Advantages | Disadvantages |
|---|---|---|
| bpftrace | Flexible, high-performance, easy to use | Requires root privileges |
| strace | Simple, no compilation required | High performance overhead |
| perf | Comprehensive features, low overhead | Steep learning curve |
| SystemTap | Powerful | Requires compilation, complex configuration |
Recommended Learning Resources
bpftrace is a powerful tool for system performance analysis and troubleshooting. By practicing these examples and mastering its core concepts, you will be able to gain a deeper understanding of Linux system runtime behavior and optimize it.
Other Extensions
Linux Command Reference