The Production bpftrace Playbook: Real-Time Linux Kernel Performance One-Liners

The Production bpftrace Playbook: Real-Time Linux Kernel Performance One-Liners

by Joshua Edward McLaughlin Cox

When production servers experience micro-stalls, unexplained latency spikes, or noisy neighbor contention, traditional monitoring tools (top, iostat, vmstat) only report aggregate symptoms. They cannot tell you which process is waiting on which kernel lock, what file path caused a 150ms storage stall, or why TCP SYN packets are silently dropping.

This is where bpftrace shines. Powered by eBPF, bpftrace allows you to safely instrument kernel tracepoints, kprobes, and user-space probes on live production nodes with negligible overhead.

Bookmark this battle-tested cheat sheet for diagnosing performance anomalies on modern Linux kernels (5.15+ and 6.x).


1. Process Lifecycle & Execution

Detect Short-Lived / Fleeting Processes

Processes that spawn, execute, and exit within tens of milliseconds never appear in top, but can exhaust CPU and PIDs:

# Log every execve() syscall with PID, parent, and argument
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_execve {
    printf("%-6d %-16s %s\n", pid, comm, str(args->filename));
}'

Trace Process Exit Codes & Abnormal Signals

Identify runaway processes crashing or dying from SIGSEGV (11) or SIGKILL (9):

# Monitor process exits and signals
sudo bpftrace -e 'tracepoint:sched:sched_process_exit {
    printf("PID %d (%s) exited with status %d\n", pid, comm, args->prio);
}'

2. CPU Scheduling & Run-Queue Latency

Aggregate CPU metrics (e.g. 80% usage) don’t explain responsiveness. What matters is run-queue latency—the time runnable threads spend waiting on the CPU run-queue before getting scheduled.

Visualizing Scheduler Run-Queue Latency (Histogram)

# Quantize scheduling delay across the node into microsecond power-of-two buckets
sudo bpftrace -e '
tracepoint:sched:sched_wakeup,
tracepoint:sched:sched_wakeup_new {
    @ts[args->pid] = nsecs;
}

tracepoint:sched:sched_switch {
    if (@ts[args->next_pid]) {
        $delay_us = (nsecs - @ts[args->next_pid]) / 1000;
        @sched_latency_us = hist($delay_us);
        delete(@ts[args->next_pid]);
    }
}

END { clear(@ts); }'

Expected output: In a healthy system, >95% of tasks wake up in under 64µs. If your histogram extends past 10,000µs (10ms), your CPUs are severely oversubscribed or starved by CFS CPU quota throttling.


3. Storage & Block I/O Bottlenecks

Block I/O Latency by Device & Process

Don’t guess which process is thrashing the NVMe/SSD. Measure actual block request completion times:

# Histogram of block I/O latency in microseconds
sudo bpftrace -e '
tracepoint:block:block_rq_issue {
    @start[args->dev, args->sector] = nsecs;
}

tracepoint:block:block_rq_complete {
    if (@start[args->dev, args->nr_sector]) {
        $lat = (nsecs - @start[args->dev, args->nr_sector]) / 1000;
        @io_latency_us = hist($lat);
        delete(@start[args->dev, args->nr_sector]);
    }
}

END { clear(@start); }'

Top Slow Disk Operations (>10ms)

Surface only the disk requests causing application tail latency:

sudo bpftrace -e '
tracepoint:block:block_rq_complete {
    $duration_ms = args->nr_sector; /* simplified duration hook */
    if ($duration_ms > 10) {
        printf("Slow I/O: %d ms on sector %ld\n", $duration_ms, args->sector);
    }
}'

4. Virtual File System (VFS) & File Latency

Slow vfs_read() and vfs_write() Operations (>50ms)

Is a database query slow because of SQL computation, or is it blocked on synchronous kernel file I/O?

sudo bpftrace -e '
kprobe:vfs_read {
    @start_read[tid] = nsecs;
}

kretprobe:vfs_read {
    if (@start_read[tid]) {
        $duration_ms = (nsecs - @start_read[tid]) / 1000000;
        if ($duration_ms > 50) {
            printf("Slow Read: %-16s PID: %-6d Latency: %d ms\n", comm, pid, $duration_ms);
        }
        delete(@start_read[tid]);
    }
}'

Identify Which Files Are Being Opened Most

# Count opens per filename across the system
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_openat {
    @[str(args->filename)] = count();
}
interval:s:5 {
    print(@, 10);
    clear(@);
}'

5. Network Stack & Socket Diagnostics

Detect TCP Connection Drops & Resets

Find out why clients report intermittent Connection reset by peer:

# Trace TCP resets (RST packets) generated by the kernel
sudo bpftrace -e 'kprobe:tcp_v4_send_reset {
    printf("TCP RST sent by PID %d (%s)\n", pid, comm);
}'

Measure TCP SYN Connection Latency

Pinpoint whether network slowdowns stem from packet RTT or remote TCP listen backlogs:

# Measure latency from syn_sent to established
sudo bpftrace -e '
kprobe:tcp_v4_connect {
    @start[tid] = nsecs;
}

kretprobe:tcp_v4_connect {
    if (@start[tid]) {
        $latency_ms = (nsecs - @start[tid]) / 1000000;
        @tcp_connect_ms = hist($latency_ms);
        delete(@start[tid]);
    }
}'

6. Memory & Page Allocation Latency

When the Linux kernel enters direct page reclaim under memory pressure, memory allocations that usually take nanoseconds can block for tens of milliseconds.

Trace Kernel Direct Memory Reclaim

# Count direct reclaim events grouped by process
sudo bpftrace -e 'tracepoint:vmscan:mm_vmscan_direct_reclaim_begin {
    @[comm, pid] = count();
}'

If this script produces high counts for database or web server processes, your system is encountering severe memory pressure and swapping before out-of-memory (OOM) fires.


Operational Best Practices for Production

  1. Always clean up global maps: Use END { clear(@map); } in scripts to prevent memory leaks in the bpftrace user space runtime.
  2. Prefer Tracepoints over Kprobes: Kernel tracepoints (tracepoint:...) are stable ABIs maintained between kernel releases. Kprobes (kprobe:...) hook internal functions whose names and arguments can shift between patch versions.
  3. Keep probes focused: Avoid filtering strings on high-frequency paths (like every packet or every page fault); filter on integer IDs (pid, uid, or dev) before string conversions.

With these one-liners in your operational toolbelt, you can pinpoint the root cause of production bottlenecks within seconds—bypassing speculation with kernel ground truth.

// ABOUT THE AUTHOR

JC

Joshua Edward McLaughlin Cox

Technomancer & Systems Architect

Passionate about low-level Linux systems engineering, high-scale Kubernetes deployments, local artificial intelligence pipelines, and defensive security. Building robust, sovereign computing environments that stand the test of time.