How To Check Cpu Utilization In Linux

17 min read

How to Check CPU Utilization in Linux

Understanding how much of your processor’s capacity is being used is essential for troubleshooting performance issues, planning capacity upgrades, or simply keeping an eye on system health. Linux provides a rich set of command‑line tools that report CPU utilization in different granularities—from overall system usage to per‑core statistics and per‑process breakdowns. This guide walks you through the most common utilities, explains what each metric means, and shows practical examples you can run on any modern distribution Took long enough..


Why Monitor CPU Utilization?

CPU utilization tells you the percentage of time the processor spends executing non‑idle tasks. High utilization can indicate:

  • A CPU‑bound workload that may need optimization or more cores.
  • A runaway process consuming excessive cycles.
  • Insufficient hardware for the current workload, prompting a scale‑up or load‑balancing decision.

Conversely, consistently low utilization may suggest over‑provisioned resources or an I/O‑bound bottleneck elsewhere Easy to understand, harder to ignore..


Core Concepts Before the Commands

  • User time (%us) – Time spent running user‑space applications.
  • System time (%sy) – Time spent in the kernel handling system calls and interrupts.
  • Nice time (%ni) – CPU time for low‑priority (nice) processes.
  • Idle time (%id) – Percentage of time the CPU is doing nothing.
  • Wait time (%wa) – Time spent waiting for I/O operations to complete.
  • Interrupt (%hi) and Soft‑interrupt (%si) – Time servicing hardware and software interrupts.

Most tools display these fields as percentages that add up to 100 % (or close, depending on rounding).


Essential Tools for Checking CPU Utilization

Tool Primary Use Key Options Typical Output
top Real‑time, interactive view of processes and overall CPU load -b (batch), -d delay, -n iterations Overall %CPU, per‑process %CPU, memory
htop Enhanced, colorized top with mouse support None needed for basic use Visual bars per core, process tree
mpstat Per‑core CPU statistics from the sysstat package -P ALL, interval, count %usr, %sys, %idle per CPU
vmstat Virtual memory statistics, includes CPU columns interval, count %usr, %sys, %id, %wa, etc.
sar Historical and real‑time system activity reporting -u for CPU, interval, count Average CPU usage over time
pidstat Per‑process CPU utilization over time -p ALL, interval, count %CPU per PID
glances All‑in‑one monitoring dashboard (curses/web) -w for web mode CPU, memory, disk, network, alerts
iostat Primarily I/O but includes %cpu column -c for CPU only, interval %user, %nice, %system, %iowait, %idle

All of these utilities are either part of the base system (top, vmstat, iostat) or available via common repositories (htop, sysstat for mpstat/sar/pidstat, glances).


Using top for a Quick Snapshot

top is pre‑installed on virtually every Linux distribution and provides an immediate overview.

top

When you run it, the first few lines look like:

top - 14:23:01 up  2:15,  2 users,  load average: 0.31, 0.28, 0.25
Tasks: 124 total,   2 running, 122 sleeping,   0 stopped,   0 zombie
%Cpu(s): 12.3 us,  3.1 sy,  0.0 ni, 84.2 id,  0.4 wa,  0.0 hi,  0.0 si,  0.0 st
KiB Mem :  8034560 total,  3124560 free,  2856000 used,  2055400 buff/cache
KiB Swap:  2097148 total,  2097148 free,      0 used.  4987600 avail Mem

The line beginning with %Cpu(s): shows the aggregate utilization.

  • 12.3 us → 12.3 % user time
  • 3.1 sy → 3.1 % system time
  • 84.2 id → 84.2 % idle

Below that, each process lists its %CPU share, letting you spot hogs instantly.

Batch mode (useful for scripting):

top -b -d 2 -n 5
  • -b runs in batch (non‑interactive) mode.
  • -d 2 sets a 2‑second delay between samples.
  • -n 5 collects five iterations then exits.

Getting a Friendlier View with htop

If you prefer colors, scrollable process trees, and the ability to kill processes with a mouse click, install htop:

sudo apt-get install htop   # Debian/Ubuntu
sudo yum install htop       # RHEL/CentOS
sudo dnf install htop       # Fedora

Launch it:

htop

You’ll see a CPU usage bar for each core at the top, making it trivial to spot imbalance (e.g.So , one core at 90 % while others sit at 5 %). Press F5 to toggle tree view, F6 to sort by column, and F9 to send a signal to a selected process It's one of those things that adds up..


Per‑Core Details with mpstat

Part of the sysstat package, mpstat reports statistics for each logical CPU.

mpstat -P ALL 1 3

Explanation:

  • -P ALL → show every CPU (including ALL for the aggregate).
  • 1 → sample interval in seconds.
  • 3 → number of reports (after which the command exits).

Sample output:

Linux 5.15.0-78-generic (myhost)   11/0

08/2025   _x86_64_ (8 CPU)

09:15:32 AM   CPU   %usr   %nice   %sys   %iowait   %steal   %idle
           All  11.That's why 20      0. 10    0.80    0.In practice, 10    0. Now, 85
             5  13. 60    0.And 20    0. 00    2.00    84.10
             1  12.Which means 25      0. 20      0.40      0.But 00    85. 50    0.00    82.15      0.90
             3  10.In practice, 40    0. 20      0.00    3.40    0.20    0.Consider this: 50    0. 25
             7  12.Even so, 60
             6  10. Think about it: 00    0. 00    79.00    3.40    0.00    5.00    3.In practice, 00    83. 15      0.20    0.That said, 80    0. On top of that, 90    0. Practically speaking, 60
             2   9. Worth adding: 00    3. 00    3.00    86.00    4.That's why 10
             0  15. 80    0.10      0.30      0.In practice, 55
             4  11. 20    0.60    0.Still, 00    86. But 00    86. Because of that, 00    3. 00    83.

The aggregate row (`All`) mirrors the values you’d see in `/proc/loadavg`, while individual CPU lines expose core‑specific behavior. A persistent high `%sys` on a single core often points to hardware interrupts or poorly tuned drivers, whereas a system‑wide spike in `%usr` usually indicates genuine application load.

---

## Tracking a Single Process Over Time with `pidstat`

`pidstat` (also from `sysstat`) lets you monitor one or more processes over successive intervals, showing CPU, memory, I/O, and context switches.

```bash
pidstat -p 1234,5678 2 5
  • -p 1234,5678 selects the target PIDs (comma‑separated).
  • 2 is the sampling interval in seconds.
  • 5 gives the number of reports.

Typical output:

Linux 5.15.0-78-generic (myhost)   11/08/2025   _x86_64_ (8 CPU)

11:45:34 AM   UID       PID    %usr %system %guest %wait   CPU  Command
              0      1234    8.20    2.10    0.00   0.00   3   java
              0      5678    5.50    1.So 80    0. 00   0.

For long‑running daemons, you can capture this data to a file and replay it later:

```bash
pidstat -p ALL 60 1440 >> /var/log/pidstat.log

This appends a 60‑second cadence for 24 hours, perfect for post‑mortem analysis or feeding into log‑aggregation tools Small thing, real impact..


Kernel‑Level Insights with vmstat

When you need a high‑level pulse of the entire system—processes, memory, swapping, and I/O—vmstat delivers in a single line per interval Small thing, real impact..

vmstat 3 10
  • 3 → wait three seconds between reports.
  • 10 → stop after ten iterations.

Sample rows:

procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
 1  0      0 2456784 123456 3456784   0    0     0     5    0    0  5  2 92  1  0
 0  0      0 2451234 124000 3457890   0    0     2     8    2    1  4  1 94  1  0

Key columns:

  • r – processes waiting for run time; values >

Digging Deeper into the vmstat Output

The single‑line snapshot that vmstat prints is a compact ledger of the system’s health. Once you know what each column represents, the line becomes a diagnostic map you can read at a glance.

Column Meaning What to watch for
r Runnable processes (ready to execute) A persistent r count above the number of CPUs suggests a backed‑up run queue and possible CPU starvation. Values below 20 % on a multi‑core system usually merit investigation. In practice,
free Free physical memory (kB) Low free combined with high buff+cache may still be healthy (cache can be reclaimed), but a sustained drop warns of tight RAM. g.Still,
id CPU idle time The complement of us+sy+wa+st.
bi Blocks received from a block device (e.Practically speaking,
cs Context switches per second High cs can be a proxy for CPU‑bound workloads that frequently yield, or for systems with many lightweight threads.
so Kibibytes swapped out per second High so indicates the kernel is aggressively moving pages to disk, again a sign of insufficient RAM. Still,
sy CPU time spent in kernel space Spikes here often betray system calls, interrupts, or driver overhead. Still, , disk reads) per second
si Kibibytes swapped in per second Frequent si activity points to memory pressure; each swap‑in incurs a performance hit. So naturally,
us CPU time spent in user space Direct indicator of application CPU consumption.
buff Memory used by block buffers (kernel page cache) Large buff is a good sign – the kernel is caching disk reads for faster later access.
b Processes in uninterruptible sleep (TASK_UNINTERRUPTIBLE) – often waiting for I/O High b values usually indicate blocked I/O, such as a hung filesystem operation or a device that isn’t delivering data. Now,
bo Blocks sent to a block device (writes) per second Elevated bo can reflect write‑intensive apps; watch for a mismatch between bi/bo that could hint at I/O imbalance. And
swpd Amount of swapped memory (kilobytes) Non‑zero values mean the system is paging; a steady rise can signal memory exhaustion. And
in Interrupts per second A sudden jump often traces to hardware events (network cards, timers) or poorly tuned drivers.
wa Time waiting for I/O (percentage of total CPU) High wa signals storage latency – think slow SSDs, network‑mounted filesystems, or overloaded storage arrays.
cache Page cache + dentries/inodes Similar to buff; a healthy system keeps this high until memory pressure forces reclamation.
st Steal time (when a virtual CPU waits for the hypervisor to give it CPU) Relevant only in virtualized environments; elevated st means the host is contending for physical CPU.

Putting It All Together

A healthy snapshot on a modest server might look like:

 0  0      0 2450000 130000 3600000   0    0    10    8   30   22  2  1 95  0  0

Interpretation: a few idle processes, no swapping, ample free memory, modest I/O, and CPUs mostly idle Simple, but easy to overlook..

A troublesome snapshot could be:

 8  2      5 1200000  80000 2100000   2    0    150  30  210   85 15 10 60  5  0

Here you see a busy run queue (r=8 on, say, 4 CPUs), non‑trivial blocking (b=2), active swapping (swpd=5), low free memory, and high

Decoding the “Troublesome” Snapshot

The second example shows a system that is far from the ideal state:

 8  2      5 1200000  80000 2100000   2    0    150  30  210   85 15 10 60  5  0

Below is a field‑by‑field walk‑through of what those numbers are whispering about the host:

Field Value What it screams
r (run queue) 8 The scheduler has eight threads/completion‑queues waiting for CPU. Which means
us 60 60 % of CPU time is spent in user space – applications are doing real work, not just waiting.
so 30 30 KiB/s are being swapped out – the kernel is actively off‑loading pages, further adding to I/O pressure. Consider this:
sy 5 5 % in kernel – normal driver/interrupt overhead. Which means
swpd (swap used) 5 Five MiB of swap are active, meaning the kernel is already paging out memory. Worth adding:
cs 10 10 context switches per second – low, meaning the CPU is not thrashing with excessive thread switching. On top of that,
cache 2100000 About 2 GiB of page‑cache/dentry/inode cache. Think about it: on a 4‑core machine this is already a double‑loaded run queue – processes are spending time context‑switching rather than making progress. This is the classic “disk stall” indicator; the storage subsystem is struggling to keep up.
bi 210 210 blocks/s are being read from the storage device.
free 1200000 Roughly **1.Day to day, the asymmetry (more reads than writes) points toward a data‑intensive service (caching, analytics, log processing) that is memory‑bound. Day to day,
si 150 150 KiB/s are being swapped in. The processor is fully occupied, which explains the long run queue. In real terms, even a modest amount of swapping can throttle applications because each page‑in/out incurs a ~10‑30 ms latency. 2 GiB** of completely free RAM.
buff 80000 ~78 MiB of block‑buffer cache – decent, but the accompanying high wa (see below) suggests the data is not being served quickly enough. This is a clear sign that the working set no longer fits in RAM and the system is compensating with disk‑based paging.
id 0 Zero idle CPU. Still, the cache is healthy, yet the I/O wait time tells us the cache isn’t preventing disk‑level latency. Which means combined with bo (writes) this hints at a read‑heavy workload that is not being satisfied from the page cache. That said,
b (blocked) 2 Two processes are in uninterruptible sleep (often waiting for I/O). That's why
bo 85 85 blocks/s are being written.
in 15 15 interrupts per second – modest, but when paired with high bi it suggests that the storage controller is generating interrupts for each block completion, adding overhead.
wa 60 60 % of total CPU time is waiting for I/O. That said, while not critical, it is far less than the 3‑4 GiB that a typical workload of this size would prefer to have free. This is the smoking gun: the storage subsystem cannot keep up with the read demand, and the high wa drags down id.

Synthesizing the Signal: What the Numbers Are Really Saying

Taken individually, each metric is a data point; together, they form a coherent narrative of a system starved for storage throughput. The run queue (r: 12) is long not because the CPU lacks cycles—the us: 60% proves the application logic is hungry for work—but because the majority of those threads are stuck in D state, reflected by the b: 4 and the crushing wa: 60% Worth knowing..

The memory subsystem is simultaneously a victim and an accomplice. 3. The presence of active swap (swpd: 5, si: 150, so: 30) indicates the working set has exceeded physical RAM, forcing the kernel to evict pages to the very block device that is already saturated (bi: 210, bo: 85). Also, memory pressure increases → kswapd wakes, writes dirty pages (bo), reads swapped pages (si). That's why this creates a positive feedback loop of degradation:

  1. Storage latency spikes → wa climbs, id hits zero.
  2. Now, 5. Application requests data not in cache → triggers read I/O (bi). Threads pile up on the run queue (r grows).
  3. Additional I/O contends with application reads → latency worsens.

The healthy cache (2 GiB) and buff (78 MiB) are effectively red herrings here; they represent capacity, not velocity. The storage device cannot drain the read queue fast enough to keep the CPU fed, rendering the cache hit rate effectively irrelevant for the active working set Small thing, real impact..

Real talk — this step gets skipped all the time.

Immediate Triage: Stop the Bleeding

Before provisioning new hardware, apply these kernel-level mitigations to buy breathing room:

  1. Throttle Writeback Aggressively
    The bo: 85 combined with high wa suggests dirty page writeback is fighting application reads for disk head time (or SSD controller queues).

    # Reduce the dirty page thresholds so writeback starts earlier, in smaller batches
    sysctl -w vm.dirty_background_ratio=5
    sysctl -w vm.dirty_ratio=10
    # Or, for absolute control on large RAM systems, use bytes:
    # sysctl -w vm.dirty_background_bytes=536870912   # 512 MiB
    # sysctl -w vm.dirty_bytes=1073741824             # 1 GiB
    

    This prevents massive, latency-spiking flush storms That's the part that actually makes a difference..

  2. Bias Reclaim Toward File Pages (Protect the Working Set)
    Since swpd is non-zero and si is active, the kernel is scanning anonymous pages. If this is a database or JVM workload, anonymous pages are the working set.

    # Favor reclaiming page-cache (file-backed) over anonymous (heap/stack)
    sysctl -w vm.swappiness=10
    # On kernels 5.8+, use memory pressure stall information (PSI) for finer control
    
  3. Elevate I/O Scheduler Priority for Critical Processes
    If a specific PID owns the bi load, ionice it into the real-time class (requires BFQ or Kyber scheduler, or none/mq-deadline with ioprio):

    ionice -c 1 -n 0 -p   # RT class, highest priority
    
  4. Disable Transparent Huge Pages (THP) if Latency-Sensitive
    THP compaction (khugepaged) can spike wa and sy unpredictably.

    echo never > /sys/kernel/mm/transparent_hugepage/enabled
    echo never > /sys/kernel/mm/transparent_hugepage/defrag
    

Strategic Remediation: Fix the Bottleneck

Kernel tuning is a bandage. The wa: 60% with id: 0 and a saturated run queue demands architectural changes:

Strategy When to Apply Expected Impact
Add RAM si/so persist after swappiness tuning; working set > RAM. Highest ROI. Even so, Eliminates swap I/O (si/so → 0), reduces bo (less reclaim pressure), frees I/O bandwidth for bi.
Tiered Storage (Fast Cache Tier) Read-heavy (bi >> bo), random access pattern, dataset > RAM but "hot set" < Fast Tier.

No fluff here — just what actually works.

| Tiered Storage (Fast Cache Tier) | Read-heavy (bi >> bo), random access pattern, dataset > RAM but "hot set" < Fast Tier. | | I/O Path Optimization | Kernel 5.Requires code changes but can eliminate the I/O bottleneck entirely. Here's the thing — | NVMe/Optane cache (bcache, dm-cache, LVM VDo) absorbs bi and bo for hot data, offloading the slow tier. g.Here's the thing — 10+, NVMe/SSD storage, latency-sensitive workload. That's why | Shifts the working set from disk to memory, directly reducing bi/bo for cached keys. , Redis, Memcached) can cache hot data in RAM. | Tune noop/none scheduler, enable noop for SSDs, or use Kyber for fine-grained latency control. Which means reduces wa by moving the I/O bottleneck to a faster device. | | Application-Level Caching | Specific services (e.Reduces scheduler overhead and queuing delays.

Validation and Iterative Monitoring

After applying kernel mitigations and strategic changes, avoid assuming the problem is solved. Re-measure using the same vmstat 2 or iostat interval:

  • Success Indicators: wa drops below 20%, id rises above 30%, si/so approach zero, and runq aligns with CPU cores.
  • Persistent High wa: If wa remains high after adding RAM or fast storage, investigate application I/O patterns (e.g., synchronous writes, inefficient queries) using iotop or strace.
  • Swap Re-engagement: If swpd grows again despite swappiness=10, the working set still exceeds RAM—re-evaluate memory allocation or consider container-level memory limits.

Conclusion

High wa with active swapping is a critical symptom of a system overwhelmed by I/O and memory pressure. But immediate kernel tuning—throttling writeback, protecting the working set via swappiness, and optimizing I/O priorities—can stabilize production environments temporarily. That said, the root cause demands architectural solutions: adding RAM to eliminate swapping, tiering storage to accelerate I/O, or refactoring applications to reduce disk dependence. On top of that, monitoring remains the cornerstone; validate each change iteratively, and let empirical data guide further adjustments. A system with wa under control is not just faster—it is resilient, predictable, and ready for growth.

Fresh Picks

Fresh Reads

Worth the Next Click

You're Not Done Yet

Thank you for reading about How To Check Cpu Utilization In Linux. We hope the information has been useful. Feel free to contact us if you have any questions. See you next time — don't forget to bookmark!
⌂ Back to Home