How to Check CPU Utilization in Linux
Understanding how much of your processor’s capacity is being used is essential for troubleshooting performance issues, planning capacity upgrades, or simply keeping an eye on system health. Linux provides a rich set of command‑line tools that report CPU utilization in different granularities—from overall system usage to per‑core statistics and per‑process breakdowns. This guide walks you through the most common utilities, explains what each metric means, and shows practical examples you can run on any modern distribution Took long enough..
Why Monitor CPU Utilization?
CPU utilization tells you the percentage of time the processor spends executing non‑idle tasks. High utilization can indicate:
- A CPU‑bound workload that may need optimization or more cores.
- A runaway process consuming excessive cycles.
- Insufficient hardware for the current workload, prompting a scale‑up or load‑balancing decision.
Conversely, consistently low utilization may suggest over‑provisioned resources or an I/O‑bound bottleneck elsewhere Easy to understand, harder to ignore..
Core Concepts Before the Commands
- User time (%us) – Time spent running user‑space applications.
- System time (%sy) – Time spent in the kernel handling system calls and interrupts.
- Nice time (%ni) – CPU time for low‑priority (nice) processes.
- Idle time (%id) – Percentage of time the CPU is doing nothing.
- Wait time (%wa) – Time spent waiting for I/O operations to complete.
- Interrupt (%hi) and Soft‑interrupt (%si) – Time servicing hardware and software interrupts.
Most tools display these fields as percentages that add up to 100 % (or close, depending on rounding).
Essential Tools for Checking CPU Utilization
| Tool | Primary Use | Key Options | Typical Output |
|---|---|---|---|
top |
Real‑time, interactive view of processes and overall CPU load | -b (batch), -d delay, -n iterations |
Overall %CPU, per‑process %CPU, memory |
htop |
Enhanced, colorized top with mouse support |
None needed for basic use | Visual bars per core, process tree |
mpstat |
Per‑core CPU statistics from the sysstat package |
-P ALL, interval, count |
%usr, %sys, %idle per CPU |
vmstat |
Virtual memory statistics, includes CPU columns | interval, count | %usr, %sys, %id, %wa, etc. |
sar |
Historical and real‑time system activity reporting | -u for CPU, interval, count |
Average CPU usage over time |
pidstat |
Per‑process CPU utilization over time | -p ALL, interval, count |
%CPU per PID |
glances |
All‑in‑one monitoring dashboard (curses/web) | -w for web mode |
CPU, memory, disk, network, alerts |
iostat |
Primarily I/O but includes %cpu column | -c for CPU only, interval |
%user, %nice, %system, %iowait, %idle |
All of these utilities are either part of the base system (top, vmstat, iostat) or available via common repositories (htop, sysstat for mpstat/sar/pidstat, glances).
Using top for a Quick Snapshot
top is pre‑installed on virtually every Linux distribution and provides an immediate overview.
top
When you run it, the first few lines look like:
top - 14:23:01 up 2:15, 2 users, load average: 0.31, 0.28, 0.25
Tasks: 124 total, 2 running, 122 sleeping, 0 stopped, 0 zombie
%Cpu(s): 12.3 us, 3.1 sy, 0.0 ni, 84.2 id, 0.4 wa, 0.0 hi, 0.0 si, 0.0 st
KiB Mem : 8034560 total, 3124560 free, 2856000 used, 2055400 buff/cache
KiB Swap: 2097148 total, 2097148 free, 0 used. 4987600 avail Mem
The line beginning with %Cpu(s): shows the aggregate utilization.
12.3 us→ 12.3 % user time3.1 sy→ 3.1 % system time84.2 id→ 84.2 % idle
Below that, each process lists its %CPU share, letting you spot hogs instantly.
Batch mode (useful for scripting):
top -b -d 2 -n 5
-bruns in batch (non‑interactive) mode.-d 2sets a 2‑second delay between samples.-n 5collects five iterations then exits.
Getting a Friendlier View with htop
If you prefer colors, scrollable process trees, and the ability to kill processes with a mouse click, install htop:
sudo apt-get install htop # Debian/Ubuntu
sudo yum install htop # RHEL/CentOS
sudo dnf install htop # Fedora
Launch it:
htop
You’ll see a CPU usage bar for each core at the top, making it trivial to spot imbalance (e.g.So , one core at 90 % while others sit at 5 %). Press F5 to toggle tree view, F6 to sort by column, and F9 to send a signal to a selected process It's one of those things that adds up..
Per‑Core Details with mpstat
Part of the sysstat package, mpstat reports statistics for each logical CPU.
mpstat -P ALL 1 3
Explanation:
-P ALL→ show every CPU (includingALLfor the aggregate).1→ sample interval in seconds.3→ number of reports (after which the command exits).
Sample output:
Linux 5.15.0-78-generic (myhost) 11/0
08/2025 _x86_64_ (8 CPU)
09:15:32 AM CPU %usr %nice %sys %iowait %steal %idle
All 11.That's why 20 0. 10 0.80 0.In practice, 10 0. Now, 85
5 13. 60 0.And 20 0. 00 2.00 84.10
1 12.Which means 25 0. 20 0.40 0.But 00 85. 50 0.00 82.15 0.90
3 10.In practice, 40 0. 20 0.00 3.40 0.20 0.Consider this: 50 0. 25
7 12.Even so, 60
6 10. Think about it: 00 0. 00 79.00 3.40 0.00 5.00 3.In practice, 00 83. 15 0.20 0.That said, 80 0. On top of that, 90 0. Practically speaking, 60
2 9. Worth adding: 00 3. 00 3.00 86.00 4.That's why 10
0 15. 80 0.10 0.30 0.In practice, 55
4 11. 20 0.60 0.Still, 00 86. But 00 86. Because of that, 00 3. 00 83.
The aggregate row (`All`) mirrors the values you’d see in `/proc/loadavg`, while individual CPU lines expose core‑specific behavior. A persistent high `%sys` on a single core often points to hardware interrupts or poorly tuned drivers, whereas a system‑wide spike in `%usr` usually indicates genuine application load.
---
## Tracking a Single Process Over Time with `pidstat`
`pidstat` (also from `sysstat`) lets you monitor one or more processes over successive intervals, showing CPU, memory, I/O, and context switches.
```bash
pidstat -p 1234,5678 2 5
-p 1234,5678selects the target PIDs (comma‑separated).2is the sampling interval in seconds.5gives the number of reports.
Typical output:
Linux 5.15.0-78-generic (myhost) 11/08/2025 _x86_64_ (8 CPU)
11:45:34 AM UID PID %usr %system %guest %wait CPU Command
0 1234 8.20 2.10 0.00 0.00 3 java
0 5678 5.50 1.So 80 0. 00 0.
For long‑running daemons, you can capture this data to a file and replay it later:
```bash
pidstat -p ALL 60 1440 >> /var/log/pidstat.log
This appends a 60‑second cadence for 24 hours, perfect for post‑mortem analysis or feeding into log‑aggregation tools Small thing, real impact..
Kernel‑Level Insights with vmstat
When you need a high‑level pulse of the entire system—processes, memory, swapping, and I/O—vmstat delivers in a single line per interval Small thing, real impact..
vmstat 3 10
3→ wait three seconds between reports.10→ stop after ten iterations.
Sample rows:
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
r b swpd free buff cache si so bi bo in cs us sy id wa st
1 0 0 2456784 123456 3456784 0 0 0 5 0 0 5 2 92 1 0
0 0 0 2451234 124000 3457890 0 0 2 8 2 1 4 1 94 1 0
Key columns:
- r – processes waiting for run time; values >
Digging Deeper into the vmstat Output
The single‑line snapshot that vmstat prints is a compact ledger of the system’s health. Once you know what each column represents, the line becomes a diagnostic map you can read at a glance.
| Column | Meaning | What to watch for |
|---|---|---|
| r | Runnable processes (ready to execute) | A persistent r count above the number of CPUs suggests a backed‑up run queue and possible CPU starvation. Values below 20 % on a multi‑core system usually merit investigation. In practice, |
| free | Free physical memory (kB) | Low free combined with high buff+cache may still be healthy (cache can be reclaimed), but a sustained drop warns of tight RAM. g.Still, |
| id | CPU idle time | The complement of us+sy+wa+st. |
| bi | Blocks received from a block device (e.Practically speaking, | |
| cs | Context switches per second | High cs can be a proxy for CPU‑bound workloads that frequently yield, or for systems with many lightweight threads. |
| so | Kibibytes swapped out per second | High so indicates the kernel is aggressively moving pages to disk, again a sign of insufficient RAM. Still, |
| sy | CPU time spent in kernel space | Spikes here often betray system calls, interrupts, or driver overhead. Still, , disk reads) per second |
| si | Kibibytes swapped in per second | Frequent si activity points to memory pressure; each swap‑in incurs a performance hit. So naturally, |
| us | CPU time spent in user space | Direct indicator of application CPU consumption. |
| buff | Memory used by block buffers (kernel page cache) | Large buff is a good sign – the kernel is caching disk reads for faster later access. |
| b | Processes in uninterruptible sleep (TASK_UNINTERRUPTIBLE) – often waiting for I/O |
High b values usually indicate blocked I/O, such as a hung filesystem operation or a device that isn’t delivering data. Now, |
| bo | Blocks sent to a block device (writes) per second | Elevated bo can reflect write‑intensive apps; watch for a mismatch between bi/bo that could hint at I/O imbalance. And |
| swpd | Amount of swapped memory (kilobytes) | Non‑zero values mean the system is paging; a steady rise can signal memory exhaustion. And |
| in | Interrupts per second | A sudden jump often traces to hardware events (network cards, timers) or poorly tuned drivers. |
| wa | Time waiting for I/O (percentage of total CPU) | High wa signals storage latency – think slow SSDs, network‑mounted filesystems, or overloaded storage arrays. |
| cache | Page cache + dentries/inodes | Similar to buff; a healthy system keeps this high until memory pressure forces reclamation. |
| st | Steal time (when a virtual CPU waits for the hypervisor to give it CPU) | Relevant only in virtualized environments; elevated st means the host is contending for physical CPU. |
Putting It All Together
A healthy snapshot on a modest server might look like:
0 0 0 2450000 130000 3600000 0 0 10 8 30 22 2 1 95 0 0
Interpretation: a few idle processes, no swapping, ample free memory, modest I/O, and CPUs mostly idle Simple, but easy to overlook..
A troublesome snapshot could be:
8 2 5 1200000 80000 2100000 2 0 150 30 210 85 15 10 60 5 0
Here you see a busy run queue (r=8 on, say, 4 CPUs), non‑trivial blocking (b=2), active swapping (swpd=5), low free memory, and high
Decoding the “Troublesome” Snapshot
The second example shows a system that is far from the ideal state:
8 2 5 1200000 80000 2100000 2 0 150 30 210 85 15 10 60 5 0
Below is a field‑by‑field walk‑through of what those numbers are whispering about the host:
| Field | Value | What it screams |
|---|---|---|
| r (run queue) | 8 | The scheduler has eight threads/completion‑queues waiting for CPU. Which means |
| us | 60 | 60 % of CPU time is spent in user space – applications are doing real work, not just waiting. |
| so | 30 | 30 KiB/s are being swapped out – the kernel is actively off‑loading pages, further adding to I/O pressure. Consider this: |
| sy | 5 | 5 % in kernel – normal driver/interrupt overhead. Which means |
| swpd (swap used) | 5 | Five MiB of swap are active, meaning the kernel is already paging out memory. Worth adding: |
| cs | 10 | 10 context switches per second – low, meaning the CPU is not thrashing with excessive thread switching. On top of that, |
| cache | 2100000 | About 2 GiB of page‑cache/dentry/inode cache. Think about it: on a 4‑core machine this is already a double‑loaded run queue – processes are spending time context‑switching rather than making progress. This is the classic “disk stall” indicator; the storage subsystem is struggling to keep up. |
| bi | 210 | 210 blocks/s are being read from the storage device. |
| free | 1200000 | Roughly **1.Day to day, the asymmetry (more reads than writes) points toward a data‑intensive service (caching, analytics, log processing) that is memory‑bound. Day to day, |
| si | 150 | 150 KiB/s are being swapped in. The processor is fully occupied, which explains the long run queue. In real terms, even a modest amount of swapping can throttle applications because each page‑in/out incurs a ~10‑30 ms latency. 2 GiB** of completely free RAM. |
| buff | 80000 | ~78 MiB of block‑buffer cache – decent, but the accompanying high wa (see below) suggests the data is not being served quickly enough. This is a clear sign that the working set no longer fits in RAM and the system is compensating with disk‑based paging. |
| id | 0 | Zero idle CPU. Still, the cache is healthy, yet the I/O wait time tells us the cache isn’t preventing disk‑level latency. Which means combined with bo (writes) this hints at a read‑heavy workload that is not being satisfied from the page cache. That said, |
| b (blocked) | 2 | Two processes are in uninterruptible sleep (often waiting for I/O). That's why |
| bo | 85 | 85 blocks/s are being written. |
| in | 15 | 15 interrupts per second – modest, but when paired with high bi it suggests that the storage controller is generating interrupts for each block completion, adding overhead. |
| wa | 60 | 60 % of total CPU time is waiting for I/O. That said, while not critical, it is far less than the 3‑4 GiB that a typical workload of this size would prefer to have free. This is the smoking gun: the storage subsystem cannot keep up with the read demand, and the high wa drags down id. |
Synthesizing the Signal: What the Numbers Are Really Saying
Taken individually, each metric is a data point; together, they form a coherent narrative of a system starved for storage throughput. The run queue (r: 12) is long not because the CPU lacks cycles—the us: 60% proves the application logic is hungry for work—but because the majority of those threads are stuck in D state, reflected by the b: 4 and the crushing wa: 60% Worth knowing..
The memory subsystem is simultaneously a victim and an accomplice. 3. The presence of active swap (swpd: 5, si: 150, so: 30) indicates the working set has exceeded physical RAM, forcing the kernel to evict pages to the very block device that is already saturated (bi: 210, bo: 85). Also, memory pressure increases → kswapd wakes, writes dirty pages (bo), reads swapped pages (si). That's why this creates a positive feedback loop of degradation:
- Storage latency spikes →
waclimbs,idhits zero. - Now, 5. Application requests data not in cache → triggers read I/O (
bi). Threads pile up on the run queue (rgrows). - Additional I/O contends with application reads → latency worsens.
The healthy cache (2 GiB) and buff (78 MiB) are effectively red herrings here; they represent capacity, not velocity. The storage device cannot drain the read queue fast enough to keep the CPU fed, rendering the cache hit rate effectively irrelevant for the active working set Small thing, real impact..
Real talk — this step gets skipped all the time.
Immediate Triage: Stop the Bleeding
Before provisioning new hardware, apply these kernel-level mitigations to buy breathing room:
-
Throttle Writeback Aggressively
Thebo: 85combined with highwasuggests dirty page writeback is fighting application reads for disk head time (or SSD controller queues).# Reduce the dirty page thresholds so writeback starts earlier, in smaller batches sysctl -w vm.dirty_background_ratio=5 sysctl -w vm.dirty_ratio=10 # Or, for absolute control on large RAM systems, use bytes: # sysctl -w vm.dirty_background_bytes=536870912 # 512 MiB # sysctl -w vm.dirty_bytes=1073741824 # 1 GiBThis prevents massive, latency-spiking flush storms That's the part that actually makes a difference..
-
Bias Reclaim Toward File Pages (Protect the Working Set)
Sinceswpdis non-zero andsiis active, the kernel is scanning anonymous pages. If this is a database or JVM workload, anonymous pages are the working set.# Favor reclaiming page-cache (file-backed) over anonymous (heap/stack) sysctl -w vm.swappiness=10 # On kernels 5.8+, use memory pressure stall information (PSI) for finer control -
Elevate I/O Scheduler Priority for Critical Processes
If a specific PID owns thebiload, ionice it into the real-time class (requiresBFQorKyberscheduler, ornone/mq-deadlinewithioprio):ionice -c 1 -n 0 -p# RT class, highest priority -
Disable Transparent Huge Pages (THP) if Latency-Sensitive
THP compaction (khugepaged) can spikewaandsyunpredictably.echo never > /sys/kernel/mm/transparent_hugepage/enabled echo never > /sys/kernel/mm/transparent_hugepage/defrag
Strategic Remediation: Fix the Bottleneck
Kernel tuning is a bandage. The wa: 60% with id: 0 and a saturated run queue demands architectural changes:
| Strategy | When to Apply | Expected Impact |
|---|---|---|
| Add RAM | si/so persist after swappiness tuning; working set > RAM. Highest ROI. Even so, |
Eliminates swap I/O (si/so → 0), reduces bo (less reclaim pressure), frees I/O bandwidth for bi. |
| Tiered Storage (Fast Cache Tier) | Read-heavy (bi >> bo), random access pattern, dataset > RAM but "hot set" < Fast Tier. |
No fluff here — just what actually works.
| Tiered Storage (Fast Cache Tier) | Read-heavy (bi >> bo), random access pattern, dataset > RAM but "hot set" < Fast Tier. |
| I/O Path Optimization | Kernel 5.Requires code changes but can eliminate the I/O bottleneck entirely. Here's the thing — | NVMe/Optane cache (bcache, dm-cache, LVM VDo) absorbs bi and bo for hot data, offloading the slow tier. g.Here's the thing — 10+, NVMe/SSD storage, latency-sensitive workload. That's why | Shifts the working set from disk to memory, directly reducing bi/bo for cached keys. , Redis, Memcached) can cache hot data in RAM. | Tune noop/none scheduler, enable noop for SSDs, or use Kyber for fine-grained latency control. Which means reduces wa by moving the I/O bottleneck to a faster device. |
| Application-Level Caching | Specific services (e.Reduces scheduler overhead and queuing delays.
Validation and Iterative Monitoring
After applying kernel mitigations and strategic changes, avoid assuming the problem is solved. Re-measure using the same vmstat 2 or iostat interval:
- Success Indicators:
wadrops below 20%,idrises above 30%,si/soapproach zero, andrunqaligns with CPU cores. - Persistent High
wa: Ifwaremains high after adding RAM or fast storage, investigate application I/O patterns (e.g., synchronous writes, inefficient queries) usingiotoporstrace. - Swap Re-engagement: If
swpdgrows again despiteswappiness=10, the working set still exceeds RAM—re-evaluate memory allocation or consider container-level memory limits.
Conclusion
High wa with active swapping is a critical symptom of a system overwhelmed by I/O and memory pressure. But immediate kernel tuning—throttling writeback, protecting the working set via swappiness, and optimizing I/O priorities—can stabilize production environments temporarily. That said, the root cause demands architectural solutions: adding RAM to eliminate swapping, tiering storage to accelerate I/O, or refactoring applications to reduce disk dependence. On top of that, monitoring remains the cornerstone; validate each change iteratively, and let empirical data guide further adjustments. A system with wa under control is not just faster—it is resilient, predictable, and ready for growth.