{"id":958,"date":"2026-10-04T06:02:59","date_gmt":"2026-10-04T00:32:59","guid":{"rendered":"https:\/\/merahost.org\/blog\/how-to-monitor-vps-resource-usage-using-htop-and-netdata\/"},"modified":"2026-10-04T06:02:59","modified_gmt":"2026-10-04T00:32:59","slug":"how-to-monitor-vps-resource-usage-using-htop-and-netdata","status":"publish","type":"post","link":"https:\/\/merahost.org\/blog\/how-to-monitor-vps-resource-usage-using-htop-and-netdata\/","title":{"rendered":"How to Monitor VPS Resource Usage Using htop and Netdata"},"content":{"rendered":"<p>Managing high-concurrency production workloads on virtualized infrastructure requires real-time observability into how the Linux kernel schedules threads, manages memory page caches, and services storage I\/O queues. When unexpected latency spikes degrade response times on unmonitored systems, systems engineers often scramble between disconnected command-line tools without clear visibility into hypervisor contention, CPU steal time, or memory pressure. Deploying high-availability infrastructure on <a href=\"https:\/\/merahost.org\">MeraHost<\/a> provides guaranteed bare-metal NVMe throughput, but maintaining optimal application performance still demands an authoritative, dual-layered monitoring strategy that pairs low-overhead terminal inspection with granular, continuous time-series telemetry.<\/p>\n<p><!-- more --><\/p>\n<h2>How to Monitor VPS Resource Usage Using htop and Netdata<\/h2>\n<div class=\"wp-block-group\" style=\"background:#f9f9f9;border-left:4px solid #001b41;padding:16px 20px;margin:20px 0;border-radius:4px\">\n<p style=\"margin:0;font-size:15px;line-height:1.6;color:#333\"><strong>Direct Answer:<\/strong> To <strong>monitor VPS resources<\/strong> effectively, combine real-time interactive CLI inspection using <strong>htop<\/strong> for immediate process triage and thread profiling with distributed, continuous telemetry collection via <strong>Netdata<\/strong> for per-second metric resolution, historical anomaly detection, and automated alerting. This hybrid approach enables sub-second bottleneck identification across CPU, memory, disk I\/O, and network stacks without degrading server throughput.<\/p>\n<\/div>\n<h2>Understanding VPS Virtualization Telemetry: vCPUs, Memory Slices, and Storage Queues<\/h2>\n<p>Virtual Private Servers (VPS) operate under Kernel-based Virtual Machine (KVM) or containerized cgroups environments where hardware resources are partitioned, scheduled, and shared. Unlike dedicated bare-metal servers where telemetry directly reflects physical silicon execution, monitoring virtualized instances requires isolating hypervisor-induced overhead from application-level bottlenecks.<\/p>\n<p>When observing Linux kernel metrics on a VPS instance, four core virtualization vectors must be monitored continuously:<\/p>\n<ul>\n<li><strong>CPU Steal Time (<code>%st<\/code>):<\/strong> Measures the percentage of time the virtual CPU had runnable threads ready for execution but was prevented from running by the hypervisor because the physical CPU cores were servicing other neighboring virtual machines. Any sustained steal time above 3% indicates severe hypervisor oversubscription.<\/li>\n<li><strong>Uninterruptible Sleep State (<code>D<\/code> State) and I\/O Wait (<code>%wa<\/code>):<\/strong> Processes blocked waiting for storage I\/O or kernel locks enter the <code>TASK_UNINTERRUPTIBLE<\/code> state. High <code>iowait<\/code> signifies that CPU cycles are idle solely because disk operations\u2014such as random SQLite\/MySQL writes or swap operations\u2014are backpressured.<\/li>\n<li><strong>Memory Cgroup Boundaries &amp; OOM Killer Invocations:<\/strong> The Linux kernel aggressively uses inactive RAM for disk page caching (Buffers and Cached). In virtualized environments, failure to differentiate between dirty application allocations (Active\/Anon) and reclaimable filesystem cache leads administrators to misdiagnose healthy caching as an imminent Out-Of-Memory (OOM) event.<\/li>\n<li><strong>Network Socket Backlog &amp; TCP Retransmissions:<\/strong> Sudden connection timeouts on web servers are frequently caused by ephemeral port exhaustion, syn-backlog overflows, or conntrack table saturation rather than raw bandwidth saturation.<\/li>\n<\/ul>\n<blockquote class=\"wp-block-quote\" style=\"background:#f9f9f9;border-left:4px solid #001b41;padding:16px 20px;margin:24px 0\">\n<p><strong style=\"color:#001b41\">Architecture Note:<\/strong> Commodity cloud hosting providers frequently oversubscribe physical CPU cores at ratios exceeding 4:1, resulting in unpredictable CPU steal spikes during peak business hours. On <a href=\"https:\/\/merahost.org\">MeraHost Enterprise Cloud<\/a>, all VPS compute nodes are provisioned on dedicated enterprise AMD EPYC and Intel Xeon processors with non-oversubscribed vCPU pinning and native PCIe Gen4 NVMe arrays, ensuring consistent zero-steal execution and sub-millisecond I\/O latency.<\/p>\n<\/blockquote>\n<h2>Mastering htop: Real-Time Interactive Process and Thread Triage<\/h2>\n<p>While the traditional UNIX <code>top<\/code> utility has served system administrators for decades, its monochromatic display, lack of hierarchical thread visualization, and clunky process management make it inefficient during high-stress operational outages. <code>htop<\/code> is an enhanced, interactive ncurses-based process viewer designed for rapid diagnostics, dynamic sorting, and granular thread accounting.<\/p>\n<h3>Navigating the htop Interface and Color Encoding<\/h3>\n<p>The top section of the <code>htop<\/code> terminal interface provides immediate graphical metering across all provisioned vCPU cores, physical memory, and swap space. Understanding the specific color-coded segments allows engineers to assess subsystem health in fractions of a second:<\/p>\n<ul>\n<li><strong>CPU Meters:<\/strong>\n<ul>\n<li><strong style=\"color:#20B038\">Green:<\/strong> Normal-priority user-space execution threads.<\/li>\n<li><strong style=\"color:#0055b3\">Blue:<\/strong> Low-priority (&#8220;niced&#8221;) user processes.<\/li>\n<li><strong style=\"color:#001b41\">Black \/ Dark Blue:<\/strong> Kernel-space system routines and syscall handling.<\/li>\n<li><strong>Cyan:<\/strong> Steal time (vCPU waiting for hypervisor allocation).<\/li>\n<li><strong>Orange \/ Amber:<\/strong> SoftIRQ and hard hardware interrupt processing.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Memory Meters:<\/strong>\n<ul>\n<li><strong style=\"color:#20B038\">Green:<\/strong> Used memory allocated by active user processes and resident sets.<\/li>\n<li><strong style=\"color:#0055b3\">Blue:<\/strong> Buffer memory (temporary metadata buffers for block devices).<\/li>\n<li><strong>Orange:<\/strong> Cache memory (page cache holding disk files in RAM for instantaneous read access).<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<h3>Essential Hotkeys for SysAdmin Rapid Incident Response<\/h3>\n<p>Under operational triage, mastering <code>htop<\/code> keyboard shortcuts eliminates the need to run multiple diagnostic commands:<\/p>\n<ul>\n<li><code>F5<\/code> \/ <code>t<\/code> (Tree View): Toggles hierarchical process parent-child relationships. This reveals instantly whether an Nginx worker, PHP-FPM pool, or Node.js cluster is spawning runaway child processes.<\/li>\n<li><code>F4<\/code> \/ <code>\\<\/code> (Incremental Filter): Filters the live process table by executable name, user account, or argument string without stopping live updates.<\/li>\n<li><code>F6<\/code> \/ <code>&gt;<\/code> (Sort Column Selector): Instantly shifts the sorting priority between CPU percentage (<code>PERCENT_CPU<\/code>), Resident Memory (<code>RES<\/code>), Virtual Memory (<code>VIRT<\/code>), and Disk I\/O rates.<\/li>\n<li><code>H<\/code>: Toggles visibility of user-space threads. Hiding threads reduces visual clutter when auditing large multithreaded applications such as Java JVMs or MySQL.<\/li>\n<li><code>K<\/code>: Toggles kernel thread visibility, hiding background kworkers and kswapd daemons to focus exclusively on application workloads.<\/li>\n<li><code>F9<\/code> \/ <code>k<\/code> (Kill Signal): Sends standard POSIX signals (<code>SIGTERM (15)<\/code>, <code>SIGKILL (9)<\/code>, <code>SIGHUP (1)<\/code>) directly to selected processes without requiring manual PID lookup.<\/li>\n<\/ul>\n<h3>Production htop Configuration Profile<\/h3>\n<p>By default, <code>htop<\/code> ships with generic settings that hide valuable columns such as detailed I\/O rates and normalized load averages. Deploying a tuned <code>htoprc<\/code> configuration file to <code>\/root\/.config\/htop\/htoprc<\/code> provides an enterprise-ready dashboard layout from the moment you establish an SSH session:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code># \/root\/.config\/htop\/htoprc - Enterprise Production Configuration\nfields=0 48 17 18 38 39 40 2 46 47 49 1\nsort_key=46\nsort_direction=-1\ntree_sort_key=0\ntree_sort_direction=1\nhide_kernel_threads=1\nhide_userland_threads=0\nshadow_other_users=0\nshow_thread_names=1\nshow_program_path=1\nhighlight_base_name=1\nhighlight_megabytes=1\nhighlight_threads=1\nhighlight_changes=1\nhighlight_changes_delay_secs=5\nfind_comm_in_cmdline=1\nstrip_exe_from_cmdline=1\nshow_merged_command=0\nheader_margin=1\nscreen_tabs=1\ndetailed_cpu_time=1\ncpu_count_from_one=1\nshow_cpu_usage=1\nshow_cpu_frequency=0\nshow_cpu_temperature=0\ndegree_fahrenheit=0\nupdate_process_names=0\naccount_guest_in_cpu_meter=1\ncolor_scheme=0\nenable_mouse=1\ndelay=15\nhide_function_bar=0\nheader_layout=two_50_50\ncolumn_meters_0=AllCPUs2 Memory Swap\ncolumn_meter_modes_0=1 1 1\ncolumn_meters_1=Tasks LoadAverage Uptime DiskIO NetworkIO\ncolumn_meter_modes_1=2 2 2 2 2<\/code><\/pre>\n<h2>Architectural Benchmarks: Comparing top, htop, and Netdata<\/h2>\n<p>Choosing the correct monitoring layer depends on whether you are conducting immediate interactive debugging during an outage or tracking longitudinal trends to forecast hardware capacity. The comparative matrix below highlights the operational trade-offs across the three dominant Linux monitoring tools:<\/p>\n<figure class=\"wp-block-table is-style-regular\">\n<table style=\"width:100%;border-collapse:collapse;margin:24px 0;font-size:15px;text-align:left\">\n<thead style=\"background:#001b41;color:#ffffff\">\n<tr>\n<th style=\"padding:12px 16px;border-bottom:2px solid #001b41\">Feature \/ Metric<\/th>\n<th style=\"padding:12px 16px;border-bottom:2px solid #001b41\">Standard \/ Default (top)<\/th>\n<th style=\"padding:12px 16px;border-bottom:2px solid #001b41\">Tuned \/ Production (htop + Netdata)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Metric Resolution<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">3.0s polling interval<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7;color:#20B038;font-weight:600\">1.0s real-time per-second resolution<\/td>\n<\/tr>\n<tr>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Historical Retention<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">None (transient snapshot only)<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7;color:#20B038;font-weight:600\">Days to months via dbengine compression<\/td>\n<\/tr>\n<tr>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Process Tree Hierarchy<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Unsupported \/ flat PID listing<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7;color:#20B038;font-weight:600\">Full parent-child tree visualization (htop F5)<\/td>\n<\/tr>\n<tr>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Daemon CPU Overhead<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">0% (runs only when invoked)<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7;color:#20B038;font-weight:600\">&lt; 1.5% single-core CPU utilization (Netdata)<\/td>\n<\/tr>\n<tr>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Automated Alerting Engine<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">None<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7;color:#20B038;font-weight:600\">Built-in threshold alerts (Slack, Discord, Email)<\/td>\n<\/tr>\n<tr>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Kernel eBPF &amp; Socket Tracing<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Unavailable<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7;color:#20B038;font-weight:600\">Native eBPF plugin for filesystem and network I\/O<\/td>\n<\/tr>\n<tr>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Visualization Interface<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Monochrome CLI text<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7;color:#20B038;font-weight:600\">Interactive CLI + Modern Web Dashboard (port 19999)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n<h2>Continuous Observability with Netdata: Per-Second High-Fidelity Telemetry<\/h2>\n<p>While <code>htop<\/code> excels at immediate manual inspection, it cannot capture intermittent micro-bursts, transient memory spikes that trigger the OOM killer at midnight, or creeping disk space exhaustion while you are away from the terminal. This is where <strong>Netdata<\/strong> establishes a production-grade telemetry foundation.<\/p>\n<p>Netdata is an autonomous, ultra-lightweight observability agent written in optimized C. Unlike heavy APM solutions that require gigabytes of RAM and complex Java runtimes, Netdata collects thousands of metrics per second while consuming less than 1.5% of a single CPU core and approximately 150 MB of memory.<\/p>\n<h3>Core Architectural Advantages of Netdata on VPS Instances<\/h3>\n<ul>\n<li><strong>Sub-Second Resolution:<\/strong> Netdata collects, stores, and evaluates metrics every single second (1s granularity). Standard monitoring systems like Prometheus or Datadog typically poll at 15-second, 30-second, or 60-second intervals, completely missing micro-bursts that degrade web application responsiveness.<\/li>\n<li><strong>Lockless Ring Buffer Time-Series Database (dbengine):<\/strong> Netdata stores time-series data using its custom <code>dbengine<\/code>, which implements zero-copy in-memory caching coupled with highly compressed on-disk tiering. This allows a VPS with modest NVMe storage to maintain weeks of per-second telemetry data without I\/O contention.<\/li>\n<li><strong>Automated Anomaly Detection:<\/strong> Netdata includes embedded machine learning models that continuously model baseline distributions for every individual metric, immediately flagging statistically anomalous behavior before it manifests as hard service failure.<\/li>\n<li><strong>Auto-Discovery of Application Stacks:<\/strong> Netdata automatically detects and instruments local Linux services including Nginx, LiteSpeed, Apache, MySQL\/MariaDB, PostgreSQL, Redis, and PHP-FPM without manual configuration.<\/li>\n<\/ul>\n<blockquote class=\"wp-block-quote\" style=\"background:#f9f9f9;border-left:4px solid #001b41;padding:16px 20px;margin:24px 0\">\n<p><strong style=\"color:#001b41\">SysAdmin Security Hardening:<\/strong> By default, Netdata binds its embedded web server to <code>0.0.0.0:19999<\/code>, exposing server telemetry to the public internet if firewall rules are unconfigured. In enterprise production environments, always bind Netdata strictly to <code>127.0.0.1<\/code> and proxy the web dashboard through Nginx with TLS encryption and HTTP Basic Authentication, or restrict access via SSH port forwarding: <code>ssh -L 19999:localhost:19999 user@vps-ip<\/code>.<\/p>\n<\/blockquote>\n<h3>Hardened Production Netdata Configuration (\/etc\/netdata\/netdata.conf)<\/h3>\n<p>Deploy the following optimized configuration file to <code>\/etc\/netdata\/netdata.conf<\/code>. This configuration restricts network binding to localhost, disables unnecessary cloud reporting telemetry for privacy, optimizes the <code>dbengine<\/code> storage tier for low-memory VPS instances, and pins collector intervals to 1 second:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code># \/etc\/netdata\/netdata.conf - Hardened Low-Footprint Production Profile\n[global]\n    run as user = netdata\n    history = 86400\n    update every = 1\n    memory mode = dbengine\n    page cache size = 32\n    dbengine multihost disk space = 512\n    dbengine disk space = 512\n    disconnect idle web clients after seconds = 60\n    enable web elevation = no\n    glibc malloc arena max for plugins = 1\n    glibc malloc arena max for netdata = 2\n\n[web]\n    bind to = 127.0.0.1:19999\n    default port = 19999\n    mode = static-threaded\n    web files owner = root\n    web files group = netdata\n    disconnect idle web clients after seconds = 60\n    respect do not track header = yes\n    x-frame-options response header = SAMEORIGIN\n\n[cloud]\n    conversation agent = no\n    statistics = no\n    anonymous statistics = no\n\n[plugins]\n    proc = yes\n    diskspace = yes\n    cgroups = yes\n    tc = no\n    idlejitter = yes\n    charts.d = no\n    python.d = yes\n    go.d = yes\n    node.d = no\n    apps = yes\n    ebpf = yes\n\n[health]\n    enabled = yes\n    in memory max health log entries = 1000\n    script to execute on alarm = \/usr\/libexec\/netdata\/plugins.d\/alarm-notify.sh<\/code><\/pre>\n<h3>Enforcing Systemd Cgroup Resource Quotas for Monitoring Daemons<\/h3>\n<p>Monitoring software should never become the cause of an outage. To prevent Netdata or its child collectors from exhausting CPU or memory during unexpected workload spikes, enforce deterministic resource constraints using a systemd drop-in override at <code>\/etc\/systemd\/system\/netdata.service.d\/override.conf<\/code>:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code># \/etc\/systemd\/system\/netdata.service.d\/override.conf\n[Service]\n# Hard resource boundaries using cgroup v2\nCPUAccounting=true\nCPUQuota=25%\nMemoryAccounting=true\nMemoryHigh=384M\nMemoryMax=512M\nMemorySwapMax=0M\n\n# Process scheduling and priority\nNice=19\nCPUSchedulingPolicy=idle\nIOSchedulingClass=idle\nIOSchedulingPriority=7\n\n# Security sandboxing\nProtectSystem=full\nProtectHome=true\nPrivateTmp=true\nCapabilityBoundingSet=CAP_SYS_PTRACE CAP_DAC_READ_SEARCH CAP_NET_ADMIN<\/code><\/pre>\n<p>After creating the override file, reload the systemd daemon and restart Netdata:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code>systemctl daemon-reload\nsystemctl restart netdata.service\nsystemctl status netdata.service<\/code><\/pre>\n<h3>Linux Kernel Sysctl Tuning for Telemetry and Memory Protection<\/h3>\n<p>High-resolution monitoring requires proper kernel tunables to ensure unprivileged profiling tools can extract process metadata without risking kernel panics or swap thrashing. Apply the following parameters to <code>\/etc\/sysctl.d\/99-vps-monitoring.conf<\/code>:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code># \/etc\/sysctl.d\/99-vps-monitoring.conf - Kernel Observability &amp; Virtual Memory Tuning\n\n# Minimize aggressive swap thrashing on low-memory VPS instances\nvm.swappiness = 10\nvm.vfs_cache_pressure = 50\n\n# Ensure memory overcommit does not crash critical database processes\nvm.overcommit_memory = 0\nvm.overcommit_ratio = 50\n\n# Increase maximum process ID allocation for high-concurrency threading\nkernel.pid_max = 65536\n\n# Enable unprivileged eBPF for Netdata kernel tracepoints\nkernel.unprivileged_bpf_disabled = 0\nkernel.perf_event_paranoid = 1\n\n# Protect dmesg output while allowing monitoring service access\nkernel.dmesg_restrict = 0\n\n# Prevent panic on OOM; prioritize kill order\nvm.panic_on_oom = 0<\/code><\/pre>\n<p>Load the updated kernel parameters immediately without rebooting:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code>sysctl -p \/etc\/sysctl.d\/99-vps-monitoring.conf<\/code><\/pre>\n<h2>Operational Playbook: Diagnosing the 4 Critical VPS Bottlenecks<\/h2>\n<p>When alerting triggers during production operations, follow this deterministic four-step triage playbook using <code>htop<\/code> and <code>Netdata<\/code> to identify and remediate the underlying failure mode:<\/p>\n<h3>1. Triaging High CPU Steal Time (%st &gt; 5%)<\/h3>\n<p><strong>Symptom:<\/strong> System load average spikes dramatically, but the sum of user (<code>%us<\/code>) and system (<code>%sy<\/code>) CPU utilization remains low. Web response latency degrades across all endpoints.<\/p>\n<p><strong>Diagnostic Flow:<\/strong><\/p>\n<ol>\n<li>Open <code>htop<\/code> and observe the CPU meter color bar. Look for the cyan segments representing steal time.<\/li>\n<li>Navigate to the Netdata dashboard under <strong>CPU &rarr; CPU Steal<\/strong>. Examine whether the steal spikes occur at regular periodic intervals (indicating a neighboring VPS running intensive cron jobs) or continuous saturation.<\/li>\n<li>Verify hypervisor scheduling latency by executing: <code>grep \"steal\" \/proc\/stat<\/code> over a 10-second interval.<\/li>\n<li><strong>Remediation:<\/strong> CPU steal cannot be resolved by application tuning because the bottleneck exists at the physical hypervisor layer. The only permanent resolution is migrating workloads to an infrastructure provider like <a href=\"https:\/\/merahost.org\">MeraHost<\/a> that enforces strict anti-oversubscription policies and guarantees dedicated hardware execution.<\/li>\n<\/ol>\n<h3>2. Identifying Uninterruptible Disk I\/O Saturation (D State Processes)<\/h3>\n<p><strong>Symptom:<\/strong> Commands hang when attempting directory traversal, database queries queue indefinitely, and <code>iowait<\/code> exceeds 20%.<\/p>\n<p><strong>Diagnostic Flow:<\/strong><\/p>\n<ol>\n<li>Launch <code>htop<\/code> and press <code>F6<\/code> to sort by <strong>IO_RATE<\/strong> or press <code>Shift + P<\/code> to inspect process states.<\/li>\n<li>Locate processes marked with the state flag <code>D<\/code>. Unlike sleeping processes (<code>S<\/code>), processes in <code>D<\/code> state are waiting directly on synchronous disk reads\/writes or filesystem journal commits.<\/li>\n<li>Check Netdata under <strong>Disks &rarr; Disk I\/O<\/strong> and <strong>Disk Backlog<\/strong>. If disk queue depth exceeds 4 requests continuously, the underlying storage tier is saturated.<\/li>\n<li><strong>Remediation:<\/strong> Identify the write-heavy process (e.g., unindexed MySQL slow queries, unbuffered log writes). Enable write buffering in application configs, partition logging to separate tmpfs mounts, or migrate to pure enterprise NVMe storage.<\/li>\n<\/ol>\n<h3>3. Diagnosing Memory Leaks and Swap Thrashing<\/h3>\n<p><strong>Symptom:<\/strong> Free memory approaches zero, swap usage rises steadily, and database daemons unexpectedly restart due to kernel OOM kills.<\/p>\n<p><strong>Diagnostic Flow:<\/strong><\/p>\n<ol>\n<li>Open <code>htop<\/code> and examine the memory bar. Distinguish between green (active allocations) and orange (reclaimable cache). If orange dominates, memory is not depleted.<\/li>\n<li>Press <code>Shift + M<\/code> in <code>htop<\/code> to sort processes strictly by resident memory consumption (<code>RES<\/code>). Look for memory growth over time in worker pools.<\/li>\n<li>In Netdata, open <strong>Memory &rarr; System Memory<\/strong> and cross-reference <strong>Page Faults<\/strong>. A rapid rise in major page faults indicates the kernel is actively reading pages from slow swap disk, severely degrading execution speed.<\/li>\n<li>Review recent kernel OOM termination events: <code>dmesg -T | grep -i \"oom-killer\"<\/code>.<\/li>\n<li><strong>Remediation:<\/strong> Cap process worker lifecycles (e.g., <code>pm.max_requests = 500<\/code> in PHP-FPM) to flush memory leaks, and tune <code>vm.swappiness = 10<\/code> to keep memory pages in physical RAM.<\/li>\n<\/ol>\n<h3>4. Detecting Ephemeral Socket and Network Backlog Saturation<\/h3>\n<p><strong>Symptom:<\/strong> Clients report intermittent &#8220;Connection Refused&#8221; or SSL handshake timeouts while CPU and memory metrics appear completely normal.<\/p>\n<p><strong>Diagnostic Flow:<\/strong><\/p>\n<ol>\n<li>In Netdata, navigate to <strong>IP &rarr; TCP Sockets<\/strong> and <strong>TCP Errors<\/strong>.<\/li>\n<li>Inspect the count of sockets in <code>TIME_WAIT<\/code> and <code>CLOSE_WAIT<\/code> states. A massive accumulation of <code>TIME_WAIT<\/code> sockets indicates backend reverse proxies are closing connections without persistent HTTP keep-alive.<\/li>\n<li>Check for TCP backlog drops: <code>netstat -s | grep \"listen queue\"<\/code>.<\/li>\n<li><strong>Remediation:<\/strong> Increase <code>net.core.somaxconn = 65535<\/code> and enable TCP socket reuse via <code>net.ipv4.tcp_tw_reuse = 1<\/code> in sysctl.<\/li>\n<\/ol>\n<h2>Frequently Asked Questions: VPS Resource Monitoring<\/h2>\n<details class=\"wp-block-group\" style=\"background:#f9f9f9;border:1px solid #e7e7e7;border-radius:4px;padding:14px;margin-bottom:12px\">\n<summary style=\"cursor:pointer;font-weight:600;color:#001b41\">Does running Netdata 24\/7 impact VPS performance on low-spec instances?<\/summary>\n<p style=\"margin-top:10px;color:#444\">No. Netdata is written in highly optimized C with lockless circular ring buffers and asynchronous multi-tier storage engines. On a standard 1-vCPU \/ 1GB RAM virtual server, Netdata typically consumes between 0.8% and 1.5% of single-core CPU capacity and less than 150 MB of memory. By applying the systemd cgroup resource slice override provided in this guide, you can strictly constrain Netdata to a maximum of 25% CPU quota and 384 MB memory ceiling, guaranteeing zero interference with your production web applications.<\/p>\n<\/details>\n<details class=\"wp-block-group\" style=\"background:#f9f9f9;border:1px solid #e7e7e7;border-radius:4px;padding:14px;margin-bottom:12px\">\n<summary style=\"cursor:pointer;font-weight:600;color:#001b41\">Why does htop report high memory usage when free -m shows ample available RAM?<\/summary>\n<p style=\"margin-top:10px;color:#444\">This is a common point of confusion in Linux memory management. The Linux kernel follows the philosophy that unused RAM is wasted RAM. When physical memory is not required by active applications, the kernel allocates it to buffer cache and page cache to accelerate disk read operations. <code>htop<\/code> represents this cache visually as orange segments on the memory meter. This cached memory is instantly reclaimable by the kernel whenever an application requests new memory allocations. As long as resident application memory (green) is within safe limits and swap is unused, high cache utilization is normal and beneficial.<\/p>\n<\/details>\n<details class=\"wp-block-group\" style=\"background:#f9f9f9;border:1px solid #e7e7e7;border-radius:4px;padding:14px;margin-bottom:12px\">\n<summary style=\"cursor:pointer;font-weight:600;color:#001b41\">How do I securely access the Netdata web dashboard without exposing port 19999?<\/summary>\n<p style=\"margin-top:10px;color:#444\">Never expose port 19999 directly to the public internet without authentication. The recommended enterprise approach is to bind Netdata strictly to <code>127.0.0.1:19999<\/code> in <code>\/etc\/netdata\/netdata.conf<\/code>. To view the dashboard securely, create an encrypted SSH tunnel from your local workstation: <code>ssh -L 19999:localhost:19999 root@your-vps-ip<\/code>. Once connected, access <code>http:\/\/localhost:19999<\/code> in your local browser. Alternatively, configure an Nginx reverse proxy with Let&#8217;s Encrypt SSL, IP whitelisting, and HTTP Basic Authentication.<\/p>\n<\/details>\n<details class=\"wp-block-group\" style=\"background:#f9f9f9;border:1px solid #e7e7e7;border-radius:4px;padding:14px;margin-bottom:12px\">\n<summary style=\"cursor:pointer;font-weight:600;color:#001b41\">What is the practical difference between CPU steal (%st) and I\/O wait (%wa)?<\/summary>\n<p style=\"margin-top:10px;color:#444\">While both metrics represent waiting states, their underlying causes are completely different. CPU steal (<code>%st<\/code>) occurs when your virtual machine has executable code ready to run, but the host hypervisor is unable to grant execution time because the physical CPU is busy running tasks for other virtual servers on the same node. Conversely, I\/O wait (<code>%wa<\/code>) means your virtual CPU is intentionally idling because your own processes are blocked waiting for disk storage or network filesystem operations to complete. High steal indicates host hardware oversubscription, whereas high iowait indicates disk throughput saturation or inefficient application queries.<\/p>\n<\/details>\n<p>When architecting business-critical infrastructure, continuous monitoring is only half the equation; the underlying hardware platform must provide guaranteed compute cycles and deterministic storage performance. If your monitoring metrics frequently reveal CPU steal spikes or I\/O wait bottlenecks on legacy cloud providers, consider migrating your production instances to <a href=\"https:\/\/merahost.org\">MeraHost Enterprise Cloud<\/a>. With dedicated NVMe arrays, unthrottled gigabit networking, and zero price hikes since 2012, MeraHost delivers the stable baseline your telemetry deserves.<\/p>\n<div class=\"wp-block-group has-background\" style=\"background:#f9f9f9;border:1px solid #e7e7e7;border-radius:8px;padding:32px;margin:40px 0;text-align:center\">\n<h3 style=\"color:#001b41;margin-top:0;font-size:24px;font-weight:700\">Deploy Enterprise-Grade Production Infrastructure<\/h3>\n<p style=\"color:#444;font-size:16px;line-height:1.6;max-width:680px;margin:12px auto 24px auto\">Need guaranteed performance with zero price hikes? Host mission-critical workloads on <strong style=\"color:#001b41\">MeraHost<\/strong> with pure Enterprise NVMe, LiteSpeed Web Server, and Same Renewal Price, Always (starting at \u20b999\/mo).<\/p>\n<div class=\"wp-block-buttons\" style=\"display:flex;gap:16px;justify-content:center;flex-wrap:wrap\">\n<div class=\"wp-block-button\"><a class=\"wp-block-button__link\" href=\"https:\/\/merahost.org\" style=\"background:#001b41;color:#ffffff;font-weight:700;padding:12px 28px;border-radius:4px;text-decoration:none;display:inline-block;font-size:15px\">Explore MeraHost NVMe Cloud &rarr;<\/a><\/div>\n<div class=\"wp-block-button is-style-outline\"><a class=\"wp-block-button__link\" href=\"https:\/\/cpanelfree.com\" style=\"background:transparent;color:#001b41;font-weight:600;padding:12px 24px;border:2px solid #001b41;border-radius:4px;text-decoration:none;display:inline-block;font-size:15px\" rel=\"nofollow noopener\" target=\"_blank\">Deploy Free Staging on CpanelFree<\/a><\/div>\n<\/div>\n<\/div>\n\n\n<div class=\"kk-star-ratings kksr-auto kksr-align-left kksr-valign-bottom\"\n    data-payload='{&quot;align&quot;:&quot;left&quot;,&quot;id&quot;:&quot;958&quot;,&quot;slug&quot;:&quot;default&quot;,&quot;valign&quot;:&quot;bottom&quot;,&quot;ignore&quot;:&quot;&quot;,&quot;reference&quot;:&quot;auto&quot;,&quot;class&quot;:&quot;&quot;,&quot;count&quot;:&quot;0&quot;,&quot;legendonly&quot;:&quot;&quot;,&quot;readonly&quot;:&quot;&quot;,&quot;score&quot;:&quot;0&quot;,&quot;starsonly&quot;:&quot;&quot;,&quot;best&quot;:&quot;5&quot;,&quot;gap&quot;:&quot;5&quot;,&quot;greet&quot;:&quot;Rate this post&quot;,&quot;legend&quot;:&quot;0\\\/5 - (0 votes)&quot;,&quot;size&quot;:&quot;20&quot;,&quot;title&quot;:&quot;How to Monitor VPS Resource Usage Using htop and Netdata&quot;,&quot;width&quot;:&quot;0&quot;,&quot;_legend&quot;:&quot;{score}\\\/{best} - ({count} {votes})&quot;,&quot;font_factor&quot;:&quot;1.25&quot;}'>\n            \n<div class=\"kksr-stars\">\n    \n<div class=\"kksr-stars-inactive\">\n            <div class=\"kksr-star\" data-star=\"1\" style=\"padding-right: 5px\">\n            \n\n<div class=\"kksr-icon\" style=\"width: 20px; height: 20px;\"><\/div>\n        <\/div>\n            <div class=\"kksr-star\" data-star=\"2\" style=\"padding-right: 5px\">\n            \n\n<div class=\"kksr-icon\" style=\"width: 20px; height: 20px;\"><\/div>\n        <\/div>\n            <div class=\"kksr-star\" data-star=\"3\" style=\"padding-right: 5px\">\n            \n\n<div class=\"kksr-icon\" style=\"width: 20px; height: 20px;\"><\/div>\n        <\/div>\n            <div class=\"kksr-star\" data-star=\"4\" style=\"padding-right: 5px\">\n            \n\n<div class=\"kksr-icon\" style=\"width: 20px; height: 20px;\"><\/div>\n        <\/div>\n            <div class=\"kksr-star\" data-star=\"5\" style=\"padding-right: 5px\">\n            \n\n<div class=\"kksr-icon\" style=\"width: 20px; height: 20px;\"><\/div>\n        <\/div>\n    <\/div>\n    \n<div class=\"kksr-stars-active\" style=\"width: 0px;\">\n            <div class=\"kksr-star\" style=\"padding-right: 5px\">\n            \n\n<div class=\"kksr-icon\" style=\"width: 20px; height: 20px;\"><\/div>\n        <\/div>\n            <div class=\"kksr-star\" style=\"padding-right: 5px\">\n            \n\n<div class=\"kksr-icon\" style=\"width: 20px; height: 20px;\"><\/div>\n        <\/div>\n            <div class=\"kksr-star\" style=\"padding-right: 5px\">\n            \n\n<div class=\"kksr-icon\" style=\"width: 20px; height: 20px;\"><\/div>\n        <\/div>\n            <div class=\"kksr-star\" style=\"padding-right: 5px\">\n            \n\n<div class=\"kksr-icon\" style=\"width: 20px; height: 20px;\"><\/div>\n        <\/div>\n            <div class=\"kksr-star\" style=\"padding-right: 5px\">\n            \n\n<div class=\"kksr-icon\" style=\"width: 20px; height: 20px;\"><\/div>\n        <\/div>\n    <\/div>\n<\/div>\n                \n\n<div class=\"kksr-legend\" style=\"font-size: 16px;\">\n            <span class=\"kksr-muted\">Rate this post<\/span>\n    <\/div>\n    <\/div>\n","protected":false},"excerpt":{"rendered":"<p>Master Linux VPS telemetry with htop and Netdata. Track CPU, RAM, and I\/O bottlenecks using real-time metrics and low-overhead collectors.<\/p>\n","protected":false},"author":1,"featured_media":957,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[149],"tags":[126,125,129,150,127],"class_list":["post-958","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-server-admin","tag-devops","tag-linux","tag-performance","tag-server-admin","tag-sysadmin"],"views":0,"_links":{"self":[{"href":"https:\/\/merahost.org\/blog\/wp-json\/wp\/v2\/posts\/958","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/merahost.org\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/merahost.org\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/merahost.org\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/merahost.org\/blog\/wp-json\/wp\/v2\/comments?post=958"}],"version-history":[{"count":0,"href":"https:\/\/merahost.org\/blog\/wp-json\/wp\/v2\/posts\/958\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/merahost.org\/blog\/wp-json\/wp\/v2\/media\/957"}],"wp:attachment":[{"href":"https:\/\/merahost.org\/blog\/wp-json\/wp\/v2\/media?parent=958"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/merahost.org\/blog\/wp-json\/wp\/v2\/categories?post=958"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/merahost.org\/blog\/wp-json\/wp\/v2\/tags?post=958"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}