Skip to content

September 2026

0.18.0 Fixed some exotic passive checks, and security hardening

Most of 0.18.0 is about things that reported success while doing nothing. check_and_forward built its submission in a form no channel could read, check_nscp had counted zero crash reports since 0.4.2, and a module enabled by a fleet bundle was never loaded until the service restarted — all three are fixed. Security reviews of the SMTP, NRPE, NRDP/NSCA and WEB modules landed alongside them, and the bundled OpenSSL moves to 3.5.8.

It also adds run_schedules for submitting a passive result without waiting out the interval, lets the MSI install your own TLS certificates, and gives real-time filters a way to prime their destination at startup.

✨ Highlights

  • 📤 check_and_forward submits again. The command ran the check, answered Message submitted and delivered nothing: the submission was built as a query message, which no channel can read. NSCA, NRDP, Graphite and every other client module were equally affected. It also gained channel, alias, destination and source. (#1452)
  • 📅 New: run_schedules. Run the configured schedules now instead of waiting out their interval, and submit the results on their normal channel — nscp client --boot --query run_schedules, optionally --argument schedule=<alias>. Works over NRPE, REST and nscp test too. (#1450, #1452)
  • 💥 check_nscp is a filter check, and its crash count works. It has read 0 crashes since 0.4.2 (it matched the extension txt against .txt), and since 0.6.10 there were no .txt reports to find. It now recognises .crash, reads the configured archive folder again, and exposes crashes, errors, uptime, crash_age, last_crash, last_error, version and date as filter keywords. (#1451)
  • 🛡️ Four security reviews — SMTP, NRPE, NRDP/NSCA and WEB — closed a set of defense-in-depth gaps: STARTTLS response injection, an unvalidated EHLO name, certificate verification that could not work on Windows, a metachar guard that ran before decoding, secrets in the trace log, and session tokens from a non-cryptographic generator.
  • 🔐 The bundled OpenSSL moves from 3.5.4 to 3.5.8 in the Windows builds, picking up four upstream security releases — most relevantly CVE-2025-11187, a stack overflow parsing a hostile PKCS#12 file, reachable through check_certificate. (#1445)
  • 🪟 The MSI can install your own TLS certificates. CERTIFICATE, CERTIFICATE_KEY and CERTIFICATE_CA place your files where every server module reads them, so the self-signed fallback is never generated. (#568)
  • ⏱️ Real-time filters can prime their destination at startup. A new run on startup key submits the filter’s empty message once when the agent starts, so check_cache stops answering “Entry not found” after a restart. (#584)
  • 📡 New ${address_ipv4} / ${address_ipv6*} hostname placeholders for every passive client, so a host can report itself by address instead of name. (#349)

🔍 Detailed changes

📤 CheckHelpers — check_and_forward delivers, and takes arguments properly

The command handed the raw QueryResponseMessage to the submission path, but channels parse a SubmitRequestMessage and the two are not wire compatible — the payload sits in a different field, and field 2 of a submit message is the channel string. The channel received a message with zero payloads and cheerfully reported success. It now converts the query result into a submission first and checks the reply.

Option Meaning
channel Where to submit (default NSCA); target remains a synonym
alias Service description; defaults to the wrapped command’s name
destination Destination host for the submission
source Source host for the submission

target previously defaulted to the empty string, for which no handler exists, so even a correctly built message had nowhere to go.

📅 Scheduler — run_schedules

run_schedules executes the schedules under [/settings/scheduler/schedules] immediately and submits each result on its own channel, target, source and alias with the same report filter, so the monitoring server cannot tell it from a timed run. The timers are untouched.

nscp client --boot --query run_schedules
nscp client --boot --query run_schedules --argument schedule=cpu

schedule= is repeatable and defaults to every schedule; an unknown alias is an error that names the ones you have. A schedule whose command is run_schedules is refused, and a reentrancy guard catches the indirect case (via check_timeout, for instance) — previously that recursed until the agent died.

The caller’s identity is forwarded to the checks it runs, so REST and NRPE permissions apply to them; the scheduler’s own timed runs stay attributed to Scheduler.

💥 CheckNSCP — check_nscp rewritten as a filter check

Three independent bugs kept the crash count at zero: the extension comparison (txt vs the .txt the helper returns) has been false since 0.4.2; 0.6.10 replaced breakpad’s <guid>.dmp + .dmp.txt pair with a single <timestamp>.crash file, so even a fixed match found nothing; and 0.4.3 stopped reading [/settings/crash] archive folder, hardcoding the compile-time default. last_crash was never populated either — the newest-file watermark started at the current time.

New filter keywords: crashes, errors, uptime, crash_age, last_crash, last_error, version, date. Thresholds accept duration units (crit=uptime < 5m, crit=crash_age < 7d), and a new max-unit option (default w) caps the largest unit rendered. Crash reports are a Windows concept; on Linux crashes is always 0.

⚙️ Core — reloads, channel verdicts and denied checks

  • A reload re-reads the included files before deciding which modules should run, so a module enabled in one since the last load is picked up. Only the includes are refreshed: clearing the whole store would discard configuration held in memory, which is exactly how nscp unit and nscp client set themselves up. One unreadable include no longer aborts the whole reload. (#1455)
  • Every channel’s verdict is reported for a channel list. All handlers were handed the same response buffer, so with channel=NSCA,GRAPHITE a failing NSCA was masked by a succeeding GRAPHITE.
  • A denied check is no longer submitted. The permission layer answers a denied query as a successful query carrying an UNKNOWN “Permission denied” payload; run_schedules and check_and_forward forwarded it, overwriting the last real result on the server while reporting success to the caller.
  • nscp client --query <cmd> no longer appends No module was specified… to every result.

🛡️ Security reviews

🔒 SMTPClient. Data pipelined across the STARTTLS handshake is refused (RFC 3207 §4) — a prepared run of 2xx replies could otherwise walk the client through MAIL/RCPT/DATA and have it report an alert delivered while nothing was sent. The EHLO name is validated before connect, closing command injection on a relayed submission. Certificates are verified against a CA bundle through a new ca target setting and --ca argument (default ${ca-path}): the client previously used OpenSSL’s default verify paths only, which on Windows excludes the certificate store, so security=starttls failed against Gmail and M365 and operators simply turned verification off. EHLO capabilities are matched per reply line rather than by substring — a greeting naming host starttls.example.com used to satisfy the STARTTLS lookup.

🔒 NRPE. The allow nasty characters = false guard now also runs on the decoded command and arguments; with a non-UTF-8 encoding, a multi-byte sequence could decode into a metacharacter that was never literally on the wire. A new expose version server setting (default true) lets the unauthenticated _NRPE_CHECK reply stop naming the exact build. nscp nrpe install reads the stored verify mode again instead of silently resetting it on every re-run.

🔒 NRDP / NSCA. A malformed <status></status> response no longer null-derefs the agent. The NRDP token and any proxy-URL credentials are redacted from the trace log — including from the Target configuration: dump, which printed the raw settings map and defeated the redaction elsewhere (NSCA’s password leaked the same way). An https:// submission made through nscp client or REST with no verify mode now defaults to peer rather than trusting any certificate.

🔒 WEBServer. Session tokens and generated admin passwords come from OpenSSL’s CSPRNG with unbiased rejection sampling, and the server now fails closed if that RNG fails (HTTP 500 and a SECURITY: log line) rather than falling back to a weaker generator. Cookie lookups require a name boundary, so eviltoken no longer satisfies a lookup for token. Session validity and identity are read in one locked observation, closing an expiry race that could drop a request onto the anonymous grant. The web installer refuses an HTTPS→HTTP redirect on the bundle download path.

🔒 OpenSSL 3.5.8. Windows builds only; Linux packages link the distribution’s OpenSSL. See Security notices.

📡 Clients — --source-host names the sender

--source-host / --sender-host were registered against the destination container, where the well-known host key is routed into the typed address field — so naming a source host silently redirected the connection to it, and the sender the handler reads was never set. SMTPClient and NRDPClient had each worked around this with their own copies, which made the option ambiguous and therefore unusable on exactly the two modules most likely to need it.

🔧 Settings

  • settings --update --add-defaults --use-samples writes the sample objects; the flag was parsed and never read. --remove-defaults now enumerates samples too, making the two exact inverses. (#233)
  • A [/includes] entry naming a directory no longer breaks saving. The directory path reached the INI writer, and the resulting “Is a directory” error aborted the save — so nscp settings --set and the web UI failed and the main file was never written. (#636)
  • New ${address_ipv4} and ${address_ipv6*} placeholders resolve in the hostname setting of NSCA, NSCANg, NRDP, Graphite, Syslog, Op5, Icinga, Elastic and Collectd clients. The address is the source address of the default route, falling back to the first non-loopback address the host name resolves to. (#349)

⏱️ Real-time filters — run on startup

A new boolean filter key on the shared filter object, so it applies to CheckLogFile, CheckEventLog and the CheckSystem/CheckSystemUnix real-time filters. When true the filter submits its empty message with OK status once at startup, priming the destination. It is registered without a default on purpose: an absent key keeps the inherited value, while an explicit false overrides an inherited true. Delivery is retried across shortened waits while later modules (such as SimpleCache) are still loading. (#584)

🪟 Windows installer — install your own TLS certificates

Three new silent-install properties place your own files under the default names every server module reads, so the self-signed fallback is never generated:

Property Installed as
CERTIFICATE certificate.pem
CERTIFICATE_KEY certificate_key.pem
CERTIFICATE_CA ca.pem

The install fails if a named file is missing or is not PEM, or if a certificate is given with no key anywhere. CERTIFICATE_KEY also writes certificate key = … for the NRPE and WEB servers, so it is incompatible with ALLOW_CONFIGURATION=0; other servers need the setting added by hand. (#568)

🐛 Bug fixes

  • SMTP: the timeout error now says what happened in the client’s own words (timed out after 30s (the budget for the whole submission)) rather than reporting a platform error that blamed the connected party; a failed connect raises the exception callers actually catch; insecure-skip-verify is accepted over REST and no longer resets a configured target’s value on every submission; the reference no longer shows its default as N/A.
  • NRPE: the arguments-case rejection says “arguments” instead of “command”.
  • check_nscp: a crash report whose timestamp cannot be read still counts but takes no part in newest-wins, so crash_age no longer reports ~56 years.
  • Configuring with -DNSCP_BUILD_TESTS=OFF no longer aborts on the mongoose_wrapper_test target.

⚠️ Upgrade notes

  • check_and_forward starts delivering, and no longer takes positional arguments. Anything that treated Message submitted as success will now see real channel failures. check_and_forward command=check_cpu warn=load>80 now fails to parse — pass one arguments= per wrapped argument instead: check_and_forward command=check_cpu "arguments=warn=load>80". The positional form had to go because its parser also swallowed the CLI’s own --argument key=value tokens, which is what fed the wrapped command garbage.
  • check_nscp may start reporting CRITICAL. Its crash count works now, so an agent with an old report still in the archive folder will report a crash where it read 0. Threshold on crash_age ("crit=crash_age < 7d") if you only care about recent crashes, or clean the folder out. The message also loses its last crash: / last error: fragments — put them back with an explicit detail-syntax if you match on the text.
  • A denied check now fails instead of submitting. If you restrict what a REST or NRPE identity may run, expect an error where a stale UNKNOWN previously appeared on the monitoring server.
  • NRDP over HTTPS verifies by default on the nscp client / REST path. Pass --verify none (or point --ca at the certificate) to keep submitting to a self-signed endpoint that way. Configured targets already defaulted to peer. 🔒
  • SMTP verifies the server certificate, against ${ca-path} by default. A target that relied on verification being effectively off needs ca pointed at the right bundle, ca = none for OpenSSL’s defaults, or insecure-skip-verify = true. 🔒
  • The SMTP timeout is now a budget for the whole submission, not a fresh deadline per operation — a target that only completed by consuming several multiples of it will now give up.
  • --source-host no longer redirects the connection. If you used it to choose where to connect, use --host / --address.
  • settings --add-defaults --use-samples now writes samples. Drop the flag if you were relying on today’s sample-free output; --remove-defaults strips them again.
  • nscp client --query output loses its trailing No module was specified… line. Scripts that stripped it can stop.
  • Settings writes work again when [/includes] names a directory. If you removed a directory include to get saving working, you can put it back.
  • WEBServer: a script or module name beginning with - is rejected (rename it; interior dashes are fine), and POST /auth/logout now enforces allowed hosts. 🔒
  • NRPE: allow nasty characters = false now also inspects decoded input, so a request the guard was always meant to block may now be rejected. Set expose version = false to stop the _NRPE_CHECK ping naming your build. 🔒

Download

You can download the new version from GitHub

// Michael Medin

0.17.0 Windows server roles get their own checks, and check messages finally read like numbers

0.17.0 adds twelve new checks — eight for IIS and Remote Desktop Services on Windows, four for the status pages of the common web servers — and gives every filter check control over how it renders numbers, so 140.293GB/0.983TB can become 141.09GB/1006.85GB (or 141,09GB/1.006,85GB). Alongside that, the filter engine stops quietly doing the wrong thing: text-versus-number comparisons are numeric, fractional thresholds mean what they say, and an error inside a syntax template is reported instead of rendering a blank.

Highlights

  • Eight new Windows checks for IIS and Remote Desktop Services. The new CheckWindowsApps module covers IIS sites, application pools, worker processes and HTTP.sys request queues, plus RDS CAL licensing, session counts, per-session load and the Connection Broker counterset.
  • Four new web-server status checks. check_apache_status, check_nginx_status, check_phpfpm_status and check_tomcat_status read the vendors’ machine-readable status endpoints over HTTP(S), sharing check_http’s auth and TLS handling.
  • Check messages can be told how to render their numbers. Four new options — decimals, byte-unit, decimal-separator, thousands-separator — on every filter check and every real-time filter (#1428). Perfdata and thresholds are untouched.
  • Filter comparisons between text and a bare number are now numeric. filter=value > 90 no longer matches value=100 as false because “100” sorts before “90”, and 90 > value evaluates at all.
  • Fractional thresholds stop being truncated. count > 2.5 meant count > 3 and working_set > 1.5g meant 1g; both now mean what they say.
  • Host name placeholders resolve across the whole settings subsystem — including attachment targets and [/includes] (#458) — and are sanitized before they land in a local path 🔒.
  • Syslog submission works again after ten years. SyslogClient read its connection settings from the wrong place and sent nothing at all; a configured syslog target will start receiving traffic on upgrade.
  • Check-specific filter keywords that shadowed the generic summary keywords are renamed, with the old names kept as deprecated aliases.

Detailed changes

CheckWindowsApps — a new module for Windows server roles

A new Windows-only module carrying IIS and Remote Desktop Services checks, built on the performance counter sets and enriched from WMI where the role’s provider is installed. The two roles share one module deliberately: every check module statically links Boost and the filter engine, so a role earns its own DLL only when it drags in a heavy or optional dependency (the way CheckMySQL carries libmariadb.dll).

Command Reports
check_iis_app_pools Per-pool state, uptime and recycles. CRITICAL by default when an auto-start pool is not running; a pool that has never started since boot surfaces as unknown rather than hiding.
check_iis_sites Per-site state, connections and uptime, plus requests_per_sec / bytes_per_sec behind averages=true. CRITICAL by default when an auto-start site is stopped.
check_iis_worker_processes Per-w3wp active and served requests, with the <pid>_<pool> instance name split into keywords. An empty set is OK — idle pools spin their workers down.
check_iis_request_queues Per-queue length, rejections and age, defaulting to HTTP.sys’ 1000-request limit (warn > 800, critical > 1000).
check_rds_licenses One record per CAL key pack from Win32_TSLicenseKeyPack: total, issued and available licences. Warns at available < 10 and total > 0, critical at available = 0 and total > 0.
check_rds_sessions Active, inactive and total session counts, all three as perfdata. No default thresholds — the interesting limits are per-farm.
check_rds_session_load One record per session (Console, Services, RDP-Tcp <n>) with CPU, working set and, on session hosts, RDP protocol bytes — the per-user attribution check_process cannot give. sessions-only=true skips the session-0 aggregate.
check_rds_broker The Connection Broker counterset. Counter names vary between Windows Server versions, so the check enumerates whatever the counterset exposes and reports one record per counter instead of hard-coding names.

A host without the role gets a clean UNKNOWN naming the missing role, not a WMI or PDH error dump. The counter plumbing landed as a reusable gather helper that collects a set of English counter names for every instance of an object in one query, keeping the existing localized/English/index resolution fallback — it was verified against live Swedish-localized counters.

CheckNet — status-page checks for the common web servers

Command Endpoint Keywords
check_apache_status mod_status (?auto appended automatically) workers, requests/s, scoreboard
check_nginx_status stub_status active/reading/writing/waiting, cumulative accepts/handled/requests, derived dropped count
check_phpfpm_status FPM status page processes, listen queue, max_children_reached, slow requests; warns by default when requests queue up
check_tomcat_status manager status?XML=true (appended automatically) per-connector thread pool, request/error counters, JVM heap; defaults fire at 75%/90% pool usage

All four share check_http’s connection handling — Basic auth, TLS version / verify / CA, timeout — and go CRITICAL by default when the endpoint is unreachable, answers non-2xx, or serves something that is not the expected status format. Numeric parsing pins the classic locale, so a host with a decimal comma no longer truncates ReqPerSec at the decimal point.

Filter messages — configurable number rendering

check_drivesize reported 140.293GB/0.983TB used: two units and six decimals in one line, with no way to change either (#1428). Every filter check now takes four options, and real-time filters take the same values as settings keys (decimals, byte unit, decimal separator, thousands separator), inheritable from the default template.

Option Effect
decimals Exactly N decimals. Default -1 keeps the historical “up to three, trailing zeros stripped”. Capped at 15.
byte-unit Pin every byte value to one unit, B…EB.
decimal-separator Radix character — , for the European rendering.
thousands-separator Digit grouping for the integer part.
check_drivesize drive=/ show-all=true decimals=2 byte-unit=GB
OK /: 141.09GB/1006.85GB used

check_drivesize drive=/ show-all=true decimals=2 byte-unit=GB \
  decimal-separator=, thousands-separator=.
OK /: 141,09GB/1.006,85GB used

The format lives on the evaluation context, so it reaches the message only: performance data is built from the raw values and keeps its full precision and its . radix, and so does every number the filter grammar parses out of a threshold — warning=used>1.5g means the same thing with a decimal comma in force.

Three defects in the byte formatter came out of this work:

  • format_bytes(used, 'gb') rendered 1.27055e-10, because the unit comparison was case sensitive against an uppercase table. Units are now case insensitive everywhere.
  • A unit that matched nothing fell out of the comparison having divided seven times, rendering value/1024^7. An unknown unit is now reported — Filter processing failed: format_bytes failed: Unknown byte unit: ZB — and the same check applies inside real-time filters.
  • format_bytes(value, '') failed to parse at all; the empty string literal is now accepted.

The filter/where engine — comparisons that mean what they say

  • Text keyword versus bare number is numeric. A string-typed keyword compared against an unquoted number used to order lexically, or — with the operands reversed — fail to evaluate. Both sides now compare as numbers. This covers value/warn/crit/min/max (filter_perf, render_perf), speed (check_network), string_value (check_registry_value) and column() (check_logfile). A value that is not a number never matches; the check logs one warning naming it and stays a certain non-match, not UNKNOWN. Quoted literals keep the lexical comparison, as do like, regexp, in, keyword-specific converters (state = 'running', age > 30m) and the = 'unknown' / = 'never' sentinels.
  • Fractional numbers survive. count > 2.5 used to be rounded into the counter’s integer domain, and unit literals lost their fraction entirely (working_set > 1.5g meant 1g, uptime < 2.5h meant 2h).
  • filter_perf/render_perf/xform_perf: max and min were swapped. max read the perf-data minimum bound and min the maximum; they now read the bounds they name.
  • Template errors are reported. A function that failed inside detail-syntax or top-syntax left the placeholder empty and said nothing; the check now returns UNKNOWN with Filter processing failed: ….
  • perf-config’s unit: converts instead of relabelling. On byte series that do not auto-scale, unit:KB used to change the label only, shipping =1536KB for 1536 bytes. The value and the warn/crit bounds now convert. An unrecognised unit leaves the value alone rather than dividing it by 1024⁷.

Filter keywords — the clash with the generic summary keywords is resolved

A handful of checks registered a keyword named status, count or total — the same names as the built-in summary keywords. The check-specific value won in filter/warning/critical and detail-syntax, while top-syntax and the reference documentation showed the generic one. Each now has a distinct name:

Check Old New
check_cpu, check_cpu_utilization total usage
check_battery status battery_status
check_network status, total link_status, throughput
check_os_updates count updates
check_patch_age count patches
check_pending_reboot count signals
check_printjobs status job_status
check_printqueue status printer_status
check_installed_software (Linux) status package_status
check_activation status activation_status
check_docker status container_status
check_connections count, total connections, total_connections
check_dns count records
check_http status status_message
check_shadowcopy count copies
check_disk_health total size

The old names remain as undocumented deprecated aliases with unchanged behaviour, so check_cpu "warn=total > 80" still works.

Settings — host name placeholders, and where they may land

${host}, ${hostname}, ${hostname_lc}, ${hostname_uc} and ${domain} now resolve in attachment target paths and in [/includes], not only in settings urls and the url an attachment is fetched from (#458). An unknown ${...} token in a path is not an error — it resolves to the installation directory — so a configuration like [/attachments] ${shared-path}/${host}.ini = … never failed, it quietly wrote one file with the installation directory in its name.

🔒 Because the host name is not fully under the operator’s control (DHCP, or any local privileged process can set it), a value substituted into a path is reduced to the characters a legal RFC-952 host name can contain: anything else becomes _, and a dots-only value becomes _. Settings urls and the submit clients’ host name specs are unaffected. See Security notices.

nscp settings --migrate-to (and the REST migrate) now keeps a placeholder you pass it as-is in boot.ini while migrating into the expanded per-host file, the way --switch already did, so the template survives on a fleet-managed machine.

Clients — submission paths that were quietly dead

  • Syslog. SyslogClient read its connection settings from the sender rather than the target, so address, port, facility, severity and templates were all ignored: the agent logged Undefined facility: and sent nothing. Broken since 0.4.3 (2015). CheckMKClient had the same defect on its query path.
  • SMTP. The sender’s host name was read from the wrong place, so it was always empty and the EHLO fell back to localhost. Set ehlo-hostname on the target if your mail server applies HELO/EHLO policy.
  • Short command names. A client command shorter than eight characters — cpu, run — answered Exception processing command line: basic_string::substr … instead of running, in every module built on the shared client machinery (NRPE, NSCA, NRDP, Graphite, …).

CheckSystem — check_pending_reboot says since when

The CBS and Windows Update reboot keys exist only while their reboot is queued, so their last-write time is when the signal appeared. Two new keywords follow check_registry’s naming: written (type_date, plus written_s) and age, duration-typed so warning=pending = 1 and age > 7d reads as seven days. The default message gains (pending since <time>) when the time is known (#1415). The file-rename, computer-rename and domain-join signals carry no timestamp, so both keywords are optional: they render as unknown, compare false against every number and emit no perfdata rather than reporting a misleading value.

Data collection — one bad field no longer sinks the cycle

  • WMI. Win32_Processor.LoadPercentage is occasionally NULL, and row::get_int had no case for it: the type-mismatch exception escaped half-way through the row and the collector threw away the entire cycle’s clock speeds and core counts (#1391). Optional fields can now opt into boost::none, mandatory ones fail with a clear <col> is NULL instead of localized COM text, and check_cpu_frequency renders a missing sample as no load sample rather than a fabricated 0.
  • PDH. MaxQueueItemAge on an idle, freshly started queue returns PDH_CALC_NEGATIVE_DENOMINATOR, which failed a whole gather even when the caller passed ignore_errors — making check_iis_request_queues misreport the object as missing. With ignore_errors the counter is now skipped for that tick (#642, #906); the background collector and every single-counter check still throw, so they hear about an uncomputable counter instead of silently reading a default.

Windows installer and file layout

A round of fixes to the modern (ProgramData) layout introduced in 0.16.2: upgrading an enrolled host resolves every path token instead of failing; ReadLayout gets the install folder before directories resolve; CURRENT_LAYOUT is set through the public property setter; an upgrade of a modern host no longer re-creates nsclient.ini in Program Files; a %ProgramData% that cannot be resolved fails outright instead of half-applying the layout; migrated files get the destination’s ACL rather than the one they came with; resetting a renamed tree to inherited strips the explicit ACEs; and --migrate-layout legacy migrates to legacy instead of silently to modern.

Bug fixes

  • nscp settings --show --path … without a --key used to print nothing and exit 0; it now reports Invalid command line please use --path and --key with show and exits non-zero.
  • The settings diff behind the REST diff endpoint kept listing an edit for the lifetime of the process after it had been written, reporting a modified entry whose old value equalled its new one.
  • Only the pending markers the backend confirms are dropped on save, and staged deletions are masked in has_key.
  • Boolean option defaults render as true/false in the generated reference instead of garbage.

Documentation

Options shared by every filter check (filter, warning, top-syntax, …) and the generic filter keywords are now single-sourced: they fold out of each command’s reference page into one shared page, so a command’s documentation shows only what is specific to it. Runtime-stubbed Windows-only checks are marked as Windows only.

Upgrade notes

  • Syslog starts delivering. If you have a syslog target configured, check it still points where you want before upgrading — it has not been delivering, and it will now. The same applies to SMTP targets, which will start announcing this host in EHLO instead of localhost.
  • Host name placeholders in paths now resolve. Check any ${host}, ${hostname} or ${domain} under [/attachments] or [/includes] and remove workarounds — such a file lands somewhere new after upgrade. 🔒 The value is sanitized when it lands in a local path. Configurations without a host name placeholder are unaffected.
  • Number rendering is opt-in, but it is all-or-nothing per check. Leave all four options unset and messages are byte-for-byte unchanged. Set any of them and plain float keywords move onto the number format too: with decimals unset they render with up to three decimals instead of the legacy 6-significant-digit form (2.71094 → 2.711), and large values stop rendering scientific. A pipeline that matches float text in the message may need its pattern relaxed.
  • An unknown unit in format_bytes() now returns UNKNOWN instead of a quietly wrong number. A syntax string with a typo’d unit will fail until the unit is fixed.
  • perf-config unit: on plain byte series changes the metric’s magnitude. A dashboard that compensated for the old mislabelling will see the metric drop by the unit ratio; a graph flat at a near-zero value because of a misspelled unit: will jump to its real magnitude.
  • Filter comparisons against a bare number are numeric. Review any filter that deliberately relied on text ordering — quote the number to keep the old behaviour.
  • Fractional thresholds change meaning. Whole-number thresholds are unchanged; expressions that already used a decimal point can behave differently.
  • max and min in filter_perf/render_perf/xform_perf were swapped. A filter that compensated needs the two names exchanged back.
  • Renamed filter keywords keep working through deprecated aliases, but three default perfdata keys change because the default perf-config names the renamed keyword: check_cpu_utilization (Linux) cpu_total → cpu_usage, check_patch_age patch_count → patch_patches, check_pending_reboot reboot_count → reboot_signals. Pass your own perf-config=extra(...) with the old name to keep the old key. check_os_updates’ default output now reports the actual number of updates instead of the matched-row count.
  • check_pending_reboot’s default message gains a suffix — Reboot required: Windows Update (pending since 2026-08-16 09:41:12). Notification pipelines matching the exact message text need their pattern relaxed.
  • nscp settings --show without --key now fails. Scripts relying on the silent success need the missing --key added.

Download

You can download the new version from GitHub

// Michael Medin

0.16.4 Security release: request-smuggling fixes in the bundled web server

0.16.4 upgrades the Cesanta Mongoose web server bundled in the Windows builds to 7.23, closing two critical HTTP request-smuggling vulnerabilities in its HTTP parser. If NSClient++’s web server is reachable through a reverse proxy or WAF, upgrade promptly.

Highlights

  • 🔒 Bundled Mongoose upgraded from 7.20 to 7.23. Fixes two critical (CVSS 9.1) HTTP request-smuggling vulnerabilities, CVE-2026-73256 and CVE-2026-73257, fixed upstream in Mongoose 7.22.
  • Windows builds only. The Windows WEBServer module (REST API and web UI) uses the Mongoose backend; the Linux DEB/RPM packages build on Boost.Beast and never contained the vulnerable code.
  • Exploitable behind an intermediary. Both flaws let an unauthenticated attacker smuggle requests past a reverse proxy, WAF or load balancer in front of NSClient++ — bypassing proxy-level ACLs or injecting into other clients’ reused connections. Direct client → NSClient++ deployments have no front end to desynchronize, and NSClient++’s own authentication is still enforced per request either way.

Detailed changes

WEBServer — bundled Mongoose upgraded to 7.23 (security)

Mongoose versions before 7.22 mis-parse HTTP message framing in two ways:

CVE Flaw
CVE-2026-73256 Broken HTTP/1.0 detection in http_cb() — a request combining Transfer-Encoding: chunked with conflicting HTTP/1.0 framing is parsed with different message boundaries than an HTTP/1.0 reverse proxy sees.
CVE-2026-73257 Requests carrying both Content-Length and Transfer-Encoding: chunked are accepted instead of rejected, enabling CL.TE desynchronization against a Content-Length-preferring front end.

All build pipelines now pin Mongoose 7.23 (the latest release, which also carries further upstream TLS and TCP/IP hardening): the Windows CI workflows, the Linux docker scenario images that fall back to the Mongoose backend (minimal, no-openssl), and the developer build instructions. The full advisory record is on the Security notices page.

Upgrade notes

  • 🔒 Upgrade Windows installs, promptly if behind a reverse proxy/WAF: the request-smuggling CVEs only matter when an intermediary in front of NSClient++ frames the HTTP stream differently than the built-in web server. No configuration change is needed — this is a drop-in upgrade.
  • Linux packages are unaffected (Boost.Beast web backend, no Mongoose), as are installs with the WEBServer module disabled.

Download

You can download the new version from GitHub

// Michael Medin

0.16.3 check_nt works with the real nagios-plugins client again

0.16.3 is a small bugfix release: it restores compatibility between the legacy check_nt server (NSClientServer) and the real nagios-plugins check_nt client — broken since 0.12.2 — and pins the fix with an integration suite that drives the genuine client against NSClient++ in CI. It also reorganises the reference documentation for readability.

Highlights

  • check_nt requests without a trailing newline are answered again. Buffer-cap hardening in 0.12.2 made the server wait for a newline terminator, but the real nagios-plugins check_nt sends <password>&<cmd>&<args> with no terminator — so every one of its requests has hung until the client’s socket timeout (No data was received from host!) in every release since (#1421).
  • The fix is pinned by a real-client integration suite. CI now compiles check_nt from the official nagios-plugins 2.5 release and drives it against the server, covering the protocol commands, password enforcement and the allow command gating (#1421).
  • Securing check_nt is now documented. New guidance covers the password, allowed hosts and the allow setting that limits which commands the legacy endpoint will answer.
  • Reference docs reorganised. Queries are listed first and every command carries an OS column with platform logos, so it is clear at a glance what exists on Windows vs Linux.

Detailed changes

check_nt — compatibility with the real nagios-plugins client restored

The buffer-cap hardening that shipped in 0.12.2 made the legacy check_nt server wait for a newline terminator before parsing a request. The real nagios-plugins check_nt sends its request with no terminator and waits for the reply, so every request from it has hung until the client’s own socket timeout in every release since. End-of-read is once again end-of-request, while both halves of the hardening are kept: the 4 KiB request cap, and the newline path (which consumes the terminator and leaves pipelined bytes intact) for line-oriented clients.

The behaviour is now pinned at two levels: unit tests on the request parser (no-terminator format, newline path, empty chunk, oversized-line cap), and an integration suite that compiles check_nt from the official nagios-plugins 2.5 tarball in a container and runs it against nscp test — covering CLIENTVERSION, UPTIME, CPULOAD, MEMUSE, USEDDISKSPACE and PROCSTATE, wrong-password handling, and the allow command gating including its fail-closed behaviour (#1421).

Documentation

  • New guidance on securing the legacy check_nt (NSClientServer) endpoint: set a password, restrict allowed hosts, and use the allow setting to limit which commands it answers.
  • The reference docs put queries first and add an OS column with platform logos to every command.

Packaging

  • Automatic Chocolatey publishing on release is disabled while the package onboarding with chocolatey.org is being sorted out (#1422). The workflow can still be run manually; the MSI, DEB, RPM and ZIP packages are unaffected.

Upgrade notes

  • check_nt clients that never sent a trailing newline get answers again. If you scripted around the hang (client-side timeouts, retries, or switching clients), those workarounds are no longer needed. No configuration change is required; the default install is unaffected unless NSClientServer is enabled.
  • Chocolatey: NSClient++ is not yet available from chocolatey.org; use the MSI from the release page for Windows installs.

Download

You can download the new version from GitHub

// Michael Medin

0.16.2 A locked-down modern Windows layout, new security and system checks, and safer settings and roles

This release introduces an opt-in modern Windows file layout that separates and locks down the agent’s writable state, adds a batch of Windows security and system checks, and hardens two areas of the WEB/settings surface. Default installs are unaffected until you opt in to the new layout.

Highlights

  • Modern, locked-down Windows file layout (opt-in). Configuration, the fleet identity, and writable state can now live in a dedicated %ProgramData% folder that is restricted to SYSTEM and Administrators, instead of sitting under Program Files. Switch with nscp settings --migrate-layout modern (or the MSI LAYOUT property); the classic layout remains the default and is untouched.
  • Writable state and a fleet folder are first-class. A ${fleet-folder} token and dedicated writable-state directories keep fleet/enrollment material and mutable state out of the package/program directories, with local overrides now visible in diagnostics. On Linux packages the writable state directories are created and migrated automatically.
  • New CheckSecurity checks. check_activation (Windows licensing state), check_file_security (file owner / DACL hardening), and check_firewall_rules (assert on individual firewall rules).
  • New CheckSystem checks. check_w32time (Windows Time service health) and check_printjobs (per-job print detail), plus reporting the printer device behind each queue.
  • Sensitive settings values are redacted on read. The REST settings read endpoints and the nscp settings --list / --show CLI now return *** for keys registered sensitive, matching the diff endpoint. Reported by @yagust.
  • The legacy WEB permission is flagged and no longer seeded by default. It unlocks deprecated query-dispatch endpoints that can run any registered command; fresh installs no longer create the role and a SECURITY warning is logged for any role that grants it. Reported by @yagust.
  • A documented upgrade path. New Upgrading and Security notices docs pages collect per-version operator actions and security-relevant changes in one place (#1410).

Detailed changes

Modern Windows file layout

The agent can now run in a “modern” layout where its configuration, fleet identity, and writable state live in a dedicated, ACL-restricted %ProgramData% folder rather than under %ProgramFiles%. The layout is recorded in boot.ini and resolved through a single shared path-token table used by both the service and the bundled clients, so ${shared-path}, ${log-path}, ${fleet-folder} and friends resolve consistently everywhere.

Migration is available both from the CLI (nscp settings --migrate-layout modern, with --dry-run) and from the MSI (via a LAYOUT property). The migration is defensive: it refuses to move into a populated destination on the first switch, locks the destination down before writing any secret into it, moves across volumes rather than failing, and keeps shipped program content out of the redirected shared path.

Change Effect
--migrate-layout modern / legacy Move an existing install between layouts (dry-run supported).
MSI LAYOUT property Install/upgrade directly into a chosen layout.
${fleet-folder} token Addresses the fleet/enrollment folder in the active layout.
Locked-down shared folder Restricted to SYSTEM + Administrators; ownership taken, not just the DACL.

New and updated checks

  • CheckSecurity: check_activation, check_file_security (owner + DACL hardening), and check_firewall_rules (individual rules). Corrected three check_file_security verdict paths and kept expect= assertions visible through a firewall filter.
  • CheckSystem: check_w32time for the Windows Time service, check_printjobs for per-job print detail, and the printer device is now reported behind each queue. Duration keywords keep their -1 sentinel and last_sync_age is treated as a duration.
  • CheckDocker: survives containers removed mid-check and counts image disk correctly.
  • CheckMySQL: plugin-dir / socket / defaults-file are settings-only.

Security & hardening

  • Settings redaction (reported by @yagust): values for keys registered sensitive are returned as *** on the settings read paths (REST GET /api/v2/settings/... and /descriptions, and the --list / --show CLI), matching the diff endpoint. Internal reads a module makes of its own configuration are unaffected. This is defense-in-depth, not an authorization boundary — the plaintext still lives in nsclient.ini. The web admin edit dialog now writes only changed fields so the mask cannot overwrite a stored secret.
  • Legacy WEB permission (reported by @yagust): the legacy grant unlocks the deprecated /query.pb and /query/{name} endpoints, which dispatch through the same command registry as /api/v2/queries. The built-in legacy role is no longer seeded on fresh installs, any role whose grant includes the legacy token now logs a SECURITY warning at startup (and from nscp web add-role / add-user), and the capability is documented in the securing guide.

Installer & packaging fixes

  • Repaired a self-initialised member and a clobbered boot.ini; stopped stamping [layout] into every boot.ini.
  • Remove the fleet identity on uninstall of a modern install; restore the config backup into the layout’s shared folder; honour boot.ini’s [paths] shared-path.
  • Linux packages create and migrate writable-state directories and keep them out of the package directory; adopt_owner handles root-written enrollment material and is symlink-safe.

Documentation

  • New Upgrading page collecting per-release operator actions (closes #1410), and a Security notices page tracking advisories and hardening changes.
  • Documented the Windows and Linux file layouts and the MSI LAYOUT property.

Upgrade notes

  • The modern layout is opt-in; the default install is unaffected. Switch deliberately with nscp settings --migrate-layout modern (try --dry-run first) or the MSI LAYOUT property. Run the CLI migration from an elevated prompt — the destination is locked to SYSTEM/Administrators.
  • 🔒 Sensitive settings values now read back as ***. Tooling that read a secret out of GET /api/v2/settings/... will now receive *** for keys registered sensitive. No configuration change is required.
  • 🔒 The legacy WEB role is no longer seeded on fresh installs and any role granting the legacy permission logs a SECURITY warning. Existing installs keep their role and are unaffected; only grant legacy to trusted legacy systems.
  • Linux writable-state migration is automatic. Packages create and migrate the writable state directories on upgrade; no action required.

Download

You can download the new version from GitHub

// Michael Medin

0.15.0 SQL Server monitoring and seventeen new checks

This release adds a new CheckMSSQL module for monitoring Microsoft SQL Server, a large batch of new Windows checks covering disks, security hygiene and patch state, richer keywords across many existing checks, and fixes a long-standing class of collector stalls caused by slow WMI providers.

✨ Highlights

  • 🗄️ New CheckMSSQL module. Five new commands monitor Microsoft SQL Server over ODBC: connectivity/health, arbitrary T-SQL queries, database state and log usage, backup age and SQL Agent jobs. Windows integrated authentication by default, with optional SQL authentication.
  • 🆕 Twelve more new check commands. Disk writability (check_disk_write), UNC share free space (check_uncpath), Storage Spaces (check_storagepool), VSS snapshots (check_shadowcopy), SMB shares (check_share), Microsoft Defender (check_defender), local account hygiene (check_local_accounts), group membership drift (check_group_members), pending reboot (check_pending_reboot), hotfix age (check_patch_age), print queues (check_printqueue) and paging I/O (check_swap_io).
  • ⚙️ The system collector no longer freezes on slow WMI providers. Slow every-12-second collections (network, temperature, CPU frequency, battery, OS updates) now run on their own thread, so a blocking WMI query no longer stretches check_cpu time windows or drops samples (#1378).
  • 📃 Multi-line check output. The new list-separator option on every filter-based check lets long results render one item per line, which Nagios-compatible frontends show as summary + long output (#1370).
  • 🔥 check_firewall now reports the effective, group-policy-aware state. A firewall enabled or disabled through group policy previously reported its pre-policy local state (#1351).
  • ⏱️ Per-disk I/O latency. check_disk_io and check_disk_health gain read_latency, write_latency and total_latency keywords in milliseconds, on both Windows and Linux (#1369).
  • 🐛 Fixed disable = cpu_frequency silently stalling check_cpu. Disabling CPU frequency collection also disabled CPU load sampling (#1368).
  • 🐧 Linux packages now ship executable scripts. Bundled scripts lost their execute bit when installed by DEB/RPM packages. Thanks to Fabio Fantoni for this and for REUSE/SPDX compliance fixes.

🔍 Detailed changes

🗄️ CheckMSSQL — new module for monitoring Microsoft SQL Server

A new Windows module connecting over ODBC with Windows integrated authentication by default and optional SQL authentication (password stored as a masked settings key). The ODBC driver is auto-detected, preferring the newest “ODBC Driver NN for SQL Server” and falling back to the legacy “SQL Server” driver; on modern drivers TrustServerCertificate=yes is applied by default (overridable via trust-cert/encrypt). Login and query timeouts keep checks from ever hanging the agent, and unreachable servers report UNKNOWN with the full ODBC diagnostic chain.

Command Purpose
check_mssql Connectivity and health: version, patch level, edition, uptime with time-unit thresholds
check_mssql_query Arbitrary T-SQL with returned columns exposed as filter keywords and perfdata
check_mssql_databases Database state, recovery model and sizes, plus log usage from DBCC SQLPERF(LOGSPACE)
check_mssql_backup Age of last full/diff/log backup from msdb; never-backed-up reported as -1 and critical by default
check_mssql_jobs SQL Agent job outcomes, duration and in-flight runs (is_running)

check_mssql_backup excludes COPY_ONLY and snapshot backups by default so an ad-hoc dev backup or a VSS agent cannot mask a failing backup job (include-copy-only / include-snapshot opt back in). A new end-to-end scenario, Monitoring a SQL Server host, combines the module with service, disk, memory, PDH and event log checks and documents a low-privilege monitoring login.

check_mssql_backup "critical=full_age > 26h or full_age = -1" "warn=log_age > 2h"

💾 CheckDisk — writability probes, UNC paths, Storage Spaces, VSS and SMB shares

Command Purpose
check_disk_write Verify a disk is actually writable: exclusive-create a probe file, write, read back, delete. Never touches a file it did not create; probe size capped at 1M
check_uncpath Free space on a UNC path (server share), with optional alternate credentials
check_storagepool Storage Spaces pool health and capacity
check_shadowcopy VSS snapshot recency, count and shadow-storage usage per volume
check_share List SMB shares or verify that specific required shares exist

Existing disk checks were extended as well:

  • check_disk_io and check_disk_health expose average per-I/O latency (read_latency, write_latency, total_latency, unit ms) with perfdata and metrics (#1369). On Windows the values are computed from raw PERF_AVERAGE_TIMER counters (the formatted WMI class truncates realistic latencies to 0); on Linux from /proc/diskstats. Thresholds like "warn=total_latency > 20" "crit=total_latency > 50" work regardless of workload shape.
  • check_drivesize gains require (alias mandatory-drives): the check goes CRITICAL if any listed drive is missing, even when scanning wildcards.
  • check_drivesize and check_disk_health can report physical-disk device state (health and operational status).
  • check_files gains aggregate file-size metrics and a folder count.

🛡️ CheckSecurity — Defender, local accounts and group membership

Command Purpose
check_defender Microsoft Defender status: signature/scan age, real-time and tamper protection, engine/signature versions
check_local_accounts Local account hygiene: enabled/disabled, locked, password-required/expires, built-in admin/guest
check_group_members Local group membership (default Administrators) with alerting on members not on an expected allow-list

🖥️ CheckSystem — patch state, reboot state, print queues and paging I/O

Command Purpose
check_pending_reboot Whether the system is waiting for a reboot, aggregating servicing, Windows Update, file-rename, computer-rename and domain-join signals
check_patch_age Installed-hotfix hygiene: time since the newest hotfix and presence of specific required hotfixes
check_printqueue Print queues: queue depth, oldest-job age, offline and error states per printer
check_swap_io System paging (swap) I/O rates: pages/bytes paged in and out per second

📈 check_process — background CPU sampling, owners and more memory keywords

check_process delta=true previously sampled inside the check, slept one second and sampled again — stalling every query by a second. CPU deltas are now published by an opt-in background collector (process cpu setting, mirroring process history) that diffs the process table once a second; the check overlays a rolling per-PID CPU% onto a normal no-sleep enumeration. With the collector off, delta=true fails fast with UNKNOWN naming the setting instead of reporting misleading values, and memory/handle fields now keep their real absolute values in delta mode.

Other process-check additions: process owner resolution (with user filtering), an rss alias for working set, thread count, working set and page file percentages, peak memory keywords and system-wide thread/memory totals. Also fixed: the time keyword always reported 0 unless delta sampling was on.

➕ More keywords and options for existing checks

Check Addition
check_network Per-interface packet rates, errors and discards (packets_in, packets_out, …) with perfdata and metrics; NIC team membership (team, team_status) and WMI source keywords
check_service summary option emitting aggregate state counts (running_services, stopped_services, paused_services, pending_services, service_count) for dashboard rollups
check_os_version CPU architecture, Windows build revision and inventory-only BIOS fields (serial, version, manufacturer); fixed version detection for Windows 10/11 and Vista/Server 2008
check_os_updates Support for Defender definition updates and update rollups
check_cpu_frequency Socket information and load percentage
check_tasksched Next run time and missed-run tracking, task URI and hidden properties, default perfdata for task state and missed-run counters
check_eventlog User SID retrieval and filtering; more efficient bookmark handling (plus a bookmark bug fix)
check_pdh Built-in memory_pages_sec counter (\Memory\Pages/sec); more robust resolution of localized counter names

⚙️ CheckSystem collector — no more stalls from slow WMI providers (#1378)

The background collector ran network, temperature, CPU frequency, battery and OS update collection on the same 1 Hz thread as CPU/memory/PDH sampling. The network collection queries Win32_PerfRawData_Tcpip_* via WMI with no timeout; when the WMI Performance Adapter service restarts (roughly every 16 minutes on an idle server) that query blocks for 21–24 seconds, freezing the whole collector — stretching check_cpu time windows and dropping samples. The five slow collections now run on their own thread, so a slow provider costs one stale cycle for that metric instead of a frozen collector.

The follow-up hardening fixed a subtle shared-state bug: CheckSystem, CheckEventLog and CheckLogFile all created the same named shutdown event, so stopping or reloading any one of them silently killed the others’ background threads — and the name let any co-resident process signal it and disable monitoring from outside. All three now use unnamed, per-instance events with proper cleanup, and a transient COM initialization failure at boot now retries instead of permanently disabling collection.

📃 Filters and output — multi-line lists and REST-safe booleans

  • list-separator (#1370): every filter-based check now accepts a separator for %(list), %(ok_list), %(warn_list), %(crit_list), %(problem_list) and %(detail_list), with \n, \r, \t and \\ escapes; real-time filters get a matching list separator settings key. The decoded separator is also exposed to templates as %(sep) so the line can break before the first item:
check_users "top-syntax=%(status): %(count) user(s) logged on:%(sep)%(list)" "detail-syntax=%(user) [%(state)]" "list-separator=\n"
OK: 7 user(s) logged on:
administrator [active]
user1 [active]

The default (,) is unchanged and templates pass through byte-for-byte. - Valued booleans on common options. debug, show-all and escape-html rejected the x=true form used by REST (answering with usage text instead of running); they now accept x=true/x=false while the bare CLI form keeps working. - %(problem_list) leak fixed. Real-time filters reuse one filter instance; %(problem_list) kept accumulating items from every previous event batch.

🔥 check_firewall — effective, group-policy-aware state (#1351)

check_firewall read only the local policy store, so a firewall configured through local or AD group policy reported its pre-policy state — a GP-disabled firewall showed as enabled and vice versa. The group-policy resultant values (EnableFirewall, default inbound/outbound actions) are now overlaid on the local store, matching Get-NetFirewallProfile -PolicyStore ActiveStore, including the legacy pre-Vista “Protect all network connections” StandardProfile key. A new policy keyword exposes whether a profile’s settings come from local or group policy.

🐛 Bug fixes

  • disable = cpu_frequency in the CheckSystem collector also disabled CPU load sampling, silently stalling check_cpu (#1368).
  • check_process time keyword always reported 0 without delta sampling.
  • A CheckEventLog bookmark bug could skew incremental event log scanning.

📦 Packaging and licensing

  • Linux DEB/RPM packages now install the bundled scripts with their execute permission, and check_ok.sh gained its missing shebang, so they can be invoked directly as external-script commands (thanks Fabio Fantoni).
  • REUSE/SPDX compliance: third-party attributions for bundled CMake modules and binaries are now correctly declared, and the SBOM no longer misattributes them (thanks Fabio Fantoni).

⚠️ Upgrade notes

  • check_process delta=true behaviour changed: it now requires the new process cpu collector setting to be enabled and returns UNKNOWN (naming the setting) when it is off, instead of sleeping one second inside the check. With the collector on, memory and handle fields report absolute values in delta mode rather than 1-second differences. Default installs (not using delta=true) are unaffected.
  • CheckMSSQL is a new optional module; it is not loaded by default. Enable it and see the new Monitoring a SQL Server host scenario in the docs.
  • All other changes are additive; existing configurations render byte-for-byte as before.

Download

You can download the new version from GitHub

// Michael Medin

0.14.1 Host security posture, JSON-aware HTTP checks, and a clearer licence

This release adds a brand-new CheckSecurity module for monitoring a host’s security posture — certificates, firewall, antivirus, BitLocker, Secure Boot, NLA and logged-on users — and teaches check_http to assert on values inside a JSON response body. It also fixes boolean check arguments over REST, tidies up process aggregation and module activation, relicenses the project under a clear dual licence, and reworks the documentation to handle Windows and Linux side by side.

Highlights

  • New CheckSecurity module. Seven new checks for host security posture: check_certificate, check_firewall, check_antivirus, check_bitlocker, check_secureboot, check_nla and check_users. check_certificate and check_users run everywhere; the rest are Windows-only. (#1339)
  • check_http can assert on JSON responses. New json-path=alias:path options extract values from a JSON body into filter keywords you can threshold on and emit as perfdata. (#1341)
  • CheckNet queries now emit performance data by default, so check_http, check_tcp and friends graph out of the box without an explicit perf syntax. (#1341)
  • Boolean check arguments accept values, not just flags. check_ping host=www.google.com total=true now works alongside the bare-flag form — the form REST already used. (#1338)
  • Activate several modules in one command: nscp settings --active-module CheckSystem CheckNet. (#1329-follow-up)
  • Clear dual licence. NSClient++ is now Apache-2.0 OR GPL-2.0-only, with machine-readable REUSE metadata and third-party notices. (#1343)
  • Reworked, multi-OS documentation that presents Windows and Linux options and features together instead of assuming one platform. (#1342)

Detailed changes

CheckSecurity — new host security-posture module

A new module, CheckSecurity (alias security), checks whether a host is in the security state you expect. Each check is a normal modern_filter check, so you can override the default warn/crit expressions, filter, and detail-syntax/top-syntax as usual.

Command Platforms What it checks
check_certificate All X.509 certificate expiry / validity / hygiene from files or the Windows store
check_users Windows + Linux Count and detail of logged-on / RDP sessions
check_firewall Windows only Firewall profile (Domain/Private/Public) enabled and active state
check_antivirus Windows only Registered antivirus products’ enabled / up-to-date state (Security Center)
check_bitlocker Windows only BitLocker drive-encryption protection status per volume
check_secureboot Windows only Whether UEFI Secure Boot is enabled (distinguishes “disabled” from “legacy”)
check_nla Windows only Network Location Awareness category (public/private/domain) per network

check_certificate defaults to warning when a certificate expires within 30 days and critical within 10 (matching common practice), emits expires_in (whole days until expiry) as perfdata, and can scan a whole directory:

check_certificate file=/etc/ssl/certs/mysite.pem
check_certificate file=/etc/ssl/certs recursive=true "detail-syntax=${subject}: ${expires_in}d"
check_certificate file=/etc/pki/tls/certs critical=expired=1

The Windows checks expose the raw state fields so you can tighten or relax the default posture. check_firewall adds an active flag (which profile is currently in effect) alongside enabled, so you can warn when a machine silently falls back to the Public profile after a network change:

check_firewall "warn=active = 1 and profile = 'Public'" "detail-syntax=${profile} profile is active"
check_secureboot "warn=supported = 0" "crit=supported = 1 and enabled = 0"
check_nla "crit=connected = 1 and category != 'domain'" "detail-syntax=${network}=${category}"

On a platform where a Windows-only check does not apply, the check returns UNKNOWN with a clear message rather than failing.

CheckNet — check_http JSON path extraction

check_http can now pull values out of a JSON response body and treat them as filter keywords. Each json-path=alias:path option extracts the value at a dotted path (numeric segments index into arrays; single-quote a segment that itself contains a dot) and makes it available for warning=/critical= expressions and perfdata:

check_http url=https://api.example.com/health "json-path=qlen:data.queue.length" "crit=qlen > 100"
check_http url=https://api.example.com/health "json-path=st:status" "crit=st != 'ok'"
check_http url=https://api.example.com/health "json-path=err:metrics.error_rate" "warn=err > 0.01" "crit=err > 0.05"
check_http url=https://api.example.com/health "json-path=first:items.0.name" "json-path=cfg:'a.b'.c"

Numeric values keep full precision, strings compare and render as strings, and booleans read as 1/0. A missing path — or a body that is not valid JSON — leaves the alias empty rather than failing the check, and multiple json-path options can be combined freely.

CheckNet — default performance data

CheckNet queries now attach sensible performance data by default, so check_http, check_tcp and the other network checks produce graphable perfdata without a hand-written perf syntax. check_ntp_offset threshold handling was also tidied up in the same change.

Check arguments — boolean options accept values

Boolean check options now accept an explicit value in addition to the bare-flag form:

check_ping host=www.google.com total=true

Previously the value form was rejected from the CLI even though REST always passes flags as key=true tokens, so a boolean option that worked over REST could look broken from the command line. Both forms now behave identically.

CheckSystem — process total aggregation

check_process process-total aggregation now correctly reports the started and hung states (on both Windows and Linux), so totals of these statuses match what the per-process detail shows.

Settings — activate multiple modules at once

nscp settings --active-module now accepts several module names in one invocation:

nscp settings --active-module CheckSystem CheckNet

Licensing — dual-licensed Apache-2.0 OR GPL-2.0-only

NSClient++ is now explicitly dual-licensed under Apache-2.0 OR GPL-2.0-only. Source headers were updated to SPDX identifiers, the project carries machine-readable REUSE metadata (REUSE.toml, LICENSES/), and a THIRD-PARTY-NOTICES.md / NOTICE collect the third-party licences. The installer, packaging and docs licence text were updated to match.

Build — Python library discovery

CMake now derives the default Python library name instead of hardcoding a version, and defaults it to the soname so the module loads without the Python development packages installed. This makes Linux builds far less sensitive to the exact Python version on the build and target hosts. (#1334)

Documentation

  • Multi-OS reference docs. The reference documentation was reworked to present Windows and Linux options and features together, handling checks whose options diverge by platform instead of documenting a single OS. Windows docs were regenerated. (#1342)
  • New docs/samples/ usage examples and descriptions for every new CheckSecurity command and the check_http JSON feature.
  • check_process docs cross-reference filter_perf for top-N processes. (#1330)
  • README restructured and dead files removed. (#1340)

Quality and CI

  • Spelling. A codespell GitHub workflow was added and spelling errors in log messages and settings descriptions were fixed. (#1314, #1344)
  • Live integration tests. A new opt-in test suite runs checks against a real VM in Azure, alongside the existing REST-driven integration tests. New integration tests cover CheckSecurity, the check_http JSON feature, and --active-module. (#1335)
  • Assorted build fixes for older Windows toolchains, Linux, and sanitizer runs.

Upgrade notes

  • Licence change: NSClient++ is now distributed as Apache-2.0 OR GPL-2.0-only. This is a clarification/relicensing — review it if your organisation tracks the exact licence of bundled software. No code or runtime behaviour changes as a result.
  • CheckNet perfdata is now on by default. Network checks emit performance data without an explicit perf syntax. If you were adding perfdata manually, double-check you are not now emitting it twice; graphs that previously showed nothing will start populating.
  • Boolean check arguments: option=true / option=false now work from the CLI as well as over REST. Existing bare-flag usage is unchanged.
  • CheckSecurity is not loaded by default. Enable it before using the new checks, e.g. nscp settings --active-module CheckSecurity. Windows-only checks return UNKNOWN on other platforms rather than erroring.

Download

You can download the new version from GitHub

// Michael Medin

0.14.0 Linux parity

This release brings Linux up to near-parity with Windows and completes the Linux story that began in 0.13.0. On the checks side it adds a full suite of Linux-native system checks (CheckSystemUnix) sourced directly from /proc and /sys, event-driven real-time monitoring on Linux, Linux disk / file / mount support in CheckDisk, and a round of cross-platform CheckNet improvements — TLS for check_tcp, a fuller check_http, multi-record-type DNS, and two new network checks. Around the daemon it delivers , a secure-by-default web server, one-command installs via winget / Chocolatey / Scoop and nscp web install-ui, and a broad set of security and reliability fixes. It also hardens plugin shutdown so a misbehaving module can no longer crash the service on exit.

🌟 Highlights

  • Linux system checks (CheckSystemUnix). New native checks — check_load, check_cpu_utilization, check_kernel_stats, check_swap_io, check_cpu_frequency, check_temperature, check_battery, check_network — plus overhauled check_process (with process history / delta CPU) and a systemd-aware check_service. All read /proc and /sys directly, with thresholds and syntax that match their Windows counterparts.
  • Real-time monitoring on Linux. CheckSystemUnix gains an event-driven real-time thread, so CPU, memory and process alerts can fire the moment a threshold is crossed rather than only on poll — the same real-time model previously available only on Windows.
  • Disk, file and mount checks on Linux (CheckDisk). CheckDisk is no longer Windows-only: free-space (check_drivesize), file (check_files) and disk-I/O checks now run on Linux, with per-device I/O sampling from /proc/diskstats, LVM / device-mapper mapping, inode statistics, file-integrity checksums, and a new check_mount.
  • TLS-aware network checks (CheckNet). check_tcp now speaks TLS (ssl=true) with new SPOP / SIMAP / SSMTP presets; check_http gains redirect policy, certificate-expiry reporting, Basic auth, SNI and non-GET methods; check_dns queries any record type against a custom resolver; and two new checks arrive — check_ssh and check_nsclient_web_online.
  • First-class Linux packaging. The build follows the FHS / CMAKE_INSTALL_PREFIX, with official .deb/.rpm targeting /usr, and a Boost.Beast web backend by default.
  • One-command installs everywhere. Windows via winget / Chocolatey / Scoop; the Linux web UI via nscp web install-ui.
  • Secure by default. The web server refuses to serve cleartext HTTP without an explicit opt-in, plus check_nt command allow-listing and stricter external-script argument checks.
  • Run Lua scripts straight from the CLI with nscp lua execute, backed by Lua thread-safety hardening.
  • Safer plugin shutdown. The plugin manager isolates broken plugins and tears modules down cleanly, so a module that fails to unload can no longer take the service down on shutdown.

📖 Detailed changes

🐧 CheckSystemUnix — native Linux system checks

A new family of checks reads Linux kernel state directly. Thresholds and detail-syntax keywords mirror the Windows checks so alerts port across platforms.

Command Source What it reports
check_load /proc/loadavg 1/5/15-minute run-queue averages; load shortcut; percpu=true scaling
check_cpu_utilization /proc/stat (~1s sample) Per-mode breakdown — user, system, iowait, steal, idle, total
check_kernel_stats /proc/stat, /proc/loadavg Context-switch rate, fork/process-creation rate, live thread count
check_swap_io /proc/vmstat Swap paging rates (swap_in/swap_out pages/s and bytes/s)
check_cpu_frequency /sys cpufreq Current / max / min CPU frequency
check_temperature thermal zones + hwmon Thermal-zone and hwmon sensor temperatures
check_battery /sys power_supply Charge level, power source, health
check_network /proc/net/dev + sysfs Per-interface link status and throughput
check_load "warn=load5 > 4" "crit=load5 > 8"
check_cpu_utilization "warn=iowait > 20" "crit=iowait > 50"
check_swap_io "warn=swap_out > 100" "crit=swap_out > 1000"
check_kernel_stats "warn=current > 8000" "crit=current > 10000"

⚙️ CheckSystemUnix — check_process history and check_service on systemd

  • check_process now tracks process history and computes delta CPU between samples (rather than lifetime CPU), and exposes memory keywords (rss, vms), matching the Windows process semantics.
  • check_service now inspects systemd units. The raw systemd state is mapped to a normalised state keyword so thresholds read the same as on Windows, while the raw fields (active, sub_state, preset) are exposed too. The default critical expression is ( state not in ('running', 'oneshot', 'static') or active = 'failed' ) and preset != 'disabled' — so a stopped-but-disabled unit stays OK while an enabled unit that failed is CRITICAL. Per-unit process metrics (rss, vms, cpu, tasks, age) are parsed from /proc for the unit’s main process.
check_service service=cron "detail-syntax=${name}=${state} active=${active} preset=${preset}"
check_service service=mysql "warn=rss > 1G" "crit=rss > 2G"
  • check_os_version now parses /etc/os-release and reports the distribution and kernel details.

⚡ CheckSystemUnix — real-time monitoring

CheckSystemUnix gains a real-time collection thread and real-time data model, bringing event-driven checks to Linux. CPU, memory and process real-time filters evaluate continuously and emit the moment a threshold is crossed, matching the Windows real-time behaviour. See the Real-Time System Monitoring scenario, now cross-platform.

💾 CheckDisk — now on Linux: disk metrics, inodes, checksums, and check_mount

CheckDisk is no longer Windows-only. Linux builds gain the core free-space and file checks (check_drivesize, check_files) plus disk-I/O sampling, and this release adds:

  • Linux disk I/O. check_disk_io and check_disk_health now sample per-device I/O from /proc/diskstats once per second on Linux (mirroring the Windows PDH path). LVM / device-mapper and RAID volumes are mapped back to their backing devices via sysfs, so space and I/O join correctly for /dev/mapper/… filesystems. The first query after startup can return UNKNOWN while the collector takes its first sample.
  • Inode statistics. check_drivesize exposes inodes_total, inodes_free, inodes_used, inodes_free_pct and inodes_used_pct, so you can catch inode exhaustion (free bytes but no free inodes).
  • File-integrity checksums. check_files exposes md5_checksum, sha1_checksum, sha256_checksum, sha384_checksum and sha512_checksum, computed lazily only when referenced.
  • check_mount (new). Verifies a filesystem is mounted — and optionally that it is mounted with the expected type and options — reading the live mount table (/proc/self/mounts). A path that is not mounted is CRITICAL; a fstype or missing-options mismatch is WARNING.
check_drivesize drive=/ "warn=used>80%" "crit=used>90%"
check_drivesize drive=/ "warn=inodes_used_pct > 85" "crit=inodes_used_pct > 95"
check_files path=/var/log pattern=*.log "crit=size>100M"
check_mount mount=/data fstype=ext4

(Some Windows-only legacy CheckDisk commands are not registered on Linux.)

🔐 CheckNet — TLS for check_tcp

check_tcp can now establish a TLS session over the connected socket (ssl=true), with tls-version (default tlsv1.2+), verify (default none) and ca options, and a response regex to match the server’s greeting. Three new TLS service presets ship alongside the existing plaintext ones:

Preset Port TLS Expected greeting
SPOP 995 yes ^\+OK
SIMAP 993 yes ^\* OK
SSMTP 465 yes ^220
check_tcp host=pop.example.com ssl=true "response=^\+OK"
check_tcp host=imap.example.com SIMAP

Peers that close the TLS session without a close_notify (reported by OpenSSL as stream_truncated) are now treated as a clean end-of-data rather than a read failure.

🌐 CheckNet — check_http features

check_http gains the features needed for real service checks:

  • Redirect policy — onredirect=ok|follow (default ok) with max-redirs (default 15); follows 301/302/303/307/308.
  • Certificate expiry — reports ssl_expiry_days (days until the served certificate expires) for HTTPS targets.
  • Authentication — username / password send an HTTP Basic Authorization header.
  • Methods and bodies — method= (HEAD/POST/…), post-data, content-type; supplying post-data with a GET promotes the request to POST.
  • SNI — sni= overrides the TLS server name / verification host.
check_http url=https://example.com method=HEAD
check_http url=https://example.com username=user password=secret
check_http url=http://example.com/old onredirect=follow
check_http url=https://example.com "warn=ssl_expiry_days < 30" "crit=ssl_expiry_days < 7"

🔎 CheckNet — check_dns record types and custom server

check_dns now queries any record type (type=A|AAAA|MX|TXT|NS|CNAME|SOA|PTR|SRV) and can direct the query at a specific resolver (server=), rather than only resolving an A record against the system resolver.

check_dns host=example.com type=MX server=8.8.8.8

🔑 CheckNet — check_ssh (new)

Connects to an SSH port and validates the protocol banner (implemented on top of the check_tcp service-preset machinery). Flags a server that fails to present a valid SSH-2.0 / SSH-1.x identification string.

check_ssh host=server.example.com
check_ssh host=server.example.com port=2222

📡 CheckNet — check_nsclient_web_online (new)

Verifies that a remote NSClient++ agent’s REST/WEB endpoint is reachable and that credentials authenticate — a lightweight liveness probe for the agent’s management interface. Reports the base URL and distinguishes “unreachable” from “authentication failed (HTTP 401/403)”.

check_nsclient_web_online url=https://agent:8443 password=... verify=none

This command is deliberately named _online because it only checks reachability. A future check_nsclient_web will run actual remote checks through the endpoint.

🌙 Lua — run scripts straight from the command line

nscp lua execute runs a Lua script directly from the CLI — useful for developing and debugging check scripts without wiring them into the configuration first:

nscp lua execute --script myscript.lua

Lua also got thread-safety hardening (a proper GIL), new helpers for targeted and forwarded queries, clearer errors when a script fails to load, and log lines that report the actual script line number.

🔏 TLS — outbound SNI and Op5 client options

  • SNI is now sent on outbound TLS connections (Graphite and the generic TLS client), so a TLS proxy hosting several certificates returns the right one.
  • The Op5 client gained explicit TLS settings:
[/settings/op5/client/targets/default]
tls version = 1.2+
verify mode = peer
ca = ${ca-path}

🔒 Security — secure-by-default web server and hardening

  • The web server refuses to run unencrypted by default. To stop NSClient++ from silently serving the REST API / web UI over plain HTTP, the WEB server now refuses to start without a certificate unless you explicitly opt in with allow insecure = true (see Upgrade notes).
  • check_nt can now be restricted to specific commands. The legacy check_nt protocol is password-only (and source-IP filtering is spoofable), so you can now limit which of its ten request codes are answered. The default is any (unchanged behaviour):
[/settings/NSClient/server]
# Answer only harmless system metrics; deny arbitrary counter/file reads
# and service/process enumeration:
allow = metrics, info

A request outside the list is rejected with ERROR: Command not allowed. - Stricter shell-metacharacter checks in external scripts. User-supplied argument values containing more shell metacharacters are now rejected. - Graphite metric paths are sanitized before being written to the line protocol, preventing injection of extra metrics. - Python sys.path handling hardened to prevent code-injection via path manipulation.

🪟 Windows — winget / Chocolatey / Scoop packages

NSClient++ is now published to the common Windows package managers:

winget install Mickem.NSClient
choco install nsclient   # Still pending approval
scoop install nsclient   # still pending approval

📦 Linux packaging — FHS layout and install prefix

The Linux build honours CMAKE_INSTALL_PREFIX like a normal CMake project, and the official .deb/.rpm are built for /usr. The file layout is now:

What Location
Daemon /usr/sbin/nscp
Modules /usr/lib/nsclient/modules
Private libs /usr/lib/nsclient
Config /etc/nsclient
State / logs /var/lib/nsclient · /var/log/nsclient

If you previously patched hardcoded paths to build for a custom location, that is no longer needed — pass -DCMAKE_INSTALL_PREFIX=/opt/nsclient (or the standard CMAKE_INSTALL_*DIR knobs) instead. To point an already-installed daemon at a boot.ini in a non-standard place there is a new override:

nscp service --run --path-override boot-conf=/etc/nsclient/boot.ini

🖥️ Linux — web UI is a separate download (.deb / .rpm)

The Linux packages no longer bundle the React/Vite web frontend (Debian/Fedora policy forbids npm install during package builds). The daemon, REST API, NRPE/NSCA listeners and every check module are still in the package — only the browser UI ships separately. After installing the package, fetch the matching UI bundle as root:

sudo nscp web install-ui      # downloads + verifies NSCP-Web-<version>.zip
sudo nscp web ui-status       # show installed version / source
sudo nscp web uninstall-ui    # remove only what install-ui put down

Until you do, the web port shows a small built-in placeholder page; the REST API and all listeners work normally without it. The Windows MSI still bundles the UI inline.

🧩 Core — filter summary-variable rendering

All check filters now prefer summary variables during summary rendering. Previously a keyword that exists both per-item and as a summary aggregate (notably status) could render the last item’s value in the summary line, making the overall status read incorrectly. Summary context now resolves to the summary value, so top-syntax reports the aggregate correctly.

🛡️ Service — safer plugin shutdown

The plugin manager now handles broken plugins defensively and prevents a module that misbehaves during teardown from crashing the service on shutdown. Modules get a clean teardown path so listeners and background threads stop before unload.

📈 collectd client — encoding and protocol fixes

Correct (little-endian) gauge encoding, working IPv6 multicast, a configurable send interval (default 10s), and previously dropped metric types (counter / derive / absolute) are now mapped instead of discarded.

🐛 Bug fixes

  • Fixed a Windows build break introduced during the Linux work.
  • Fixed regressions where some metrics and real-time checks stopped reporting.
  • Corrected check_ntp_offset threshold handling and improved default accuracy.
  • Improved check_connections performance-data accuracy for total connections.
  • Yet another possible fix for installer deleting config on upgrade.
  • Fix unreliable per-process CPU% from check_process delta=true. The delta calculation for per-process CPU usage produced inconsistent readings; it now returns stable, accurate values.
  • Fix perf-config=none reporting “Failed to parse syntax”. Setting perf-config=none to suppress performance-data formatting no longer fails parsing.
  • Fixed http(s) headers should be case-insensitive.
  • IPv6: listeners set IPV6_V6ONLY on Linux to avoid port conflicts with IPv4, and IPv6 address resolution was improved.
  • Thread-safety: logger subscriber management, the scheduler, and timer callbacks were made properly thread-safe; CommandClient now shuts down gracefully on POSIX signals.
  • check_mk server: fixed a memory leak.

🚚 Packaging & distribution notes

  • The bundled check_nsclient Nagios plugin moved to its own repository (mickem/check_nsclient) and is pulled in at build time. This only matters if you build from source.
  • Package/file names were normalised — double-check the exact asset name on the releases page if you script downloads.
  • Reduced Linux build dependencies: the build now uses libzip (instead of vendored Miniz), can use the system Google Test, and degrades cleanly when an optional dependency is missing. Linux uses the Boost.Beast web backend by default.

📚 Documentation and tests

  • New Linux Server Health scenario, plus updated cross-platform Network Checks, Disk Space Alerting, Service & Process Monitoring and Real-Time System Monitoring scenarios.
  • New docs/samples/ usage examples and clarifying descriptions for every new command.
  • Extensive new unit tests (CheckSystemUnix, CheckDisk unix, CheckNet) and REST-driven integration tests under tests/ covering the new system, disk and network checks.

⚠️ Upgrade notes

A few defaults were tightened for security and the Linux packaging layout changed. None of these affect a normal Windows MSI upgrade, but Linux users and anyone running the web server in cleartext should read this section.

  • The web server refuses to run unencrypted by default. If you intentionally run the web server in cleartext (e.g. behind a TLS-terminating proxy, or on an isolated network), set allow insecure = true:
[/settings/WEB/server]
allow insecure = true

Otherwise, provide a certificate (certificate = …). If you do nothing and the server has no certificate, it logs an error and does not start the listener. - The web UI is a separate download on Linux (.deb / .rpm). After installing the package, run sudo nscp web install-ui to fetch the matching UI bundle. Until then the web port serves a built-in placeholder; the REST API and all listeners work normally without it. The Windows MSI still bundles the UI inline. - Linux install layout now follows the FHS / install prefix. The official .deb/.rpm install to /usr (daemon /usr/sbin/nscp, config /etc/nsclient, state/logs under /var). If you patched hardcoded paths to build for a custom location, pass -DCMAKE_INSTALL_PREFIX (and the standard CMAKE_INSTALL_*DIR knobs) instead. - Linux check_service now targets systemd. If you previously scripted around the old behaviour, note the normalised state keyword and the default expression that keeps disabled units OK. Match units by unit name (e.g. service=ssh), and use state, active, sub_state and preset in thresholds. - Linux check_process reports delta CPU. CPU is now the usage between samples rather than lifetime CPU. Review any CPU thresholds that assumed the old semantics. - Linux disk I/O needs one collector sample. The first check_disk_io / check_disk_health query immediately after startup may return UNKNOWN (“collector still initializing”); this is invisible with a running service and normal in one-shot testing. - check_tcp / check_http boolean ssl. Enable TLS with ssl=true. When verifying certificates, set verify=peer and provide a ca= bundle; the default remains verify=none.

Download

You can download the new version from GitHub

// Michael Medin

0.12.6 New permission system

The release has three big stories — a new core permission system with optional client-cert principals on NRPE, a PDH overhaul that fixes long-standing counter-collection crashes and adds counter functions, and a WEB hardening option that lets monitoring-only deployments expose the WEB UI without seeding a privileged admin account. Everything else is bug fixes, small features, and follow-ups around those three threads.


Highlights

  • Core permission system — opt-in policy layer that gates which caller can run which command. Configured under /settings/permissions. Disabled by default; existing installs keep working. See https://nsclient.org/docs/concepts/permissions/ for the model, identity table, and rollout recipe.
  • NRPE client identity from cert CN — when client identity source = cn is set on NRPEServer and the listener verifies the client cert, the CN is stamped as the policy principal so rules can be written per-cert ( NRPEServer:icinga-master = ...). Hard guardrail at module start refuses to load the module if the TLS verify mode would let the CN be attacker-supplied.
  • Global allow exec toggle — exec is now gated by a single on/off switch under /settings/permissions. The per-command rule table applies to queries only. Default true so enabling the policy system does not break exec callers.
  • PDH (performance counter) overhaul — fixes for service crashes when PDH misbehaves (#592, #547), counter retry when temporarily unavailable (#634), reliable English counter lookup (#652, #906), a resource leak in the counter-lookup path, and a refactor to smart-buffer-based PDH enumeration. Most users running CheckSystem on Windows should see meaningfully better reliability.
  • check_pdh counter scaling and functions (#281) — details-syntax and related rendering paths can now apply scaling and other functions, e.g. '${counter}'=${value:scale(/1024)}MB.
  • check_network — human-readable strings, scaling, speed, and percentages (#329); team-network statistics (#625). See https://nsclient.org/docs/reference/check/CheckNet.
  • Nagios range syntax in performance data (#748) — 1:10, ~:5, @10:20 etc. work in perfdata thresholds, matching the Nagios plugin spec.
  • disable admin user on WEBServer — monitoring-only deployments can expose the WEB UI without ever seeding the built-in admin (and previously seeded admin entries are ignored). Pairs naturally with the new permission system to lock down reconfiguration surfaces.
  • Path overrides moved to boot.ini + new --path-override CLI flag — path tokens (module-path, certificate-path, etc.) are now declared early in boot.ini so they take effect before the main config is loaded. Per-invocation overrides via --path-override KEY=VALUE. See https://nsclient.org/docs/concepts/settings.
  • NRPE startup is no longer fatal on listener failure — bad bind address / port already in use logs a clear error and leaves the module loaded so settings and commands stay usable for diagnostics.
  • Dual-stack listening fixed (#312) — v4 and v6 acceptors no longer trample each other’s pending connection slot.
  • disable admin user, client identity source, allow exec, and the policy table are all documented in https://nsclient.org/docs/concepts/permissions/ and https://nsclient.org/docs/setup/securing. Treat those two as the starting point for any new install.

Detailed changes

Security and permissions

Core permission system A policy layer in the core decides whether a given caller may run a given command. Disabled by default; when enabled, rules form a strict allow-list.

[/settings/permissions]
enabled = true
log denials = true
log allows = false      ; noisy, only flip on while rolling out
allow exec = true       ; queries-only rule table; exec is a global toggle

[/settings/permissions/policies]
NRPEServer = CheckHelpers.*, CheckSystem.check_cpu
WEBServer:admin   = *
WEBServer:viewer  = CheckSystem.check_cpu, CheckSystem.check_drivesize
Scheduler = CheckHelpers.*, CheckSystem.*

Subject is module[:principal]; object is module.command. Wildcards (*, ?) supported. Rules combine additively. See https://nsclient.org/docs/concepts/permissions/ for the full identity model, the CheckHelpers identity-forwarding behaviour, and a step-by-step rollout recipe.

NRPE client cert CN as principal When two-way TLS is configured and verifying client certs against your CA, the Common Name is stamped as the policy principal:

[/settings/NRPE/server]
client identity source = cn        ; default: none
verify mode = peer-cert
ca = /etc/nsclient/ca.pem
[/settings/permissions/policies]
NRPEServer:icinga-master   = CheckHelpers.*, CheckSystem.*
NRPEServer:metrics-shipper = CheckSystem.check_cpu, CheckSystem.check_drivesize

Guardrails: the module refuses to start if client identity source = cn is configured without SSL, without verify_mode containing peer and fail-if-no-peer-cert (or the peer-cert alias), or without a non-empty ca path. The CN is logged at debug level on every accepted handshake for diagnostics. CN-only (not full DN) because INI key syntax uses = as the key/value separator and would corrupt DN-shaped policy keys; see the “Why CN-only” section of the permissions doc. See https://nsclient.org/docs/reference/client/NRPEServer.

Global allow exec toggle Per-command rules apply to queries only. The exec surface (WEB scripts UI, lua/python core:simple_exec(...), CLI exec) is gated by a single boolean:

[/settings/permissions]
allow exec = false   ; hard lockdown; default is true

When false and enabled = true, every exec call returns Permission denied: exec is globally disabled (/settings/permissions/allow exec = false). See “Why exec is a single toggle” in https://nsclient.org/docs/concepts/permissions/.

disable admin user on WEBServer For installations that expose the WEB UI for status/visualisation only and never want a remote-reconfiguration surface:

[/settings/WEB/server]
disable admin user = true

With this set, the built-in admin is not seeded on first boot, and any existing admin entry in the user settings is ignored at load time.

Security guide updates https://nsclient.org/docs/setup/securing was rewritten with concrete configurations for NRPE (with and without mTLS) and the WEB server. Read it before exposing either to a network you don’t fully control.


Performance counters / PDH

The PDH subsystem (the Windows performance-counter collection backbone behind CheckSystem, check_cpu, check_pdh, check_network, etc.) got a substantial reliability pass. Most users running NSClient++ as a long-running service on Windows should see fewer crashes and more consistent results.

  • Service crashes when PDH misbehaves on a particular machine (#592, #547) — root-caused and fixed. Misbehaving counter registrations no longer take the service down.
  • Counter not retried if unavailable (#634) — counters that fail to bind at first sight now get retried on subsequent collection cycles, instead of being permanently unhealthy for the lifetime of the process.
  • English counter lookup improved (#652, #906) — addresses reading of localised counters by their canonical English names on non- English Windows installs.
  • Resource leak in PDH counter lookup fixed.
  • PDH enumeration refactored to smart buffers — clearer memory ownership across the enumeration path, fewer footguns for future changes.
  • check_pdh counter scaling and functions (#281) — all the details-syntax / rendering paths can now apply functions. Examples:
    check_pdh "counter=\Processor(_Total)\% Processor Time" \
              "details-syntax=${counter} = ${value:round(2)}%"
    
    See https://nsclient.org/docs/reference/check/CheckSystem for the function reference.

check_network

  • Human-readable strings, scaling, speed, and percentages (#329) — perfdata and message output now render numbers in a way operators actually want to read:
    check_network 'filter=interface=Ethernet' \
                  'top-syntax=${list}' \
                  'detail-syntax=${interface}: ${total_rx_human}/s in, ${total_tx_human}/s out'
    
  • Team network statistics (#625) — aggregate stats across Windows NIC teams.

See https://nsclient.org/docs/check/CheckNet.


Performance data formatting

  • Nagios range syntax in performance data (#748) — the perfdata threshold fields now accept the standard Nagios range syntax: 5:10, ~:5, @10:20, etc. Brings NSClient++ into line with what Nagios consumers already expect.

Settings, paths, and CLI

  • Path overrides moved to boot.ini — path tokens (module-path, certificate-path, data-path, log-path, …) now live under [paths] in boot.ini (next to nscp.exe), not in nsclient.ini. Overrides take effect before the main config is loaded — including the bootstrap step that decides where the main config itself lives.
    ; boot.ini
    [paths]
    module-path = D:\monitoring\modules
    certificate-path = D:\monitoring\certs
    
  • --path-override CLI flag — per-invocation override, repeatable. (Renamed from --path to avoid colliding with the nscp settings --path subcommand option.)
    nscp client --path-override module-path=/build/modules --path-override log-path=. ...
    
  • See https://nsclient.org/docs/concepts/settings for the precedence rules and the migration note for installs that had a [/paths] section in nsclient.ini.

Aliases and command registration

  • CheckHelpers alias — aliases can now be defined under [/settings/check helpers/alias] and are registered by CheckHelpers directly, without requiring CheckExternalScripts to be loaded. This is the preferred place going forward; the legacy [/settings/external scripts/alias] is still honoured for backward compatibility.
  • API to list registered query aliases (#506) — programmatic introspection of the alias table, useful for tooling.
  • simple_command / simple_command_map — internal refactor that streamlines how modules register aliases. No user-visible behaviour change, but module authors may want to look at the new pattern.
  • Icinga client alias (7c49a3d3) — minor module-specific addition.

NRPEServer

  • Listener failure no longer kills the module — a bad bind to address that the resolver can’t look up, or a port already in use, used to make the whole module fail to load. Now the failure is logged clearly, the listener stays down, and the module’s settings and commands remain accessible for diagnostics and reconfiguration. Fix the config and reload — no service restart needed.
  • Dual-stack fixed (#312) — the v4 and v6 acceptors used to share a single pending-connection slot, which caused intermittent Already open errors on v6 once v4 accepted a client. Each family now owns its own slot.
  • Insecure mode produces an error-level log line — flipping insecure = true (for legacy check_nrpe interop) now surfaces as an ERROR so it shows up in monitoring dashboards, instead of silently disabling cert-based peer auth.

Plugin lifecycle

  • prepare_shutdown hook — modules can opt in to a first-phase shutdown pass before any plugin is unloaded. Used by the Scheduler and similar long-running submitters to finish in-flight work cleanly. Operators see fewer “submission failed during shutdown” lines during service stop.

Settings store

  • simpleini buffer NUL-termination fix — fixes a buffer allocation issue in the INI parser that could affect non-UTF-8 data paths.
  • cache allowed host is now a real boolean — previously parsed as a string with surprising truthiness; matches what the docs always claimed.

Modules and clean-ups

  • WMI module refactor — target handling and settings management cleaned up.
  • IcingaClient cleanup — removed unused command-handling code paths.
  • CheckLogFile config and descriptions — fixed misleading defaults and improved the help text.
  • Web UI improvements — more settings elements exposed under modules, simpler module configuration. Web dependencies refreshed.
  • Installer: UninstallString is now correct (#495) — removal via Windows “Apps & Features” works again.
  • Rust dependencies bumped.

Upgrade notes

Most installs can upgrade in place — defaults are preserved. Read the specific items below if any of them apply.

Permission system

The new policy layer is disabled by default. Existing installs continue to behave exactly as before until an operator opts in via /settings/permissions/enabled = true.

If you do opt in:

  • Per-command rules under /settings/permissions/policies apply to queries only. Any rules you might have written for exec command patterns will be silently ignored for the exec dispatch path — exec is gated by the single global allow exec boolean.
  • The default for allow exec is true, so enabling the policy will not silently break the WEB scripts UI, lua/python core:simple_exec(...), or CLI exec. Flip to false only if you want a hard exec lockdown.
  • Roll out with log allows = true first so you can inventory what your actual traffic looks like before tightening to a real allow-list. See the step-by-step recipe in https://nsclient.org/docs/concepts/permissions/.

NRPEServer

  • The new client identity source setting defaults to none, which matches the previous behaviour (subject is bare NRPEServer). Set to cn only when you want per-cert principals — and only after you’ve configured verify_mode = peer-cert and a ca path. The module will refuse to start with a clear error if you set cn without those.
  • Pin the ca path to your private monitoring CA. The system trust store (Windows root store / Linux distro bundle) accepts certs from every public CA on the planet and would let an attacker with a public cert choose their own CN. See “Pin to a private CA” in the permissions doc.

Path overrides

  • If you had a [/paths] section in nsclient.ini from an older NSClient++ install, those overrides moved to [paths] in boot.ini (note: same section name, different file). There is no automatic migration. Copy each key = value to a [paths] section in boot.ini (next to nscp.exe) and delete the old section from nsclient.ini.

WEB server

  • The new disable admin user = true setting is opt-in. Existing installs keep their admin and continue to work unchanged. Use this when you want to expose the WEB UI for status-only viewing and have no need to reconfigure the agent through the web.

NRPEServer startup robustness

  • A failed listener (bad bind address, port in use) used to make the whole NRPEServer module fail to load. It now logs an ERROR and leaves the module loaded with no active listener — so you can reconfigure via nscp settings --path /settings/NRPE/server --key ... --set ... and reload, without restarting the service. If you had monitoring on “module load failed” specifically, you may want to add “NRPE listener failed” as a separate signal.

insecure = true on NRPEServer

  • This option (for legacy check_nrpe interop) now logs at ERROR rather than DEBUG/INFO. Behaviour is unchanged; the message is louder so it shows up in dashboards. If your monitoring filters by severity, you may want to whitelist this specific message on agents that intentionally run in insecure mode.

cache allowed host

  • Previously parsed as a string with surprising truthiness; now a real boolean. If you had cache allowed host = yes or = on, switch to true. Numeric 1 / 0 still work.

Nagios range syntax in performance data

  • This is additive — existing perfdata that doesn’t use range syntax continues to work. Plain numbers still parse as before. Only consumers that previously had to special-case NSClient++’s output may need adjusting, but most Nagios-ecosystem tools handle both forms.

Download

You can download the new version from GitHub

// Michael Medin

0.12.3 Fixed almost all bugs :)

0.12.4 Fixes a few important regression issues, so please use that version.

What’s Changed

This release rolls up everything since the last stable: five pre-releases (0.11.31, 0.11.32, 0.11.33, 0.12.1, 0.12.2) plus the latest in-development changes.

The headline themes are:

  1. New monitoring scenarios — first-class Checkmk and Icinga 2 integration, plus a real check_net family.
  2. A modern Web UI and REST API — events, metadata, settings DELETE, filterable lists, dedicated widgets for PDH counters and real-time filters.
  3. Hardened by default — 0.12.2 is a security release that closes listener defaults that used to be silently permissive (empty allowed hosts, plaintext check_nt, query-string tokens, etc.).
  4. Many long-standing check fixes — check_service, check_process, check_files, check_drivesize, check_uptime, CheckLogFile, and the shared filter/threshold engine all behave correctly now.

Read the Breaking changes section before upgrading — several long-standing-but-incorrect behaviours have been corrected and a number of listener defaults are now fail-closed. If you have an existing configuration, plan to review it.


TL;DR for end users

  • New scenario: Checkmk agent integration. Point a Checkmk site at port 6556 and you get a native-looking agent dump. See scenarios/check-mk.md.
  • New scenario: Icinga 2 passive submission. A new IcingaClient module submits passive results to Icinga 2’s REST API as an alternative to NSCA / NRDP.
  • New scenario: NSCA-ng. A new hardened NSCAngClient with PSK and AEAD-first cipher selection.
  • Native cross-platform network checks: check_tcp, check_dns, check_http, check_ntp_offset, check_connections.
  • Native Windows registry checks: check_registry_key, check_registry_value.
  • HTTP proxy support for every HTTP-based client (NRDP, Elastic, Op5, Icinga, the configuration loader, …).
  • Windows ROOT trust store auto-export — HTTPS-bound checks validate certificates against the system trust store automatically.
  • A modern Web UI with filterable lists, settings diff, dashboard, and dedicated CheckSystem widgets.
  • New REST endpoints: GET/DELETE /api/v2/events, GET /api/v2/metadata, DELETE /api/v2/settings/.... Covered in api/rest/.
  • Linux real-time metrics — the same background CPU/memory/disk/ network/load sampling that Windows has had for years.
  • Many bug fixes in check_service, check_process, check_files, CheckLogFile, the filter/threshold engine and the HTTP stack.

Major new features

Checkmk agent integration

NSClient++ can now serve a Checkmk-compatible agent dump on TCP port 6556. A real Checkmk site can register the host with tag_agent = cmk-agent, discover services, and run checks — no proxy, no NSCA gateway.

Enable it:

[/modules]
CheckMKServer = enabled
LUAScript = enabled
CheckSystem = enabled
CheckDisk = enabled
CheckHelpers = enabled

[/settings/check_mk/server]
port = 6556
allowed hosts = 127.0.0.1, <checkmk-site-ip>
submission ttl = 60          ; seconds, default 60
mrpe channel = check_mk-mrpe
local channel = check_mk-local

Out-of-the-box sections (no extra config):

Section Contents
<<<check_mk>>> Version, OS, hostname
<<<systemtime>>> Unix epoch (Windows clock-skew check)
<<<uptime>>> Seconds since boot (read from internal metrics store)
<<<mem>>> MemTotal:/MemFree:/SwapTotal:/SwapFree: (from metrics store)
<<<df>>> Per-volume size/used/free/mountpoint (Windows)
<<<services>>> name state/start_type display_name per Windows service
<<<ps>>> (user,vsz_kb,rss_kb,cputime,pid) cmdline per process

Expose any nscp check as a Checkmk service under <<<local>>>:

[/settings/check_mk/server/local]
CPU Load = command=check_cpu warn=load>80 crit=load>95
Disk C = command=check_drivesize drive=C: "warn=free<20%" "crit=free<10%"

MRPE relay under <<<mrpe>>>:

[/settings/check_mk/server/mrpe]
Uptime = command=check_uptime warn=uptime<2d
Memory = command=check_memory type=committed warn=used>80% crit=used>90%

Documentation: https://nsclient.org/docs/scenarios/check-mk.md`.

IcingaClient — Icinga 2 REST API submission

A new client module submits passive check results directly to an Icinga 2 master/satellite via the /v1/actions/process-check-result REST endpoint, as an alternative to NSCA or NRDP.

[/modules]
IcingaClient = enabled

[/settings/IcingaClient/targets/default]
address = https://icinga2.example.com:5665
username = nscp
password = secret
hostname = ${hostname}
nscp client --module IcingaClient \
            --command submit_icinga \
            --address https://icinga2.example.com:5665 \
            --username nscp --password secret \
            --command heartbeat \
            --result 0 \
            --message "Hello from NSClient++" \
            --ensure-objects

NSCA-ng client

A new NSCAngClient module ships a hardened NSCA-ng submission client with PSK support, AEAD-first cipher selection, and connection retry logic.

Native support for Windows CA-store

On startup NSClient++ now exports the machine’s ROOT certificate store as a single PEM bundle, so any check that does TLS (check_http, IcingaClient, NRDP, …) can validate certificates against the trust store the rest of Windows already uses.

check_http url=https://www.ibm.com
OK: https://www.ibm.com -> 303 ok (0B in 33ms)

check_http url=https://self-signed.badssl.com/
CRITICAL: https://self-signed.badssl.com/ -> 0 error: Failed to connect ... certificate verify failed

CheckNet — five new (cross-platform) checks

CheckNet graduated from a placeholder into a full network-check module. All five commands work over NRPE as well as locally:

  • check_tcp — open a TCP socket to one or more host/port pairs, optionally send a payload and require an expected substring.
  • check_dns — resolve a hostname and optionally assert which addresses come back.
  • check_http — fetch one or more URLs, check status code, response time and body content; supports custom headers and user-agent.
  • check_ntp_offset — query one or more NTP servers and alert on offset / stratum.
  • check_connections — Windows TCP/UDP connection table inspection (counts per protocol/family/state).
check_tcp host=smtp.gmail.com port=25 send="EHLO nsclient.org" expect="250"
check_dns host=google.com expected-address=172.217.20.174
check_http url=https://nsclient.org/ expected-body="NSClient" \
    "warn=time > 500 or code >= 400" \
    "crit=time > 2000 or code >= 500 or result != 'ok'"
check_ntp_offset "servers=0.pool.ntp.org,1.pool.ntp.org" timeout=2000
check_connections "filter=protocol = 'tcp' and state = 'TIME_WAIT'" \
    "warn=count > 200" "crit=count > 1000"

CheckSystem (Windows) — registry checks

Two new commands let you monitor the Windows registry directly from NSClient++ instead of relying on external scripts. They support recursion, exclude lists, 32/64-bit (WoW64) views, custom filters and the usual warn=/crit= expression syntax.

  • check_registry_key — verify that a key exists, count sub-keys/values, watch its last-write time.
  • check_registry_value — read a single value assert its type, size or content.
check_registry_key "key=HKLM\Software\NSClient" \
    "warn=age > 7d" "crit=age > 30d or not exists"

check_registry_key "key=HKLM\Software\Microsoft\Windows\CurrentVersion\Uninstall" \
    recursive max-depth=1 exclude=KB5005463 exclude=KB5005539

check_registry_value "key=HKLM\System\CurrentControlSet\Services\W32Time\Config" \
    value=MaxPollInterval "warn=int_value > 14" "crit=int_value > 17"

CheckSystem — check_os_updates (Windows)

A new check using the Windows Update Agent (WUA) reports pending OS updates. By default any pending update returns warning; thresholds let you alert only on security/critical:

check_os_updates "warning=important > 0" "critical=security > 0 or critical > 0"

CheckSystem (Linux) — real-time metrics

The Linux build of CheckSystem now ships with the same real-time metric collection that has been available on Windows for a long time: CPU, memory, disk, network and load are sampled in the background and exposed both to dashboards/metrics and to real-time filters (filter=... rules that fire when a threshold is crossed). Existing real-time filter configuration just works on Linux now.

Real-time filter metrics

CheckSystem’s real-time filters now publish per-filter match and error counts under system.realtime.<filter_name>.fired / system.realtime.<filter_name>.errors. Visible via:

  • The metrics REST endpoint (/api/v2/metrics + filter)
  • Prometheus scrape
  • The new Metrics() Lua API in default_check_mk.lua

Useful for spotting filters that never fire (typo in the where-clause) or filters that always error (broken expression).

CheckDisk — check_single_file

A focused variant of check_files for inspecting a single, known path. Compared to using check_files for the same job:

  • Only one required argument (file=<path>).
  • A clear error when the input is empty.
  • UNKNOWN: File not found: <path> when the file is missing — instead of the empty-set / “No files found” workflow.
  • A useful default detail-syntax so a no-threshold run is informative on its own.
check_single_file file=C:/windows/WindowsUpdate.log "warn=age > 5m" "crit=age > 1h"
CRITICAL: WindowsUpdate.log (size=276, age=917)

CheckDisk — filesystem filtering for check_drivesize

check_drivesize can now filter drives by filesystem type — useful for excluding tmpfs, nfs, etc.

check_drivesize drive=* "filter=fs = 'NTFS'"

check_nscp_update

A new check command queries the GitHub releases API (with caching) and reports whether the running NSClient++ is up to date.

HTTP proxy support across every HTTP client

NSClient++ can now route HTTP and HTTPS traffic through a corporate proxy. The same surface is used by every component built on the internal http::simple_client (NRDPClient, ElasticClient, Op5Client, IcingaClient, the remote boot.ini loader, …).

For HTTPS targets the client opens a CONNECT tunnel to the proxy, validates the proxy’s response, and only then performs the TLS handshake — so a single setting covers both http:// and https:// URLs.

Two new options on every HTTP client command and target:

Option Purpose
proxy Proxy URL — scheme://[user:pass@]host[:port]/. Empty value disables the proxy.
no-proxy Comma-separated list of hosts that bypass the proxy. A leading . is a suffix match.
[/settings/NRDP/client/targets/nagios]
address = https://nagios.example.com/nrdp/
token = mytoken
proxy = http://proxy.corp.example:3128/
no proxy = localhost,127.0.0.1,.internal

Configuration loader (boot.ini):

[proxy]
url = http://proxy.corp.example:3128/
no_proxy = localhost,127.0.0.1,.internal

Notes / limits:

  • Only the http:// proxy scheme is supported. socks5:// / https:// proxies are not.
  • No automatic detection of system proxy settings (HTTP_PROXY env vars, WinINET / WPAD). The proxy must be configured explicitly.
  • On 407 Proxy Authentication Required the proxy’s response body is captured in the error message.

Web UI / REST API expansion

New web routes:

Route Method Purpose
/api/v2/events GET List buffered real-time events
/api/v2/events DELETE Drain (returns + clears) the event buffer in one call
/api/v2/metadata GET Module/setting metadata index
/api/v2/metadata/counters GET List of available PDH counters
/api/v2/metadata/channels GET List of registered submission channels
/api/v2/settings/<path> DELETE Remove a settings key or path (staged delete; survives restart)

The settings store gained staged deletion: a DELETE is recorded so that subsequent reads of the deleted key/path return “not present” until the change is saved. Stops a deleted-but-not-yet-saved key from being re-resurrected by a concurrent read.

Web UI refresh

The bundled web interface has been heavily reworked:

  • Modern theme with active-navigation highlighting and a redesigned login page.
  • Filterable lists for Modules, Queries and Settings.
  • Settings diff dialog — the “settings changed” widget can now show exactly which keys changed.
  • CheckSystem settings UI got dedicated widgets for PDH counters and real-time filters: a counter picker that hits /api/v2/metadata/counters, “Add filter” / “Add counter” dialogs, and a live preview of metric values pulled from the metrics endpoint.

If you’ve been editing real-time filters in nsclient.ini by hand, the web UI is now a much faster way to do it.

SMTPClient rewrite

The SMTPClient module has been substantially rewritten with proper SMTP handling, integration tests, and a Python-based test harness.

Smaller features

  • nscp settings --sort — produce stable, sorted output, useful for diffing exported settings between hosts.
  • Performance threshold min/max bounds — perfdata threshold expressions can declare minimum and maximum bounds, propagated into emitted perfdata:
    check_pdh "counter=\\Processor(_Total)\\% Processor Time" \
              "perf-config=*(minimum:0;maximum:100)"
    
  • Timezone-aware check_uptime and Schduler — applies a timezone cache on both Windows and Unix, so absolute boot-time output and cron expressions agree with the host’s local time.
  • WEBServer cookie attribute support — Secure, HttpOnly, SameSite, Path, Domain, Expires, Max-Age.
  • WEBServer password hashing with constant-time verification — removes the timing oracle on the previous plaintext equality check.
  • WEBServer authentication rate limiter — per-source throttling of failed authentication attempts:
    [/settings/WEB/server]
    auth rate limit max failures   = 10   ; 0 disables the limiter
    auth rate limit block seconds  = 60
    

Filter engine — stable summary thresholds

These changes touch the shared filter / threshold engine and therefore affect every modular check (check_files, check_service, check_process, check_eventlog, …).

Stable count / total / *_count in warn= / crit=

warn= / crit= were evaluated during iteration. Summary variables such as count therefore exposed their running value instead of the final post-iteration value, so a mixed expression like

crit = state = 'hung' OR count < 5

mis-fired on the very first row (count == 1 < 5) regardless of how many rows ultimately matched. Per-row evaluation is now deferred: matched rows are recorded during iteration, and the warn/crit/ok engines run once the summary state is final.

Mixed warn= / crit= evaluated when no rows match

If a filter excluded every row, mixed expressions like crit = state = 'stopped' OR count = 0 were skipped entirely — leaving the check OK in the empty case. They are now evaluated with object-bound variables defaulting to false and summary variables at their final values, so the check correctly returns CRITICAL when the service is missing.

Quieter, more predictable expression evaluation

  • Operators audited so is_unsure propagates consistently; invalid-type comparisons resolve to unsure-false instead of erroring.
  • String variables on no-object cases now return an empty string with is_unsure=true and produce a warning in the log instead of an error per row — log volume on complex queries drops dramatically.
  • Removed the misleading “most likely mutating” warnings.
  • Substantial new test coverage.

check_service and check_process fixes (Windows)

  • “Failed to enumerate service: 6f7” on busy hosts — enumeration is now properly looped until the SCM signals end-of-data.
  • perf-syntax=none actually suppresses perfdata — check_service used to emit empty perfdata aliases ( ''=4;0;1 ''=4;0;1 ...), blowing past NRPE size limits.
  • No more TODO leaking into ${desc} — check_service service=Spooler used to render as OK: Spooler: TODO. Now: OK: Spooler: Print Spooler.
  • delayed only reported for SERVICE_AUTO_START — manual / boot / system / disabled services no longer randomly show up as delayed.
  • check_process sees protected / cross-user processes as NETWORK SERVICE — a PROCESS_QUERY_LIMITED_INFORMATION fallback is now attempted, so winlogon.exe, csrss.exe etc. no longer report CRITICAL: <name>=stopped when the agent runs unprivileged.
  • Realtime check_process is now case-insensitive, matching the active path and Windows itself.

check_files fixes

  • #730 — max-depth=0 now scans the top directory only (was: bail out before scanning anything, returning “no files found”).
  • #598 — Non-ASCII paths (accented letters, CJK, …) are no longer silently mangled by mismatched codepage conversions.
  • #613 — Top-level paths that cannot be opened now produce UNKNOWN: Path was not found: <path> instead of being hidden behind the configured empty-state.
  • #605 — NTFS junctions / symlinks / mount points are now skipped during recursion, preventing double-counts on self-referential trees.
  • #717 — The legacy CheckFiles shim now sets empty-state=ok when translating, restoring 0.4-era behaviour for legacy calls that find zero files.

Other check / module fixes

  • CheckDisk resilience — an error on a single unavailable volume no longer aborts the entire check_drivesize run.
  • #581 — CheckLogFile honours the line-split argument (previously hard-coded to \n); multi-character delimiters such as \r\n are handled correctly. Real-time seek behaviour fixed; CRLF handling harmonised.
  • #589 — Time/duration arguments such as time=3000foobar or time=3000mfoobar are no longer silently accepted; malformed inputs are rejected with a clear error.
  • #669 — The literal U (Nagios “undefined”) in performance data is preserved end-to-end instead of being coerced to 0. Only an exact U, u, U% or u% token matches.
  • NSCA wire timestamps are now correctly built in UTC. Both server (IV packet) and client (data packet) used to derive seconds-since-epoch from second_clock::local_time(), which drifted by the host’s TZ offset. A timezone setting on both ends allows legacy interop with agents that emit local-clock-as-Unix-time stamps.
  • Metrics collection regression fixed (some metrics were silently dropped).
  • Op5Client / ElasticClient unified on the new HTTP client; 401 path fixed; reponse → response typos corrected.
  • Gracefully handle non-numeric NSClient command codes.
  • TLS support fixes; better randomness for encryption; race condition fixes; boundary checks for various network payloads and reading certificates.
  • NRDP integration tests added; new nrdp client alias.

HTTP refactor

  • HTTP request and response are now distinct types instead of one shared bag.
  • Chunked transfer-encoding is decoded properly. check_http against servers using Transfer-Encoding: chunked ( most modern reverse proxies, Icinga 2, Kubernetes ingress, …) now returns the full body instead of a truncated/garbled one. The IcingaClient module relies on this.
  • Header storage is normalised — case-insensitive lookup, no more duplicate-header surprises.

Security hardening

The 0.12.2 release is a security-focused pass. These do not change documented behaviour for well-formed traffic but close down attacker-controlled edge cases.

DoS / resource-exhaustion limits

  • Authorization header capped at 8 KiB to mitigate amplification.
  • Per-connection parser buffer cap to prevent memory pinning from oversized or never-completed requests.
  • Session token cap with eviction to prevent unbounded memory growth.
  • Payload lengths below the protocol minimum are rejected before allocation.
  • Path expansion now detects cycles and refuses to recurse, preventing stack overflow on pathological configurations.

NSCA hardening

  • Packet version is checked.
  • Timestamp validation tightened to mitigate replay attacks.

Log/output injection prevention

Control characters are stripped from values before they are written to external sinks, removing log/protocol-injection vectors:

  • Log file entries (file names and messages)
  • syslog messages (CR/LF/NUL stripped)
  • Graphite metric paths and values
  • HTTP response headers (header keys and values)
  • log_status is now JSON-serialised so attacker-controlled fields cannot inject extra structured fields.

Filesystem / process safety

  • PID file creation hardened against symlink attacks; exclusive access enforced.
  • Archive extraction has a zip-slip guard that validates entry paths and refuses traversal sequences.
  • Module and script names are validated to prevent path traversal at load time.
  • Argument substitution in external scripts is isolated to prevent command injection through user-controlled tokens.

Cryptography / TLS

  • HTTPS now logs explicitly when no certificate is present and warns on HTTP fallback in production.
  • SSL connections enable hostname verification by default.
  • Auto-generated passwords use OpenSSL RAND_bytes (cryptographically secure) instead of the previous predictable generator.
  • Sensitive values are no longer logged at debug level.
  • check_nt password compare is constant-time.

Breaking changes

Read this section carefully. Some changes are listener defaults that are now fail-closed; some are corrections to long-standing buggy behaviour; some are internal API changes for out-of-tree modules.

Listeners default to safer behaviour

  1. Empty allowed hosts now rejects all connections. Previously treated as “allow any source”. To genuinely expose the agent to any source, set it explicitly:
    allowed hosts = 0.0.0.0/0,::/0
    
  2. check_nt (NSClientServer) defaults to ssl = true. The legacy check_nt protocol carries the password in every request. The listener will not refuse to start if TLS is off, but it will log a warning. To keep the old plaintext behaviour for legacy clients, set ssl = false explicitly in [/settings/NSClient/server].
  3. check_nt: the literal password None no longer authenticates. Empty server passwords now reject all requests. Errors are also genericised (ERROR: Bad request.) to remove the online password-guessing oracle.
  4. WEBServer: /auth/token and /auth/logout are removed (HTTP 410). They accepted the password and session token as URL query parameters, leaking credentials into browser history and proxy logs. Migrate to:
    • POST /api/v2/login with Authorization: Basic to obtain a token
    • DELETE /api/v2/login with Authorization: Bearer to log out
  5. WEBServer: ?TOKEN= / ?__TOKEN= query-string token auth removed. Send the token in a header instead: Authorization: Bearer <token>, TOKEN: <token>, or X-Auth-Token: <token>.
  6. WEBServer: anonymous access is now opt-in. A role named anonymous registered in settings is silently ignored unless the new allow_anonymous flag is enabled.
  7. WEBServer: existing admin user is no longer overwritten on restart. Deployments that relied on the password being reset to the default at boot must adapt.

Scheduler — cron expressions evaluate in local time by default (#570)

The Scheduler module previously used UTC, so 40 15 * * * fired at 15:40 UTC regardless of host TZ. The default has changed to local time, matching standard cron semantics. Hour and minute fields will shift accordingly on non-UTC hosts.

A new timezone setting under [/settings/scheduler] controls the reference clock:

[/settings/scheduler]
timezone = local                          ; default — standard cron semantics
; timezone = utc                          ; restore the pre-0.12 behaviour
; timezone = EST-05EDT,M3.2.0,M11.1.0     ; any POSIX TZ string is honoured

IANA names such as Europe/Stockholm are not supported — use the POSIX form. Unparseable values fall back to UTC and surface as UTC? in any timezone label.

Filter / threshold engine

  1. warn= / crit= no longer fire mid-iteration on running counts. Configurations “tuned” against the buggy early-fire will produce different results.
    crit = state = 'hung' OR count < 5
    # Old: CRITICAL on the very first row (count == 1).
    # New: CRITICAL only if any row is 'hung' OR final count < 5.
    
  2. Mixed warn= / crit= now evaluate when no rows match.
    crit = state = 'stopped' OR count = 0
    # Old: OK when nothing matched (count = 0).
    # New: CRITICAL when nothing matched (count = 0).
    
    If your old config implicitly treated “empty” as OK, add a count > 0 AND ... guard or move the empty-case logic into a dedicated check.

Check-specific corrections

  1. check_service: delayed is no longer reported for non-auto services. Filters that matched start_type = 'delayed' on Manual / Boot / System / Disabled services will stop matching. To alert on “any non-running service that isn’t disabled”:
    filter = start_type IN ('auto','delayed','boot','system') AND state != 'running'
    
  2. Realtime check_process is now case-insensitive. A rule that intentionally matched only an exact casing will now match all variants (almost certainly the desired behaviour).
  3. check_service: ${desc} no longer returns the literal TODO. Use the real display name.
  4. check_service: perf-syntax=none actually suppresses perfdata. Backends that consumed the empty-aliased entries (highly unlikely) will see them disappear.

check_files — corner cases changed

  1. max-depth=0 now scans the top directory instead of returning empty (#730).
  2. Missing paths now return UNKNOWN instead of OK / empty (#613).
  3. NTFS junction loops are no longer double-counted (#605).
  4. Legacy CheckFiles calls that previously returned UNKNOWN on empty results will now return OK (#717).

Configuration / startup

  1. CheckExternalScripts: malformed alias commands are refused at startup. The fallback “split-on-space” parser has been removed. Aliases whose command line does not parse cleanly are refused with an error in the log instead of being silently registered with surprising tokenisation. Review your logs after upgrading.

Internal API (out-of-tree module authors)

  1. HTTP request/response API changed. Internal C++ types http::request / http::response are now distinct, headers are case-insensitive, and chunked decoding happens transparently. Out-of-tree modules linked against the old shared bag type need a small adjustment:
    // before
    http::packet pkt = client.send(...);
    auto body = pkt.body;
    
    // after
    http::response resp = client.send(http::request{...});
    auto body = resp.body();   // chunked decoding already applied
    

Documentation reorganisation

  1. The documentation tree was restructured (concepts/, checks-in-depth/, scenarios/, tutorial/, reference/ are now clearly separated). Bookmarks and external links may need updating.

Upgrade checklist

  1. Audit allowed hosts on every node — empty values now reject everything.
  2. check_nt (NSClientServer) now defaults to ssl = true. If your clients don’t speak TLS, set ssl = false explicitly. Either way the listener will log a warning at startup if TLS is off or a password is configured, recommending a switch to REST or NRPE.
  3. Replace any client that calls /auth/token or /auth/logout with the /api/v2/login flow.
  4. Replace any client that passes ?TOKEN= / ?__TOKEN= in the query string with a header-based token.
  5. Scheduler cron expressions on non-UTC hosts will shift to local time. Either update them or set [/settings/scheduler] timezone = utc to restore the previous behaviour.
  6. Review check_service / check_process / check_files filters that may have relied on the corrected behaviours listed above.
  7. Restart the service and review the log for new “refused alias” or “rejected connection” warnings — these flag configurations that were previously silently accepted.

No configuration migration is required for the new HTTP proxy keys, the Checkmk server, the Icinga client, the NSCA-ng client, or the new checks — they are all opt-in.

Download

You can download the new version from GitHub

// Michael Medin