CheckKubernetes¶
Experimental
This module is experimental: it works, but its options, filter keywords and output may change in a future release. Please try it and report anything that does not behave the way you expect.
Check a Kubernetes cluster through its API server: cluster health, pods, nodes and workloads.
Enable module¶
To enable this module and allow using the commands you need to add CheckKubernetes = enabled to the [/modules] section in nsclient.ini:
[/modules]
CheckKubernetes = enabled
Queries¶
A quick reference for all available queries (check commands) in the CheckKubernetes module.
List of commands:
A list of all available queries (check commands)
| Command | Description |
|---|---|
| check_kubernetes (experimental) | Check that the Kubernetes API server is reachable and ready: version, readiness and node counts. |
| check_nodes (experimental) | Check node readiness, pressure conditions, schedulability and capacity. |
| check_pods (experimental) | Check the state of pods: phase, the kubectl STATUS column, readiness and restarts. |
| check_workloads (experimental) | Check that deployments, statefulsets and daemonsets have their desired replicas available. |
check_kubernetes¶
Check that the Kubernetes API server is reachable and ready: version, readiness and node counts.
About check_kubernetes¶
check_kubernetes is the “is the cluster there” check: it reads /version,
/readyz and /api/v1/nodes from the API server configured under
[/settings/kubernetes] and reports the Kubernetes version, whether the API
server considers itself ready, and how many nodes are Ready.
Reaching the API is the health signal. The default critical is
api_ready = 0, which trips when /readyz answers with a failing check (the
readyz keyword then names it, e.g. failed: etcd); an unreachable server, a
rejected token or a missing RBAC rule are reported as UNKNOWN with the reason
(see below). That holds for /readyz too: a service account without the
nonResourceURLs: ["/readyz"] rule gets an UNKNOWN naming the rule, not a
CRITICAL about a healthy cluster. Only a readiness report counts as a verdict
(ok, or the API server’s own 5xx listing its checks); a 404 or 502 from an
ingress that does not forward the path leaves api_ready at 1 and shows
readyz as unavailable (HTTP 404), since /version did answer. Node counts carry no default threshold, so add
warning=nodes_not_ready > 0 when a NotReady node should show up here rather
than in check_nodes.
The cluster is chosen by the operator, once, in the settings. No command in
this module takes a url=, host= or token= argument: a caller who can run
checks over REST (anyone holding queries.execute) must not be able to point
the agent - and the bearer token it sends - at a server of their choosing. The
token itself never appears in check output, error text or the log.
Error messages are meant to be acted on:
| Situation | Result | Message starts with |
|---|---|---|
| Nothing configured | UNKNOWN | No Kubernetes API server configured: setapi serverandtoken... |
| TCP/TLS failure | UNKNOWN | Failed to connect to Kubernetes API server at 'https://...' |
| HTTP 401 | UNKNOWN | ... rejected the credentials (HTTP 401 ...), naming the credential sent |
| HTTP 403 | UNKNOWN | ... denied GET /api/v1/... (HTTP 403: <the API server's own message>) |
Jump to section:
Sample Commands¶
Check that the API server is reachable and ready:
check_kubernetes
OK: Kubernetes v1.30.4 at https://127.0.0.1:6443: API ready, 3/4 nodes ready
Alert on nodes that are not Ready, with perf data:
check_kubernetes "warning=nodes_not_ready > 0" "critical=nodes_ready < 2"
WARNING: Kubernetes v1.30.4 at https://127.0.0.1:6443: API ready, 3/4 nodes ready|'https://127.0.0.1:6443 not ready nodes'=1;0;0 'https://127.0.0.1:6443 ready nodes'=3;0;2
Use the keywords in the output:
check_kubernetes "detail-syntax=%(version) on %(platform), source %(source), readyz %(readyz)"
OK: v1.30.4 on linux/amd64, source settings, readyz ok
An API server that is down is clearly reported (UNKNOWN):
check_kubernetes
Failed to connect to Kubernetes API server at 'https://127.0.0.1:1' (settings): Failed to connect to 127.0.0.1:1: Connection refused
Command-line Arguments¶
| Option | Default Value | Description |
|---|---|---|
| timeout | 30 | Timeout in seconds for each network step of an API server request: the connect, the TLS handshake and each read. It is not a total for the check (a server that keeps trickling data takes longer). A positive number. |
timeout:
Timeout in seconds for each network step of an API server request: the connect, the TLS handshake and each read. It is not a total for the check (a server that keeps trickling data takes longer). A positive number.
Default Value: 30
Common options:
These options are shared by all filter based commands and are described on the common options page; the default values below are specific to this command.
| Option | Default Value |
|---|---|
| filter | |
| warning | |
| warn | |
| critical | api_ready = 0 |
| crit | |
| ok | |
| debug | false |
| show-all | false |
| empty-state | unknown |
| perf-config | |
| escape-html | false |
| list-separator | , |
| top-syntax | ${status}: ${list} |
| ok-syntax | |
| empty-syntax | %(status): No cluster information returned |
| detail-syntax | Kubernetes ${version} at ${server}: API ${api_ready}, ${nodes_ready}/${nodes} nodes ready |
| perf-syntax | ${server} |
| byte-unit | |
| decimal-separator | |
| decimals | -1 |
| thousands-separator |
This command also accepts the standard help options: help, help-pb, show-default, help-short.
Filter keywords¶
| Option | Description |
|---|---|
| api_ready | 1 unless /readyz carries a readiness report saying the API server is not ready (renders as ready / not ready); a path that is unavailable (a 404 or 502 from an ingress) leaves it 1 - /version answered |
| nodes | Number of nodes in the cluster |
| nodes_not_ready | Number of nodes whose Ready condition is False or Unknown |
| nodes_ready | Number of nodes whose Ready condition is True |
| platform | Platform the API server runs on (e.g. linux/amd64) |
| readyz | What /readyz answered: ok, the failing checks it listed, or unavailable (HTTP n) when the path did not answer with a readiness report |
| server | The API server address (never the credential) |
| source | Where the cluster configuration came from: settings, kubeconfig |
| version | Kubernetes version reported by /version (e.g. v1.30.2) |
This command also supports the common filter keywords: count, total, ok_count, warn_count, crit_count, problem_count, list, ok_list, warn_list, crit_list, problem_list, detail_list, sep, status.
check_nodes¶
Check node readiness, pressure conditions, schedulability and capacity.
About check_nodes¶
check_nodes reads /api/v1/nodes and evaluates each node: the Ready
condition, the pressure conditions (memory_pressure, disk_pressure,
pid_pressure, network_unavailable), whether it is cordoned
(schedulable), and what it offers - cpu_capacity / cpu_allocatable in
millicores, memory_capacity / memory_allocatable in bytes (thresholds take
units: memory_allocatable < 4G) and pods_capacity.
node_status is the STATUS column of kubectl get nodes: Ready,
NotReady or Unknown, with ,SchedulingDisabled appended for a cordoned
node. The defaults go critical on ready != 'True' and warn on any pressure
condition or a cordon; a drained node during maintenance therefore shows up as
a warning, which is usually what you want - add filter=schedulable = 1 to
hide it.
node=<name> (repeatable) restricts the check to the named nodes, and a node
the API server does not return is reported as missing: it left the cluster,
which is exactly what a per-node check is there to tell you. label-selector=
and field-selector= are passed to the API server, e.g.
label-selector=node-role.kubernetes.io/worker=.
Jump to section:
Sample Commands¶
Check every node with the default thresholds (NotReady is critical, pressure or a cordon a warning):
check_nodes
CRITICAL: worker-2=NotReady, worker-3=Ready,SchedulingDisabled
Only alert on readiness, ignoring cordoned nodes:
check_nodes warning=none "critical=ready != 'True'"
CRITICAL: worker-2=NotReady
Require that specific nodes are still in the cluster:
check_nodes node=worker-1 node=worker-9
CRITICAL: worker-9=missing
Show what a node offers, using the node keywords:
check_nodes node=worker-1 "detail-syntax=%(name): %(node_status), kubelet %(kubelet_version), %(os) %(arch), cpu %(cpu_allocatable)m of %(cpu_capacity)m, memory %(memory_allocatable) of %(memory_capacity) bytes, %(pods_capacity) pods, roles=%(roles) taints=%(taints)" "top-syntax=${list}" ok-syntax=
worker-1: Ready, kubelet v1.30.4, Ubuntu 22.04.4 LTS amd64, cpu 7900m of 8000m, memory 33285996544 of 34359738368 bytes, 110 pods, roles= taints=
List cordoned nodes without alerting:
check_nodes "filter=schedulable = 0" warning=none critical=none "detail-syntax=%(name)=%(node_status)" "top-syntax=${status}: ${list}" ok-syntax=
OK: worker-3=Ready,SchedulingDisabled
Command-line Arguments¶
| Option | Default Value | Description |
|---|---|---|
| label-selector | Label selector passed to the API server, e.g. node-role.kubernetes.io/worker= or topology.kubernetes.io/zone=eu-1a. | |
| field-selector | Field selector passed to the API server, e.g. metadata.name=worker-1. | |
| node | Name of a node that must exist (repeatable). Only the named nodes take part in the check; a name the API server does not return gets node_status ‘missing’. | |
| timeout | 30 | Timeout in seconds for each network step of an API server request: the connect, the TLS handshake and each read. It is not a total for the check (a server that keeps trickling data takes longer). A positive number. |
timeout:
Timeout in seconds for each network step of an API server request: the connect, the TLS handshake and each read. It is not a total for the check (a server that keeps trickling data takes longer). A positive number.
Default Value: 30
Common options:
These options are shared by all filter based commands and are described on the common options page; the default values below are specific to this command.
| Option | Default Value |
|---|---|
| filter | |
| warning | memory_pressure = 1 or disk_pressure = 1 or pid_pressure = 1 or network_unavailable = 1 or schedulable = 0 |
| warn | |
| critical | ready != ‘True’ |
| crit | |
| ok | |
| debug | false |
| show-all | false |
| empty-state | critical |
| perf-config | |
| escape-html | false |
| list-separator | , |
| top-syntax | ${status}: ${problem_list} |
| ok-syntax | %(status): All %(count) nodes are ready |
| empty-syntax | %(status): No nodes found |
| detail-syntax | ${name}=${node_status} |
| perf-syntax | ${name} |
| byte-unit | |
| decimal-separator | |
| decimals | -1 |
| thousands-separator |
This command also accepts the standard help options: help, help-pb, show-default, help-short.
Filter keywords¶
| Option | Description |
|---|---|
| age | Seconds since the node joined the cluster, -1 when unknown (supports units, e.g. age < 1h) |
| arch | CPU architecture (amd64, arm64, …) |
| cpu_allocatable | CPU available to pods in millicores (-1 when not reported) |
| cpu_capacity | CPU capacity in millicores (-1 when not reported) |
| created | When the node joined the cluster (date) |
| disk_pressure | 1 when the DiskPressure condition is True, else 0 |
| internal_ip | The node’s InternalIP address |
| kubelet_version | Kubelet version (e.g. v1.30.2) |
| memory_allocatable | Memory available to pods in bytes, -1 when not reported (thresholds take units, e.g. memory_allocatable < 4G) |
| memory_capacity | Memory capacity in bytes, -1 when not reported (thresholds take units, e.g. memory_capacity < 8G) |
| memory_pressure | 1 when the MemoryPressure condition is True, else 0 |
| name | Node name |
| network_unavailable | 1 when the NetworkUnavailable condition is True, else 0 |
| node_status | The kubectl STATUS column: Ready, NotReady or Unknown, with ,SchedulingDisabled appended for a cordoned node; missing for a requested node the API server does not know about |
| os | OS image the node runs (e.g. Ubuntu 22.04.4 LTS) |
| pid_pressure | 1 when the PIDPressure condition is True, else 0 |
| pods_capacity | Maximum number of pods the node accepts (-1 when not reported) |
| ready | The Ready condition: True, False or Unknown (empty for a requested node the API server does not know about) |
| roles | Node roles from the node-role.kubernetes.io/ |
| schedulable | 1 unless the node is cordoned (spec.unschedulable), then 0 |
| taints | Taints as key=value:effect, comma separated |
This command also supports the common filter keywords: count, total, ok_count, warn_count, crit_count, problem_count, list, ok_list, warn_list, crit_list, problem_list, detail_list, sep, status.
check_pods¶
Check the state of pods: phase, the kubectl STATUS column, readiness and restarts.
About check_pods¶
check_pods lists pods - across the cluster or in the given namespace=
(repeatable) - and evaluates each one. The list is fetched page by page
(limit=500, following metadata.continue), so it works on large clusters;
label-selector= and field-selector= are handed to the API server, so the
filtering happens before the payload is built.
pod_status reproduces the STATUS column of kubectl get pods (Running,
Completed, CrashLoopBackOff, ImagePullBackOff, OOMKilled,
Terminating, Init:1/2, ExitCode:3, …) from the container statuses and
deletionTimestamp. Use it rather than phase for alerting: a crash-looping
pod has phase Running.
Three ways to use it:
- No arguments - every pod that has not finished (
filter=phase != 'Succeeded'). The defaults warn onPendingpods and more than five restarts, and go critical onFailed,Unknown, any*BackOff,OOMKilledor amissingpod. pod=<name>orpod=<namespace>/<name>(repeatable) - only the named pods take part, and a pod the API server does not return is reported withpod_statusmissing, so it trips the default critical instead of silently disappearing from the listing. A bare name is matched in every namespace listed and, when missing, reported as*/<name>(or under the onenamespace=the check was scoped to).- A selector -
label-selector=app=weborfield-selector=spec.nodeName=worker-1for the pods of one application or one node.
age takes duration units (age < 10m), created is a date, and oom_killed
is 1 when a container’s current or last termination was an out-of-memory kill,
even if the pod has since restarted and shows Running. Native sidecars (init
containers with restartPolicy: Always) count as running once started and are
included in containers / ready_containers, as in kubectl; restarts is
the kubectl RESTARTS column, which counts the main containers and running
sidecars once the pod is initialised and the init containers only while it is
not, and ready_containers counts a container only while it is actually
running; a finished job
pod being cleaned up keeps Completed rather than turning into Terminating,
while terminating still reports the pending deletion.
Jump to section:
Sample Commands¶
Check every pod with the default thresholds (finished job pods are ignored):
check_pods
CRITICAL: shop/api-5f6c7d8b9-xyz12=CrashLoopBackOff, shop/db-0=Pending|'shop/web-7d4b9c6f8-k2xqz restarts'=0;5;0 'shop/web-7d4b9c6f8-p9lmn restarts'=0;5;0 'shop/api-5f6c7d8b9-xyz12 restarts'=7;5;0 'shop/db-0 restarts'=0;5;0 'kube-system/coredns-76f75df574-abcde restarts'=0;5;0 'kube-system/coredns-76f75df574-fghij restarts'=0;5;0
Only one namespace, with your own thresholds:
check_pods namespace=shop "warning=restarts > 3" "critical=pod_status like 'BackOff'"
CRITICAL: shop/api-5f6c7d8b9-xyz12=CrashLoopBackOff|'shop/web-7d4b9c6f8-k2xqz restarts'=0;3;0 'shop/web-7d4b9c6f8-p9lmn restarts'=0;3;0 'shop/api-5f6c7d8b9-xyz12 restarts'=7;3;0 'shop/db-0 restarts'=0;3;0
Require that specific pods exist (a pod the API server does not know is CRITICAL):
check_pods pod=shop/web-7d4b9c6f8-k2xqz pod=shop/db-0 pod=shop/search-0
CRITICAL: shop/db-0=Pending, shop/search-0=missing|'shop/web-7d4b9c6f8-k2xqz restarts'=0;5;0 'shop/db-0 restarts'=0;5;0 'shop/search-0 restarts'=0;5;0
Select by label and show the pod keywords:
check_pods label-selector=app=web "detail-syntax=%(namespace)/%(name) on %(node): %(pod_status) %(ready_containers)/%(containers) ready, %(restarts) restarts" "top-syntax=${status}: ${list}" ok-syntax=
OK: shop/web-7d4b9c6f8-k2xqz on worker-1: Running 1/1 ready, 0 restarts, shop/web-7d4b9c6f8-p9lmn on worker-3: Running 1/1 ready, 0 restarts|'shop/web-7d4b9c6f8-k2xqz restarts'=0;5;0 'shop/web-7d4b9c6f8-p9lmn restarts'=0;5;0
A namespace the service account may not read is reported with the rule to grant (UNKNOWN):
check_pods namespace=secret
Kubernetes API server at 'https://127.0.0.1:6443' denied GET /api/v1/namespaces/secret/pods?limit=500 (HTTP 403: pods is forbidden: User "system:serviceaccount:monitoring:nscp" cannot list resource "pods" in API group "" in the namespace "secret"): grant the agent's service account get and list on the resource (see the CheckKubernetes documentation for the ClusterRole)
Command-line Arguments¶
| Option | Default Value | Description |
|---|---|---|
| namespace | Only list pods in this namespace (repeatable). Default: all namespaces. | |
| label-selector | Label selector passed to the API server, e.g. app=web,tier!=cache: filtering happens before the payload is built. | |
| field-selector | Field selector passed to the API server, e.g. spec.nodeName=worker-1 or status.phase!=Succeeded. | |
| pod | Name of a pod that must exist (repeatable). Only the named pods take part in the check; a name the API server does not return gets pod_status ‘missing’. | |
| timeout | 30 | Timeout in seconds for each network step of an API server request: the connect, the TLS handshake and each read. It is not a total for the check (a server that keeps trickling data takes longer). A positive number. |
timeout:
Timeout in seconds for each network step of an API server request: the connect, the TLS handshake and each read. It is not a total for the check (a server that keeps trickling data takes longer). A positive number.
Default Value: 30
Common options:
These options are shared by all filter based commands and are described on the common options page; the default values below are specific to this command.
| Option | Default Value |
|---|---|
| filter | phase != ‘Succeeded’ |
| warning | phase = ‘Pending’ or restarts > 5 |
| warn | |
| critical | phase = ‘Failed’ or phase = ‘Unknown’ or pod_status like ‘BackOff’ or pod_status = ‘OOMKilled’ or pod_status = ‘missing’ |
| crit | |
| ok | |
| debug | false |
| show-all | false |
| empty-state | ok |
| perf-config | |
| escape-html | false |
| list-separator | , |
| top-syntax | ${status}: ${problem_list} |
| ok-syntax | %(status): All %(count) pods are fine |
| empty-syntax | %(status): No pods found |
| detail-syntax | ${namespace}/${name}=${pod_status} |
| perf-syntax | ${namespace}/${name} |
| byte-unit | |
| decimal-separator | |
| decimals | -1 |
| thousands-separator |
This command also accepts the standard help options: help, help-pb, show-default, help-short.
Filter keywords¶
| Option | Description |
|---|---|
| age | Seconds since the pod was created, -1 when unknown (supports units, e.g. age < 1h) |
| containers | Number of containers in the pod spec |
| created | When the pod was created (date) |
| ip | Pod IP address |
| labels | Pod labels as key=value, comma separated |
| name | Pod name |
| namespace | Namespace the pod lives in |
| node | Node the pod is scheduled on (empty while Pending) |
| oom_killed | 1 when a container’s current or last termination was an out-of-memory kill, else 0 |
| owner | Name of the controller owning the pod |
| owner_kind | Kind of the controller owning the pod: ReplicaSet, StatefulSet, DaemonSet, Job, Node, … |
| phase | Pod phase: Pending, Running, Succeeded, Failed or Unknown |
| pod_status | The kubectl STATUS column: Running, Completed, CrashLoopBackOff, ImagePullBackOff, OOMKilled, Terminating, Init:1/2, … or missing for a requested pod the API server does not know about |
| qos | QoS class: Guaranteed, Burstable or BestEffort |
| ready | 1 when the pod’s Ready condition is True, else 0 |
| ready_containers | Number of containers reporting ready |
| restarts | The kubectl RESTARTS column: restarts of the containers and native sidecars once the pod is initialised, of the init containers while it is not |
| terminating | 1 when the pod is being deleted (deletionTimestamp is set), else 0 |
This command also supports the common filter keywords: count, total, ok_count, warn_count, crit_count, problem_count, list, ok_list, warn_list, crit_list, problem_list, detail_list, sep, status.
check_workloads¶
Check that deployments, statefulsets and daemonsets have their desired replicas available.
About check_workloads¶
check_workloads reads the deployments, statefulsets and daemonsets under
/apis/apps/v1 and normalises the three status shapes into one record:
desired, ready, available, updated, unavailable and missing
(desired minus available), plus paused for a deployment whose rollout is
paused. For a daemonset desired is the number of nodes it should run on.
The defaults go critical when nothing is available of a non-zero desired
(available = 0 and desired > 0) and warn when replicas are missing or a
rollout has not finished (missing > 0 or updated < desired). A workload
scaled to zero wants nothing and is fine.
kind=deployment|statefulset|daemonset (repeatable, singular or plural, any
case) restricts the check to one type and saves the other list calls;
namespace=, label-selector= and field-selector= are passed to the API
server. workload=<name> or workload=<namespace>/<name> (repeatable) names
workloads that must exist; one the API server does not return is reported
with 0/1 available under the kind missing, which trips the default
critical.
Jump to section:
Sample Commands¶
Check every deployment, statefulset and daemonset with the default thresholds:
check_workloads
CRITICAL: Deployment shop/api=1/3, Deployment ops/legacy=0/2, StatefulSet shop/db=2/3, DaemonSet monitoring/node-exporter=3/4|'shop/web available'=2;0;0 'shop/web desired'=2;0;0 'shop/web missing'=0;0;0 'shop/api available'=1;0;0 'shop/api desired'=3;0;0 'shop/api missing'=2;0;0 'ops/legacy available'=0;0;0 'ops/legacy desired'=2;0;0 'ops/legacy missing'=2;0;0 'kube-system/coredns available'=2;0;0 'kube-system/coredns desired'=2;0;0 'kube-system/coredns missing'=0;0;0 'shop/db available'=2;0;0 'shop/db desired'=3;0;0 'shop/db missing'=1;0;0 'monitoring/node-exporter available'=3;0;0 'monitoring/node-exporter desired'=4;0;0 'monitoring/node-exporter missing'=1;0;0
One kind in one namespace:
check_workloads kind=deployment namespace=shop
WARNING: Deployment shop/api=1/3|'shop/web available'=2;0;0 'shop/web desired'=2;0;0 'shop/web missing'=0;0;0 'shop/api available'=1;0;0 'shop/api desired'=3;0;0 'shop/api missing'=2;0;0
Require that specific workloads are fully available:
check_workloads workload=shop/web workload=kube-system/coredns
OK: All 2 workloads are available|'shop/web available'=2;0;0 'shop/web desired'=2;0;0 'shop/web missing'=0;0;0 'kube-system/coredns available'=2;0;0 'kube-system/coredns desired'=2;0;0 'kube-system/coredns missing'=0;0;0
Use the workload keywords in the output:
check_workloads kind=daemonset "detail-syntax=%(kind) %(namespace)/%(name): %(available)/%(desired) available, %(updated) updated, %(missing) missing" "top-syntax=${status}: ${list}" ok-syntax=
WARNING: DaemonSet monitoring/node-exporter: 3/4 available, 4 updated, 1 missing|'monitoring/node-exporter available'=3;0;0 'monitoring/node-exporter desired'=4;0;0 'monitoring/node-exporter missing'=1;0;0
Command-line Arguments¶
| Option | Default Value | Description |
|---|---|---|
| namespace | Only list workloads in this namespace (repeatable). Default: all namespaces. | |
| kind | Only check this workload kind: deployment, statefulset or daemonset (repeatable). Default: all three. | |
| label-selector | Label selector passed to the API server, e.g. app.kubernetes.io/part-of=shop. | |
| field-selector | Field selector passed to the API server, e.g. metadata.name=web. | |
| workload | Name of a workload that must exist, as name or namespace/name (repeatable). Only the named workloads take part in the check; one the API server does not return is reported with 0 available of 1 desired. | |
| timeout | 30 | Timeout in seconds for each network step of an API server request: the connect, the TLS handshake and each read. It is not a total for the check (a server that keeps trickling data takes longer). A positive number. |
timeout:
Timeout in seconds for each network step of an API server request: the connect, the TLS handshake and each read. It is not a total for the check (a server that keeps trickling data takes longer). A positive number.
Default Value: 30
Common options:
These options are shared by all filter based commands and are described on the common options page; the default values below are specific to this command.
| Option | Default Value |
|---|---|
| filter | |
| warning | missing > 0 or updated < desired |
| warn | |
| critical | available = 0 and desired > 0 |
| crit | |
| ok | |
| debug | false |
| show-all | false |
| empty-state | ok |
| perf-config | |
| escape-html | false |
| list-separator | , |
| top-syntax | ${status}: ${problem_list} |
| ok-syntax | %(status): All %(count) workloads are available |
| empty-syntax | %(status): No workloads found |
| detail-syntax | ${kind} ${namespace}/${name}=${available}/${desired} |
| perf-syntax | ${namespace}/${name} |
| byte-unit | |
| decimal-separator | |
| decimals | -1 |
| thousands-separator |
This command also accepts the standard help options: help, help-pb, show-default, help-short.
Filter keywords¶
| Option | Description |
|---|---|
| age | Seconds since the workload was created, -1 when unknown (supports units, e.g. age < 1h) |
| available | Replicas available (Ready for at least minReadySeconds) |
| created | When the workload was created (date) |
| desired | Replicas wanted (spec.replicas, or the nodes a daemonset should run on) |
| kind | Workload kind: Deployment, StatefulSet or DaemonSet |
| labels | Workload labels as key=value, comma separated |
| missing | desired minus available, never below 0 |
| name | Workload name |
| namespace | Namespace the workload lives in |
| paused | 1 when a deployment’s rollout is paused, else 0 |
| ready | Replicas whose pod is Ready |
| unavailable | Replicas the controller reports unavailable |
| updated | Replicas running the current template (less than desired during a rollout) |
This command also supports the common filter keywords: count, total, ok_count, warn_count, crit_count, problem_count, list, ok_list, warn_list, crit_list, problem_list, detail_list, sep, status.
Configuration¶
| Path / Section | Description |
|---|---|
| /settings/kubernetes |
/settings/kubernetes ¶
| Key | Default Value | Description |
|---|---|---|
| api server | API SERVER | |
| ca | ${ca-path} | CERTIFICATE AUTHORITY |
| context | KUBECONFIG CONTEXT | |
| kubeconfig | KUBECONFIG (JSON) | |
| max response size | 64 | MAX RESPONSE SIZE |
| timeout | 30 | TIMEOUT |
| tls version | tlsv1.2+ | TLS VERSION |
| token | BEARER TOKEN | |
| token file | TOKEN FILE | |
| verify mode | peer | TLS PEER VERIFY MODE |
#
[/settings/kubernetes]
ca=${ca-path}
max response size=64
timeout=30
tls version=tlsv1.2+
verify mode=peer
API SERVER ¶
The Kubernetes API server, e.g. https://k8s.example.com:6443. Leave empty to auto-detect when the agent runs inside the cluster (KUBERNETES_SERVICE_HOST plus the mounted service account token and CA). Which cluster the agent talks to is an operator decision: a check request cannot choose another server. Use https: an http:// server receives the bearer token in cleartext, and the agent logs a warning when it connects to one.
| Key | Description |
|---|---|
| Path: | /settings/kubernetes |
| Key: | api server |
| Default value: | N/A |
Sample:
[/settings/kubernetes]
# API SERVER
api server=
CERTIFICATE AUTHORITY ¶
CA bundle (a PEM file or a hashed directory) used to verify the API server certificate. Defaults to the trusted system store (in-cluster: the mounted service account CA, unless this is set); point it at the cluster CA (`kubectl config view –raw -o jsonpath=’{.clusters[0].cluster.certificate-authority-data}’ | base64 -d`) for a private cluster CA.
| Key | Description |
|---|---|
| Path: | /settings/kubernetes |
| Key: | ca |
| Default value: | ${ca-path} |
Sample:
[/settings/kubernetes]
# CERTIFICATE AUTHORITY
ca=${ca-path}
KUBECONFIG CONTEXT ¶
The kubeconfig context to use; empty means its current-context.
| Key | Description |
|---|---|
| Path: | /settings/kubernetes |
| Key: | context |
| Advanced: | Yes (means it is not commonly used) |
| Default value: | N/A |
Sample:
[/settings/kubernetes]
# KUBECONFIG CONTEXT
context=
KUBECONFIG (JSON) ¶
Path to a kubeconfig in JSON form (`kubectl config view –raw –minify -o json > nscp-kubeconfig.json`), used when `api server` is empty. YAML kubeconfigs are not read. Supplies the server, CA, token or client certificate of the selected context.
| Key | Description |
|---|---|
| Path: | /settings/kubernetes |
| Key: | kubeconfig |
| Advanced: | Yes (means it is not commonly used) |
| Default value: | N/A |
Sample:
[/settings/kubernetes]
# KUBECONFIG (JSON)
kubeconfig=
MAX RESPONSE SIZE ¶
Largest API response the agent will buffer, in megabytes. A pod list in a large cluster runs to tens of megabytes; 0 removes the cap.
| Key | Description |
|---|---|
| Path: | /settings/kubernetes |
| Key: | max response size |
| Advanced: | Yes (means it is not commonly used) |
| Default value: | 64 |
Sample:
[/settings/kubernetes]
# MAX RESPONSE SIZE
max response size=64
TIMEOUT ¶
Timeout in seconds for each network step of an API server request: the connect, the TLS handshake and each read. It bounds a stalled server, not the whole check: a server that keeps trickling data, or a list of many pages, takes longer. Must be positive: the checks refuse 0 or less rather than wait forever.
| Key | Description |
|---|---|
| Path: | /settings/kubernetes |
| Key: | timeout |
| Advanced: | Yes (means it is not commonly used) |
| Default value: | 30 |
Sample:
[/settings/kubernetes]
# TIMEOUT
timeout=30
TLS VERSION ¶
Minimum TLS protocol version accepted: tlsv1.0, tlsv1.1, tlsv1.2, tlsv1.2+ (the default: TLS 1.2 and 1.3), tlsv1.3.
| Key | Description |
|---|---|
| Path: | /settings/kubernetes |
| Key: | tls version |
| Advanced: | Yes (means it is not commonly used) |
| Default value: | tlsv1.2+ |
Sample:
[/settings/kubernetes]
# TLS VERSION
tls version=tlsv1.2+
BEARER TOKEN ¶
Service account bearer token used to authenticate against the API server. Prefer `token file` for a token that is rotated.
| Key | Description |
|---|---|
| Path: | /settings/kubernetes |
| Key: | token |
| Default value: | N/A |
Sample:
[/settings/kubernetes]
# BEARER TOKEN
token=
TOKEN FILE ¶
Path to a file holding the bearer token, read at check time so a rotated (projected) token keeps working without a reload. Takes precedence over `token`.
| Key | Description |
|---|---|
| Path: | /settings/kubernetes |
| Key: | token file |
| Default value: | N/A |
Sample:
[/settings/kubernetes]
# TOKEN FILE
token file=
TLS PEER VERIFY MODE ¶
Comma separated list of options: none, peer (or certificate), peer-cert, fail-if-no-cert (or fail-if-no-peer-cert, client-certificate). For a self signed certificate use peer-cert and point `ca` at that certificate; none disables verification entirely and sends the token to an unverified peer.
| Key | Description |
|---|---|
| Path: | /settings/kubernetes |
| Key: | verify mode |
| Default value: | peer |
Sample:
[/settings/kubernetes]
# TLS PEER VERIFY MODE
verify mode=peer