Configure Kubernetes Monitoring using Prometheus Mode in Applications Manager

Configure Kubernetes Monitoring using Prometheus Mode in Applications Manager

Overview

Applications Manager can monitor a Kubernetes cluster without an agent by reading its metrics from a Prometheus server. Prometheus scrapes the cluster exporters, and Applications Manager queries Prometheus over its HTTP API on every poll.

Each Kubernetes monitor is discovered from kube_node_info. By default, Applications Manager then collects data by matching the job, instance and endpoint labels of that series. This breaks when node-exporter, cAdvisor and kubelet metrics carry different job/instance values than kube-state-metrics, which is common in Rancher, kube-prometheus-stack and federated setups.

The Data Collection Filter Label (custom label) solves this. You attach one label, with the same value, to every exporter's metrics. Applications Manager then discovers one monitor per label value and filters every query by that label alone.

ExporterWhat it providesDefault port
kube-state-metricsNodes, pods, deployments, services, PVs, PVCs and their status8080
node-exporterNode filesystem size and free space9100
cAdvisor (kubelet /metrics/cadvisor)Container CPU, memory, throttling, network and filesystem usage10250
kubelet (/metrics)Kubelet version and volume stats10250

Prerequisites

  • A running Kubernetes cluster, version 1.22 or later.
  • kubectl configured against that cluster with cluster-admin rights, needed to create the ClusterRoles.
  • Either an existing Prometheus that already scrapes kube-state-metrics, node-exporter and the kubelet/cAdvisor, or permission to deploy one using the steps below.
  • Network access from the Applications Manager server to the Prometheus API (default port 9090, or the NodePort you expose). For SSH mode, the SSH host must reach Prometheus and have curl installed.
  • The following metric families must be present in Prometheus: kube_* (kube-state-metrics), node_filesystem_* (node-exporter), container_* (cAdvisor) and kubelet_* (kubelet).

Note: If Prometheus is already running and scraping these exporters, skip to Step 2 and only add the custom label.

Step 1: Deploy the exporters and Prometheus

This step deploys a plain Prometheus, with no Helm and no Operator, into a new monitoring namespace. Download the manifest files attached to this article and apply them in order from the Kubernetes control-plane node.

FileCreates
00-namespace.yamlNamespace monitoring
01-node-exporter.yamlServiceAccount, DaemonSet (host port 9100) and headless Service node-exporter
02-kube-state-metrics.yamlServiceAccount, ClusterRole, ClusterRoleBinding, Deployment and Service kube-state-metrics (port 8080)
03-prometheus-rbac.yamlServiceAccount prometheus with read access to nodes, nodes/proxy, nodes/metrics, services, endpoints and pods
04-prometheus-config.yamlConfigMap prometheus-config holding prometheus.yml (edited in Step 2)
05-prometheus-deployment.yamlPrometheus Deployment and NodePort Service (port 30090)
  1. Create the namespace and deploy the exporters:
    kubectl apply -f 00-namespace.yaml
    kubectl apply -f 01-node-exporter.yaml
    kubectl apply -f 02-kube-state-metrics.yaml
  2. Create the Prometheus RBAC. Prometheus reads kubelet and cAdvisor metrics through the API server proxy, so it needs no direct access to port 10250 on the nodes:
    kubectl apply -f 03-prometheus-rbac.yaml
  3. The attached 04-prometheus-config.yaml already includes the custom label. Change its value to your cluster name as described in Step 2, then create the ConfigMap and deploy Prometheus:
    kubectl apply -f 04-prometheus-config.yaml
    kubectl apply -f 05-prometheus-deployment.yaml
  4. Wait for the rollouts to finish:
    kubectl -n monitoring rollout status deploy/kube-state-metrics --timeout=180s
    kubectl -n monitoring rollout status ds/node-exporter --timeout=180s
    kubectl -n monitoring rollout status deploy/prometheus --timeout=180s
    kubectl -n monitoring get pods,svc

Note: The Prometheus pod uses emptyDir storage, so data is lost on restart. For production, replace the storage volume in 05-prometheus-deployment.yaml with a PersistentVolumeClaim.

Step 2: Add the custom label to every scrape job

Add one label, with one fixed value per cluster, to every job that scrapes a Kubernetes exporter. Applications Manager uses this label to tie the kube-state-metrics, node-exporter, cAdvisor and kubelet series of one cluster together.

How Applications Manager uses the label

  • Discovery runs sum by (job, instance, endpoint, <label>) (kube_node_info). Each distinct value of the label becomes a separate Kubernetes monitor.
  • Data collection filters every query by <label>="<value>" alone, instead of job, instance and endpoint. So a series that lacks the label, or carries a different value, is not collected.
SettingRuleExample
Label nameStarts with a letter or _; then letters, digits or _; up to 100 characters. Default is apm_filter.apm_filter
Label valueLetters, digits, space and - . _ | : / @; up to 100 characters. Use a different value for each cluster.prod-cluster-01

2a. Plain Prometheus (the setup from Step 1)

The attached 04-prometheus-config.yaml already includes the label. The two lines below close the relabel_configs of each of the five Kubernetes jobs: kubernetes-nodes, kubernetes-cadvisor, kubernetes-node-resource, kube-state-metrics and node-exporter. Change prod-cluster-01 to your cluster name in all five places.

          - target_label: apm_filter
replacement: prod-cluster-01

The complete prometheus.yml section of the ConfigMap then reads:

global:
scrape_interval: 30s
scrape_timeout: 15s
evaluation_interval: 30s

scrape_configs:

- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']

# Kubelet /metrics (via API server proxy)
- job_name: 'kubernetes-nodes'
scheme: https
tls_config:
ca_file: /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
insecure_skip_verify: true
bearer_token_file: /var/run/secrets/kubernetes.io/serviceaccount/token
kubernetes_sd_configs:
- role: node
relabel_configs:
- action: labelmap
regex: __meta_kubernetes_node_label_(.+)
- target_label: __address__
replacement: kubernetes.default.svc:443
- source_labels: [__meta_kubernetes_node_name]
regex: (.+)
target_label: __metrics_path__
replacement: /api/v1/nodes/${1}/proxy/metrics
- source_labels: [__meta_kubernetes_node_name]
target_label: instance
- target_label: apm_filter
replacement: prod-cluster-01

# cAdvisor container metrics
- job_name: 'kubernetes-cadvisor'
scheme: https
tls_config:
ca_file: /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
insecure_skip_verify: true
bearer_token_file: /var/run/secrets/kubernetes.io/serviceaccount/token
kubernetes_sd_configs:
- role: node
relabel_configs:
- action: labelmap
regex: __meta_kubernetes_node_label_(.+)
- target_label: __address__
replacement: kubernetes.default.svc:443
- source_labels: [__meta_kubernetes_node_name]
regex: (.+)
target_label: __metrics_path__
replacement: /api/v1/nodes/${1}/proxy/metrics/cadvisor
- source_labels: [__meta_kubernetes_node_name]
target_label: instance
- target_label: apm_filter
replacement: prod-cluster-01

# Kubelet resource metrics
- job_name: 'kubernetes-node-resource'
scheme: https
tls_config:
ca_file: /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
insecure_skip_verify: true
bearer_token_file: /var/run/secrets/kubernetes.io/serviceaccount/token
kubernetes_sd_configs:
- role: node
relabel_configs:
- action: labelmap
regex: __meta_kubernetes_node_label_(.+)
- target_label: __address__
replacement: kubernetes.default.svc:443
- source_labels: [__meta_kubernetes_node_name]
regex: (.+)
target_label: __metrics_path__
replacement: /api/v1/nodes/${1}/proxy/metrics/resource
- source_labels: [__meta_kubernetes_node_name]
target_label: instance
- target_label: apm_filter
replacement: prod-cluster-01

# kube-state-metrics
- job_name: 'kube-state-metrics'
honor_labels: true
kubernetes_sd_configs:
- role: endpoints
namespaces:
names: [monitoring]
relabel_configs:
- source_labels: [__meta_kubernetes_service_name, __meta_kubernetes_endpoint_port_name]
action: keep
regex: kube-state-metrics;http-metrics
- source_labels: [__meta_kubernetes_pod_node_name]
target_label: instance
- target_label: apm_filter
replacement: prod-cluster-01

# node-exporter
- job_name: 'node-exporter'
kubernetes_sd_configs:
- role: endpoints
namespaces:
names: [monitoring]
relabel_configs:
- source_labels: [__meta_kubernetes_service_name, __meta_kubernetes_endpoint_port_name]
action: keep
regex: node-exporter;metrics
- source_labels: [__meta_kubernetes_pod_node_name]
target_label: instance
- target_label: apm_filter
replacement: prod-cluster-01

If Prometheus is already running, apply the change and reload it:

kubectl -n monitoring apply -f 04-prometheus-config.yaml
kubectl -n monitoring exec deploy/prometheus -- wget -qO- --post-data='' http://localhost:9090/-/reload

2b. Prometheus Operator, kube-prometheus-stack or Rancher Monitoring

These stacks generate scrape jobs from ServiceMonitors, so edit each ServiceMonitor (or the matching Helm values) instead of prometheus.yml. Add the relabeling to the kube-state-metrics, node-exporter, kubelet and cAdvisor endpoints:

spec:
endpoints:
- port: http-metrics
relabelings:
- action: replace
targetLabel: apm_filter
replacement: prod-cluster-01

Note: Do not use global.external_labels for this label. External labels are added only to federated, remote-write and alert data, so the Prometheus query API that Applications Manager calls never returns them.

Note: Rancher's default cAdvisor relabel rules drop container_cpu_cfs_throttled_seconds_total and node-level container_fs_* series. Remove those metrics from the drop rules, or the related container metrics stay empty in Applications Manager.

Monitoring several clusters from one Prometheus

If one Prometheus, Thanos or federated server holds metrics for several clusters, give each cluster its own value, for example prod-cluster-01 and qa-cluster-01. Discovery then creates one Kubernetes monitor per value, and each monitor collects only its own cluster's series.

Step 3: Add the Prometheus integration in Applications Manager

  1. Get the Prometheus address. The Service from Step 1 is already a NodePort on port 30090:
    kubectl get nodes -o wide
    kubectl -n monitoring get svc prometheus
    Prometheus is then reachable at http://<node-ip>:30090. For a Service of another type, convert it with kubectl -n monitoring patch svc prometheus -p '{"spec":{"type":"NodePort"}}'.
  2. In Applications Manager, go to Settings → Product Settings → Prometheus and click Add New.
  3. Enter the General Specifications:
    • Display Name: a name for the integration, for example Prod Kubernetes Prometheus.
    • Polling Interval: in minutes; the minimum is 5.
    • Mode of Data collection: Exposed API if Prometheus is reachable over the network, or SSH if it is reachable only on localhost of a Linux host.
  4. Enter the Prometheus Server Details: Host Name / IP Address (<node-ip>), Port (30090), Endpoint, Protocol (HTTP or HTTPS) and the authentication type (No Authentication, Basic Authentication or Service Account token).
  5. Under Discovery specifications:
    • Monitor Types: select Kubernetes.
    • Data Collection Filter Label: enter the label name you added in Step 2, not its value. Leave the default apm_filter if you used that name.
    • Alert action if monitor data not available: choose the action to take when a monitor's data is missing.
  6. Click Save. Applications Manager connects to Prometheus and runs discovery.
  7. Choose All Instances, or Select Instances for monitoring and pick the clusters you want. Each discovered instance corresponds to one label value, such as prod-cluster-01.

Note: If you leave the Data Collection Filter Label empty, Applications Manager uses apm_filter. If kube_node_info does not carry the configured label at all, Applications Manager falls back to matching job, instance and endpoint. After you add or change the label in Prometheus, run discovery again from the integration so the monitor picks up the new value.

Verification and troubleshooting

Confirm the targets and the label in Prometheus before you add the integration. Open http://<node-ip>:30090, go to Status → Targets, and check that these jobs are UP: kubernetes-nodes, kubernetes-cadvisor, kubernetes-node-resource, kube-state-metrics and node-exporter.

Run these queries in the Prometheus Graph page. Each should return data, and each series should carry apm_filter="prod-cluster-01":

sum by (job, instance, endpoint, apm_filter) (kube_node_info)
count by (job) ({apm_filter="prod-cluster-01"})
container_cpu_usage_seconds_total{apm_filter="prod-cluster-01"}
node_filesystem_size_bytes{apm_filter="prod-cluster-01"}
kubelet_volume_stats_used_bytes{apm_filter="prod-cluster-01"}

The second query should list all five Kubernetes jobs. A missing job means its relabel_configs lacks the label.

IssueCheckFix
No Kubernetes monitor discoveredsum by (job, instance, endpoint, apm_filter) (kube_node_info) returns nothingCheck kube-state-metrics is UP and its job carries the label. Confirm Kubernetes is selected under Monitor Types.
Monitor discovered but node or container metrics are emptycount by (job) ({apm_filter="<value>"}) misses node-exporter, kubernetes-cadvisor or kubernetes-nodesAdd the label, with the same value, to the missing jobs and reload Prometheus.
Two clusters show up as one monitorBoth clusters use the same label valueGive each cluster a unique value, then run discovery again.
Wrong label name in Applications ManagerThe name in Data Collection Filter Label differs from the one in prometheus.ymlEdit the integration so the names match, then run discovery again.
Kubelet or cAdvisor targets show 403 Forbiddenkubectl auth can-i get nodes/proxy --as=system:serviceaccount:monitoring:prometheusRe-apply 03-prometheus-rbac.yaml, or bind the system:kubelet-api-admin ClusterRole to the monitoring:prometheus ServiceAccount.
kube-state-metrics shows 0 serieskubectl -n monitoring logs deploy/kube-state-metricsRe-apply 02-kube-state-metrics.yaml and check its ClusterRoleBinding.
Applications Manager cannot connectcurl http://<node-ip>:30090/api/v1/status/buildinfo from the Applications Manager serverOpen the NodePort in the firewall, or use SSH mode.
Some container metrics are missing on RancherQuery container_cpu_cfs_throttled_seconds_total and container_fs_usage_bytesRemove them from Rancher's cAdvisor drop rules (see Step 2b).