Configure Weaviate Monitoring using Prometheus Mode in Applications Manager

Configure Weaviate Monitoring using Prometheus Mode in Applications Manager

Overview

Applications Manager monitors Weaviate only through Prometheus. Weaviate exposes its own metrics on port 2112, Prometheus scrapes them, and Applications Manager queries Prometheus over its HTTP API on every poll.

Each Weaviate monitor is discovered from the object_count metric. By default, Applications Manager creates one monitor per job, instance and endpoint combination, and filters every query by those three labels. In a multi-node Weaviate cluster, this gives one monitor per node. On Kubernetes, the instance value is the pod IP, so it changes on every restart and the monitor stops matching.

The Data Collection Filter Label (custom label) solves this. You attach one label with a fixed value to the Weaviate scrape job. Applications Manager then discovers one monitor per label value, filters every query by that label alone, and sums the metrics across all nodes that carry it.

Prerequisites

  • A running Weaviate instance or cluster, deployed with Docker, Docker Compose or Kubernetes (Helm).
  • A Prometheus server that can reach port 2112 on every Weaviate node.
  • Network access from the Applications Manager server to the Prometheus API (default port 9090). For SSH mode, the SSH host must reach Prometheus and have curl installed.
  • At least one collection (class) in Weaviate. Discovery runs on object_count, which Weaviate reports only for existing collections, so an empty instance is not discovered.
  • For the async indexing queue metrics, Weaviate must run with ASYNC_INDEXING=true. Without it, the queue_* metrics stay empty.

Step 1: Enable the Weaviate metrics endpoint

Weaviate exposes Prometheus metrics at http://<weaviate-host>:2112/metrics once PROMETHEUS_MONITORING_ENABLED is set to true. Set it on every Weaviate node.

Docker Compose

Add the variable to the Weaviate service and publish port 2112:

services:
weaviate:
image: cr.weaviate.io/semitechnologies/weaviate:<version>
ports:
- "8080:8080"
- "50051:50051"
- "2112:2112"
environment:
PROMETHEUS_MONITORING_ENABLED: 'true'
ASYNC_INDEXING: 'true' # optional, needed for the queue metrics

Then restart the service:

docker compose up -d weaviate

Docker

docker run -d --name weaviate -p 8080:8080 -p 50051:50051 -p 2112:2112 -e PROMETHEUS_MONITORING_ENABLED=true cr.weaviate.io/semitechnologies/weaviate:<version>

Kubernetes (Helm)

Set the variable in the Weaviate Helm values and upgrade the release:

env:
PROMETHEUS_MONITORING_ENABLED: true
helm upgrade weaviate weaviate/weaviate -n weaviate -f values.yaml

Check the endpoint

curl -s http://<weaviate-host>:2112/metrics | grep -E '^(object_count|go_memstats_sys_bytes)'

The output should contain object_count lines, one per collection and shard, and a go_memstats_sys_bytes line.

Note: Do not set PROMETHEUS_MONITORING_GROUP=true. It merges all collections and shards into a single series, so the collection and shard counts in Applications Manager show 1.

Step 2: Configure the Prometheus scrape job with the custom label

Add a weaviate scrape job to Prometheus and attach the custom label to it. Every node of one Weaviate cluster must carry the same label value.

How Applications Manager uses the label

  • Discovery runs sum by (job, instance, endpoint, <label>) (object_count). Each distinct value of the label becomes a separate Weaviate monitor.
  • Data collection filters every query by <label>="<value>" alone, instead of job, instance and endpoint, and sums the results across nodes. So a node that lacks the label, or carries a different value, is not counted.
SettingRuleExample
Label nameStarts with a letter or _; then letters, digits or _; up to 100 characters. Default is apm_filter.apm_filter
Label valueLetters, digits, space and - . _ | : / @; up to 100 characters. Use a different value for each Weaviate cluster.prod-weaviate-01

2a. Docker or VM hosts (static targets)

List every Weaviate node under targets and set the label under labels:

scrape_configs:
- job_name: 'weaviate'
metrics_path: /metrics
static_configs:
- targets:
- '<weaviate-node1-ip>:2112'
- '<weaviate-node2-ip>:2112'
- '<weaviate-node3-ip>:2112'
labels:
apm_filter: prod-weaviate-01

2b. Kubernetes (pod discovery)

This job finds the Weaviate pods by their app=weaviate label, scrapes port 2112 on each pod and adds the custom label:

scrape_configs:
- job_name: 'weaviate'
metrics_path: /metrics
kubernetes_sd_configs:
- role: pod
namespaces:
names: [weaviate]
relabel_configs:
- source_labels: [__meta_kubernetes_pod_label_app]
action: keep
regex: weaviate
- source_labels: [__meta_kubernetes_pod_ip]
target_label: __address__
replacement: '${1}:2112'
- source_labels: [__meta_kubernetes_pod_name]
target_label: pod
- target_label: apm_filter
replacement: prod-weaviate-01

Prometheus needs get, list and watch on pods in the weaviate namespace for this job.

2c. Prometheus Operator or kube-prometheus-stack

Create a PodMonitor instead of editing prometheus.yml:

apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: weaviate
namespace: weaviate
spec:
selector:
matchLabels:
app: weaviate
podMetricsEndpoints:
- targetPort: 2112
path: /metrics
relabelings:
- action: replace
targetLabel: apm_filter
replacement: prod-weaviate-01

The PodMonitor must match the podMonitorSelector of your Prometheus resource, for example a release: <prometheus-release> label.

Reload Prometheus

After editing prometheus.yml, reload Prometheus:

curl -X POST http://<prometheus-host>:9090/-/reload

This needs Prometheus to run with --web.enable-lifecycle; otherwise restart it.

Note: Do not use global.external_labels for this label. External labels are added only to federated, remote-write and alert data, so the Prometheus query API that Applications Manager calls never returns them.

Step 3: Add the Prometheus integration in Applications Manager

  1. In Applications Manager, go to Settings → Product Settings → Prometheus and click Add New. If an integration for this Prometheus server already exists, edit it instead and add Weaviate to its monitor types.
  2. Enter the General Specifications:
    • Display Name: a name for the integration, for example Prod Weaviate Prometheus.
    • Polling Interval: in minutes; the minimum is 5.
    • Mode of Data collection: Exposed API if Prometheus is reachable over the network, or SSH if it is reachable only on localhost of a Linux host.
  3. Enter the Prometheus Server Details: Host Name / IP Address, Port (default 9090), Endpoint, Protocol (HTTP or HTTPS) and the authentication type (No Authentication or Basic Authentication).
  4. Under Discovery specifications:
    • Monitor Types: select Weaviate.
    • Data Collection Filter Label: enter the label name you added in Step 2, not its value. Leave the default apm_filter if you used that name.
    • Alert action if monitor data not available: choose the action to take when a monitor's data is missing.
  5. Click Save. Applications Manager connects to Prometheus and runs discovery.
  6. Choose All Instances, or Select Instances for monitoring and pick the Weaviate clusters you want. Each discovered instance corresponds to one label value, such as prod-weaviate-01.

Note: If you leave the Data Collection Filter Label empty, Applications Manager uses apm_filter. If object_count does not carry the configured label at all, Applications Manager falls back to matching job, instance and endpoint, and creates one monitor per Weaviate node. After you add or change the label in Prometheus, run discovery again from the integration so the monitor picks up the new value.

Verification and troubleshooting

Confirm the target and the label in Prometheus before you add the integration. Open http://<prometheus-host>:9090, go to Status → Targets, and check that every Weaviate node under the weaviate job is UP.

Run these queries in the Prometheus Graph page. Each should return data:

sum by (job, instance, endpoint, apm_filter) (object_count)
count by (instance) (go_memstats_sys_bytes{apm_filter="prod-weaviate-01"})
sum(max by (class_name, shard_name) (object_count{apm_filter="prod-weaviate-01"}))
sum(rate(queries_durations_ms_count{apm_filter="prod-weaviate-01"}[5m]))

The first query should show one row per node, all with the same apm_filter value. The second should list every Weaviate node. The query rate is 0 until the instance receives traffic.

IssueCheckFix
No Weaviate monitor discoveredsum by (job, instance, endpoint, apm_filter) (object_count) returns nothingCreate at least one collection in Weaviate, and confirm the weaviate target is UP and Weaviate is selected under Monitor Types.
One monitor per node instead of one per clusterThe first query shows rows without apm_filterAdd the label to the scrape job as in Step 2, reload Prometheus and run discovery again.
Target shows connection refusedcurl http://<weaviate-host>:2112/metricsSet PROMETHEUS_MONITORING_ENABLED=true and publish or open port 2112.
Monitor stops collecting after a pod restartThe monitor was discovered before the label was added, so it matches the old pod IPAdd the label, run discovery again and remove the old monitor.
Collection and shard counts show 1object_count has class_name="n/a"Remove PROMETHEUS_MONITORING_GROUP=true from the Weaviate settings.
Queue metrics are emptyqueue_size returns nothingSet ASYNC_INDEXING=true in Weaviate if you use async indexing. Otherwise these metrics are expected to be empty.
Two clusters show up as one monitorBoth clusters use the same label valueGive each cluster a unique value, then run discovery again.
Wrong label name in Applications ManagerThe name in Data Collection Filter Label differs from the one in the scrape jobEdit the integration so the names match, then run discovery again.
Applications Manager cannot connectcurl http://<prometheus-host>:9090/api/v1/status/buildinfo from the Applications Manager serverOpen the Prometheus port in the firewall, or use SSH mode.