New to KubeDB? Please start here.

Kafka Alerting with Prometheus

This tutorial shows you how to configure Prometheus-based alerting for a KubeDB-managed Kafka cluster using the kafka-alerts Helm chart. This chart also bundles a Grafana dashboard that it imports automatically through a post-install Job — no separate dashboard chart is required.

Before You Begin

  • Ensure you have a Kubernetes cluster and that kubectl is configured to communicate with it. If you do not already have a cluster, you can create one using kind.

  • Install the KubeDB operator by following the steps here.

  • Deploy the database in the alert-kafka namespace:

    $ kubectl create ns alert-kafka
    namespace/alert-kafka created
    
  • To learn more about how Prometheus monitoring works with KubeDB, see the overview here.

  • You will also need a Grafana API key / token with Editor permission so the chart’s dashboard-import Job can push the dashboard. See Step 1 below.

Note: YAML files used in this tutorial are stored in docs/examples/kafka folder in GitHub repository kubedb/docs.

Configuration

Step 1 (kube-prometheus-stack) is required to follow this tutorial. Step 2 (Panopticon) is required for the Provisioner Group alerts below (KubeDBKafkaPhase...) — skip it only if you just want the exporter-based Database Group alerts. If you have already completed the step(s) you need in another guide, skip ahead.

Step 1: Deploy kube-prometheus-stack

kube-prometheus-stack installs Prometheus, Prometheus Operator, Alertmanager, and Grafana together. This is the recommended way to get the full monitoring stack on Kubernetes.

Add the prometheus-community Helm repo and install:

$ helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
$ helm repo update

$ helm upgrade --install prometheus prometheus-community/kube-prometheus-stack \
  --namespace monitoring --create-namespace \
  --set grafana.image.tag=7.5.5

Wait for all pods to be ready:

$ kubectl get pods -n monitoring
NAME                                                   READY   STATUS    RESTARTS   AGE
alertmanager-prometheus-kube-prometheus-alertmanager-0 2/2     Running   0          2m
prometheus-grafana-xxxx                                3/3     Running   0          2m
prometheus-kube-prometheus-operator-xxxx               1/1     Running   0          2m
prometheus-kube-prometheus-prometheus-0                2/2     Running   0          2m
prometheus-kube-state-metrics-xxxx                     1/1     Running   0          2m

Find the serviceMonitorSelector/ruleSelector labels that Prometheus uses to pick up ServiceMonitor/PrometheusRule objects — this is the release: prometheus label used throughout this tutorial.

$ kubectl get prometheus -n monitoring -o jsonpath='{.items[0].spec.ruleSelector}'
{"matchLabels":{"release":"prometheus"}}

$ kubectl get prometheus -n monitoring -o jsonpath='{.items[0].spec.serviceMonitorSelector}'
{"matchLabels":{"release":"prometheus"}}

Step 2: Install Panopticon (required for the Provisioner Group alerts)

Panopticon is the Appscode operator that exports the KubeDB operator’s own view of every resource — kubedb_com_kafka_status_phase and related metrics. It’s what powers the Provisioner Group alerts below (KubeDBKafkaPhaseNotReady/KubeDBKafkaPhaseCritical). Skip this step if you only need the exporter-based Database Group alerts.

$ helm repo add appscode https://charts.appscode.com/stable/
$ helm repo update

$ helm upgrade --install panopticon appscode/panopticon \
  --version v2026.4.30 \
  --namespace kubeops --create-namespace \
  --set monitoring.enabled=true \
  --set monitoring.agent=prometheus.io/operator \
  --set monitoring.serviceMonitor.labels.release=prometheus \
  --set-file license=/path/to/kubedb-license.txt \
  --wait --timeout 5m0s

Verify Panopticon is running:

$ kubectl get pods -n kubeops
NAME                          READY   STATUS    RESTARTS   AGE
panopticon-xxxx               1/1     Running   0          1m

Overview

Kafka Alerting Architecture

  • KubeDB deploys Kafka with metrics exposed by a JMX Exporter running as a Java agent inside the kafka container itself — not a separate sidecar container. KubeDB uses the JMX agent because the officially recognized Kafka exporter image does not yet expose metrics for the KRaft-mode versions KubeDB supports.
  • ServiceMonitor (named {kafka-name}-stats) is created automatically by KubeDB and tells Prometheus to scrape the JMX agent’s HTTP endpoint every 10 seconds.
  • PrometheusRule is created by the kafka-alerts chart and contains alert definitions grouped by concern: database health (which also embeds KubeDB-operator-sourced KafkaDown/KafkaPhaseCritical alerts) and provisioner.
  • Dashboard-import Job — when grafana.enabled is true, the chart also creates a one-shot Job that POSTs a bundled dashboard JSON straight to your Grafana instance’s /api/dashboards/import endpoint.
  • Prometheus Operator evaluates every rule expression every 30 seconds and fires matching alerts to AlertManager.
  • AlertManager groups, inhibits, and silences alerts, then routes them to configured receivers (Slack, email, PagerDuty, webhook, etc.).

Deploy Kafka with Monitoring Enabled

Below is a single-broker Kafka object for this tutorial (a production cluster would use spec.topology for separate broker/controller roles).

apiVersion: kubedb.com/v1
kind: Kafka
metadata:
  name: kafka-alert-demo
  namespace: alert-kafka
spec:
  replicas: 1
  version: "3.9.0"
  storageType: Durable
  storage:
    storageClassName: "local-path"
    accessModes:
      - ReadWriteOnce
    resources:
      requests:
        storage: 1Gi
  deletionPolicy: WipeOut
  monitor:
    agent: prometheus.io/operator
    prometheus:
      serviceMonitor:
        labels:
          release: prometheus
        interval: 10s

Here,

  • spec.replicas: 1 deploys a single-broker Kafka instance for this tutorial.
  • spec.monitor.agent: prometheus.io/operator tells KubeDB to create a ServiceMonitor resource managed by the Prometheus operator.
  • spec.monitor.prometheus.serviceMonitor.labels.release: prometheus adds the release: prometheus label to the created ServiceMonitor, matching the Prometheus serviceMonitorSelector so the target is discovered automatically.

Let’s create the Kafka resource.

$ kubectl apply -f https://github.com/kubedb/docs/raw/v2026.7.10/docs/examples/kafka/monitoring/kafka-alert-demo.yaml
kafka.kubedb.com/kafka-alert-demo created

Wait for the cluster to go into Ready state.

$ kubectl get kafka -n alert-kafka kafka-alert-demo
NAME                VERSION   STATUS   AGE
kafka-alert-demo    3.9.0     Ready    3m

KubeDB creates a dedicated stats service with the -stats suffix for monitoring.

$ kubectl get svc -n alert-kafka --selector="app.kubernetes.io/instance=kafka-alert-demo"
NAME                        TYPE        CLUSTER-IP     EXTERNAL-IP   PORT(S)     AGE
kafka-alert-demo            ClusterIP   10.43.10.20    <none>        9092/TCP    3m
kafka-alert-demo-pods       ClusterIP   None           <none>        9092/TCP    3m
kafka-alert-demo-stats      ClusterIP   10.43.10.21    <none>        56790/TCP   3m

KubeDB also creates a ServiceMonitor that tells Prometheus where to scrape.

$ kubectl get servicemonitor -n alert-kafka
NAME                     AGE
kafka-alert-demo-stats   3m

Verify that the ServiceMonitor carries the release: prometheus label so Prometheus discovers it.

$ kubectl get servicemonitor -n alert-kafka kafka-alert-demo-stats \
    -o jsonpath='{.metadata.labels.release}'
prometheus

Step 1 — Create a Grafana API Key

The chart’s dashboard-import Job authenticates to Grafana with a bearer token, so create one first.

  • Grafana 9+: Administration → Service accounts → Add service account → role EditorAdd token. Copy the token.

  • Grafana 8.x and earlier (no Service Accounts UI, e.g. the bundled kube-prometheus-stack Grafana 7.5.5): use the legacy API Keys endpoint instead:

    # Port-forward Grafana
    $ kubectl port-forward -n monitoring svc/prometheus-grafana 3000:80&
    
    # Retrieve the admin password
    $ kubectl get secret -n monitoring prometheus-grafana \
        -o jsonpath='{.data.admin-password}' | base64 -d && echo
    
    # Create an API key with Editor role
    $ curl -s -X POST -H "Content-Type: application/json" \
        -u admin:<grafana_password> \
        http://localhost:3000/api/auth/keys \
        -d '{"name":"kafka-alerts-demo","role":"Editor"}'
    # Note the returned "key"
    
    # Stop the port-forward
    $ kill %1
    

Either way, you end up with a bearer token to use as grafana.apikey below.

Step 2 — Install kafka-alerts

Why the Helm release name matters

The chart derives the PrometheusRule name and scopes every PromQL expression from the Helm release name — so the release name must match the Kafka object’s name (kafka-alert-demo).

Install

$ helm upgrade -i kafka-alert-demo oci://ghcr.io/appscode-charts/kafka-alerts \
    -n alert-kafka \
    --create-namespace \
    --version=v2026.7.14 \
    --set form.alert.labels.release=prometheus \
    --set grafana.enabled=true \
    --set grafana.url="http://prometheus-grafana.monitoring.svc:80" \
    --set grafana.apikey="<token-from-above>" \
    --set form.alert.appSuffix=kf-grafana-demo
FlagValuePurpose
kafka-alert-demo (release name)Scopes every PromQL expression to this instance (job="kafka-alert-demo-stats"). This must exactly match the Kafka object’s name.
form.alert.labels.releaseprometheusMatches the Prometheus ruleSelector so the rules are loaded
grafana.urlin-cluster Grafana URLThe dashboard-import Job runs inside the cluster, so this must be a cluster-internal address, not localhost
grafana.apikeytoken from Step 1Authenticates the dashboard-import POST request

To install alerts only, without the dashboard, omit the grafana.* flags (or set --set grafana.enabled=false).

Verify the PrometheusRule is created

$ kubectl get prometheusrule -n alert-kafka
NAME                 AGE
kafka-alert-demo     30s

Confirm the release: prometheus label is present.

$ kubectl get prometheusrule -n alert-kafka kafka-alert-demo \
    -o jsonpath='{.metadata.labels.release}'
prometheus

Verify the dashboard-import Job

$ kubectl get job -n alert-kafka
NAME                        STATUS     COMPLETIONS   AGE
kafka-alert-demo-post-job   Complete   1/1           17s

$ kubectl logs -n alert-kafka job/kafka-alert-demo-post-job
{"pluginId":"","title":"kubedb.com / Kafka / alert-kafka / kafka-alert-demo","imported":true, ...}

A "imported":true response confirms the dashboard kubedb.com / Kafka / alert-kafka / kafka-alert-demo now exists in Grafana.

Confirm Prometheus loaded the rules

$ kubectl port-forward -n monitoring \
    svc/prometheus-kube-prometheus-prometheus 9090:9090

Open http://localhost:9090/rules and locate the kafka.database and kafka.provisioner groups.

Prometheus Rule Health

Both groups should show OK. kafka-alerts v2026.7.14 has no opsManager/stash/kubeStash groups — only database and provisioner.

Note the overlap: the database group’s KafkaDown (for: 30s) and KafkaPhaseCritical (for: 3m) key off the same kubedb_com_kafka_status_phase metric as the provisioner group’s KubeDBKafkaPhaseNotReady/KubeDBKafkaPhaseCritical (for: 1m/15m) — expect both pairs to eventually fire together during a real outage, at different times.


Verify End-to-End

1. Check the Prometheus target is UP

Open http://localhost:9090/query?g0.expr=up%7Bnamespace%3D%22alert-kafka%22%7D&g0.tab=1.

Prometheus up query — kafka-alert-demo-0 UP

2. Confirm all Kafka alerts are inactive

Open http://localhost:9090/alerts.

Prometheus Alerts — Kafka groups inactive

All rules should show INACTIVE. KafkaTopicCount and the replication-related alerts (KafkaUnderReplicatedPartitions, KafkaUnderMinIsrPartitionCount, KafkaISRExpandRate/KafkaISRShrinkRate) are naturally quiet on a single-broker cluster with no topics yet.

3. Check AlertManager

$ kubectl port-forward -n monitoring \
    svc/prometheus-kube-prometheus-alertmanager 9093:9093

Open http://localhost:9093.

AlertManager


Simulating a Firing Alert

This section deliberately triggers KafkaDown (for: 30s, the fastest down-signal) by crashing the main Kafka JVM process.

1. Crash the Kafka process

$ kubectl exec -n alert-kafka kafka-alert-demo-0 -c kafka -- sh -c '
    end=$(( $(date +%s) + 60 ));
    while [ $(date +%s) -lt $end ]; do
      pid=$(pgrep -f "kafka.Kafka" | head -1);
      [ -n "$pid" ] && kill -9 "$pid" 2>/dev/null;
      sleep 1;
    done'

2. Watch the alert fire in Prometheus

Open http://localhost:9090/alerts.

Prometheus Alerts — KafkaDown Firing

KafkaDown (kubedb_com_kafka_status_phase{phase!="Ready"} == 1, for: 30s) should transition to FIRING first.

3. Check the AlertManager dashboard

Open http://localhost:9093.

AlertManager — KafkaDown Firing

4. Restore Kafka

Stop the loop from step 1.

$ kubectl get kafka -n alert-kafka kafka-alert-demo -w
NAME               VERSION   STATUS   AGE
kafka-alert-demo   3.9.0     Ready    24m

If Kafka does not recover on its own within a minute or two, force a clean restart: kubectl delete pod -n alert-kafka kafka-alert-demo-0.


Alert Reference

All alerts are scoped to the kafka-alert-demo instance in the alert-kafka namespace via job="kafka-alert-demo-stats" / namespace="alert-kafka" (database group), or app="kafka-alert-demo" / namespace="alert-kafka" (provisioner group and the two operator-phase alerts embedded in the database group).

Database Group

AlertSeverityForWhat It Means
KafkaUnderReplicatedPartitionswarning10sPartitions have fewer in-sync replicas than expected.
KafkaAbnormalControllerStatewarning10sMore or fewer than one active controller in the cluster.
KafkaOfflinePartitionswarning10sOne or more partitions have no leader.
KafkaUnderMinIsrPartitionCountwarning10sPartitions below the minimum in-sync-replica count.
KafkaOfflineLogDirectoryCountwarning10sA log directory has gone offline (likely a disk issue).
KafkaISRExpandRatewarning1mISR set is expanding frequently — sign of flakiness.
KafkaISRShrinkRatewarning1mISR set is shrinking frequently — sign of flakiness.
KafkaBrokerCountcritical1mBroker count has dropped.
KafkaNetworkProcessorIdlePercentcritical1mNetwork processor threads are saturated.
KafkaRequestHandlerIdlePercentcritical1mRequest handler threads are saturated.
KafkaReplicaFetcherManagerMaxLagcritical1mReplica fetcher lag is high.
KafkaTopicCountwarning1mTopic count changed unexpectedly.
KafkaPhaseCriticalwarning3mKubeDB operator view: resource Critical (duplicates the provisioner group’s own version at a different for).
KafkaDowncritical30sKubeDB operator view: resource not Ready. Fastest down-signal available.
DiskUsageHighwarning1mPersistent volume usage exceeds 80%.
DiskAlmostFullcritical1mPersistent volume usage exceeds 95%.

Provisioner Group

AlertSeverityForWhat It Means
KubeDBKafkaPhaseNotReadycritical1mKubeDB marked the Kafka resource NotReady.
KubeDBKafkaPhaseCriticalwarning15mKafka is degraded but not fully unavailable.

Customising Alerts

# custom-alerts.yaml
form:
  alert:
    labels:
      release: prometheus
    groups:
      database:
        enabled: warning
        rules:
          kafkaBrokerCount:
            enabled: true
            duration: "2m"
            severity: critical
$ helm upgrade kafka-alert-demo oci://ghcr.io/appscode-charts/kafka-alerts \
    -n alert-kafka \
    --version=v2026.7.14 \
    -f custom-alerts.yaml

Cleaning up

To remove all resources created in this tutorial, run the following commands.

# Remove the kafka-alerts release (PrometheusRule + dashboard-import Job)
$ helm uninstall kafka-alert-demo -n alert-kafka

# Remove the imported Grafana dashboard (it is not removed by helm uninstall)
$ curl -s -X DELETE -H "Authorization: Bearer <grafana-token>" \
    http://localhost:3000/api/dashboards/uid/<uid>

$ kubectl delete kafka -n alert-kafka kafka-alert-demo
$ kubectl delete ns alert-kafka

# Uninstall monitoring stack (optional — skip if other tutorials on this cluster still need them)
$ helm uninstall panopticon -n kubeops
$ helm uninstall prometheus -n monitoring

Next Steps