New to KubeDB? Please start here.

Hazelcast Alerting with Prometheus

This tutorial shows you how to configure Prometheus-based alerting for a KubeDB-managed Hazelcast instance using the hazelcast-alerts Helm chart. This chart also bundles a Grafana dashboard that it imports automatically through a post-install Job — no separate dashboard chart is required.

Before You Begin

  • Ensure you have a Kubernetes cluster and that kubectl is configured to communicate with it. If you do not already have a cluster, you can create one using kind.

  • Install the KubeDB operator by following the steps here.

  • Deploy the database in the alert-hazelcast namespace:

    $ kubectl create ns alert-hazelcast
    namespace/alert-hazelcast created
    
  • Hazelcast requires an Enterprise license. Create a secret with your license key before deploying:

    $ kubectl create secret generic hz-license-key -n alert-hazelcast \
        --from-literal=licenseKey='your hazelcast license key'
    secret/hz-license-key created
    
  • To learn more about how Prometheus monitoring works with KubeDB, see the overview here.

  • You will also need a Grafana API key / token with Editor permission so the chart’s dashboard-import Job can push the dashboard. See Step 1 below.

Note: YAML files used in this tutorial are stored in docs/examples/hazelcast folder in GitHub repository kubedb/docs.

Configuration

Step 1 (kube-prometheus-stack) is required to follow this tutorial. Step 2 (Panopticon) is required for the Provisioner Group alerts below (KubeDBHazelcastPhase...) — skip it only if you just want the exporter-based Database Group alerts. If you have already completed the step(s) you need in another guide, skip ahead.

Step 1: Deploy kube-prometheus-stack

kube-prometheus-stack installs Prometheus, Prometheus Operator, Alertmanager, and Grafana together. This is the recommended way to get the full monitoring stack on Kubernetes.

Add the prometheus-community Helm repo and install:

$ helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
$ helm repo update

$ helm upgrade --install prometheus prometheus-community/kube-prometheus-stack \
  --namespace monitoring --create-namespace \
  --set grafana.image.tag=7.5.5

Wait for all pods to be ready:

$ kubectl get pods -n monitoring
NAME                                                   READY   STATUS    RESTARTS   AGE
alertmanager-prometheus-kube-prometheus-alertmanager-0 2/2     Running   0          2m
prometheus-grafana-xxxx                                3/3     Running   0          2m
prometheus-kube-prometheus-operator-xxxx               1/1     Running   0          2m
prometheus-kube-prometheus-prometheus-0                2/2     Running   0          2m
prometheus-kube-state-metrics-xxxx                     1/1     Running   0          2m

Find the serviceMonitorSelector/ruleSelector labels that Prometheus uses to pick up ServiceMonitor/PrometheusRule objects — this is the release: prometheus label used throughout this tutorial.

$ kubectl get prometheus -n monitoring -o jsonpath='{.items[0].spec.ruleSelector}'
{"matchLabels":{"release":"prometheus"}}

$ kubectl get prometheus -n monitoring -o jsonpath='{.items[0].spec.serviceMonitorSelector}'
{"matchLabels":{"release":"prometheus"}}

Step 2: Install Panopticon (required for the Provisioner Group alerts)

Panopticon is the Appscode operator that exports the KubeDB operator’s own view of every resource — kubedb_com_hazelcast_status_phase and related metrics. It’s what powers the Provisioner Group alerts below (KubeDBHazelcastPhaseNotReady/KubeDBHazelcastPhaseCritical). Skip this step if you only need the exporter-based Database Group alerts.

$ helm repo add appscode https://charts.appscode.com/stable/
$ helm repo update

$ helm upgrade --install panopticon appscode/panopticon \
  --version v2026.4.30 \
  --namespace kubeops --create-namespace \
  --set monitoring.enabled=true \
  --set monitoring.agent=prometheus.io/operator \
  --set monitoring.serviceMonitor.labels.release=prometheus \
  --set-file license=/path/to/kubedb-license.txt \
  --wait --timeout 5m0s

Verify Panopticon is running:

$ kubectl get pods -n kubeops
NAME                          READY   STATUS    RESTARTS   AGE
panopticon-xxxx               1/1     Running   0          1m

Overview

Hazelcast Alerting Architecture

  • KubeDB deploys Hazelcast with a metrics-exporter sidecar (container exporter) that exposes Hazelcast’s own JMX-derived metrics (com_hazelcast_Metrics_*).
  • ServiceMonitor (named {hazelcast-name}-stats) is created automatically by KubeDB and tells Prometheus to scrape the exporter every 10 seconds.
  • PrometheusRule is created by the hazelcast-alerts chart and contains alert definitions grouped by concern: database health (which also embeds KubeDB-operator-sourced hazelcastDown/hazelcastPhaseCritical alerts) and provisioner.
  • Dashboard-import Job — when grafana.enabled is true, the chart also creates a one-shot Job that POSTs a bundled dashboard JSON straight to your Grafana instance’s /api/dashboards/import endpoint.
  • Prometheus Operator evaluates every rule expression every 30 seconds and fires matching alerts to AlertManager.
  • AlertManager groups, inhibits, and silences alerts, then routes them to configured receivers (Slack, email, PagerDuty, webhook, etc.).

Deploy Hazelcast with Monitoring Enabled

At first, let’s deploy a single-node Hazelcast instance with monitoring enabled. Below is the Hazelcast object we are going to create.

apiVersion: kubedb.com/v1alpha2
kind: Hazelcast
metadata:
  name: hazelcast-alert-demo
  namespace: alert-hazelcast
spec:
  version: "5.5.2"
  replicas: 1
  licenseSecret:
    name: hz-license-key
  storage:
    storageClassName: "local-path"
    accessModes:
      - ReadWriteOnce
    resources:
      requests:
        storage: 1Gi
  deletionPolicy: WipeOut
  monitor:
    agent: prometheus.io/operator
    prometheus:
      serviceMonitor:
        labels:
          release: prometheus
        interval: 10s

Here,

  • spec.replicas: 1 creates a single-node Hazelcast instance.
  • spec.licenseSecret.name: hz-license-key points to the Enterprise license secret created in Before You Begin.
  • spec.monitor.agent: prometheus.io/operator tells KubeDB to create a ServiceMonitor resource managed by the Prometheus operator.
  • spec.monitor.prometheus.serviceMonitor.labels.release: prometheus adds the release: prometheus label to the created ServiceMonitor, matching the Prometheus serviceMonitorSelector so the target is discovered automatically.

Let’s create the Hazelcast resource.

$ kubectl apply -f https://github.com/kubedb/docs/raw/v2026.7.10/docs/examples/hazelcast/monitoring/hazelcast-alert-demo.yaml
hazelcast.kubedb.com/hazelcast-alert-demo created

Wait for the database to go into Ready state.

$ kubectl get hazelcast -n alert-hazelcast hazelcast-alert-demo
NAME                    VERSION   STATUS   AGE
hazelcast-alert-demo    5.5.2     Ready    3m

KubeDB creates a dedicated stats service with the -stats suffix for monitoring.

$ kubectl get svc -n alert-hazelcast --selector="app.kubernetes.io/instance=hazelcast-alert-demo"
NAME                              TYPE        CLUSTER-IP     EXTERNAL-IP   PORT(S)     AGE
hazelcast-alert-demo              ClusterIP   10.43.10.20    <none>        5701/TCP    3m
hazelcast-alert-demo-pods         ClusterIP   None           <none>        5701/TCP    3m
hazelcast-alert-demo-stats        ClusterIP   10.43.10.21    <none>        8080/TCP    3m

KubeDB also creates a ServiceMonitor that tells Prometheus where to scrape.

$ kubectl get servicemonitor -n alert-hazelcast
NAME                           AGE
hazelcast-alert-demo-stats     3m

Verify that the ServiceMonitor carries the release: prometheus label so Prometheus discovers it.

$ kubectl get servicemonitor -n alert-hazelcast hazelcast-alert-demo-stats \
    -o jsonpath='{.metadata.labels.release}'
prometheus

Step 1 — Create a Grafana API Key

The chart’s dashboard-import Job authenticates to Grafana with a bearer token, so create one first.

  • Grafana 9+: Administration → Service accounts → Add service account → role EditorAdd token. Copy the token.

  • Grafana 8.x and earlier (no Service Accounts UI, e.g. the bundled kube-prometheus-stack Grafana 7.5.5): use the legacy API Keys endpoint instead:

    # Port-forward Grafana
    $ kubectl port-forward -n monitoring svc/prometheus-grafana 3000:80&
    
    # Retrieve the admin password
    $ kubectl get secret -n monitoring prometheus-grafana \
        -o jsonpath='{.data.admin-password}' | base64 -d && echo
    
    # Create an API key with Editor role
    $ curl -s -X POST -H "Content-Type: application/json" \
        -u admin:<grafana_password> \
        http://localhost:3000/api/auth/keys \
        -d '{"name":"hazelcast-alerts-demo","role":"Editor"}'
    # Note the returned "key"
    
    # Stop the port-forward
    $ kill %1
    

Either way, you end up with a bearer token to use as grafana.apikey below.

Step 2 — Install hazelcast-alerts

Why the Helm release name matters

The chart derives the PrometheusRule name and scopes every PromQL expression from the Helm release name — so the release name must match the Hazelcast object’s name (hazelcast-alert-demo).

Install

$ helm upgrade -i hazelcast-alert-demo oci://ghcr.io/appscode-charts/hazelcast-alerts \
    -n alert-hazelcast \
    --create-namespace \
    --version=v2026.7.14 \
    --set form.alert.labels.release=prometheus \
    --set grafana.enabled=true \
    --set grafana.url="http://prometheus-grafana.monitoring.svc:80" \
    --set grafana.apikey="<token-from-above>" \
    --set grafana.jobName=hazelcast-alert-demo-stats \
    --set form.alert.appSuffix=hz-grafana-demo
FlagValuePurpose
grafana.urlin-cluster Grafana URLThe dashboard-import Job runs inside the cluster, so this must be a cluster-internal address, not localhost
grafana.apikeytoken from Step 1Authenticates the dashboard-import POST request
grafana.jobNamehazelcast-alert-demo-statsRequired — the chart’s default (kubedb-databases) doesn’t match any real Prometheus job, so most of the dashboard’s panels show “No data” unless you override it to your instance’s actual stats-service name

To install alerts only, without the dashboard, omit the grafana.* flags (or set --set grafana.enabled=false).

Verify the PrometheusRule is created

$ kubectl get prometheusrule -n alert-hazelcast
NAME                     AGE
hazelcast-alert-demo     30s

$ kubectl get prometheusrule -n alert-hazelcast hazelcast-alert-demo \
    -o jsonpath='{.metadata.labels.release}'
prometheus

Verify the dashboard-import Job

$ kubectl get job -n alert-hazelcast
NAME                            STATUS     COMPLETIONS   AGE
hazelcast-alert-demo-post-job   Complete   1/1           17s

$ kubectl logs -n alert-hazelcast job/hazelcast-alert-demo-post-job
{"pluginId":"","title":"kubedb.com / Hazelcast / alert-hazelcast / hazelcast-alert-demo","imported":true, ...}

A "imported":true response confirms the dashboard kubedb.com / Hazelcast / alert-hazelcast / hazelcast-alert-demo now exists in Grafana.

Confirm Prometheus loaded the rules

$ kubectl port-forward -n monitoring \
    svc/prometheus-kube-prometheus-prometheus 9090:9090

Open http://localhost:9090/rules and locate the hazelcast.database and hazelcast.provisioner groups.

Prometheus Rule Health

Both groups should show OK. hazelcast-alerts v2026.7.14 has no opsManager/stash/kubeStash groups at all — only database and provisioner.

Note the overlap: the database group’s hazelcastDown (for: 30s) and hazelcastPhaseCritical (for: 3m) key off the exact same kubedb_com_hazelcast_status_phase metric as the provisioner group’s KubeDBhazelcastPhaseNotReady/KubeDBhazelcastPhaseCritical (for: 1m/15m) — a real outage fires both pairs of alerts (at different times, since the for windows differ). Worth knowing so you don’t mistake it for two independent problems.


Verify End-to-End

1. Check the Prometheus target is UP

Open http://localhost:9090/query?g0.expr=up%7Bnamespace%3D%22alert-hazelcast%22%7D&g0.tab=1.

Prometheus up query — hazelcast-alert-demo-0 UP

2. Confirm the Hazelcast alerts are inactive

Open http://localhost:9090/alerts.

Prometheus Alerts — Hazelcast groups inactive

All rules should show INACTIVE.

3. Check AlertManager

$ kubectl port-forward -n monitoring \
    svc/prometheus-kube-prometheus-alertmanager 9093:9093

Open http://localhost:9093.

AlertManager


Simulating a Firing Alert

This section deliberately triggers hazelcastDown (for: 30s, the fastest down-signal) by crashing the main Hazelcast JVM process.

1. Crash the Hazelcast process

$ kubectl exec -n alert-hazelcast hazelcast-alert-demo-0 -c hazelcast -- sh -c '
    end=$(( $(date +%s) + 60 ));
    while [ $(date +%s) -lt $end ]; do
      pid=$(pgrep -f "java.*hazelcast" | head -1);
      [ -n "$pid" ] && kill -9 "$pid" 2>/dev/null;
      sleep 1;
    done'

2. Watch the alert fire in Prometheus

Open http://localhost:9090/alerts.

Prometheus Alerts — hazelcastDown Firing

hazelcastDown (kubedb_com_hazelcast_status_phase{phase!="Ready"} == 1, for: 30s) should transition to FIRING first; if the crash loop runs long enough, KubeDBhazelcastPhaseNotReady (for: 1m, provisioner group) fires shortly after.

3. Check the AlertManager dashboard

Open http://localhost:9093.

AlertManager — hazelcastDown Firing

4. Restore Hazelcast

Stop the loop from step 1.

$ kubectl get hazelcast -n alert-hazelcast hazelcast-alert-demo -w
NAME                    VERSION   STATUS   AGE
hazelcast-alert-demo    5.5.2     Ready    24m

If Hazelcast does not recover on its own within a minute or two, force a clean restart: kubectl delete pod -n alert-hazelcast hazelcast-alert-demo-0.


Alert Reference

All alerts are scoped to the hazelcast-alert-demo instance in the alert-hazelcast namespace, mostly via namespace/service label filters matching $app-stats (database group), or app="hazelcast-alert-demo" / namespace="alert-hazelcast" (provisioner group and the two operator-phase alerts embedded in the database group).

Database Group

Fired based on live metrics from the Hazelcast exporter sidecar (JMX-derived) and, for the hazelcastDown/hazelcastPhaseCritical pair, the KubeDB operator’s own view of the resource phase.

AlertSeverityForWhat It Means
hazelcastPartitionCountExceedwarning30sActive partition count is unusually high.
hazelcastHighHeapPercentagewarning30sJVM heap usage is high.
hazelcastHighMemoryUsagewarning30sHazelcast memory usage is high.
hazelcastHighPhysicalMemoryUsagewarning30sPhysical memory usage is high relative to total.
hazelcastHighLatencywarning30sGet-operation latency is elevated.
hazelcastSystemCPULoadExceedwarning30sSystem CPU load is high.
hazelcastPhaseCriticalwarning3mKubeDB operator view: resource Critical (duplicates the provisioner group’s own version at a different for).
hazelcastDowncritical30sKubeDB operator view: resource not Ready. Fastest down-signal available.
DiskUsageHighwarning1mPersistent volume usage exceeds 80%.
DiskAlmostFullcritical1mPersistent volume usage exceeds 95%.

Provisioner Group

Monitors the KubeDB operator’s view of the Hazelcast resource phase (sourced from Panopticon, not the Hazelcast metrics endpoint).

AlertSeverityForWhat It Means
KubeDBhazelcastPhaseNotReadycritical1mKubeDB marked the Hazelcast resource NotReady.
KubeDBhazelcastPhaseCriticalwarning15mHazelcast is degraded but not fully unavailable.

Customising Alerts

To override thresholds or disable specific alert groups, create a custom values file and upgrade the chart.

# custom-alerts.yaml
form:
  alert:
    labels:
      release: prometheus
    groups:
      database:
        enabled: warning
        rules:
          hazelcastHighHeapPercentage:
            enabled: true
            duration: "2m"
            severity: warning
$ helm upgrade hazelcast-alert-demo oci://ghcr.io/appscode-charts/hazelcast-alerts \
    -n alert-hazelcast \
    --version=v2026.7.14 \
    -f custom-alerts.yaml

Cleaning up

To remove all resources created in this tutorial, run the following commands.

# Remove the hazelcast-alerts release (PrometheusRule + dashboard-import Job)
$ helm uninstall hazelcast-alert-demo -n alert-hazelcast

# Remove the imported Grafana dashboard (it is not removed by helm uninstall)
$ curl -s -X DELETE -H "Authorization: Bearer <grafana-token>" \
    http://localhost:3000/api/dashboards/uid/<uid>

$ kubectl delete hazelcast -n alert-hazelcast hazelcast-alert-demo
$ kubectl delete secret -n alert-hazelcast hz-license-key
$ kubectl delete ns alert-hazelcast

# Uninstall monitoring stack (optional — skip if other tutorials on this cluster still need them)
$ helm uninstall panopticon -n kubeops
$ helm uninstall prometheus -n monitoring

Next Steps