Skip to content

Repository files navigation

JSM Home Assistant Notifier

CI GHCR image License: Apache 2.0

A lightweight Docker service that bridges Jira Service Management (JSM / OpsGenie) alerts to Home Assistant — with smart on-call routing, escalation detection, rich TTS announcements, and persistent dashboard notifications.

Jump to Quick Start ↓


Contents


How It Works

JSM alert created / escalated
         │
         ▼
  jsm-ha-notifier (Docker)
         │
         ├─ Parse alert payload
         ├─ Deduplicate (suppress retries within 60 s)
         ├─ Route decision:
         │    always_notify mode?   → NOTIFY
         │    escalated to me?      → NOTIFY
         │    I'm on-call?          → NOTIFY   (JSM API, cached 5 min)
         │    none of the above     → DROP
         │
         ▼
  Home Assistant REST API
    ├─ media_player.play_media  (TTS with rich metadata / real alert title)
    └─ persistent_notification  (visible in HA dashboard)

  On Acknowledge / Close → dismiss the persistent notification automatically

Two webhook URLs, one for each routing mode:

JSM Webhook URL Behaviour
https://your-host:8080/alert?key=YOUR_KEY Notify only when on-call
https://your-host:8080/alert?mode=always&key=YOUR_KEY Always notify regardless of schedule

Features

  • On-call aware — queries JSM in real time and caches results; only wakes you when you are actually on-call
  • Escalation detection — EscalateNext events always notify regardless of on-call status or dedup window
  • Always-notify mode — a separate webhook path for schedules that should always page you (e.g. infrastructure monitors)
  • Rich TTS — spoken announcements include priority, alert title, system name, and a description excerpt
  • Real media player title — uses extra.metadata so HA shows the actual alert title instead of "Playing Default Media Receiver"
  • Persistent HA notifications — created on alert, auto-dismissed on Acknowledge or Close
  • Configurable announcement formats — customise the detailed and terse TTS templates with placeholders
  • Time-based quiet hours — silent windows (no TTS) and terse windows (short format), with cross-midnight support
  • Priority override for silent mode — P1/P2 alerts can bypass silent windows so critical incidents always wake you
  • After-hours suppression — weekday-aware business hours (Mon-Fri 09:00-17:00) keep low-priority alerts from speaking outside office hours, including weekends. Fills the gap left by your alerting platform's quiet hours, which do not apply to outgoing webhooks — see Quiet Hours & After-Hours Suppression
  • Per-media-player routing — route TTS to different speakers by time of day (e.g. bedroom at night, office during the day)
  • Volume control — set media player volume before TTS playback, with separate levels for full and terse modes
  • Alert batching — combine multiple alerts arriving within a configurable window into one TTS announcement
  • TTS repeat (pager mode) — repeat TTS at intervals for critical alerts until acknowledged or max repeats hit
  • Acknowledge from HA — POST /alert/{id}/acknowledge endpoint lets HA automations ack alerts without opening JSM
  • Token health check — daily background job verifies the Atlassian API token; fires a HA TTS warning if expired (TTS suppressed during quiet hours; persistent notification still created)
  • Deep health check — GET /healthz verifies both JSM and HA API connectivity (returns 503 if either fails)
  • Startup connectivity checks — verifies JSM and HA reachability at boot, logs warnings if unreachable
  • HA automation webhooks — fire HA webhook triggers on Create, Escalate, Acknowledge, Close, Update, and SLA Breach events to control lights, scenes, scripts
  • Incident state dashboard — optional SQLite-backed GET /incidents API with status/priority filters, summary endpoint, and Grafana JSON datasource compatibility
  • JSM incident sync — optional background task to poll JSM for open alerts and keep the incident dashboard current
  • Emoji toggle — ENABLE_EMOJIS=false strips all emojis from notifications, metadata, and incoming alert text
  • Generic webhook support — any system that sends HTTP POST (Grafana, Uptime Kuma, shell scripts, HA automations) can trigger HA alerts
  • API key authentication — optional API key via query parameter (?key=), HTTP header (X-API-Key), or URL path prefix (/KEY/endpoint)
  • Webhook signature verification — optional HMAC-SHA256 validation via X-Hub-Signature-256
  • Request body size limit — rejects payloads over 1 MB to prevent memory exhaustion
  • Safe format templates — user-configurable announcement formats use a restricted formatter that blocks attribute/index access
  • Prometheus metrics — GET /metrics exposes alert counters, credential check stats, rate limit hits, and uptime for Grafana/Prometheus dashboards
  • Structured JSON logging — LOG_FORMAT=json for Datadog, Loki, CloudWatch, ELK; default is human-readable text
  • Hot config reload — POST /reload re-reads .env and applies changes without container restart
  • Per-IP rate limiting — 60 requests/minute on /alert to prevent webhook abuse
  • Secure container — non-root user, read-only filesystem, tmpfs at /tmp, localhost-only port binding

Prerequisites

  • Docker + Docker Compose
  • A server or device accessible from the internet (or from JSM's webhook delivery IPs)
  • Home Assistant with a Long-Lived Access Token and a TTS service configured
  • An Atlassian API token with access to JSM Ops (OpsGenie) schedules

Quick Start

git clone https://github.com/RealDougEubanks/JSM-HomeAssistant-Notifier.git
cd JSM-HomeAssistant-Notifier

cp .env.example .env
# Edit .env and fill in all required values (see Configuration below)

docker compose up -d
docker compose logs -f

Verify the service is running:

curl http://localhost:8080/health
# {"status":"ok"}

Local build testing

To build the Docker image locally instead of pulling from GHCR:

docker build -t ghcr.io/realdougeubanks/jsm-ha-notifier:latest .
docker compose up -d
docker compose logs -f

This tags the local build with the same image name the compose file expects.

If you've made code changes and Docker serves a cached layer, force a clean rebuild:

docker build --no-cache -t ghcr.io/realdougeubanks/jsm-ha-notifier:latest .
docker compose up -d
docker compose logs -f

Note: .env changes do not require a rebuild — just restart the container with docker compose up -d.


Configuration

Step 1 — Copy and edit .env

cp .env.example .env

Open .env and fill in each value. The file is fully commented with instructions for finding each value. The sections below expand on the key ones.

Step 2 — Find your Atlassian Cloud ID

Your Cloud ID is a UUID that identifies your Atlassian organisation. Retrieve it with:

curl -s -u "you@yourcompany.com:YOUR_API_TOKEN" \
  https://your-org.atlassian.net/_edge/tenant_info \
  | python3 -m json.tool

Look for the "cloudId" field. Copy it into JSM_CLOUD_ID.

Step 3 — Find your Atlassian Account ID

Your account ID (JSM_MY_USER_ID) is the UUID Atlassian uses internally for your user. The easiest way to find it:

curl -s -u "you@yourcompany.com:YOUR_API_TOKEN" \
  "https://api.atlassian.com/jsm/ops/api/YOUR_CLOUD_ID/v1/schedules/YOUR_SCHEDULE_ID/on-calls" \
  | python3 -m json.tool

Find your name in the onCallParticipants array; the "id" field is your account ID.

Step 4 — Find your exact schedule names

Schedule names are case-sensitive. List all schedules visible to your token:

curl -s -u "you@yourcompany.com:YOUR_API_TOKEN" \
  "https://api.atlassian.com/jsm/ops/api/YOUR_CLOUD_ID/v1/schedules" \
  | python3 -m json.tool | grep '"name"'

Copy the exact names into ALWAYS_NOTIFY_SCHEDULE_NAMES and/or CHECK_ONCALL_SCHEDULE_NAMES in .env.

Step 5 — Create a Home Assistant Long-Lived Access Token

  1. In Home Assistant, click your profile picture (bottom-left)
  2. Scroll to Security → Long-Lived Access Tokens
  3. Click Create token, give it a descriptive name (e.g. JSM Notifier)
  4. Copy the token into HA_TOKEN in .env — it is only shown once

Step 6 — Find your media player entity ID

In Home Assistant go to Developer Tools → States, filter by media_player. Copy the entity_id (e.g. media_player.living_room) into HA_MEDIA_PLAYER_ENTITY.

Step 7 — Verify everything works

Once the container is running, check on-call status directly:

curl http://localhost:8080/status | python3 -m json.tool

You should see your schedules listed and an on_call field. If a schedule shows "error": "not found", the name in .env doesn't match — compare carefully against the output of the schedule listing curl above.


Making the Service Externally Accessible

JSM's servers need to reach your webhook URL over the internet.

WARNING: Do not expose this container directly to the internet. The service runs plain HTTP without TLS. All traffic — including API keys, webhook payloads, and authentication tokens — is transmitted in cleartext. Always place a TLS-terminating proxy or tunnel in front of the service. Direct internet exposure risks credential interception, replay attacks, and unauthorized access to your Home Assistant instance.

Option 1 — Cloudflare Tunnel (recommended — no open ports)

A Cloudflare Tunnel creates an encrypted outbound connection from your network to Cloudflare's edge, with no inbound ports to open on your router or firewall. Cloudflare handles TLS termination and DDoS protection automatically.

For setup instructions, see the Cloudflare Tunnel documentation.

Option 2 — Reverse proxy with TLS (NGINX / Traefik / Caddy)

Run a TLS-terminating reverse proxy on the same host and forward traffic to http://127.0.0.1:8080. Detailed reverse proxy configuration is outside the scope of this README — consult your proxy's documentation for TLS certificate setup (e.g. Let's Encrypt via Certbot or Caddy's automatic HTTPS).


JSM Webhook Configuration

Configure two outgoing webhooks in JSM Ops — one for on-call schedules and one for always-notify schedules.

Go to JSM Ops Settings

JSM project → Settings → Integrations → Add Integration → choose Webhook (under "Outgoing").

Webhook for On-Call Schedule(s)

Field Value
Name HA Notifier — On-Call
Webhook URL https://your-host/alert?key=YOUR_API_KEY
Method POST
Send alert payload ✅ Enabled
Alert actions Create, EscalateNext, Acknowledge, Close
Teams / Schedules filter Your on-call schedule's team

Webhook for Always-Notify Schedule(s)

Field Value
Name HA Notifier — Always Notify
Webhook URL https://your-host/alert?mode=always&key=YOUR_API_KEY
Method POST
Send alert payload ✅ Enabled
Alert actions Create, EscalateNext, Acknowledge, Close
Teams / Schedules filter Your always-notify team/schedule

Optional — API Key Authentication (recommended)

The simplest way to secure your webhook endpoints. Set WEBHOOK_API_KEY in .env and pass the key using any of these methods:

Method Example Best for
Query parameter https://your-host/alert?key=YOUR_KEY JSM webhooks (URL-only config)
Path prefix https://your-host/YOUR_KEY/alert Tools that can't add headers or query params
HTTP header X-API-Key: YOUR_KEY Scripts, HA automations, Grafana

All three methods work on every authenticated endpoint. Generate a key: openssl rand -hex 32

Requests without a valid key receive a stealth 404 (not 401), so attackers cannot confirm that authenticated endpoints exist.

⚠️ Keys in URLs leak into logs

The query-parameter and path-prefix methods exist because JSM webhooks can only be configured with a bare URL — but any URL-borne secret ends up in reverse-proxy access logs, browser history, and intermediate proxies. Prefer the X-API-Key header wherever the caller supports it (scripts, HA automations, Grafana), and if you front this service with a reverse proxy:

  • Scrub or disable access logging for the key. nginx example:

    # Redact the key query param / path segment before logging
    map $request_uri $loggable_uri {
        ~^(?<pre>.*[?&]key=)[^&]+(?<post>.*)$  "${pre}REDACTED${post}";
        default                                 $request_uri;
    }
    log_format scrubbed '$remote_addr - [$time_local] "$request_method $loggable_uri" $status';
    access_log /var/log/nginx/access.log scrubbed;

    Caddy: add log { format filter { request>uri query { replace key REDACTED } } } to the site block.

  • Rotate WEBHOOK_API_KEY (and update the JSM webhook URLs) if a proxy has already logged it.

The service itself never logs request URLs, and uvicorn access logging is disabled in the container.

Optional — HMAC Webhook Signature

If set, WEBHOOK_SECRET requires every POST /alert to carry an X-Hub-Signature-256: sha256=<hex> header containing the HMAC-SHA256 of the raw request body keyed with the secret.

⚠️ JSM cannot sign requests itself. JSM / OpsGenie outgoing-webhook custom headers are static strings — there is no template function to compute a per-request HMAC over the body. If JSM posts directly to this service, leave WEBHOOK_SECRET empty and rely on WEBHOOK_API_KEY.

Use WEBHOOK_SECRET when a signing-capable caller sits in front of the service:

Caller How
Forwarding proxy you control (Cloudflare Worker, AWS Lambda, nginx+Lua) Receive the JSM webhook, compute sha256= HMAC of the body, forward with the header
Your own scripts / automations Compute the HMAC at send time (e.g. openssl dgst -sha256 -hmac "$SECRET")

You can use both WEBHOOK_API_KEY and WEBHOOK_SECRET together for defense in depth when your caller supports signing.


Testing With curl

Send a test alert (on-call path)

curl -X POST http://localhost:8080/alert \
  -H "Content-Type: application/json" \
  -d '{
    "action": "Create",
    "alert": {
      "alertId": "test-001",
      "message": "Test Alert — please ignore",
      "priority": "P3",
      "entity": "dev-server",
      "description": "This is a test alert sent manually."
    }
  }'

If you are currently on-call, this will trigger a TTS announcement and create a persistent notification in HA.

Send a test alert (always-notify path)

curl -X POST "http://localhost:8080/alert?mode=always" \
  -H "Content-Type: application/json" \
  -d '{
    "action": "Create",
    "alert": {
      "alertId": "always-test-001",
      "message": "Infrastructure Monitor Test",
      "priority": "P2",
      "entity": "prod-server-01"
    }
  }'

This path always notifies regardless of on-call status.

Send a test escalation

curl -X POST "http://localhost:8080/alert?mode=always" \
  -H "Content-Type: application/json" \
  -d '{
    "action": "EscalateNext",
    "alert": {
      "alertId": "test-001",
      "message": "Test Alert — please ignore",
      "priority": "P1",
      "entity": "prod-db-01"
    }
  }'

Check on-call status

curl http://localhost:8080/status | python3 -m json.tool

Invalidate on-call cache

curl -X POST http://localhost:8080/cache/invalidate

Test with webhook signature

If WEBHOOK_SECRET is set, generate the signature before sending:

SECRET="your-webhook-secret"
BODY='{"action":"Create","alert":{"alertId":"sig-test","message":"Signed test","priority":"P3"}}'
SIG="sha256=$(echo -n "$BODY" | openssl dgst -sha256 -hmac "$SECRET" | awk '{print $2}')"

curl -X POST http://localhost:8080/alert \
  -H "Content-Type: application/json" \
  -H "X-Hub-Signature-256: $SIG" \
  -d "$BODY"

Using With Other Webhook Sources

The /alert endpoint accepts any JSON payload matching the OpsGenie webhook format. You don't need JSM — any monitoring system, script, or automation that can send HTTP POST requests can trigger HA alerts.

Required Payload Format

{
  "action": "Create",
  "alert": {
    "alertId": "unique-id-123",
    "message": "Your alert title here",
    "priority": "P1",
    "entity": "optional-system-name",
    "description": "Optional longer description text"
  }
}
Field Required Description
action Yes Create, EscalateNext, Acknowledge, or Close
alert.alertId Yes Unique identifier (used for dedup and notification tracking)
alert.message Yes Alert title / summary (spoken by TTS)
alert.priority No P1–P5 (default: P3)
alert.entity No System / host name
alert.description No Longer details (first 200 chars used in TTS)

Example: Uptime Kuma

Configure a webhook notification in Uptime Kuma with the Notification Type set to "Webhook" / custom JSON:

# Uptime Kuma → Settings → Notifications → Add → Webhook
# URL: http://your-notifier:8080/alert?mode=always&key=YOUR_KEY
# Method: POST
# Body:
{
  "action": "Create",
  "alert": {
    "alertId": "uptime-kuma-{{ monitorJSON.id }}",
    "message": "{{ monitorJSON.name }} is {{ heartbeatJSON.status == 1 ? 'UP' : 'DOWN' }}",
    "priority": "P2",
    "entity": "{{ monitorJSON.hostname }}"
  }
}

Example: Grafana Alerting

Use a Grafana "webhook" contact point with the OpsGenie payload format:

# Grafana → Alerting → Contact Points → New → Webhook
# URL: http://your-notifier:8080/alert?mode=always&key=YOUR_KEY
# Method: POST
#
# Or use curl to forward Grafana alerts via a script:
curl -X POST "http://your-notifier:8080/alert?mode=always&key=YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "action": "Create",
    "alert": {
      "alertId": "grafana-cpu-alert-prod01",
      "message": "CPU usage above 95% on prod-01",
      "priority": "P1",
      "entity": "prod-01",
      "description": "CPU has been above 95% for the last 5 minutes. Current: 98.2%."
    }
  }'

Example: Prometheus Alertmanager

Use Alertmanager's webhook receiver to POST to the notifier:

# alertmanager.yml
receivers:
  - name: ha-notifier
    webhook_configs:
      - url: "http://your-notifier:8080/alert?mode=always&key=YOUR_KEY"
        send_resolved: true

Then use a small relay script or Alertmanager template to transform alerts into the expected format.

Example: Home Assistant Automation

Trigger an alert from HA itself (e.g. a sensor threshold):

# HA automation action
service: rest_command.trigger_notifier_alert
data:
  alert_id: "ha-temp-alert-{{ now().isoformat() }}"
  message: "Temperature sensor above threshold"
  priority: "P2"
  entity: "sensor.living_room_temperature"
  description: "Current temperature: {{ states('sensor.living_room_temperature') }}°C"
# configuration.yaml
rest_command:
  trigger_notifier_alert:
    url: "http://your-notifier:8080/alert?mode=always&key=YOUR_KEY"
    method: POST
    content_type: "application/json"
    payload: >
      {"action":"Create","alert":{"alertId":"{{ alert_id }}","message":"{{ message }}","priority":"{{ priority }}","entity":"{{ entity }}","description":"{{ description }}"}}

Example: Simple Shell Script

Trigger an alert from any script or cron job:

#!/bin/bash
# notify-ha.sh — send an alert to the JSM-HA Notifier
NOTIFIER_URL="http://your-notifier:8080/alert?mode=always&key=YOUR_KEY"

curl -s -X POST "$NOTIFIER_URL" \
  -H "Content-Type: application/json" \
  -d "{
    \"action\": \"Create\",
    \"alert\": {
      \"alertId\": \"script-$(date +%s)\",
      \"message\": \"$1\",
      \"priority\": \"${2:-P3}\",
      \"entity\": \"$(hostname)\"
    }
  }"

Usage: ./notify-ha.sh "Backup failed on NAS" P2

Closing / Acknowledging Alerts

To dismiss the persistent HA notification and stop TTS repeats, send a Close or Acknowledge action with the same alertId:

curl -X POST "http://your-notifier:8080/alert?mode=always&key=YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"action": "Close", "alert": {"alertId": "the-original-alert-id", "message": "resolved"}}'

Or use the dedicated acknowledge endpoint:

curl -X POST "http://your-notifier:8080/alert/the-original-alert-id/acknowledge?key=YOUR_KEY"

JSM Webhook Events Reference

JSM sends different action values for each alert lifecycle event. The notifier handles all of them:

JSM Action Notifier Behaviour HA Webhook Config
Create TTS + persistent notification (if on-call/always-notify) HA_WEBHOOK_ON_CREATE
EscalateNext TTS + persistent notification (always if targeted at you) HA_WEBHOOK_ON_ESCALATE
Acknowledge Dismiss HA notification, cancel TTS repeat HA_WEBHOOK_ON_ACKNOWLEDGE
Close Dismiss HA notification, cancel TTS repeat HA_WEBHOOK_ON_CLOSE
AddNote Fire HA webhook only (no TTS) HA_WEBHOOK_ON_UPDATE
UnAcknowledge Fire HA webhook only HA_WEBHOOK_ON_UPDATE
AssignOwnership Fire HA webhook only HA_WEBHOOK_ON_UPDATE
SlaBreached Fire HA webhook only HA_WEBHOOK_ON_SLA_BREACH

JSM Webhook Configuration for All Events

In JSM, configure your outgoing webhook to include all action types you want to handle:

Field Value
Alert actions Create, EscalateNext, Acknowledge, Close, AddNote, UnAcknowledge, AssignOwnership
Webhook URL https://your-host/alert?mode=always&key=YOUR_API_KEY

Generic Webhook Payloads for Each Event Type

These work with JSM, Uptime Kuma, Grafana, scripts, or any HTTP client:

New alert:

{"action": "Create", "alert": {"alertId": "inc-001", "message": "Server down", "priority": "P1", "entity": "prod-01"}}

Escalation:

{"action": "EscalateNext", "alert": {"alertId": "inc-001", "message": "Server down", "priority": "P1", "entity": "prod-01"}}

Acknowledged:

{"action": "Acknowledge", "alert": {"alertId": "inc-001", "message": "Server down"}}

Resolved / Closed:

{"action": "Close", "alert": {"alertId": "inc-001", "message": "Server down"}}

Updated (note added):

{"action": "AddNote", "alert": {"alertId": "inc-001", "message": "Server down", "description": "Restarting services..."}}

SLA Breached:

{"action": "SlaBreached", "alert": {"alertId": "inc-001", "message": "Server down", "priority": "P1"}}

Complete Lifecycle Script

Send a full alert lifecycle from the command line for testing:

URL="http://localhost:8080/alert?mode=always&key=YOUR_KEY"
ID="test-lifecycle-$(date +%s)"

# Create
curl -s -X POST "$URL" -H "Content-Type: application/json" \
  -d "{\"action\":\"Create\",\"alert\":{\"alertId\":\"$ID\",\"message\":\"Test lifecycle alert\",\"priority\":\"P2\",\"entity\":\"test-server\"}}"

sleep 5

# Acknowledge
curl -s -X POST "$URL" -H "Content-Type: application/json" \
  -d "{\"action\":\"Acknowledge\",\"alert\":{\"alertId\":\"$ID\",\"message\":\"Test lifecycle alert\"}}"

sleep 5

# Close
curl -s -X POST "$URL" -H "Content-Type: application/json" \
  -d "{\"action\":\"Close\",\"alert\":{\"alertId\":\"$ID\",\"message\":\"Test lifecycle alert\"}}"

Quiet Hours & After-Hours Suppression

Read this before you go to production. By default this service announces every alert that reaches it, aloud, at full volume — including a P4 disk-space warning at 3am. If you are wired to a real on-call rotation, configure BUSINESS_HOURS_WINDOW below.

Why your alerting platform's quiet hours do not protect you

This is the single most common way people get burned by this project, so it is worth being explicit.

Most alerting platforms (JSM/Opsgenie, PagerDuty, Grafana OnCall) let you defer or suppress low-priority notifications outside office hours. Those rules apply to the platform's own delivery channels — mobile push, email, SMS.

Outgoing webhooks are not one of those channels. A webhook fires the instant the alert is created, before and independent of any notification policy, deferral, or quiet-hours rule.

The result is a trap:

 3:00am   P3 alert created
          ├─ platform notification policy → "defer to 09:00" → teammates sleep ✅
          └─ outgoing webhook            → fires immediately  → YOUR SPEAKER 🔊

Your whole team is protected by the platform rule. You are not, because you are the only one on the webhook. It looks like the alert is "only alerting here" — and it is, by design.

So the suppression has to be re-implemented on this side. That is what this section configures.

Configuration

Variable Default Purpose
BUSINESS_HOURS_WINDOW (empty — feature off) Master switch. Weekday-aware office hours, e.g. Mon-Fri 09:00-17:00
AFTER_HOURS_AUDIBLE_PRIORITIES P1,P2 Outside those hours, only these speak aloud
AFTER_HOURS_SILENT_TAGS (empty) Optional. Tags that force silence even for an audible priority

Minimal setup — this is all most people need:

TZ=America/New_York
BUSINESS_HOURS_WINDOW=Mon-Fri 09:00-17:00
AFTER_HOURS_AUDIBLE_PRIORITIES=P1,P2

Outside Mon–Fri 09:00–17:00, a P3/P4/P5 still posts the persistent HA notification and still updates the status light — it just makes no sound. Nothing is lost; you see it when you wake up. P1 and P2 still wake you, because they should.

TZ matters. Windows are evaluated in the container's local time, and the container defaults to UTC. Without TZ your "business hours" will be silently wrong by your UTC offset.

BUSINESS_HOURS_WINDOW vs SILENT_WINDOW — pick the right one

These look similar and are easy to confuse. The difference has bitten people.

SILENT_WINDOW BUSINESS_HOURS_WINDOW
Knows the day of week ❌ no ✅ yes
Knows alert priority via SILENT_WINDOW_OVERRIDE_PRIORITIES via AFTER_HOURS_AUDIBLE_PRIORITIES
Mutes 02:00 Tuesday ✅ ✅
Mutes 14:00 Saturday ❌ fully audible ✅

SILENT_WINDOW=22:00-06:00 only understands clock time, so a Saturday afternoon P4 sails straight through. If you want "outside office hours", you need BUSINESS_HOURS_WINDOW.

Also beware: TERSE_WINDOW does not make anything quiet. It only shortens the spoken text — a terse alert is still spoken at full volume. Setting TERSE_WINDOW overnight and expecting silence is a common mistake.

You can use both settings together. After-hours suppression is applied after SILENT_WINDOW_OVERRIDE_PRIORITIES, so that override can never resurrect an alert suppression has muted.

Window format

Mon-Fri 09:00-17:00                    # weekdays, 9 to 5
Mon-Fri 08:00-18:00, Sat 10:00-14:00   # plus Saturday mornings
Fri 17:00-09:00                        # crosses midnight into Saturday
Mon-Sun 00:00-23:59                    # always "business hours" (feature effectively off)

Day names are Mon–Sun. Ranges wrap the week, so Fri-Mon means Fri, Sat, Sun and Mon. For windows crossing midnight the weekday matches the day the window started, so Fri 17:00-09:00 covers early Saturday morning.

Optional — honouring tags from your alerting platform

If your platform already decides which alerts are deferrable, you can honour that decision instead of duplicating the rule here.

For example, a JSM alert policy can tag matching alerts business-hours-only, and a notification policy defers anything with that tag to the next weekday. Name the tag here and this service follows the same decision:

AFTER_HOURS_SILENT_TAGS=business-hours-only

Your alerting platform stays the single source of truth — widen or narrow the policy there and this service tracks it with no config change. Tags are matched case-insensitively, and any tag listed here forces silence even if the alert's priority appears in AFTER_HOURS_AUDIBLE_PRIORITIES.

Most users should leave this empty and rely on priority alone.

Verifying it works

Send an off-hours test and confirm nothing is spoken:

curl -s -X POST "http://localhost:8080/alert?mode=always&key=$WEBHOOK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"action":"Create","alert":{"alertId":"quiet-test-1",
       "message":"After-hours suppression test","priority":"P3"}}' | jq

With suppression active you get "announcement_mode": "silent" and "suppressed_after_hours": true, and the logs show:

After-hours suppression — alert quiet-test-1 (priority=P3) arrived outside
BUSINESS_HOURS_WINDOW; announcing silently

The HA persistent notification should still appear. Repeat with "priority":"P1" and you should get "announcement_mode": "full" plus audible TTS.


HA Automation Webhooks

The notifier fires Home Assistant webhook triggers when alert state changes, letting you control lights, scenes, scripts, or any HA automation in response to incidents.

State Machine — How Webhook Firing Works

⚠️ Prerequisite: set INCIDENT_DASHBOARD_ENABLED=true. The aggregate state below is computed by counting rows in the incident store. With the dashboard disabled there is nothing to count, so ON_CREATE / ON_ACKNOWLEDGE / ON_CLOSE degrade to firing once per matching event — which shows the wrong colour whenever more than one alert is open. The service logs a warning at startup if you configure state webhooks without the dashboard. You should also mount a volume for the DB so state survives a container restart.

The three main webhooks (ON_CREATE, ON_ACKNOWLEDGE, ON_CLOSE) reflect the aggregate state of all open incidents, not just the individual event that arrived. The notifier updates its incident store first, then fires the webhook that matches the current totals:

Aggregate state Webhook fired Example light color
Any open, unacknowledged alert HA_WEBHOOK_ON_CREATE 🔴 Red
All open alerts are acknowledged (none unacked) HA_WEBHOOK_ON_ACKNOWLEDGE 🟡 Yellow
No open alerts at all HA_WEBHOOK_ON_CLOSE 🟢 Green / off

Key implications:

  • UnAcknowledge automatically drives the light back to red — no extra config is needed.
  • If alert A is acknowledged but alert B is still unacked, the light stays red until both are acked.
  • If alert A closes while alert B is still acked-open, the light stays yellow until B closes too.
  • The payload always includes unacked_count, acked_count, and total_open so your automations can show counts if you want.

Per-event webhooks (ON_ESCALATE, ON_UPDATE, ON_SLA_BREACH) fire for that specific action regardless of aggregate state — useful for flashing a light on escalation without changing the color it holds.

EscalateNext and UnAcknowledge fire both: their per-event webhook and a state webhook. So an escalation can flash the light (ON_ESCALATE) while the state webhook keeps it red, and an un-acknowledge fires ON_UPDATE while the state webhook drives the colour from yellow back to red.

Reconciliation — surviving downtime

State webhooks are edge-triggered: they fire when an alert event arrives. If the service is offline while an alert is created, acknowledged or closed, that edge is missed permanently — the webhook was simply never delivered — and your light stays wrong until an unrelated alert happens to arrive.

To close that gap the state is re-derived from the incident store and re-fired:

When reason sent with the payload
Service startup startup
Scheduled sync, if INCIDENT_SYNC_INTERVAL_MINUTES > 0 scheduled-sync
POST /incidents/sync manual-sync

Startup reconciliation runs regardless of INCIDENT_SYNC_INTERVAL_MINUTES. JSM is treated as authoritative — the store is refreshed from the API first, then the webhook fires from the resulting counts.

Reconcile payloads carry event: "Reconcile" and a reason, with the per-alert fields (alert_id, priority, tags) deliberately blank, since no single alert triggered them. Automations that key off state / unacked_count need no changes; one that reads alert_id should tolerate an empty value.

Force one at any time:

curl -X POST -H "X-API-Key: $WEBHOOK_API_KEY" \
  http://localhost:8080/incidents/sync
# {"status":"synced","alerts_upserted":3,"state_webhook_fired":"ha_webhook_on_create"}

Trigger Variables

Every webhook POST includes these fields accessible as trigger.json.* in your HA templates:

Variable Description
trigger.json.event JSM action that triggered the state change (Create, Acknowledge, UnAcknowledge, Close, …)
trigger.json.state Aggregate state: create, acknowledge, or close
trigger.json.alert_id Unique alert identifier
trigger.json.message Alert title / summary
trigger.json.priority P1–P5
trigger.json.entity System / host name
trigger.json.description First 200 chars of description
trigger.json.source Alert source
trigger.json.tags List of tags
trigger.json.unacked_count Number of unacknowledged open alerts
trigger.json.acked_count Number of acknowledged open alerts
trigger.json.total_open Total open alerts (unacked + acked)

state, unacked_count, acked_count, and total_open are only present on state webhooks (ON_CREATE / ON_ACKNOWLEDGE / ON_CLOSE), not on per-event webhooks.

.env Configuration

# State webhooks — drive your status light
HA_WEBHOOK_ON_CREATE=jsm_state_unacked       # any unacked alert → red
HA_WEBHOOK_ON_ACKNOWLEDGE=jsm_state_acked    # all open alerts acked → yellow
HA_WEBHOOK_ON_CLOSE=jsm_state_clear         # no open alerts → green/off

# Per-event webhooks — optional, fire regardless of aggregate state
HA_WEBHOOK_ON_ESCALATE=jsm_escalation
HA_WEBHOOK_ON_UPDATE=jsm_alert_updated
HA_WEBHOOK_ON_SLA_BREACH=jsm_sla_breached

Comma-separate multiple IDs to fire several automations from one event:

HA_WEBHOOK_ON_CREATE=jsm_state_unacked,flash_office_lights,mobile_critical_push

Example: Bias Light Behind Your Monitor

This pattern sets a wall bias light red/yellow/green based on incident state. Three automations share one webhook each.

Step 1 — Create an HA helper to track state

In HA, go to Settings → Devices & Services → Helpers → Create → Dropdown and create:

  • Name: JSM Alert State
  • Entity ID: input_select.jsm_alert_state
  • Options: clear, acknowledged, unacked

This helper tracks the true incident state independently of whether the light is on or off. It is the source of truth used when the light is turned back on after being manually switched off.

Step 2 — Three automations, one per state

# automations.yaml

# ── Unacked alert → red ───────────────────────────────────────────────────────
- alias: "JSM — unacked alert, bias light red"
  trigger:
    - platform: webhook
      webhook_id: "jsm_state_unacked"
      allowed_methods: [POST]
      local_only: true
  action:
    # Always update the helper so light-on restore works correctly.
    - service: input_select.select_option
      target:
        entity_id: input_select.jsm_alert_state
      data:
        option: unacked
    # Only change the light if it is already on — avoids turning it on
    # in a dark room during off-hours. Remove this condition if you always
    # want the light to turn on for new alerts.
    - condition: state
      entity_id: light.bias_light
      state: "on"
    - service: light.turn_on
      target:
        entity_id: light.bias_light
      data:
        rgb_color: [255, 0, 0]
        brightness_pct: 100

# ── All open alerts acked → yellow ───────────────────────────────────────────
- alias: "JSM — all alerts acked, bias light yellow"
  trigger:
    - platform: webhook
      webhook_id: "jsm_state_acked"
      allowed_methods: [POST]
      local_only: true
  action:
    - service: input_select.select_option
      target:
        entity_id: input_select.jsm_alert_state
      data:
        option: acknowledged
    - condition: state
      entity_id: light.bias_light
      state: "on"
    - service: light.turn_on
      target:
        entity_id: light.bias_light
      data:
        rgb_color: [255, 200, 0]
        brightness_pct: 80

# ── No open alerts → green then off ──────────────────────────────────────────
- alias: "JSM — no open alerts, bias light green then off"
  trigger:
    - platform: webhook
      webhook_id: "jsm_state_clear"
      allowed_methods: [POST]
      local_only: true
  action:
    - service: input_select.select_option
      target:
        entity_id: input_select.jsm_alert_state
      data:
        option: clear
    - condition: state
      entity_id: light.bias_light
      state: "on"
    - service: light.turn_on
      target:
        entity_id: light.bias_light
      data:
        rgb_color: [0, 200, 0]
        brightness_pct: 60
    - delay: "00:00:30"
    - service: light.turn_off
      target:
        entity_id: light.bias_light

Step 3 — Restore light state when turned on

When someone manually switches the light off during an active incident, the helper still holds the correct state. This automation re-applies the right color whenever the light is turned on:

- alias: "JSM — restore bias light color on turn-on"
  trigger:
    - platform: state
      entity_id: light.bias_light
      to: "on"
  action:
    - choose:
        - conditions:
            - condition: state
              entity_id: input_select.jsm_alert_state
              state: unacked
          sequence:
            - service: light.turn_on
              target:
                entity_id: light.bias_light
              data:
                rgb_color: [255, 0, 0]
                brightness_pct: 100
        - conditions:
            - condition: state
              entity_id: input_select.jsm_alert_state
              state: acknowledged
          sequence:
            - service: light.turn_on
              target:
                entity_id: light.bias_light
              data:
                rgb_color: [255, 200, 0]
                brightness_pct: 80
      # No default — if state is "clear", leave the light at whatever
      # color/brightness it was turned on with (manual use).

How the restore works: The webhook automations always write to input_select.jsm_alert_state before touching the light, even when the "light is on" condition prevents a color change. So the helper always reflects reality. When the light is switched on later, the restore automation reads the helper and sets the correct color immediately.

Variations

Always turn the light on for new alerts (no "if on" condition):

Remove or comment out the condition: block in the red automation. The light will turn on in a dark room whenever a new alert arrives.

Show the open alert count on a display:

- alias: "JSM — update alert count badge"
  trigger:
    - platform: webhook
      webhook_id: "jsm_state_unacked"
      allowed_methods: [POST]
      local_only: true
  action:
    - service: input_number.set_value
      target:
        entity_id: input_number.jsm_open_count
      data:
        value: "{{ trigger.json.total_open }}"

Flash on escalation without changing the color the light is holding:

- alias: "JSM — flash bias light on escalation"
  trigger:
    - platform: webhook
      webhook_id: "jsm_escalation"
      allowed_methods: [POST]
      local_only: true
  action:
    - condition: state
      entity_id: light.bias_light
      state: "on"
    - service: light.turn_on
      target:
        entity_id: light.bias_light
      data:
        flash: long
HA_WEBHOOK_ON_ESCALATE=jsm_escalation

Priority-based colors on the ON_CREATE state:

- alias: "JSM — unacked alert, priority-colored bias light"
  trigger:
    - platform: webhook
      webhook_id: "jsm_state_unacked"
      allowed_methods: [POST]
      local_only: true
  action:
    - service: input_select.select_option
      target:
        entity_id: input_select.jsm_alert_state
      data:
        option: unacked
    - condition: state
      entity_id: light.bias_light
      state: "on"
    - service: light.turn_on
      target:
        entity_id: light.bias_light
      data:
        brightness_pct: 100
        rgb_color: >
          {% if trigger.json.priority == 'P1' %}
            [255, 0, 0]
          {% elif trigger.json.priority == 'P2' %}
            [255, 80, 0]
          {% else %}
            [255, 200, 0]
          {% endif %}

Incident State Dashboard

An optional SQLite-backed incident tracker that exposes a JSON API at /incidents. Useful for building Grafana dashboards, monitoring tools, or just quickly checking what's open.

Enabling the Dashboard

INCIDENT_DASHBOARD_ENABLED=true
INCIDENT_DB_PATH=/data/incidents.db
INCIDENT_SYNC_INTERVAL_MINUTES=5

Note: the default INCIDENT_DB_PATH=/tmp/incidents.db lives on the container's tmpfs mount, so incident history is intentionally wiped on every restart. Mount a volume to persist it.

For persistent storage, mount a volume in docker-compose.yml (a named-volume variant ships commented-out in the file):

services:
  jsm-ha-notifier:
    volumes:
      - ./data:/data   # host dir must be writable by uid 1000 (non-root user)

API Endpoints

GET /incidents

List all incidents with optional filters:

# All incidents
curl http://localhost:8080/incidents

# Only open incidents
curl "http://localhost:8080/incidents?status=open"

# Only P1 incidents
curl "http://localhost:8080/incidents?priority=P1"

# Open P1 incidents
curl "http://localhost:8080/incidents?status=open&priority=P1"

Response:

{
  "incidents": [
    {
      "alert_id": "abc-123",
      "message": "Database connection pool exhausted",
      "priority": "P1",
      "entity": "prod-db-01",
      "description": "All 200 connections in use...",
      "source": "Datadog",
      "status": "open",
      "action": "Create",
      "created_at": "2026-03-22T10:30:00+00:00",
      "updated_at": "2026-03-22T10:30:00+00:00",
      "acknowledged_at": null,
      "closed_at": null
    }
  ],
  "count": 1
}

GET /incidents/summary

Aggregate counts:

curl http://localhost:8080/incidents/summary
{
  "total_open": 3,
  "total_closed": 12,
  "by_status": {"open": 2, "escalated": 1, "closed": 12},
  "by_priority": {"P1": 1, "P2": 2}
}

GET /incidents/{alert_id}

Single incident detail:

curl http://localhost:8080/incidents/abc-123

POST /incidents/sync

Force an immediate sync from JSM:

curl -X POST http://localhost:8080/incidents/sync

Grafana Integration

The /incidents endpoint is compatible with Grafana's JSON or Infinity datasource plugins:

  1. Install the Infinity datasource plugin in Grafana
  2. Add a new datasource:
    • Type: Infinity
    • URL: http://your-notifier:8080
  3. Create a dashboard panel:
    • Source: Infinity
    • Type: JSON
    • URL: /incidents?status=open
    • Root selector: $.incidents
  4. Add columns: alert_id, message, priority, status, entity, created_at

For the summary endpoint, use /incidents/summary to build gauge or stat panels showing open incident counts by priority.

Pre-built Dashboard

A ready-to-import Grafana dashboard is included in this repo:

grafana/incident-dashboard.json

To import:

  1. In Grafana, go to Dashboards > Import
  2. Upload grafana/incident-dashboard.json
  3. Select your Infinity datasource
  4. Set the api_key variable to your WEBHOOK_API_KEY value (under dashboard Settings > Variables)

The dashboard includes:

  • Stat panels: Total Open, Total Closed, Open P1, Open P2, Open P3
  • Full incident table with priority/status color coding and column filters
  • Pie charts: By Status, By Priority (open only)
  • Auto-refresh every 30 seconds

Force-Close Incidents

Close a stale incident directly from the API (without waiting for JSM):

curl -X POST "http://localhost:8080/incidents/the-alert-id/close?key=YOUR_KEY"

This sets the status to closed, dismisses the HA persistent notification, and cancels any TTS repeats.

Retention Policy

Automatically clean up old incidents to prevent unbounded database growth:

INCIDENT_RETENTION_OPEN_DAYS=30       # Delete stale open incidents after 30 days
INCIDENT_RETENTION_CLOSED_DAYS=90     # Delete resolved incidents after 90 days

Retention runs during each sync cycle (INCIDENT_SYNC_INTERVAL_MINUTES). Set to 0 to keep everything forever (default).

Alert Enrichment

When the incident dashboard is enabled, the notifier automatically enriches new alerts by fetching full details from the JSM API on Create events. This adds:

  • Tags — alert tags from JSM
  • Teams — team assignments
  • Responders — who the alert was sent to
  • Custom details — any key/value pairs from the alert's details field

All enrichment data is stored in the SQLite database and returned in the /incidents API response.


Running on unRAID (or any Docker host)

Option A — Docker Compose (recommended)

# Build locally (until you push to GHCR)
docker compose up -d --build

# Or pull the pre-built image after the first CI release
docker compose pull
docker compose up -d

# Watch logs
docker compose logs -f jsm-ha-notifier

Option B — docker run

docker run -d \
  --name jsm-ha-notifier \
  --restart unless-stopped \
  -p 8080:8080 \
  --env-file /path/to/.env \
  --read-only \
  --tmpfs /tmp \
  ghcr.io/realdougeubanks/jsm-ha-notifier:latest

CI/CD — GitHub Actions & Container Registries

The repository includes two workflows. Images are published to both GitHub Container Registry (GHCR) and Docker Hub.

CI (.github/workflows/ci.yml)

Triggers on every push to main or develop and on pull requests. Runs:

  • ruff (lint)
  • black (format check)
  • mypy (type check, advisory)
  • pip-audit (dependency CVE scan)
  • bandit (Python SAST, advisory)
  • pytest with coverage (fails below 70%)
  • Tests against Python 3.11, 3.12, 3.13

Release (.github/workflows/release.yml)

Triggers on push to main or any version tag (v*). Builds a multi-arch Docker image (linux/amd64 + linux/arm64) and pushes to both GHCR and Docker Hub. Also runs Trivy container vulnerability scanning and uploads results to GitHub Security.

No personal access tokens or manual secrets are needed — the workflow uses the built-in GITHUB_TOKEN that GitHub provides automatically to every Actions run, which already has packages: write permission as configured in the workflow.

Image tags produced

Git event Image tags
Push to main latest, main, <short-sha>
Push tag v1.2.3 v1.2.3, <short-sha>

First-time GHCR setup

After the first successful release workflow run, your container image is private by default. To make it public so others (and your unRAID server) can pull it without authentication:

  1. Go to https://github.com/RealDougEubanks?tab=packages
  2. Click the jsm-ha-notifier package
  3. Click Package settings (right side)
  4. Under Danger Zone, click Change visibility → Public

Alternatively, link the package to your repository:

  1. On the package page, click Connect repository and select your repo
  2. The package inherits the repository's visibility

Once public, docker pull ghcr.io/realdougeubanks/jsm-ha-notifier:latest works without login from any machine.

Docker Hub setup

The release workflow also pushes to Docker Hub if credentials are configured. To enable:

  1. Create a Docker Hub access token at https://hub.docker.com/settings/security
  2. In your GitHub repo, go to Settings > Secrets and variables > Actions
  3. Add these secrets/variables:
    • Variable DOCKERHUB_USERNAME = your Docker Hub username (e.g. dougeubanks)
    • Secret DOCKERHUB_TOKEN = the access token from step 1

Once configured, every release pushes to both registries:

docker pull ghcr.io/realdougeubanks/jsm-ha-notifier:latest
docker pull dougeubanks/jsm-ha-notifier:latest

If Docker Hub credentials are not set, the workflow gracefully skips the Docker Hub login and only pushes to GHCR.

Image signing and provenance

Every image pushed by the release workflow is:

  1. Signed with cosign (Sigstore keyless / OIDC) — proves the image was built by this GitHub Actions workflow
  2. Attested with SLSA Build Level 2 — GitHub's native build provenance attestation

To verify an image before pulling:

# Install cosign: https://docs.sigstore.dev/cosign/system_config/installation/
cosign verify ghcr.io/realdougeubanks/jsm-ha-notifier:latest \
  --certificate-oidc-issuer https://token.actions.githubusercontent.com \
  --certificate-identity-regexp github.com/RealDougEubanks/JSM-HomeAssistant-Notifier

This confirms the image was built from this repository's GitHub Actions — not tampered with after the fact.

Updating docker-compose.yml to use a published image

After the image has been published, edit docker-compose.yml. Pin a specific release tag so docker compose pull never silently swaps the code under you:

services:
  jsm-ha-notifier:
    image: ghcr.io/realdougeubanks/jsm-ha-notifier:v3.1.0
    # build: .   ← comment out or remove this line

To update, bump the tag and re-run docker compose up -d. Floating tags (:3.1, :3, :latest) auto-update on pull — convenient, but you are trusting every future release; :latest is not recommended for unattended deployments.


API Reference

POST /alert

Receives JSM webhook payloads.

Query param Values Behaviour
mode always Skip on-call check; always notify
(absent) — Check on-call status before notifying

Expected payload: standard OpsGenie / JSM Ops outgoing webhook JSON.

POST /alert/{alert_id}/acknowledge

Acknowledges a JSM alert, dismisses the HA notification, and cancels TTS repeats. Intended for use from HA automations (see .env.example for a ready-to-use rest_command snippet).

Returns {"alert_id": "...", "acknowledged": true} on success, 502 if JSM rejects the request.

GET /health

Returns {"status": "ok"}. Used by Docker health-check and external monitors.

GET /healthz

Deep health check — verifies JSM and HA connectivity, validates configured schedules, and reports operational state. Returns 200 if core checks pass, 503 if any fail. Gated by API key (query param, header, or path prefix) when WEBHOOK_API_KEY is set.

{
  "healthy": true,
  "timestamp": "2026-03-25T14:32:01+00:00",
  "started_at": "2026-03-25T14:00:00+00:00",
  "uptime_seconds": 1921.0,
  "version": "2.0.0",
  "checks": { "jsm_api": "ok", "ha_api": "ok" },
  "schedules": {
    "check_oncall": {
      "Cloud Engineering On-Call Schedule": {
        "schedule_id": "abc-123",
        "exists_in_jsm": true,
        "on_call": true
      }
    },
    "always_notify": ["Internal Systems_schedule"]
  },
  "cache": {
    "schedule_id_entries": 2,
    "oncall_entries": 1,
    "dedup_entries": 0
  },
  "background_tasks": {
    "batch_queue_size": 0,
    "active_tts_repeats": 0,
    "tts_repeat_alert_ids": []
  },
  "incident_dashboard": { "enabled": false },
  "configuration": {
    "oncall_cache_ttl_seconds": 300,
    "alert_dedup_ttl_seconds": 60,
    "token_check_interval_hours": 24,
    "alert_batch_window_seconds": 0,
    "tts_repeat_interval_seconds": 0,
    "tts_repeat_max": 5,
    "tts_repeat_priorities": "P1",
    "silent_window": "(none)",
    "terse_window": "(none)",
    "webhook_secret_configured": true,
    "webhook_api_key_configured": true,
    "emojis_enabled": true
  }
}

No tokens, secrets, URLs, or user IDs are included in the response.

GET /status

Returns current on-call status for all watched schedules (forces a fresh JSM API lookup, bypasses cache).

{
  "on_call_schedules": {
    "Your On-Call Schedule": {
      "schedule_id": "abc-123",
      "on_call": true
    }
  },
  "always_notify_schedules": ["Your Always-Notify Schedule"]
}

GET /metrics

Prometheus-compatible metrics in text exposition format. Gated by API key when configured.

jsm_notifier_alerts_received_total 42
jsm_notifier_alerts_notified_total 38
jsm_notifier_alerts_deduplicated_total 3
jsm_notifier_alerts_dismissed_total 12
jsm_notifier_alerts_rate_limited_total 0
jsm_notifier_credential_checks_total 7
jsm_notifier_credential_checks_failed_total 0
jsm_notifier_healthz_requests_total 15
jsm_notifier_uptime_seconds 86412.3

POST /reload

Re-reads .env and applies configuration changes without restarting the container. Clears all caches (schedule ID, on-call, dedup) on reload. Gated by API key when configured.

Returns {"status": "reloaded"} on success, 500 if the new config is invalid (previous config remains active).

POST /cache/invalidate

Clears the cached on-call status so the next alert forces a fresh JSM API check. Useful immediately after a rotation hand-off.

GET /incidents (dashboard)

List incidents with optional ?status= and ?priority= filters. Requires INCIDENT_DASHBOARD_ENABLED=true.

GET /incidents/summary (dashboard)

Aggregate incident counts by status and priority.

GET /incidents/{alert_id} (dashboard)

Single incident detail by alert ID.

POST /incidents/sync (dashboard)

Force an immediate sync of open alerts from JSM into the incident store.


Notification Details

TTS Announcement

The spoken message includes:

  • Escalation prefix ("Escalated alert!") when applicable
  • Priority level in plain English ("Priority 1, Critical")
  • Alert message / title
  • System / entity name
  • Truncated description (first 200 characters)

Example: "Attention! Priority 1, Critical alert from Jira Service Management. Alert: Database connection lost. System: prod-db-01. Details: All connections exhausted..."

Media Player Display

Instead of "Playing Default Media Receiver", the HA media player will show:

🔴 P1: Database connection lost
Your Notifier Label
prod-db-01

This is set via the extra.metadata block in the media_player.play_media service call. The label shown as the artist is configurable via HA_NOTIFIER_LABEL in .env.

Persistent Notification

A persistent notification is created in the HA dashboard with the full alert details. It is automatically dismissed when JSM sends an Acknowledge or Close action for that alert.


Local Development

# Create a virtual environment
python3 -m venv .venv
source .venv/bin/activate

# Install all dependencies
pip install -r requirements-dev.txt

# Copy and edit config
cp .env.example .env
# Fill in .env before running tests or the server

# Run tests
pytest tests/ -v

# Run the service locally
uvicorn src.main:app --reload --port 8080

Troubleshooting

"Schedule not found" error

Schedule names in .env must match JSM exactly (case-sensitive). List your schedules:

curl -s -u "you@yourcompany.com:YOUR_API_TOKEN" \
  "https://api.atlassian.com/jsm/ops/api/YOUR_CLOUD_ID/v1/schedules" \
  | python3 -m json.tool | grep '"name"'

Copy the exact name into .env and restart.

No audio / TTS not playing

Check first: was it silenced on purpose? Several settings suppress TTS by design, and each still posts the HA notification — so a missing announcement is often correct behaviour, not a fault. Rule these out before debugging tokens.

Response field / log line Meaning This is
"suppressed_after_hours": true Outside BUSINESS_HOURS_WINDOW and the priority is not in AFTER_HOURS_AUDIBLE_PRIORITIES working as configured
"announcement_mode": "silent" Inside a SILENT_WINDOW working as configured
"announcement_mode": "terse" Inside a TERSE_WINDOW — audio still plays, just shorter working as configured
"reason": "not on-call for any watched schedule" You are not on-call, so no announcement working as configured
"batched": true Queued for up to ALERT_BATCH_WINDOW_SECONDS before speaking working as configured
# Grep for a deliberate suppression
docker compose logs jsm-ha-notifier | grep -E "After-hours suppression|Silent window|No notification"

If BUSINESS_HOURS_WINDOW is set but TZ is not, the container evaluates windows in UTC and your office hours will be wrong by your UTC offset. TZ is read at process start, so POST /reload will not pick up a change — recreate the container.

If none of the above applies, then debug the connection:

  1. Verify the HA token is valid:
    curl -H "Authorization: Bearer YOUR_HA_TOKEN" https://your-ha-url/api/
  2. Verify the media player entity ID:
    curl -H "Authorization: Bearer YOUR_HA_TOKEN" \
      https://your-ha-url/api/states \
      | python3 -m json.tool | grep media_player
  3. Check service logs: docker compose logs -f jsm-ha-notifier

An alert announced overnight that should not have

Your alerting platform's quiet hours do not cover this service. Notification policies govern the platform's own channels (mobile, email, SMS); outgoing webhooks fire at alert creation and bypass them entirely — so your whole team can be protected by a deferral rule while you are not.

Fix: set BUSINESS_HOURS_WINDOW. See Quiet Hours & After-Hours Suppression for the full explanation.

Setting TERSE_WINDOW will not fix this. Terse only shortens the spoken text — it still plays at full volume.

HA shows "Playing Default Media Receiver"

Your HA media player integration may not support the extra.metadata block. This is normal for some Google Cast / Chromecast firmware versions. The TTS audio itself will still play correctly — only the display label is affected.

On-call check returns true when I'm not on-call

The on-call cache may be stale. Force a refresh:

curl -X POST http://localhost:8080/cache/invalidate

Getting 404 on endpoints that should exist

When WEBHOOK_API_KEY is set, endpoints return 404 Not Found (not 401) if the API key is missing or wrong. This is intentional — it prevents attackers from discovering that authenticated endpoints exist. If you're getting unexpected 404s:

  1. Confirm WEBHOOK_API_KEY is set in your .env
  2. Confirm your request includes the key via one of:
    • Query parameter: ?key=YOUR_KEY
    • Path prefix: /YOUR_KEY/endpoint
    • HTTP header: X-API-Key: YOUR_KEY
  3. Verify the key value matches exactly (no extra whitespace)

The /health and /robots.txt endpoints are always unauthenticated.

Invalid webhook signature

Confirm that the WEBHOOK_SECRET in .env matches the secret configured in JSM exactly. Remember: the HMAC is computed over the raw request body, not the parsed JSON.

JSM is not calling the webhook

  • Confirm the URL is reachable from the internet: curl https://your-host/health
  • Check JSM webhook delivery logs: JSM → Settings → Integrations → your webhook → Logs

Token expiry notification in HA dashboard

The service checks token validity every TOKEN_CHECK_INTERVAL_HOURS hours (default: 24). If your token has expired:

  1. Create a new token at https://id.atlassian.com/manage-profile/security/api-tokens
  2. Update JSM_API_TOKEN in .env
  3. Restart the container: docker compose restart

The persistent HA notification will be dismissed automatically on the next successful token check (within 30 seconds of startup).

Quiet hours: If the credential check fails during a SILENT_WINDOW, the TTS announcement is suppressed — only the persistent dashboard notification is created. This prevents the service from waking you up at night for a non-urgent token issue that can wait until morning.


Production Recommendations

External uptime monitoring

The service is designed to wake you up — but nothing wakes you up if the service itself is down. Configure an external uptime monitor to poll GET /health and alert you if it stops responding.

NodePing (recommended):

  1. Create a new HTTP check
  2. URL: https://your-host/health (or your Cloudflare Tunnel URL)
  3. Expected status: 200
  4. Expected body contains: "ok"
  5. Check interval: 1 minute
  6. Notification contacts: your email, SMS, or PagerDuty

Other options: Uptime Kuma (self-hosted), UptimeRobot, Pingdom, or Cloudflare Health Checks.

The /health endpoint is always unauthenticated and has no external dependencies — it returns {"status": "ok"} as long as the process is alive.

Single-worker resilience

The service runs a single uvicorn worker. This is sufficient for typical JSM webhook volume (a few alerts per hour), but be aware:

  • If the process crashes, Docker's restart: unless-stopped will restart it automatically, but alerts arriving during the ~5s restart window will be lost (JSM retries, so they'll typically arrive again)
  • If the event loop blocks on a slow JSM/HA API call, other webhooks queue behind it — the async architecture mitigates this, but a truly hung connection could stall processing

For higher availability, run the service on a host with reliable uptime and use the external uptime monitor above to detect outages quickly.

Persistent incident storage

By default, the incident database uses /tmp/incidents.db which is lost on container restart (tmpfs). For production use, mount a Docker volume:

# docker-compose.yml
services:
  jsm-ha-notifier:
    volumes:
      - ./data:/data
# .env
INCIDENT_DB_PATH=/data/incidents.db

Security Checklist

  • Atlassian API token created with minimum necessary permissions (JSM Ops schedule access)
  • .env is in .gitignore and was never committed
  • WEBHOOK_API_KEY is set (openssl rand -hex 32) — or — WEBHOOK_SECRET is set (or both)
  • Service runs as non-root user (handled in Dockerfile)
  • Container filesystem is read-only (read_only: true in compose)
  • Port 8080 is behind a TLS-terminating reverse proxy or Cloudflare Tunnel before reaching the internet
  • HA long-lived token was created specifically for this service (not shared with other integrations)

Project Structure

jsm-ha-notifier/
├── .github/
│   └── workflows/
│       ├── ci.yml          # Lint, test, coverage
│       └── release.yml     # Build & push multi-arch Docker image to GHCR
├── src/
│   ├── __init__.py
│   ├── main.py             # Logging, app composition, lifespan
│   ├── security.py         # Middleware, rate limiting, signature / API key auth
│   ├── metrics.py          # Prometheus-compatible counters
│   ├── routes/
│   │   ├── ops.py          # /health, /healthz, /metrics, /status, /reload, /cache
│   │   ├── incidents.py    # /incidents dashboard endpoints
│   │   └── webhook.py      # /alert and /alert/{id}/acknowledge
│   ├── config.py           # Pydantic settings (all from .env)
│   ├── models.py           # JSM webhook payload models
│   ├── jsm_client.py       # Async JSM Ops API client with caching
│   ├── ha_client.py        # Async Home Assistant REST API client
│   ├── alert_processor.py  # Core routing / dedup / notification logic
│   ├── incident_store.py   # SQLite-backed incident state tracker
│   └── time_windows.py     # Time-window parsing and media player routing
├── tests/
│   ├── conftest.py                     # Shared fixtures
│   ├── test_models.py
│   ├── test_config.py
│   ├── test_ha_client.py
│   ├── test_ha_client_coverage.py
│   ├── test_jsm_client.py
│   ├── test_alert_processor.py
│   ├── test_alert_processor_coverage.py # Batch, repeat, state webhooks
│   ├── test_main_routes.py             # Route, reload, and rate-limit tests
│   ├── test_security_coverage.py       # Auth, signature, middleware
│   ├── test_announcement_format.py     # Format, time windows, priority override, repeat
│   ├── test_robustness.py              # Security: sanitization, safe formatter, emoji toggle
│   ├── test_incident_store.py          # Incident store, webhooks, force-close, retention
│   ├── test_after_hours.py             # Business-hours windows, after-hours suppression
│   └── test_time_windows.py            # Window parsing, player routing
├── docs/
│   ├── RUNBOOK.md          # 2am operational guide — start here when paged
│   └── ENV_VARS.md         # Every environment variable, with rotation steps
├── grafana/
│   └── incident-dashboard.json      # Pre-built Grafana dashboard (import-ready)
├── .env.example            # Template — copy to .env and fill in values
├── .gitignore
├── CHANGELOG.md
├── CONTRIBUTING.md         # Dev setup, CI gates, PR checklist
├── SECURITY.md             # Vulnerability reporting and security controls
├── LICENSE                 # Apache License 2.0
├── docker-compose.yml
├── Dockerfile
├── pyproject.toml          # black, ruff, pytest, mypy config
├── requirements.txt
├── requirements-dev.txt
└── README.md

Further Reading

Document Read it when
docs/RUNBOOK.md You were paged and need to diagnose or restart the service
docs/ENV_VARS.md You need to know what a variable does or how to rotate a credential
SECURITY.md You are reporting a vulnerability or reviewing the security controls
CONTRIBUTING.md You are setting up a dev environment or opening a pull request
CHANGELOG.md You want to know what changed between releases

AI Disclosure

This project was designed and built by Doug Eubanks to solve a real on-call alerting problem. The architecture, requirements, testing, and deployment decisions were driven by him throughout.

Claude (Anthropic's AI assistant) was used as a collaborative engineering tool during development — writing and iterating on code, debugging issues, and helping document the project. All code was reviewed, tested in a live environment, and validated by the author before use.

This disclosure is provided in the spirit of transparency. The use of AI assistance does not diminish the engineering decisions, debugging work, or operational responsibility that went into this project.


License

Apache License 2.0 — see LICENSE for details.

About

Docker service that bridges Jira Service Management (JSM/OpsGenie) alerts to Home Assistant — smart on-call routing, escalation detection, rich Google Home TTS announcements, and persistent dashboard notifications.

Topics

Resources

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages