"부하가 임계치를 넘으면 스크립트를 돌리고 싶다" 는 요구는 흔하지만 Alertmanager 는 셸 명령을 직접 실행하지 못한다. 수신자로 지원하는 것은 이메일이나 Slack 같은 기성 연동과, 임의의 HTTP 엔드포인트로 POST 를 보내는 webhook_configs 뿐이다.
따라서 구조는 이렇게 된다.
Prometheus (규칙 평가)
-> Alertmanager (그룹핑 · 억제 · 라우팅)
-> webhook 수신기 (직접 만든 작은 HTTP 서버)
-> 스크립트 실행
사내 메일 API 로 보내는 경우도 마찬가지다. Alertmanager 가 보내는 JSON 은 형식이 고정돼 있어 임의의 API 스펙에 맞출 수 없으므로 중간에 변환기를 둔다.
groups:
- name: node_alerts
rules:
- alert: HighLoad
expr: node_load1 > 16
for: 1m
labels:
severity: critical
annotations:
summary: "High load on {{ $labels.instance }}"
description: "load1 = {{ $value }}"
rule_files:
- /etc/prometheus/rules/*.yml
alerting:
alertmanagers:
- static_configs:
- targets: ['localhost:9093']
for: 1m 은 조건이 1분 동안 계속 참일 때만 발화한다는 뜻이다. 순간적인 스파이크로 스크립트가 도는 것을 막는다.
문법을 먼저 검사하고 반영한다.
promtool check rules /etc/prometheus/rules/node_alerts.yml
promtool check config /etc/prometheus/prometheus.yml
curl -X POST http://localhost:9090/-/reload
리로드 엔드포인트는 Prometheus 를 --web.enable-lifecycle 로 띄웠을 때만 열린다.
global:
resolve_timeout: 5m
route:
group_by: ['alertname', 'instance']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
receiver: script-runner
receivers:
- name: script-runner
webhook_configs:
- url: 'http://127.0.0.1:5001/alert'
send_resolved: false
amtool check-config /etc/alertmanager/alertmanager.yml
sudo systemctl restart alertmanager
repeat_interval 을 짧게 두면 같은 경보로 스크립트가 반복 실행된다. 스크립트가 여러 번 돌아도 문제없는 형태가 아니라면 값을 넉넉히 준다.
Alertmanager 가 보내는 본문은 다음 형태다.
{
"status": "firing",
"alerts": [
{
"status": "firing",
"labels": {"alertname": "HighLoad", "instance": "node1:9100"},
"annotations": {"summary": "High load on node1:9100"},
"startsAt": "2026-09-20T10:00:00Z"
}
]
}
이를 받아 스크립트를 돌리는 최소 구현이다. 표준 라이브러리만 쓰므로 폐쇄망에서도 추가 설치가 필요 없다.
import json
import subprocess
from http.server import BaseHTTPRequestHandler, HTTPServer
SCRIPT = "/opt/alert-hook/on_alert.sh"
ALLOWED = {"HighLoad"}
class Handler(BaseHTTPRequestHandler):
def do_POST(self):
length = int(self.headers.get("Content-Length", 0))
payload = json.loads(self.rfile.read(length) or b"{}")
for alert in payload.get("alerts", []):
if alert.get("status") != "firing":
continue
labels = alert.get("labels", {})
name = labels.get("alertname", "")
if name not in ALLOWED:
continue
subprocess.Popen([SCRIPT, name, labels.get("instance", "")])
self.send_response(200)
self.end_headers()
self.wfile.write(b"ok")
if __name__ == "__main__":
HTTPServer(("127.0.0.1", 5001), Handler).serve_forever()
#!/usr/bin/env bash
set -euo pipefail
echo "$(date -Is) fired: $1 $2" >> /var/log/alert-hook.log
백그라운드로 띄워 두면 재부팅 후 사라지고 로그도 흩어진다. 유닛으로 만든다.
[Unit]
Description=Alertmanager webhook runner
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
User=alerthook
Group=alerthook
ExecStart=/usr/bin/python3 /opt/alert-hook/webhook_runner.py
Restart=on-failure
RestartSec=5
[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable --now alert-hook
임의의 스크립트를 HTTP 요청으로 실행하는 구조는 그 자체가 위험하다. 다음을 지킨다.
alertname 화이트리스트로 거른다경보를 직접 밀어 넣어 경로 전체를 확인한다.
curl -XPOST http://localhost:9093/api/v2/alerts \
-H 'Content-Type: application/json' \
-d '[{"labels":{"alertname":"HighLoad","instance":"test:9100"}}]'
amtool alert
tail -f /var/log/alert-hook.log