- Python 100%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
Remove the dead MeshSettings.brain_max_concurrent field: it parsed WARDEN_BRAIN_MAX_CONCURRENT with a different default (1) than the live orchestrator, which reads the same env var directly in spine.py (default 2) and never consulted this field. Verified with a full-repo grep that nothing reads settings.brain_max_concurrent. Removing it collapses the dual-reader trap onto the one path actually used. Also close four coverage gaps flagged in review of the Tier-2 edge-brain build: - test_mesh_brain: decode_job/decode_result reject malformed/non-JSON bytes via Reject (already handled by _decode_dict; adds the assertion). - test_mesh_grant: mutating jti inside a signed grant while keeping the original signature is caught as a bad signature, proving the HMAC covers jti. - test_brain_dispatch: a publish that raises AFTER insert_brain_job commits still leaves the lease row live (a sweepable orphan) - persist-before-publish holds under a publish fault, and the exception propagates rather than being swallowed. - config.py: note on _is_loopback_url that a scheme-less WARDEN_ORCHESTRATOR_URL (e.g. orch.prod:8899) parses as hostless and is therefore also treated as loopback -> fail-closed. Full suite: 1062 passed, 1 skipped. Co-Authored-By: Claude Opus 4.8 <[email protected]> Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af |
||
| docs | ||
| src/warden | ||
| tests | ||
| .gitignore | ||
| LICENSE | ||
| pyproject.toml | ||
| README.md | ||
warden
A small, deterministic operational incident engine for devops. warden turns any check — an error rate, a latency SLO, a full disk, an expiring certificate, a downed endpoint, config drift, a crash-looping container, and the failure-by-absence class error-based tools miss — into deduped incidents, notifies, and (opt-in, later) auto-resolves them via runbook-driven agents behind a hard allowlist.
It is plugin-based and domain-agnostic: the core knows nothing about your systems or
the kind of problem. You point it at your world with detector plugins and a declarative
checks.yaml; each detector decides what "failing" means for its own check.
Why
One engine for the whole operational surface, not another single-purpose alerter. Point-tools each watch one slice (metrics, logs, uptime, certs, jobs) and route nowhere useful; warden gives every kind of check the same lifecycle — stable-fingerprint dedupe, transition-only notification, crash-resumable incidents, and a path to automated remediation. It also covers one class the others structurally can't: failure-by-absence (a run that didn't happen, a feed that stopped) throws no error, so warden treats "did the expected thing happen, and is its output current?" as a first-class check — and reports a check that can't run as loudly as one that fails, never silently OK.
Design in one breath
- Deterministic spine, no LLM in the hot path. Routing, dedupe, and severity are
plain code keyed on a stable
fingerprint; untrusted evidence text never decides an action. - Detector → Finding → Incident. A detector is pure:
check(ctx) -> [Finding]. The store diffs findings per fingerprint and emits only transitions (ok→problem opens one incident, problem→ok resolves it, a severity change re-notifies). All state is in sqlite, so the process is crash-resumable. - Everything is a plugin. Built-in generic detectors ship in this repo; third-party
add-ons register via the
warden.detectorsentry-point group — no fork required.
Install
pip install warden # core (light: pyyaml + structlog)
pip install "warden[delta]" # + the delta_freshness detector's deps (deltalake/pyarrow)
Quickstart
# checks.yaml — your detection surface (reviewed like code)
checks:
nightly-etl:
plugin: heartbeat
trigger: {every: 5m}
params: {key: nightly-etl, max_age: 26h} # the job POSTs /heartbeat on success
api-up:
plugin: probe
trigger: {every: 1m}
params: {kind: http, url: https://api.example.com/health}
WARDEN_CHECKS=checks.yaml WARDEN_DB=warden.db warden
# GET :8891/health · GET :8891/incidents · GET :8891/metrics (Prometheus) · POST :8891/heartbeat {"key":"nightly-etl"}
# SIGHUP reloads checks.yaml live (a bad file is refused, keeping the running config).
# WARDEN_RESOLVE_DWELL=5m suppresses flapping (resolve only after OK holds that long).
Built-in detectors
| plugin | detects |
|---|---|
heartbeat |
failure-by-absence — a keyed heartbeat older than max_age (or never seen) |
probe |
an HTTP endpoint (status/body), a TCP port, or an SSH/SFTP banner not answering |
promql |
a Prometheus expression that crosses a threshold (turns any recording rule into a warden check) |
prom_scrape |
a metric read straight from an exporter's /metrics crosses a threshold (when the Prometheus server isn't reachable but the exporter is) |
disk |
free space on a path below a floor (severity scales with how far below) |
cert_expiry |
a TLS certificate within warn_days/crit_days of expiry — or already expired |
systemd |
a systemd unit not in the active state |
json_check |
polls a JSON health/readiness endpoint and maps each record ({name,status,detail}) to a finding — adopts a service's own checklist |
dagster |
a Dagster sensor not RUNNING / stale last-tick, a job run FAILURE, or a run-storm (retry loop), via GraphQL |
delta_freshness |
a Delta table stale beyond a business-day threshold (weekend/holiday-aware) |
Every built-in is pure and capability-injected, so it unit-tests without touching
the network or a real host. Since 0.26.0 warden also ships the agent_audit drift
sweep and ingests external error/business events (error_events, WARDEN_EVENTS_INGEST,
on warden.events.<env>.<source>); container-health remains a roadmap plugin. Bring
your own with the warden.detectors entry point (docs/adding-a-plugin.md) — no fork
required.
Notifications
Transitions fan out to every configured channel (one failing channel never silences
the others): a structured log (always on), an HTTP push (WARDEN_NOTIFY_URL —
ntfy/webhook), Matrix (WARDEN_MATRIX_* — posts an m.notice to a room), and a
Gitea/Forgejo issue tracker as the durable case record (WARDEN_FORGEJO_URL +
_TOKEN + _OWNER + _REPO): an incident opens an issue, severity changes and the
resolution append comments (the timeline), and resolving closes it — idempotent per
incident. Wire the channels your deployment has; secrets stay in your private overlay.
(For an internal tracker behind an expired/self-signed cert, WARDEN_TLS_INSECURE_HOSTS
skips TLS verification for those exact hosts only.)
Telemetry
GET /metrics is Prometheus exposition. Beyond warden's own health (open incidents,
last-tick age) it exposes durable counters for the whole incident lifecycle and — the
part that matters for an agentic system — per-fix token and cost accounting parsed
from each resolver run's stream-json result:
warden_transitions_total{kind,check}— opened / resolved / severity_changedwarden_incidents_resolved_total{check,mode}·_escalated_total·_waiting_totalwarden_agent_runs_total{check,mode}·warden_agent_cost_usd_total{check}warden_agent_{input,output,cache_read_input,cache_creation_input}_tokens_total{check}warden_agent_turns_total{check}·warden_agent_duration_ms_total{check}
Every agent run also writes a usage event onto the incident timeline (tokens, cost,
turns, duration), so you can see exactly what a given fix cost. Point Prometheus at
/metrics for the metrics half.
For logs, set WARDEN_OTLP_LOGS_URL to an OTLP/HTTP collector endpoint and warden ships
its structured logs there (batched, best-effort, zero extra dependencies — a plain
OTLP-JSON POST to /v1/logs), which forwards them to Loki. A dead collector never blocks
the loop.
Configuration & secrets
warden reads its detection surface from WARDEN_CHECKS and its state db from WARDEN_DB.
Keep your checks.yaml, runbooks, and secrets in your own private repo/overlay — they
are deployment-specific and never belong in this open-source repo. See docs/checks.example.yaml.
Writing a plugin
A detector is ~30 lines. Ship it in this repo (generic) or as your own package via the
warden.detectors entry point (proprietary/add-on). See docs/adding-a-plugin.md.
Status
Phases 1–3 are implemented and tested (0.26.0 landed the last roadmap items:
error_events NATS ingest, the agent_audit drift sweep, the fleet-hub
first-connect retry, and sandbox.run mounted on /invoke):
-
Detect + notify — the detectors above, a Prometheus self-exporter (
/metrics), and log / webhook / Matrix channels; SIGHUP config reload and flap suppression. -
Runbook resolution — on an incident, warden matches a runbook and composes an exact, allowlist-scoped
claude -p …remediation command (seedocs/runbooks.example/). -
Auto-resolution (opt-in,
WARDEN_AUTO_RESOLVE=1) — for runbooks that setauto: true, warden launches that command itself behind hard guardrails: deny-by-default allowlist, per-attempt wall-clock timeout, 2-strikes, spine-sideverify(the agent never self-certifies), a full action timeline on the incident, and escalation carrying the resumable agent session id. A runbook withoutauto: trueis only ever composed for a human. Untrusted evidence can never widen the allowlist or inject instructions. -
Guide-the-agent (Phase 4, with auto-resolve + Matrix) — when auto-resolution can't fix an incident but has a resumable session, warden hands off to
WAITING_HUMANand listens in the room. A human's reply resumes the same agent session under the same allowlist (a chat message can't widen the agent's tools), then the spine re-verifies. Escalation carries theclaude --resume <session>id. -
Chat issue-reporting (0.19.0) — an allowlisted Matrix sender can open a tracked incident directly with
@warden report .../!warden report ...on the same room the guidance loop already listens in; gated by its own allowlist + per-sender rate limit, source-scoped so it can't collide with a detector finding, and escalates to a human unless a runbook is explicitly authored for it. Seedocs/mesh.md. -
Auto-draft Tier-0 runbooks (0.20.0) — the "border-collie" learning loop: after the LLM (Tier 1) resolves a novel incident with exactly one clean, reproducible mesh invoke, warden drafts a
DRAFTinvoke:runbook (auto: false) and attaches it to the Forgejo case; a human reviews, merges, and flipsauto: trueto make it a $0, no-LLM Tier-0 fix next time. Free-form/ambiguous fixes are never drafted. Seedocs/mesh.md. -
NATS TLS (0.21.0) — the ops bus can run over verified TLS: an opt-in
tls://mode with mandatory server-cert + hostname verification (no way to disable it) and optional mutual TLS, configured with three file-path env vars; a missing/unreadable CA or cert/key fails loud instead of falling back to plaintext. Seedocs/mesh.md. -
Fleet across envs (0.22.0) — a prod "hub" orchestrator can aggregate a read-only, view-only pane of sibling per-env orchestrators' incidents (
GET /fleet,warden_fleet_incidents{env,kind}): a remote orchestrator best-effort mirrors its own incidents to the hub over a second, publish-only, TLS-required connection, and the hub'sFleetViewis data-only — no actuation collaborator, so a remote incident can never trigger a local (or remote) exec, route, grant, or remediation. Gated on both sides; unset ⇒ byte-for-byte the prior single-env behavior. Seedocs/mesh.md. -
MCP bridge (read-only, 0.23.0) —
warden mcp(the optionalpip install "warden[mcp]"extra) runs a stdio MCP server so an operator's Claude Code/Desktop session can query live warden state during triage — open incidents, an incident's detail+timeline, topology, fleet — through four tools that proxy the orchestrator's own read-only HTTP endpoints; it holds no NATS credential, no DB handle, and no invoker/grant/exec collaborator, so it cannot mutate state or invoke a capability. Seedocs/mesh.md. -
sandbox.runephemeral fixers (0.24.0; mounted on/invokein 0.26.0) — a grant-gated capability to run a bounded fix in a TTL'd, resource-capped, one-shot, hardened container (pinned-digest image allowlist only, read-only rootfs,CapDrop: ALL, no host mounts ever, no fallback to the restart-proxy) that the caller cannot widen.POST /invokenow dispatchescapability: sandbox.runto it (grant-gated exactly like the mesh exec path); it stays inert without a configured spawn-proxy (WARDEN_SANDBOX_DOCKER_HOST). Seedocs/mesh.md. -
Roadmap sweep (0.26.0) — the last roadmap items: the fleet-hub first-connect retry/backoff (a hub down at startup is picked up later with no orchestrator restart), the
error_eventsexternal-event NATS ingest (opt-in, source-scoped,WARDEN_EVENTS_INGEST), theagent_auditpure-Python cross-source freshness/drift sweep, and mountingsandbox.runon/invoke(above). -
Off-box fate-sharing twin (Phase 5) — the primary warden reports its own liveness (
WARDEN_SELF_HEARTBEAT_URL→ a twin's/heartbeateach loop). A second warden on another host watches that key with the heartbeat detector, so if the primary — or its whole host — goes dark, the twin pages. The watcher is watched; they share fate only if both hosts die at once. Seedocs/twin.example.yaml.
All five phases are implemented and tested (0.26.0 landed agent_audit and the
error_events ingest detector; container-health remains a roadmap plugin — see
docs/design.md). Auto-resolution, guidance, and the twin are
opt-in and gated, so warden is safe to run in detect-and-notify mode from day one and graduate
toward autonomy at your pace.
Fleet / mesh (optional, NATS; M1–M6 through 0.26.0 — the 0.21–0.26 items (NATS TLS,
cross-env fleet, MCP bridge, sandbox.run + its /invoke mount, hub retry,
error_events) are detailed in docs/mesh.md). warden-core stays
the deterministic orchestrator: one orchestrator per environment, over a dedicated ops bus (the
nats extra), never a bus that carries business traffic. Agents (WARDEN_ROLE=sensor|actuator| node) now actually run: a sensor runs detectors and publishes findings, an actuator executes
typed capabilities from its manifest in a locked-down subprocess and replies, and both
self-register and re-announce on a lease. The orchestrator tracks them in a topology registry,
diffs an optional topology.yaml (desired state) against what actually announced, and serves it
at GET /topology plus fleet metrics on /metrics
(warden_agent_up/_info/_lease_age_seconds/_capability, warden_agents_registered,
warden_mesh_connected). A bus outage is always one SKIPPED incident (mesh-bus), never a
storm of per-agent pages. WARDEN_NATS_FINDINGS_SUBJECT (a single hardcoded subject) is
deprecated in favour of the versioned per-agent subject warden.findings.<env>.<agent_id>
with the {"v":1,...} envelope, which every sensor is expected to use going forward. Ingested
findings are untrusted, subject to the same fenced, allowlist-bounded handling as local
detectors. An operator (and, since 0.16.0, the Tier-1 LLM fallback) drives a capability through
warden invoke or a grant-gated POST /invoke on the orchestrator — a per-attempt HMAC grant
bound to {incident, capabilities, expiry} so the caller never needs a NATS credential. Full wire
protocol, manifest shape, the docker.restart capability, and the /invoke contract are in
docs/mesh.md.
Tier-0 deterministic self-heal is now live (0.15.0). A runbook can declare invoke: — a
capability call with literal params, routed to a fresh actuator and confirmed by a polling
verify check — that the spine tries before any LLM is ever consulted. When it resolves the
incident, it's a $0, no-LLM fix; only if it can't (or isn't configured) does the existing LLM
auto-resolver run, and only if neither can does the incident escalate to a human. A warden with
no mesh, or a runbook with no invoke:, behaves exactly as before — this is opt-in per runbook.
See Remediation tiers in docs/mesh.md for the full tier
sequence, the invoke: frontmatter, and the metrics it emits.
The LLM fallback can now act through the mesh too (0.16.0). When Tier 0 fails, Tier 1's LLM
is no longer stuck with a runbook author's hand-written allowed_tools — it's handed an exact,
fully-enumerated warden invoke <cap> --target <id> --param k=v toolset composed from the
routed actuator's own registered capabilities, plus a short-lived, per-attempt HMAC grant
bound to {incident, capabilities, expiry}. The child gets that grant and the orchestrator's URL,
never a NATS credential; a capability whose param space can't be safely and exactly enumerated
declines the whole composition and escalates to a human rather than handing the LLM a wildcard
tool. See Tier-1 composition in docs/mesh.md.
Verify can now nudge a remote sensor, and routing can be pinned ad hoc (0.17.0). A
sensor/node serves POST /recheck — re-run one of its own checks and re-publish
immediately — and Tier 0's polling verify can best-effort trigger it once, on its remote-observe
path only, via the mesh warden.recheck capability, instead of only ever waiting out that
sensor's own scheduled cadence; a trigger failure never fails or skips the observe, and with
nothing wired the verify step is byte-for-byte the pre-0.17.0 observe-only behavior. Separately,
warden invoke and POST /invoke can now pin a --host/--service locus directly (server-side
precedence: explicit target, then host/service, then env-wide) — the same routing a runbook's
invoke: scope: already used, now reachable from an ad hoc operator or Tier-1-composed call, not
only a declared runbook. See Remote recheck and Locus
routing in docs/mesh.md.
License
Apache-2.0.