AI-powered autonomous incident response

Incidents investigated while you read the alert.

Elpis is an AI SRE that correlates logs, metrics, traces, deployments and Git changes into an evidence-backed root cause, proposes a risk-classified fix, and executes only after a human says yes, with a full audit trail. Open signup, no card, no sales call, and you bring your own LLM key, so there is zero markup on your AI spend.

Why Elpis?

1
Correlate evidence across every source in one pass.

Logs, metrics, traces, deployments, Git history and past incidents are collected in parallel by specialist agents, so there is no more tab-hopping through six consoles to form one hypothesis.

collect_logs ........ done
collect_metrics ...... done
collect_traces ....... done
collect_deployments .. done
git_analysis ......... done
historical_search .... done
2
Root causes backed by evidence, not vibes.

Every conclusion ships with the log lines, metric deltas, commits and deployments that support it, plus alternative hypotheses and a confidence score you can challenge.

root_cause = connection pool exhaustion
confidence = 0.93
✓ DB connections reached 98/100
✓ timeout errors up 14×
✓ pool size changed in v1.8.3
3
Nothing touches production without a human.

A risk engine classifies every proposed action. Low-risk steps run automatically, anything riskier stops at an RBAC-gated approval screen, and destructive actions are prohibited outright.

risk_engine → HIGH
rollback_deployment
approval → SRE role required
awaiting your decision…
delete_database → PROHIBITED
4
Guardrails on tool calls, time, and spend.

Every investigation runs inside hard caps: 30 tool calls, 5 minutes, $0.10 of LLM spend, 3 retries. When a limit trips, the graph stops and says so instead of burning budget.

max_tool_calls 30
max_duration    300s
max_cost       $0.10
max_retries    3
budget spent $0.04 ▮▮▮▯▯▯▯▯▯▯
5
It remembers, and gets cheaper every time.

Investigations are embedded into pgvector memory and retrieved as incident RAG. Semantic caching reuses near-identical answers, and recurring root causes auto-generate runbooks.

rag.search_similar(k=3) → 3 hits
cache redis → postgres → pgvector
recurrence ≥ 3 → runbook draft
route classify→cheap root_cause→reasoning
compress_logs 10k lines → ≤30 patterns
Built withFastAPILangGraphNeon Postgres + pgvectorOllama (llama3.1:8b)Next.js 16JWT / RBACPrometheusOpenTelemetryLokiGitHub APIDockerKubernetes
Platform · one loop, end to end

From alert to verified recovery, with evidence at every step.

Elpis doesn't guess. Each stage produces structured output the next stage can check: classifications, evidence rows, hypotheses with confidence, risk-classified actions, approval records, and verification results.

Investigate

An 18-node LangGraph workflow fans out across six evidence sources in parallel, then assembles one correlated picture of the incident.

intake → classify → collect → assemble
Diagnose

Hypotheses are generated, tested against telemetry, and promoted to a root cause only when confidence clears the 0.8 gate. Otherwise, the graph gathers more evidence.

hypothesis → test → confidence gate
Remediate

A planner proposes the smallest safe fix, such as restart, scale, rollback, or clear cache, classified by a risk table before anyone is paged.

12 typed actions · 4 risk levels
Approve

Human-in-the-loop gates with JWT + four RBAC roles. Engineers run investigations, SREs approve fixes, admins own policy.

viewer · engineer · sre · admin
Verify

After execution, a verification step compares error rate and latency against the pre-incident baseline before declaring RECOVERED.

before/after comparison
Report

Each incident ends with a structured report covering summary, root cause, evidence, actions, and verification, then is written back into organizational memory.

report → RAG memory → runbooks
How it works · the investigation loop

Eleven stages. One shared state. Zero hidden reasoning.

Alert→Classify→Collect evidence (×6 parallel)→Guardrail check→Hypotheses→Root cause→Risk assessment→Human approval→Execute→Verify→Report

Six agents collect evidence in parallel.

Logs come from Loki, host metrics from Prometheus, spans from your trace store, and real commits from the GitHub API. Each tool is narrowly scoped, times out safely, and returns structured rows. If the telemetry stack is down, collectors return empty results and the investigation still runs.

search_logsget_metricsanalyze_tracesget_recent_deploymentsgit_analysishistorical_search

Autonomy with a hard leash.

The risk engine, not the model, decides who must approve. LOW actions run automatically, MEDIUM and HIGH require SRE approval, and four CRITICAL actions like delete_database are prohibited outright. Guardrails run as a graph node, so an over-budget investigation is blocked before the next LLM call.

risk table (12 actions)RBAC (4 roles)audit trail30s token rotation

Remembered, measured, and cost-aware.

Every finished incident is embedded into pgvector and retrievable as context for the next one. A three-tier model router sends classification to cheap models and root-cause work to reasoning models, semantic caching skips repeats, and log compression trims raw telemetry to ≤30 patterns before it ever reaches the LLM.

incident RAGmodel routingsemantic cachecontext compressioncost analytics
By the numbers · measured, not claimed
18
workflow nodes in the investigation graph
24
reproducible fault scenarios (8 base + 16 variants)
$0.10
hard LLM budget per incident
4
RBAC roles gating every action
30
tool-call cap per investigation
6
specialist agents collecting evidence in parallel
384
dimension pgvector embeddings for incident RAG
26
passing tests in the CI suite
Evaluation · the AI is graded too

Quality, safety and cost are scored on every run.

A built-in evaluator replays all 24 fault scenarios against the full graph and scores two things that matter: did it name the correct root cause, and did it propose a safe action? Prometheus counters track LLM calls and investigations, and OpenTelemetry traces export every graph transition.

root-cause accuracysafe-action rateelpis_llm_calls_totalOTLP → Jaeger
Get started · open signup, no install

Try it free in your browser.

Sign up with an email and password and you start at SRE. Inject a fault, hit investigate, and watch the agent graph, timeline and evidence fill in live. Then approve the fix and see it verify recovery. Every account gets private incidents, your own OpenAI-compatible key in Settings (or the free built-in model), and demo logins viewer / engineer / sre / admin.