Incidents investigated while you read the alert.
Elpis is an AI SRE that correlates logs, metrics, traces, deployments and Git changes into an evidence-backed root cause, proposes a risk-classified fix, and executes only after a human says yes, with a full audit trail. Open signup, no card, no sales call, and you bring your own LLM key, so there is zero markup on your AI spend.
Why Elpis?
Logs, metrics, traces, deployments, Git history and past incidents are collected in parallel by specialist agents, so there is no more tab-hopping through six consoles to form one hypothesis.
Every conclusion ships with the log lines, metric deltas, commits and deployments that support it, plus alternative hypotheses and a confidence score you can challenge.
A risk engine classifies every proposed action. Low-risk steps run automatically, anything riskier stops at an RBAC-gated approval screen, and destructive actions are prohibited outright.
Every investigation runs inside hard caps: 30 tool calls, 5 minutes, $0.10 of LLM spend, 3 retries. When a limit trips, the graph stops and says so instead of burning budget.
Investigations are embedded into pgvector memory and retrieved as incident RAG. Semantic caching reuses near-identical answers, and recurring root causes auto-generate runbooks.
From alert to verified recovery, with evidence at every step.
Elpis doesn't guess. Each stage produces structured output the next stage can check: classifications, evidence rows, hypotheses with confidence, risk-classified actions, approval records, and verification results.
An 18-node LangGraph workflow fans out across six evidence sources in parallel, then assembles one correlated picture of the incident.
Hypotheses are generated, tested against telemetry, and promoted to a root cause only when confidence clears the 0.8 gate. Otherwise, the graph gathers more evidence.
A planner proposes the smallest safe fix, such as restart, scale, rollback, or clear cache, classified by a risk table before anyone is paged.
Human-in-the-loop gates with JWT + four RBAC roles. Engineers run investigations, SREs approve fixes, admins own policy.
After execution, a verification step compares error rate and latency against the pre-incident baseline before declaring RECOVERED.
Each incident ends with a structured report covering summary, root cause, evidence, actions, and verification, then is written back into organizational memory.
Eleven stages. One shared state. Zero hidden reasoning.
Six agents collect evidence in parallel.
Logs come from Loki, host metrics from Prometheus, spans from your trace store, and real commits from the GitHub API. Each tool is narrowly scoped, times out safely, and returns structured rows. If the telemetry stack is down, collectors return empty results and the investigation still runs.
Autonomy with a hard leash.
The risk engine, not the model, decides who must approve. LOW actions run automatically, MEDIUM and HIGH require SRE approval, and four CRITICAL actions like delete_database are prohibited outright. Guardrails run as a graph node, so an over-budget investigation is blocked before the next LLM call.
Remembered, measured, and cost-aware.
Every finished incident is embedded into pgvector and retrievable as context for the next one. A three-tier model router sends classification to cheap models and root-cause work to reasoning models, semantic caching skips repeats, and log compression trims raw telemetry to ≤30 patterns before it ever reaches the LLM.
Quality, safety and cost are scored on every run.
A built-in evaluator replays all 24 fault scenarios against the full graph and scores two things that matter: did it name the correct root cause, and did it propose a safe action? Prometheus counters track LLM calls and investigations, and OpenTelemetry traces export every graph transition.
Try it free in your browser.
Sign up with an email and password and you start at SRE. Inject a fault, hit investigate, and watch the agent graph, timeline and evidence fill in live. Then approve the fix and see it verify recovery. Every account gets private incidents, your own OpenAI-compatible key in Settings (or the free built-in model), and demo logins viewer / engineer / sre / admin.