OpenSRE investigates incidents with an AI agent runtime built on the Claude Agent SDK — not a hand-rolled graph orchestration layer. A root investigator agent plans the work, loads integration skills on demand, can dispatch specialist subagents for independent sub-questions, and streams progress back to clients in real time. Episodic memory and a service topology knowledge graph both live in Neo4j; configuration and full investigation traces live in PostgreSQL.
Web UI Slack Bot Teams Bot
(Next.js) (Socket Mode) (Microsoft Teams)
\ | /
\ | /
v v v
┌─────────────────────────────────────────┐
│ sre-agent │
└──────────────────┬──────────────────────┘
│
┌───────────────────┼──────────────┬─────────────────┐
v v v v
┌──────────┐ ┌──────────┐ ┌──────────┐ ┌─────────────┐
│ config- │ │ Postgres │ │ Neo4j │ │ LiteLLM │
│ service │ │ agent │ │ memory + │ │ (optional │
│ │ │ runs + │ │ knowledge│ │ proxy) │
│ │ │ config │ │ graph │ │ │
└──────────┘ └──────────┘ └──────────┘ └─────────────┘
Entry points
POST /investigate and related thread endpoints on sre-agent, used by the clients above.InteractiveAgentSession (Claude Agent SDK).investigator, falling back to planner) plans the work, loads skills on demand, and dispatches specialists via the SDK Task tool.memory-search and infrastructure-neo4j. There is no automatic pre-injection of past episodes or topology into the first prompt; recall is agent-driven.:Episode (one per conversation) and persists the full tool trace to PostgreSQL for replay in the web UI.There is no LangGraph, no fixed pipeline of nodes. The SDK session itself is the runtime.
The investigation engine is InteractiveAgentSession, which wraps the SDK's ClaudeSDKClient and maps SDK messages onto OpenSRE's own SSE event protocol.
investigator preferred, then planner). Its system prompt comes from a nested config path (agents.{agent_id}.prompt.system) so it can be customized per team without touching other teams.AgentDefinition objects, built from the team's configured agent topology. The root delegates to them via the SDK Task tool — optionally in the background, so multiple specialists can work concurrently.SKILL.md files plus optional scripts. The SDK Skill tool loads lightweight metadata for every skill up front (roughly 100 tokens each) and the full skill content only when the agent actually chooses to use it. Skill scripts then run via Bash under that loaded skill.The session follows the SDK-recommended pattern rather than the simpler single-response API:
client.query() with a streaming user-message generator, running concurrently with the receive loop.receive_messages() — not receive_response() alone.receive_messages() doesn't stop automatically at the first result. That matters because a turn isn't necessarily over just because the SDK reports one — background subagents may still be running.
Subagents started in the background emit their own lifecycle messages on the SDK stream. While any are still outstanding:
thought event, not a final answer.background_waiting event with the pending task IDs.You can add context to a running investigation without waiting for it to finish:
POST /threads/{thread_id}/queue-messagemessage_queued event when queued text is merged into the agent's turnQueued messages are debounced for a fraction of a second and merged into a single numbered guidance block, so several quick follow-ups don't each interrupt the agent separately. Any remaining queue is flushed at the turn boundary so nothing gets silently dropped.
| Type | Purpose |
|---|---|
thought | Agent reasoning / interim narration |
tool_start / tool_end | Skill, Bash, Task, and other tool lifecycle |
task_started / task_notification | Background subagent lifecycle |
background_waiting | Parent waiting on outstanding background tasks |
message_queued | Mid-run user message accepted or consumed |
question | Agent asking a clarifying question |
result | Terminal answer for the turn |
error | Failure or timeout |
Investigations run under a wall-clock timeout independent of the SDK's own turn-count limit — the two are separate caps and raising one doesn't raise the other.
Default (make dev / self-host) | Sandbox mode | |
|---|---|---|
| Server | server_simple.py | server.py |
| Agent process | Same process as the API | Per-thread isolated pod |
| Isolation | Trusted local / single-tenant use | Filesystem and network isolation |
| Skills | Copied to a per-thread workspace | Baked into the sandbox image |
The default mode is what make dev runs and what typical self-hosting uses. Sandbox mode is a separate, production-oriented stack for stronger isolation between concurrent investigations and isn't required to run OpenSRE.
OpenSRE stores a condensed summary of every conversation as a Neo4j :Episode, and recalls similar past episodes on demand via semantic vector search — the agent decides when to search, based on evidence gathered so far, not a fixed rule fired on every alert. See Episodic Memory for the full model, including how strategy playbooks get synthesized from repeated episodes.
Neo4j also holds service topology — dependencies and blast radius — in the same database as episodic memory, queried agent-driven via a dedicated skill once an affected service is known. See Knowledge Graph for what's tracked and how it's populated today.
Skills are the primary integration surface: a SKILL.md methodology doc plus optional scripts, organized by domain (Kubernetes and cloud, observability, databases, incident tooling, version control, and more). See Investigation Skills for the full catalog and how to add your own.
config-service is the control plane:
${VAR} substitution — never hardcoded in config.See Configuration for the full reference.
cp .env.example .env # set ANTHROPIC_API_KEY
make dev # core stack
make dev-slack # core + Slack bot
make dev-teams # core + Teams bot
| Service | Port | Role |
|---|---|---|
| Web UI | 3002 | Next.js console |
| sre-agent | 8001 | Investigation API + SSE |
| config-service | 8081 | Config, tokens, agent-run storage |
| PostgreSQL | 5433 | Config DB + agent runs |
| Neo4j Browser / Bolt | 7475 / 7688 | Memory + topology graph |
| Slack bot | — | Optional (--profile slack / make dev-slack) |
| Teams bot | 3978 | Optional (--profile teams / make dev-teams) |
| LiteLLM | 4001 | Optional (--profile litellm) |
By default the agent calls Anthropic directly (ANTHROPIC_API_KEY). To use OpenRouter or another provider, start the LiteLLM profile and set ANTHROPIC_BASE_URL to the proxy.