Spots

I Designed Incident Response Around Persistent Memory

I Gave Incident Response a Memory With Hindsight The first useful question during an outage is often not “what could be wrong?” but “have we seen this before?”

I built Incident-Memory-Copilot around that question. The system

I built Incident-Memory-Copilot around that question. The system combines an incident-response workflow with persistent organizational memory so that a new incident can be investigated using what the organization learned from previous incidents.

The application is an incident operations console. It

The application is an incident operations console. It gives an engineer a place to inspect active incidents, search historical memory, review runbooks and postmortems, investigate a new incident, and explicitly teach the system what was learned after resolution.

The important architectural decision is that Hindsight is

The important architectural decision is that Hindsight is not treated as a secondary search box. It sits inside the incident lifecycle. At a high level, the flow is:

DATA SOURCES │ ┌─────────────┼──────────────┐ ▼ ▼ ▼ Rootly

DATA SOURCES │ ┌─────────────┼──────────────┐ ▼ ▼ ▼ Rootly PagerDuty PagerDuty Logs Incident Docs Postmortems │ │ │ └─────────────┼──────────────┘ │ ▼ SYNTHETIC ORGANIZATIONAL DATA │ ┌─────────────┼──────────────┐ ▼ ▼ ▼ 100–150 50–100 50–100 Incidents Runbooks Postmortems │ │ │ └─────────────┼──────────────┘ ▼ HINDSIGHT CLOUD │ ┌──────────┼──────────┐ ▼ ▼ ▼ RETAIN RECALL REFLECT │ │ │ └──────────┼──────────┘ ▼ INCIDENT RESPONSE AGENT │ ▼ EVIDENCE-BACKED ACTIONS │ ▼ HUMAN ENGINEER │ ▼ POSTMORTEM │ └──────→ RETAIN

The data foundation combines operational incident information, incident-response

The data foundation combines operational incident information, incident-response knowledge, postmortems, realistic synthetic incidents, runbooks, and postmortems. Hindsight becomes the layer that turns that accumulated information into reusable memory. The application then has a simple loop: Recall → Investigate → Human decision → Resolve → Retain That loop is more important to me than any individual UI screen. Why I wanted memory in the incident workflow An LLM can already explain an HTTP 502, list possible database problems, or suggest checking a deployment. That is not the difficult part. The difficult part is knowing what happened in this environment before.

Suppose a payment service is returning HTTP 502

Suppose a payment service is returning HTTP 502 responses. A generic assistant might recommend checking the gateway, application health, upstream dependencies, database connectivity, and recent deployments. Those are reasonable checks.

But suppose the organization previously had a related

But suppose the organization previously had a related incident where a deployment changed database connection-pool limits. The team discovered that restarting the API gateway made the situation worse because it created a traffic spike, while rolling back the deployment restored the correct connection-pool configuration. That historical experience is much more useful than a generic list of possible causes. The incident console captures exactly this kind of information. The current incident screen represents an example around a Payment API returning HTTP 502 errors.

The incident contains operational context such as the

The incident contains operational context such as the service, severity, error, current CPU and memory utilization, recent deployment, and explanatory description. The investigation then moves through three memory-oriented stages: The key point is that historical information is presented alongside the current evidence. The agent is not simply saying: “This happened before, so do the same thing.” Instead, the previous incident becomes evidence that an engineer can compare against the current situation.

That distinction matters in production systems because infrastructure

That distinction matters in production systems because infrastructure changes over time. The same symptom can have a different cause. Retain, Recall, and Reflect The integration with Hindsight follows three operations.

News

I Designed Incident Response Around Persistent Memory

I Gave Incident Response a Memory With Hindsight The first useful question during an outage is often not “what could be wrong?” but “have we seen this before?”

@spots #dev
Source: Dev.to
See more like this