Spots

INCIDEX

Production incidents rarely happen in isolation.

When a Payment API starts returning 500 errors

When a Payment API starts returning 500 errors at 2 AM, the useful information is often not in the current alert. It is buried in incidents that happened weeks or months earlier: what failed, what engineers checked, what remediation worked, and what initially looked plausible but turned out to be wrong. I wanted to build an incident-response agent that could actually use that history.

The result is an Incident Learning & Response

The result is an Incident Learning & Response Agent that combines an AI investigation layer with persistent organizational memory using Hindsight. The interesting part wasn't simply getting an LLM to suggest a root cause. It was making previous incident experience available at the right moment—and then feeding verified outcomes back into memory so future investigations could benefit from them. The problem: every incident starts too cold An LLM can be very good at reasoning about an incident description.

"Payment API is returning intermittent HTTP 500 errors

"Payment API is returning intermittent HTTP 500 errors during heavy traffic. Database connection usage is near its configured limit." It can recognize that database connection exhaustion is a plausible explanation. The model doesn't inherently know what happened the last time our Payment API experienced the same pattern.

Maybe an earlier incident showed that increasing the

Maybe an earlier incident showed that increasing the connection pool and restarting the service resolved the issue. Maybe another incident looked similar but was actually caused by something else. Maybe an engineer discovered an important operational detail that isn't present in today's alert. Without persistent memory, that context has to be manually added to every prompt. That is exactly the problem I wanted to solve.

Instead of treating every production incident as a

Instead of treating every production incident as a completely new reasoning problem, I wanted the system to reuse relevant engineering experience. The application has a web interface where an engineer can submit an incident with: The frontend communicates with two n8n workflows.

The first workflow handles investigation. It prepares the

The first workflow handles investigation. It prepares the incident, retrieves relevant historical memories from Hindsight, and passes the current incident together with that context to an LLM for analysis.

The second workflow handles the engineer's decision and

The second workflow handles the engineer's decision and outcome. Once an engineer reviews the recommendation and records what actually happened, the result is stored back in Hindsight. The important part is that these aren't two disconnected AI calls. They form a learning loop.

Current Incident ↓ Hindsight Recall ↓ Historical Experience

Current Incident ↓ Hindsight Recall ↓ Historical Experience ↓ AI Investigation ↓ Engineer Review ↓ Actual Outcome ↓ Hindsight Retain ↓ Future Investigations That last step is the reason memory matters. A successful incident isn't just closed. Its outcome becomes useful context for the next incident. I didn't want to build a simple table of previous incidents and dump the entire database into every LLM prompt. The useful question isn't: "Show me every incident we've ever had." "Which previous experiences are relevant to this incident?" That's where Hindsight fits naturally.

Hindsight provides persistent memory through operations such as

Hindsight provides persistent memory through operations such as retain and recall. The system can store durable information and later retrieve relevant context rather than requiring the entire history to be manually supplied to the model. I created a dedicated Hindsight memory bank for the incident-response system:

News

INCIDEX

Production incidents rarely happen in isolation.

@spots #dev
Source: Dev.to
See more like this