Spots

Why I Stopped Using Generic LLM Wrappers for My Agent

Building an AI agent is easy when you only need it to generate a response. Building one that remembers what happened five calls ago is a different problem. While building DealMemory, a sales intelligence agent with persistent memory, I ran into a surprisingly simple bug:

That small issue cost me hours of debugging

That small issue cost me hours of debugging and taught me an important lesson about working with LLMs, memory systems, and agent frameworks. The Problem: LLMs Don't Remember Your Deals Imagine a sales representative has spoken with Acme Corp five times. The CTO raised API latency concerns. The CFO pushed back on pricing three times. Salesforce was mentioned as a competitor. The security team requested a SOC 2 report. Legal was discussing a 10% volume discount. All this information might exist in CRM notes, but a normal LLM doesn't automatically know it. "Review the stakeholder map, identify objections, and prepare for pricing discussions." "Review the stakeholder map, identify objections, and prepare for pricing discussions." Technically correct, but not very useful. That's what we wanted to solve with DealMemory. DealMemory gives every deal its own memory bank.

Hindsight — persistent memory Streamlit — user interface

Hindsight — persistent memory Streamlit — user interface Python — application logic The important part is that the LLM doesn't start with an empty context. It receives relevant information retrieved from the deal's history. Hindsight: Retain, Recall, Reflect Hindsight gave us three important operations. Each deal gets its own memory bank, so Acme's information doesn't mix with Globex or Initech. When the rep asks for a briefing: This is where I ran into the bug. I initially assumed the retrieved memory would behave like a typical LLM response. I tried accessing the result using: But the Hindsight recall result exposed the actual stored memory through: So the context needed to be constructed like this: The frustrating part was that .content is something we commonly encounter when working with LLM responses. But a memory retrieval result isn't necessarily an LLM response.

Different layers of an AI application can have

Different layers of an AI application can have completely different object interfaces. When debugging these systems, checking the actual object is often more useful than guessing: That simple step would have saved me a lot of time. Recall gives us the relevant information. But Hindsight also provides reflect(). The difference is roughly: For example, recall might show: Reflection can turn that history into a useful pattern: The CFO consistently shows pricing sensitivity, but ROI-based positioning has improved engagement. The CFO consistently shows pricing sensitivity, but ROI-based positioning has improved engagement. That's much closer to a real sales copilot. "Schedule a discovery call and prepare for pricing objections." "Schedule a discovery call and prepare for pricing objections."

Stakeholder: CTO raised API latency concerns. Objection: CFO

Stakeholder: CTO raised API latency concerns. Objection: CFO pushed on pricing three times. Competitor: Salesforce mentioned during discovery. Open items: SOC 2 report and 10% volume discount. Next action: Send an updated proposal using ROI framing and follow up with legal.

Stakeholder: CTO raised API latency concerns. Objection: CFO

Stakeholder: CTO raised API latency concerns. Objection: CFO pushed on pricing three times. Competitor: Salesforce mentioned during discovery. Open items: SOC 2 report and 10% volume discount. Next action: Send an updated proposal using ROI framing and follow up with legal. The LLM didn't become smarter. The context became better. One More Debugging Detail: Async Indexing There was another small issue we had to handle.

Hindsight's memory retention is asynchronous, so newly stored

Hindsight's memory retention is asynchronous, so newly stored information may take a short time before it becomes searchable. Our flow therefore waits briefly after logging a call: This prevents the confusing situation where you store information and immediately wonder: "Why can't my agent find it?" "Why can't my agent find it?" The biggest lesson wasn't simply to use .text instead of .content.

Don't assume every component in an AI pipeline

Don't assume every component in an AI pipeline follows the same response format. Memory systems, LLM APIs, retrieval systems, and tools can all return different objects. Inspect the actual data before building assumptions around it. More importantly, building DealMemory changed how I think about AI agents.

"How do I make my LLM smarter?" sometimes

"How do I make my LLM smarter?" sometimes the better question is: "How do I give my LLM better context?" That's what DealMemory is built around: "How do I make my LLM smarter?" sometimes the better question is: "How do I give my LLM better context?" That's what DealMemory is built around: text Generic ↓ Specific ↓ Pattern-aware And that small .text bug was one of the debugging lessons that helped us get there. For further actions, you may consider blocking this person and/or reporting abuse

News

Why I Stopped Using Generic LLM Wrappers for My Agent

Building an AI agent is easy when you only need it to generate a response.

@spots #dev
Source: Dev.to
See more like this