
# When does an AI agent actually earn its cost? I measured it on 100 questions
Built for the TigerGraph Agentic GraphRAG Hackathon. Repo and dashboard linked at the end.
Spots 

Built for the TigerGraph Agentic GraphRAG Hackathon. Repo and dashboard linked at the end.

RAG retrieves text. GraphRAG adds structure. Agentic GraphRAG lets a model plan its own investigation. Everyone assumes the third one is better — but by how much, and on which questions? That's the question this hackathon asked, so I built six pipelines over one corpus and measured them against the same 100 questions, with the same generation model throughout.
The short version: exact match went from 67% to 99%, a 32-point jump. But the interesting part isn't that number. It's that an agent by itself only got me 3 of those 32 points.
2,951 Wikipedia articles in TigerGraph Savanna. 100 evaluation questions across five types: lookup, temporal, multi-hop, superlative and aggregation. Every pipeline uses gemini-3.1-flash-lite — a cheap model, deliberately — with local BGE embeddings and TigerGraph's native vector index. No pipeline gets a better model than any other.
Scoring is exact match against the gold answer, computed with no model in the loop. I also ran an LLM judge, and I'll come back to why I stopped trusting it. The failure that taught me the most
Plain RAG scored 67%. GraphRAG — entity linking plus one-hop traversal — also scored 67%. Identical. That was my first surprise.
Breaking it down by question type showed why. On aggregation questions — "how many cycling events had more than 30 competitors?" — RAG got 1 out of 21. GraphRAG got 0.
The reason is structural, and no amount of prompt engineering touches it. Answering that question requires every matching document; for one question, 43 of them. Retrieval fetches the top five and counts those. The model then confidently reports a number that is simply the size of what it was shown.
My LLM-extracted entity graph didn't help either. It had 12,000 entities with free-text relationship labels — but "competitors: 43" was never a property you could filter or count on. I had modelled the prose and not the facts.
Every one of those articles opens with an infobox: event, games, venue, date, competitors, nations, gold medallist. So I built a second layer in the graph directly from those boxes — OlympicEvent linked to its Games, Sport and Venue, plus a PREV_GAMES edge so "the Olympics before 2016" is one hop.
Built for the TigerGraph Agentic GraphRAG Hackathon.
