[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f39kyqizprivmm":3},{"_id":4,"slug":5,"title":6,"subtitle":7,"kind":8,"cards":9,"tags":58,"categories":60,"source":62,"lang":65,"author":66,"audioState":69,"stats":70,"publishedAt":73,"renderer":74},"6abd04fdca21c797c7ea173e","when-does-an-ai-agent-actually-earn-its-cost-i-measured-it-o-3acefa09","# When does an AI agent actually earn its cost? I measured it on 100 questions","Built for the TigerGraph Agentic GraphRAG Hackathon.","news",[10,13,18,23,28,33,38,43,48,53],{"headline":6,"body":11,"imageUrl":12,"sourceImageUrl":12},"Built for the TigerGraph Agentic GraphRAG Hackathon. Repo and dashboard linked at the end.","https:\u002F\u002Fmedia2.dev.to\u002Fdynamic\u002Fimage\u002Fwidth=1200,height=627,fit=cover,gravity=auto,format=auto\u002Fhttps%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5u5ayb5ht26a0ik3cw48.png",{"headline":14,"body":15,"imageUrl":16,"images":17},"RAG retrieves text. GraphRAG adds structure. Agentic GraphRAG","RAG retrieves text. GraphRAG adds structure. Agentic GraphRAG lets a model plan its own investigation. Everyone assumes the third one is better — but by how much, and on which questions? That's the question this hackathon asked, so I built six pipelines over one corpus and measured them against the same 100 questions, with the same generation model throughout.","\u002Fapi\u002Fmedia\u002Fposts\u002Fwhen-does-an-ai-agent-actually-earn-its-cost-i-measured-it-o-3acefa09\u002F1.webp",{"local":16},{"headline":19,"body":20,"imageUrl":21,"images":22},"The short version: exact match went from 67%","The short version: exact match went from 67% to 99%, a 32-point jump. But the interesting part isn't that number. It's that an agent by itself only got me 3 of those 32 points.","\u002Fapi\u002Fmedia\u002Fposts\u002Fwhen-does-an-ai-agent-actually-earn-its-cost-i-measured-it-o-3acefa09\u002F2.webp",{"local":21},{"headline":24,"body":25,"imageUrl":26,"images":27},"2,951 Wikipedia articles in TigerGraph Savanna. 100 evaluation","2,951 Wikipedia articles in TigerGraph Savanna. 100 evaluation questions across five types: lookup, temporal, multi-hop, superlative and aggregation. Every pipeline uses gemini-3.1-flash-lite — a cheap model, deliberately — with local BGE embeddings and TigerGraph's native vector index. No pipeline gets a better model than any other.","\u002Fapi\u002Fmedia\u002Fposts\u002Fwhen-does-an-ai-agent-actually-earn-its-cost-i-measured-it-o-3acefa09\u002F3.webp",{"local":26},{"headline":29,"body":30,"imageUrl":31,"images":32},"Scoring is exact match against the gold answer","Scoring is exact match against the gold answer, computed with no model in the loop. I also ran an LLM judge, and I'll come back to why I stopped trusting it. The failure that taught me the most","\u002Fapi\u002Fmedia\u002Fposts\u002Fwhen-does-an-ai-agent-actually-earn-its-cost-i-measured-it-o-3acefa09\u002F4.webp",{"local":31},{"headline":34,"body":35,"imageUrl":36,"images":37},"Plain RAG scored 67%. GraphRAG — entity linking","Plain RAG scored 67%. GraphRAG — entity linking plus one-hop traversal — also scored 67%. Identical. That was my first surprise.","\u002Fapi\u002Fmedia\u002Fposts\u002Fwhen-does-an-ai-agent-actually-earn-its-cost-i-measured-it-o-3acefa09\u002F5.webp",{"local":36},{"headline":39,"body":40,"imageUrl":41,"images":42},"Breaking it down by question type showed why","Breaking it down by question type showed why. On aggregation questions — \"how many cycling events had more than 30 competitors?\" — RAG got 1 out of 21. GraphRAG got 0.","\u002Fapi\u002Fmedia\u002Fposts\u002Fwhen-does-an-ai-agent-actually-earn-its-cost-i-measured-it-o-3acefa09\u002F6.webp",{"local":41},{"headline":44,"body":45,"imageUrl":46,"images":47},"The reason is structural, and no amount of","The reason is structural, and no amount of prompt engineering touches it. Answering that question requires every matching document; for one question, 43 of them. Retrieval fetches the top five and counts those. The model then confidently reports a number that is simply the size of what it was shown.","\u002Fapi\u002Fmedia\u002Fposts\u002Fwhen-does-an-ai-agent-actually-earn-its-cost-i-measured-it-o-3acefa09\u002F7.webp",{"local":46},{"headline":49,"body":50,"imageUrl":51,"images":52},"My LLM-extracted entity graph didn't help either. It","My LLM-extracted entity graph didn't help either. It had 12,000 entities with free-text relationship labels — but \"competitors: 43\" was never a property you could filter or count on. I had modelled the prose and not the facts.","\u002Fapi\u002Fmedia\u002Fposts\u002Fwhen-does-an-ai-agent-actually-earn-its-cost-i-measured-it-o-3acefa09\u002F8.webp",{"local":51},{"headline":54,"body":55,"imageUrl":56,"images":57},"Every one of those articles opens with an","Every one of those articles opens with an infobox: event, games, venue, date, competitors, nations, gold medallist. So I built a second layer in the graph directly from those boxes — OlympicEvent linked to its Games, Sport and Venue, plus a PREV_GAMES edge so \"the Olympics before 2016\" is one hop.","\u002Fapi\u002Fmedia\u002Fposts\u002Fwhen-does-an-ai-agent-actually-earn-its-cost-i-measured-it-o-3acefa09\u002F9.webp",{"local":56},[59],"dev",[61],"Technology",{"name":63,"url":64},"Dev.to","https:\u002F\u002Fdev.to\u002Fanant_kumar_bf65d0d3994d3\u002F-when-does-an-ai-agent-actually-earn-its-cost-i-measured-it-on-100-questions-4j5l","en",{"handle":67,"displayName":68},"spots","Spots","queued",{"views":71,"likes":72,"saves":72,"shares":72,"completions":72,"opens":72,"skips":72,"depthSum":72},1,0,"2026-09-30T12:47:57.034Z","local"]