[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f2duw8gfuy26vu":3},{"_id":4,"slug":5,"title":6,"subtitle":7,"kind":8,"cards":9,"tags":33,"categories":35,"source":37,"lang":40,"author":41,"audioState":44,"stats":45,"publishedAt":48,"renderer":49},"6abcac4dca21c797c7ea0434","the-day-all-our-ai-agents-stopped-working-742f5db5","The Day All Our AI Agents Stopped Working 💀","Kanvas is right in the middle of a presentation.","news",[10,13,18,23,28],{"headline":6,"body":11,"imageUrl":12,"sourceImageUrl":12},"Kanvas is right in the middle of a presentation. The tests from the day before? All successful. Our PO was happy with the results. I had my protein shake that morning. I was prepared for the presentation. Then, out of nowhere, I get a call: — What’s going on with our agents? None of them are working. Not a single one is working. 💀 What the hell is going on? The codebase hasn’t changed in the last 24 hours. Our core is still operational. The business logic is working. So I check our monitoring system and find the problem: After that unnecessarily dramatic introduction, let me explain what actually happened. Our AI models experienced a usage spike right in the middle of a presentation. And that forced us to ask ourselves a question we hadn't seriously considered before: How do we make AI agents highly available? Kanvas Agent has become one of our flagship products.","https:\u002F\u002Fmedia2.dev.to\u002Fdynamic\u002Fimage\u002Fwidth=1200,height=627,fit=cover,gravity=auto,format=auto\u002Fhttps%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0wjhed5aixldaaekw2dg.png",{"headline":14,"body":15,"imageUrl":16,"images":17},"And with Kanvas, we’ve always had something very","And with Kanvas, we’ve always had something very important: Our core is under our control. Business rules, resources, APIs, infrastructure — all of it goes through systems we manage. If something fails, we can investigate it. We can deploy another instance. We can change the infrastructure.","\u002Fapi\u002Fmedia\u002Fposts\u002Fthe-day-all-our-ai-agents-stopped-working-742f5db5\u002F1.webp",{"local":16},{"headline":19,"body":20,"imageUrl":21,"images":22},"We have redundancy, autoscaling, health checks, blue\u002Fgreen deployments","We have redundancy, autoscaling, health checks, blue\u002Fgreen deployments, and years of knowledge about building highly available systems. But now we're living through this new wave of AI. And AI introduced a new problem. You can have your infrastructure working perfectly. Your database can be healthy. Your servers can be running without a problem. But if the provider powering your agents stops responding... When part of your infrastructure depends on a third-party AI provider, there's something you simply don't control. At some point, all you can do is trust them. our faith wasn't enough. 😂 Fortunately, we managed to save the presentation thanks to our PO, Estrella. But the incident left us with a problem we needed to solve. Our Lead came up with a simple idea: If a request to an AI model fails, the request shouldn't die with that model. There should always be another route.","\u002Fapi\u002Fmedia\u002Fposts\u002Fthe-day-all-our-ai-agents-stopped-working-742f5db5\u002F2.webp",{"local":21},{"headline":24,"body":25,"imageUrl":26,"images":27},"If Model A isn't available, try Model B","If Model A isn't available, try Model B. If the entire provider is having problems, move to Provider B. And that's when we started implementing something we internally call Routing. More specifically, one of the main strategies behind our router is fallback. Our agentic framework can now receive multiple models and multiple providers. A request might start with our primary model. If that model fails, the request doesn't die. The router tries the next available model. And if the problem affects the entire provider, we can move to another provider completely. Model A ❌ → Model B ❌ → Provider B → Model C ✅ Every production problem leaves you with a lesson. You can't build highly available AI agents while depending on a single model or a single provider.","\u002Fapi\u002Fmedia\u002Fposts\u002Fthe-day-all-our-ai-agents-stopped-working-742f5db5\u002F3.webp",{"local":26},{"headline":29,"body":30,"imageUrl":31,"images":32},"Today, we can assign multiple models from the","Today, we can assign multiple models from the same provider and also configure multiple providers with different models. And that opened another interesting door For further actions, you may consider blocking this person and\u002For reporting abuse","\u002Fapi\u002Fmedia\u002Fposts\u002Fthe-day-all-our-ai-agents-stopped-working-742f5db5\u002F4.webp",{"local":31},[34],"dev",[36],"Technology",{"name":38,"url":39},"Dev.to","https:\u002F\u002Fdev.to\u002Ffrederickpeal\u002Fthe-day-all-our-ai-agents-stopped-working-1iia","en",{"handle":42,"displayName":43},"spots","Spots","queued",{"views":46,"likes":47,"saves":47,"shares":47,"completions":47,"opens":47,"skips":47,"depthSum":47},1,0,"2026-09-30T06:29:33.739Z","local"]