[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f1m9yehb6cgd5i":3},{"_id":4,"slug":5,"title":6,"subtitle":7,"kind":8,"cards":9,"tags":58,"categories":60,"source":62,"lang":65,"author":66,"audioState":69,"stats":70,"publishedAt":73,"renderer":74},"6abc8953ca21c797c7ea0144","spark-structured-streaming-cant-save-you-from-bad-architectu-492b3025","Spark Structured Streaming Can’t Save You From Bad Architecture","Roughly 80% of the \"exactly-once\" guarantees I see implemented in production financial pipelines are actually \"at-least-once\" disguised by a fragile post-processing cleanup script.","news",[10,13,18,23,28,33,38,43,48,53],{"headline":6,"body":11,"imageUrl":12,"sourceImageUrl":12},"Roughly 80% of the \"exactly-once\" guarantees I see implemented in production financial pipelines are actually \"at-least-once\" disguised by a fragile post-processing cleanup script. We obsess over the Spark checkpointing mechanism while ignoring the fact that our sink is a leaky bucket.","https:\u002F\u002Fmedia2.dev.to\u002Fdynamic\u002Fimage\u002Fwidth=1200,height=627,fit=cover,gravity=auto,format=auto\u002Fhttps%3A%2F%2Fimages.unsplash.com%2Fphoto-1680992046626-418f7e910589%3Fcrop%3Dentropy%26cs%3Dtinysrgb%26fit%3Dmax%26fm%3Djpg%26ixid%3DM3w5NzU0MjJ8MHwxfHNlYXJjaHwyfHxzZXJ2ZXIlMjByb29tJTIwY2FibGVzfGVufDB8MHx8fDE3OTA3MTk2ODB8MA%26ixlib%3Drb-4.1.0%26q%3D80%26w%3D1080",{"headline":14,"body":15,"imageUrl":16,"images":17},"Why I chose this topic: I’ve spent the","Why I chose this topic: I’ve spent the last six years cleaning up \"exactly-once\" messes that caused multi-million dollar reconciliation failures in healthcare billing systems. If you don't understand the physical limitations of your sink, your streaming job is just a very expensive way to generate duplicate records.","\u002Fapi\u002Fmedia\u002Fposts\u002Fspark-structured-streaming-cant-save-you-from-bad-architectu-492b3025\u002F1.webp",{"local":16},{"headline":19,"body":20,"imageUrl":21,"images":22},"Most of my peers argue that Spark’s checkpointLocation","Most of my peers argue that Spark’s checkpointLocation is the gold standard for fault tolerance. They treat it like a holy relic. The reality is that Spark’s exactly-once guarantee is strictly confined to the boundary between the source offset management and the state store. Once your data hits an external database or a file system that doesn't participate in a distributed transaction with Spark, your \"exactly-once\" guarantee is effectively vaporware. Why the common approach falls short","\u002Fapi\u002Fmedia\u002Fposts\u002Fspark-structured-streaming-cant-save-you-from-bad-architectu-492b3025\u002F2.webp",{"local":21},{"headline":24,"body":25,"imageUrl":26,"images":27},"The industry standard is to use foreachBatch to","The industry standard is to use foreachBatch to sink data into a relational database or a cloud object store. Developers assume that because Spark retries the batch upon failure, they are protected against duplicates. This is a dangerous fallacy.","\u002Fapi\u002Fmedia\u002Fposts\u002Fspark-structured-streaming-cant-save-you-from-bad-architectu-492b3025\u002F3.webp",{"local":26},{"headline":29,"body":30,"imageUrl":31,"images":32},"If your Spark executor writes 10,000 rows to","If your Spark executor writes 10,000 rows to a Postgres table and crashes at row 9,999, the driver will restart the task. If you aren’t using idempotent writes—where the sink itself knows how to deduplicate based on a business key—you now have a partial write. If your sink logic involves an INSERT statement without an ON CONFLICT DO UPDATE or a similar upsert mechanism, you have just injected garbage into your downstream analytical layer. Take a look at a common, broken pattern I see in Spark 3.3.x:","\u002Fapi\u002Fmedia\u002Fposts\u002Fspark-structured-streaming-cant-save-you-from-bad-architectu-492b3025\u002F4.webp",{"local":31},{"headline":34,"body":35,"imageUrl":36,"images":37},"This code is a lie. If the save()","This code is a lie. If the save() operation fails midway, Spark reruns the batch. If the database processed 5,000 rows before the crash, those 5,000 rows remain. The retry will attempt to insert all 10,000 rows again. You have just doubled your data. Unless your schema has a strict unique constraint and you’ve configured your JDBC driver to handle the resulting exception, you’ve corrupted your downstream data. The illusion of checkpointing","\u002Fapi\u002Fmedia\u002Fposts\u002Fspark-structured-streaming-cant-save-you-from-bad-architectu-492b3025\u002F5.webp",{"local":36},{"headline":39,"body":40,"imageUrl":41,"images":42},"Spark’s checkpointing is brilliant for maintaining the state","Spark’s checkpointing is brilliant for maintaining the state of your streaming query. It stores the metadata of processed offsets in HDFS or S3. But people confuse \"I know which offsets I’ve processed\" with \"I have committed the data to the sink.\"","\u002Fapi\u002Fmedia\u002Fposts\u002Fspark-structured-streaming-cant-save-you-from-bad-architectu-492b3025\u002F6.webp",{"local":41},{"headline":44,"body":45,"imageUrl":46,"images":47},"When a worker node dies, Spark rolls back","When a worker node dies, Spark rolls back the state to the last successful offset recorded in the checkpoint directory. It then re-executes the batch. This is where the physical reality of your sink takes over. If you are writing to S3 using the default file sink, Spark uses a commit protocol that involves renaming temporary files. If that rename operation isn't atomic—and on many S3-compatible object stores, it isn't—you are left with \"ghost\" files that weren't cleaned up by the failed task but aren't technically part of the final manifest.","\u002Fapi\u002Fmedia\u002Fposts\u002Fspark-structured-streaming-cant-save-you-from-bad-architectu-492b3025\u002F7.webp",{"local":46},{"headline":49,"body":50,"imageUrl":51,"images":52},"In my experience, relying on checkpointLocation without a","In my experience, relying on checkpointLocation without a high-performance, atomic storage layer is like trying to build a skyscraper on a swamp. You aren't doing exactly-once; you are doing \"at-least-once plus a prayer that the infrastructure handles the retry gracefully.\" The sink is the bottleneck","\u002Fapi\u002Fmedia\u002Fposts\u002Fspark-structured-streaming-cant-save-you-from-bad-architectu-492b3025\u002F8.webp",{"local":51},{"headline":54,"body":55,"imageUrl":56,"images":57},"True exactly-once in a streaming system requires a","True exactly-once in a streaming system requires a two-phase commit protocol or an idempotent sink. Spark Structured Streaming implements this only when both the source and the sink support it natively. For example, if you are reading from Kafka and writing to Kafka, Spark can participate in the transaction.","\u002Fapi\u002Fmedia\u002Fposts\u002Fspark-structured-streaming-cant-save-you-from-bad-architectu-492b3025\u002F9.webp",{"local":56},[59],"dev",[61],"Technology",{"name":63,"url":64},"Dev.to","https:\u002F\u002Fdev.to\u002Faniketsoni\u002Fspark-structured-streaming-cant-save-you-from-bad-architecture-173k","en",{"handle":67,"displayName":68},"spots","Spots","queued",{"views":71,"likes":72,"saves":72,"shares":72,"completions":72,"opens":72,"skips":72,"depthSum":72},1,0,"2026-09-30T04:00:19.190Z","local"]