Spots

Moving Your Local Airflow to GCP for under $150 a month

My team runs ML training and prediction pipelines on Vertex AI. For a long time, the thing telling those pipelines when to run was Apache Airflow, installed by hand on a few on-prem Linux boxes. From there it orchestrated our hybrid cloud setup from one place.

Next to Airflow sat our own configuration management

Next to Airflow sat our own configuration management system. It lets people change a pipeline's settings on demand, and every change is tracked and logged over time, so if a model behaves differently on Tuesday, we can see what changed on Monday.

Both of these lived on Linux boxes we

Both of these lived on Linux boxes we maintained ourselves, and those boxes had to go: our organization is phasing out its on-prem setup and moving fully cloud native. We are in the middle of that move right now. The on-prem pipelines are moving out to different clouds, and our own pipelines and compute are moving into GCP. With the work landing in GCP, putting the orchestrator there too was the natural answer. The goal was to make that move without ending up any less safe or less reliable than what we had.

In this post I want to share how

In this post I want to share how we did it, chapter by chapter. The code is in the self-managed-airflow-on-gcp repo, and everything below is a trimmed-down, working version of it. Chapter 1: Trying Cloud Composer

If you ask "how do I run Airflow

If you ask "how do I run Airflow on GCP?", the answer is Cloud Composer. It is managed Airflow: Google runs the scheduler, the database and the workers, and you drop DAGs into a bucket. We built it first. We stood up a Composer 3 environment with Terraform, pointed our DAGs at it, and it worked.

The cost was the problem. With the pricing

The cost was the problem. With the pricing calculator, a small Composer environment came to about $350 a month. Composer 3 bills for compute units every hour the environment is alive, whether a DAG runs or not. Our pipelines do not need a cluster awake around the clock to schedule a handful of jobs, because the heavy work happens in Vertex AI anyway. My search was not done there. Chapter 2: Self-managed Airflow on a VM

Our Airflow does not do the heavy work

Our Airflow does not do the heavy work. It decides when work happens and hands it to Vertex AI or BigQuery. A scheduler like that fits on one small VM. So we built the same thing a second time: one Compute Engine VM running a self-managed Airflow 3.x, with Postgres for the metadata database Terraform to provision the platform, per environment (dev, test, prod) a Fabric script to install Airflow, deploy code, and handle day-to-day maintenance

The estimate for this came in under $150

The estimate for this came in under $150 a month. About $20 of that is the HTTPS Load Balancer in front of our config app (GCP charges a flat $0.025/hour for the first five forwarding rules, roughly $18/month, plus data). The rest is the VM, its disk, and Cloud NAT.

That number is not the whole cost. With

That number is not the whole cost. With Composer, Google upgrades Airflow and patches the machine. On a VM, we do. We had both versions running side by side, and the team picked the self-managed one: we are a technical team that already ran Airflow ourselves, so the extra work was work we knew. Cheaper only counts if it is also safe and reliable, so the rest of this post is about how we got there. Chapter 3: How the pieces fit There are two components we deploy:

config_hq, the new home of our configuration system

config_hq, the new home of our configuration system. It is a small Flask app with a plain HTML/JS frontend, running on Cloud Run. A user types config, presses Save, and it is written to a GCS bucket. airflow-vm, the orchestrator. Its DAGs read that config and kick off the ML jobs. The part I like most in this design is that Airflow never calls config_hq. The two only share a bucket.

News

Moving Your Local Airflow to GCP for under $150 a month

My team runs ML training and prediction pipelines on Vertex AI.

@spots #dev
Source: Dev.to
See more like this