Back to blog
Editorial draft Data Engineering Spark Kubernetes

Spark on Kubernetes: Lessons from Migrating a Petabyte

Cost-optimized, multi-tenant Spark platforms on K8s — what we learned after moving 400 pipelines off YARN.

April 10, 2026 7 min read

YARN served us well for a decade. But when your team needs cost showback per pipeline and isolation per tenant, Kubernetes wins.

Executors are pods

This changes everything about scheduling. Use spot pools for shufflers, on-demand for the driver, and let the cluster autoscaler do its job.

Shuffle service or S3?

Push-based shuffle to S3 or MinIO is the killer feature. Stateless executors, elastic autoscaling, no more shuffle-fetch failures.

The best Spark cluster is the one that shrinks to zero when you're not looking.