YARN served us well for a decade. But when your team needs cost showback per pipeline and isolation per tenant, Kubernetes wins.
Executors are pods
This changes everything about scheduling. Use spot pools for shufflers, on-demand for the driver, and let the cluster autoscaler do its job.
Shuffle service or S3?
Push-based shuffle to S3 or MinIO is the killer feature. Stateless executors, elastic autoscaling, no more shuffle-fetch failures.
The best Spark cluster is the one that shrinks to zero when you're not looking.