Introduction to Celeborn

What is Celeborn?

Apache Celeborn is a Remote Shuffle Service (RSS) designed to improve the efficiency, stability, and flexibility of shuffle operations in distributed compute engines. It supports Apache Spark, Apache Flink, Apache Tez, and MapReduce.


Why Celeborn?

Traditional shuffle frameworks have significant limitations that become critical at scale:

Problem

Traditional Shuffle

Celeborn Solution

Network Efficiency

M × N connections between Mappers and Reducers

Consolidated M+N connections via Celeborn workers

Disk I/O

Random I/O on compute nodes

Sequential I/O on dedicated shuffle nodes

Dynamic Allocation

Limited by shuffle data locality

Full executor elasticity

Node Failure

Shuffle data lost, job fails or retries

Data replicated — job continues without retry

Storage

Large local disks required on compute nodes

Dedicated shuffle storage (local, HDFS, S3)


Key Benefits

  • Performance: 2–5× improvement in shuffle-heavy workloads

  • Stability: Data replication prevents job failures from node loss

  • Elasticity: Enables true dynamic resource allocation

  • Disaggregation: Separates compute from shuffle storage

  • Multi-Engine: Supports Spark, Flink, Tez, and MapReduce



  Last updated