Title
Create new category
Edit page index title
Edit category
Edit link
Introduction to Celeborn
What is Celeborn?
Apache Celeborn is a Remote Shuffle Service (RSS) designed to improve the efficiency, stability, and flexibility of shuffle operations in distributed compute engines. It supports Apache Spark, Apache Flink, Apache Tez, and MapReduce.
Why Celeborn?
Traditional shuffle frameworks have significant limitations that become critical at scale:
Problem | Traditional Shuffle | Celeborn Solution |
|---|---|---|
Network Efficiency | M × N connections between Mappers and Reducers | Consolidated M+N connections via Celeborn workers |
Disk I/O | Random I/O on compute nodes | Sequential I/O on dedicated shuffle nodes |
Dynamic Allocation | Limited by shuffle data locality | Full executor elasticity |
Node Failure | Shuffle data lost, job fails or retries | Data replicated — job continues without retry |
Storage | Large local disks required on compute nodes | Dedicated shuffle storage (local, HDFS, S3) |
Key Benefits
Performance: 2–5× improvement in shuffle-heavy workloads
Stability: Data replication prevents job failures from node loss
Elasticity: Enables true dynamic resource allocation
Disaggregation: Separates compute from shuffle storage
Multi-Engine: Supports Spark, Flink, Tez, and MapReduce