Modal Shatters Scale Limits: 1M Sandboxes in Seconds
Alps Wang
Sep 24, 2026 · 1 views
Reimagining Orchestration for Extreme Scale
The article from InfoQ presents a compelling case for Modal's innovative approach to scaling sandbox infrastructure, moving beyond traditional Kubernetes limitations. The core insight lies in decentralizing coordination and treating scheduling more like load balancing, where workers become their own sources of truth and scheduling is parallelized. This design choice directly addresses the O(containers) and O(nodes) complexity that plagues centralized systems like Kubernetes, particularly its reliance on etcd. The reported achievement of launching 1 million sandboxes in under a minute with sub-second startup times is truly remarkable and directly tackles a critical bottleneck for AI/ML workloads that demand rapid, isolated execution environments. The technical details, such as replacing global coordination with RPC calls to workers and using a Redis stream as the sole bottleneck, provide a clear picture of their architectural departure. This is a significant advancement for any platform requiring massive concurrency and low latency for ephemeral compute, especially in the rapidly evolving AI landscape.
Key Points
- Modal rebuilt its sandbox infrastructure to scale to 1 million concurrent sandboxes.
- Traditional systems like Kubernetes struggle at this scale due to centralized coordination and state.
- Modal decentralized coordination, making scheduling resemble load balancing.
- Workers became their own sources of truth, and scheduling was parallelized.
- The architecture's primary bottleneck is a single Redis stream, viable up to 100,000 workers.
- Achieved 1 million sandboxes in under a minute with sub-second startup times.

📖 Source: Beyond Kubernetes at Modal: How to Scale 1 Million Concurrent Sandboxes in Seconds
Related Articles
Comments (0)
No comments yet. Be the first to comment!
