InfoQ 中文 · 10/9/2026, 17:24:41
Modal Bypasses Kubernetes Limits with Decentralized Sandbox Architecture for Million-Scale Concurrency
Modal engineers bypassed Kubernetes' strong consistency constraints by building a decentralized sandbox infrastructure that supports one million concurrent sandboxes and tens of thousands of creations per second. The core strategy eliminates O(N) state coordination bottlenecks by treating scheduling as load balancing, making each worker node an independent source of truth and deploying parallel schedulers to avoid etcd performance limits under high pod churn.
SOURCE COVERAGEOriginal coverage
Modal engineers built custom sandbox infrastructure to support millions of concurrent sandboxes and tens of thousands of creations per second, abandoning strongly consistent orchestration systems like Kubernetes. The core strategy was to avoid O(N) state coordination bottlenecks by designing all high-complexity components for horizontal scalability and radically simplifying the sandbox creation path.
Key takeaways: Traditional container platforms struggle to scale to millions of sandboxes due to etcd’s strong consistency requirements and O(N) scheduling complexity; etcd lacks native key-space sharding capabilities, becoming a bottleneck under high pod creation or churn rates; the primary design priorities for sandbox systems are horizontal scalability and minimalism in the creation path.
This article is suitable for cloud-native platform engineers, large-scale scheduling system architects, and serverless infrastructure developers.

In a recent post, Modal engineers Colin Weld and Connor Adams detailed how they rebuilt their sandbox infrastructure from scratch to support millions of concurrent sandboxes and tens of thousands of sandbox creations per second.
According to Weld and Adams, traditional container orchestration systems like Kubernetes struggle at this scale because they rely heavily on centralized coordination and strongly consistent state.
Running 1 million sandboxes pushes any container platform to its limits, both due to the sheer number of containers and the need for tens of thousands of compute nodes to run them. Many operations have a complexity of either O(containers), O(nodes), or both, causing traditional container platforms to hit their scaling ceilings.
They explained that in Kubernetes, the load on both the scheduling algorithm and the central persistent store (etcd) grows as the number of nodes and pods increases. Additionally, pods and nodes write data to etcd multiple times, "which can cause severe issues when pod creation rates or pod churn rates are high, especially since etcd does not natively support sharding within the same key space." They noted that overcoming this limitation is possible but requires "significant work," including rewriting or replacing etcd and parallelizing the scheduling algorithm.
To optimize for scalability, we decided that any component generating O(sandbox) or O(node) level load must support horizontal scaling by default; the sandbox creation flow should be as simple as possible; everything else is secondary.
The fundamental change Modal engineers made to their platform was stopping global coordination, making scheduling more akin to load balancing. Each worker node no longer relies on a central datastore as the "single source of truth" but becomes its own "single source of truth." Similarly, instead of using a single serial scheduler, they deployed a set of parallel-running scheduling servers, enabling horizontal scaling of the scheduling layer.
Once a scheduling server decides which worker node will host a new sandbox, it contacts that worker directly via RPC to request sandbox creation. If the worker has available resources, it accepts the scheduling request; otherwise, it rejects it.
According to them, the resulting architecture has only one bottleneck: all workers publish status updates to a Redis stream. However, "load tests showed that this approach remains viable even when the number of workers far exceeds 100,000." In benchmarks, they created 1 million sandboxes in under a minute, with a median time from launch to code execution of less than 0.5 seconds.
On LinkedIn, Hopsworks CEO Jim Dowling commented on the announcement, noting that "each order-of-magnitude increase in scale introduces new technical challenges," suggesting the team had to iterate on the design multiple times to reliably achieve 50,000 sandbox creations per second. More substantively, AWS Chief AI Engineer Alex Jones pointed out that the key to Modal’s success was that they did not try to scale Kubernetes, but rather "bypassed the entire system" after understanding its limitations. Jones argued that this is "the first credible signal that Kubernetes is failing to adapt quickly enough to the actual needs of generative AI infrastructure," stating:
We are moving toward decoupling coordination from execution. The execution layer needs what Modal has implemented: establishing isolation boundaries in milliseconds. The coordination layer (where multi-agent workflows require shared memory and overlapping security boundaries) still needs the capabilities that systems like Kubernetes excel at.
Modal is a serverless compute platform designed specifically for AI workloads, offering programmatic access to CPUs, GPUs, containers, inference, training, batch jobs, and isolated sandboxes. It is not the only platform attempting to "rebuild the cloud" around goals of highly scalable infrastructure and cold starts completed within 10 milliseconds. Other projects pursuing similar goals include Unikraft, Google Substrate, and Overdrive.
Original article: https://www.infoq.com/news/2026/09/modal-scaling-sandboxes/