Explore indexBack to Terms
PD Disaggregation
An architecture that deploys prefill and decode phases of LLM inference on separate compute resources to optimize throughput and latency.
No public content is connected to this entity yet.
An architecture that deploys prefill and decode phases of LLM inference on separate compute resources to optimize throughput and latency.
No public content is connected to this entity yet.