Explore indexBack to Terms

Speculative Decoding

Speculative Decoding is an inference acceleration technique for autoregressive language models. A built-in draft head or lightweight auxiliary model rapidly drafts candidate tokens, which are verified in parallel in a single forward pass by the target model, boosting generation speed without compromising mathematical output quality.

No public content is connected to this entity yet.