Explore indexBack to Terms
Speculative Decoding
Speculative Decoding is an inference acceleration technique for autoregressive language models. A built-in draft head or lightweight auxiliary model rapidly drafts candidate tokens, which are verified in parallel in a single forward pass by the target model, boosting generation speed without compromising mathematical output quality.
No public content is connected to this entity yet.