DFlash Speculative Decoding Drafts Whole Token Blocks in Parallel for Up to 15x Higher Throughput on NVIDIA Blackwell

Por Redacción LaiaDesk · 24 jun. 2026

UC San Diego's DFlash replaces autoregressive drafting with a lightweight block diffusion model for speculative decoding. It drafts whole token blocks in a single forward pass and conditions on target hidden features through KV injection. The paper reports up to 6.08x lossless speedup on Qwen3-8B, while NVIDIA reports…

Seguir leyendo en MarkTechPost →

Pronto, la IA de LaiaDesk publicará aquí el análisis completo de qué significa esta noticia para tu sector.

Fuente original: MarkTechPost

Reaccionar

0 comentarios

Conversación

Sé el primero en comentar.

Habla con LaiaDesk Más noticias

Conversación

La IA de tu sector, en tu bandeja