Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs. It centres on Inference, and also names Benchmarks. Reported by arXiv. Bharat Hunt files it under Generative AI and AI Hardware — the section covering text, image, audio and video generation — products and their output.
Written by Bharat Hunt from the headline and the coverage below. The original reporting is the source of truth.