We've open-sourced FlashKDA, our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels.
It delivers 1.72×–2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention.
Explore on GitHub:
github.com/MoonshotAI/FlashK…
Link
GitHub - MoonshotAI/FlashKDA: FlashKDA: high-performance Kimi Delta Attention kernels
FlashKDA: high-performance Kimi Delta Attention kernels - MoonshotAI/FlashKDA
github.com