Twitter/X

MoonshotAI open-sourced FlashKDA, a CUTLASS-based implementation of Kimi Delta…

Brief

MoonshotAI released FlashKDA, a high-performance, CUTLASS-based implementation of Kimi Delta Attention kernels on GitHub (MoonshotAI/FlashKDA). The repo claims 1.72×–2.22× prefill speedups versus the flash-linear-attention baseline on H20 hardware and advertises seamless use as a drop-in backend for flash-linear-attention.

Why it matters

MoonshotAI open-sourced FlashKDA, a CUTLASS-based implementation of Kimi Delta Attention kernels available at github.com/MoonshotAI/FlashKDA

Key details

  • FlashKDA provides a 1.72×–2.22× prefill speedup over the flash-linear-attention baseline measured on H20
  • FlashKDA is compatible as a drop-in backend for the flash-linear-attention library
Source evidence

We've open-sourced FlashKDA, our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels.

It delivers 1.72×–2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention.

Explore on GitHub:
github.com/MoonshotAI/FlashK…

Link

GitHub - MoonshotAI/FlashKDA: FlashKDA: high-performance Kimi Delta Attention kernels

FlashKDA: high-performance Kimi Delta Attention kernels - MoonshotAI/FlashKDA
github.com