CODA Rewrites Transformer Kernels: FlashAttention Obsolete?
CODA compresses multi-kernel transformer computations into a single matrix multiply with a fused epilogue. Early results show competitive performance with hand-tuned libraries, raising the question of whether the era of custom CUDA kernels is ending.











