Skip to content

Pull requests: fla-org/flash-linear-attention

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

[Perf] Fuse ds in Ascend chunk_bwd_dqkwg to avoid recomputing do@v.T
#1048 opened Jul 20, 2026 by sunyi0505 Contributor Loading…
[Ops] Add triton-ascend backend for KDA kernels
#1047 opened Jul 20, 2026 by zheliuyu Contributor Draft
[Fix] Reject non-divisible GQA head counts
#1032 opened Jul 16, 2026 by morluto Contributor Loading…
[Ops] Rewrite NPU chunk_scaled_dot_kkt_fwd
#1023 opened Jul 14, 2026 by OsirisDuan Contributor Loading…
[CP] Enable context parallelism for GDP
#1022 opened Jul 13, 2026 by Mellonta Loading…
[Model] Add CAT (Compress and Attend Transformer)
#1017 opened Jul 12, 2026 by bicycleman15 Loading…
[Perf] Generalize fused q/k/v short convolution across layers
#977 opened Jun 23, 2026 by zhiyuan1i Collaborator Loading…
[GDN] Fix GDN precision on Blackwell
#948 opened Jun 14, 2026 by syeehyn Loading…
[Fix] Fix shared memory race in tilelang chunk_bwd dg_last accumulation help wanted Extra attention is needed
#890 opened May 11, 2026 by Erix025 Contributor Loading…
[SSE] Add SSE integration
#882 opened May 9, 2026 by Pan-Yuqi Contributor Loading…
ProTip! Type g p on any issue or pull request to go back to the pull request listing page.