#40033 [NVFP4][Hopper/AMD Instinct] Add Triton kernels for NVFP4 dequantization and QDQ emulation
原始 PR · 作者 fxmarty-amd · 合并时间 2026-05-01 05:35
添加Triton内核加速NVFP4反量化和QDQ模拟
值得精读: - 学习 Triton 内核优化技巧:二进制树 E2M1 查找、2D tile 批处理、interleave 合并写。 - 理解设备间功能兼容性处理:通过 `current_platform.is_cuda_alike()` 动态切换实现。 - 关注社区反馈中对类型安全的关注,建议合并后进一步放宽 `global_scale` 类型以支持 float。
参与讨论