[LLVMCPU][RISCV] Add RVV bf16 mmt4d ukernel using the Zvfbfwma widening MAC (bf16*bf16->f32) (#24695) ## Summary On RISC-V, a data-tiled bf16 matmul was computed by promoting LHS/RHS to f32, ignoring the `Zvfbfwma` bf16 widening multiply-accumulate extension (`vfwmaccbf16`) that LLVM already lowers to and that shipping silicon implements. This PR adds a native `bf16 x bf16 -> f32` `mmt4d` microkernel for RVV+`Zvfbfwma`. Naturally, this change is related to (and was implementing looking at) the existing f32, `int8` (#23734), and f16/`Zvfh` (#22231) ukernels. ## How the change was tested 1. Compiler e2e A `128x256 bf16 * 256x128 bf16 -> 128x128 f32` matmul compiled for `riscv64 +v,+zvfbfwma` with `--iree-opt-data-tiling --iree-llvmcpu-enable-ukernels=all`: Before: dispatch contains `arith.extf` f32-promotion, no ukernel. After: `Mmt4dTilingExpert` dispatch `matmul_128x128x256_bf16xbf16xf32` with `iree_uk_mmt4d` + `iree_uk_pack`; **zero `arith.extf`**; all four `iree_uk_mmt4d_tile_bf16bf16f32_*_riscv_64_zvfbfwma` symbols present; generated RISC-V asm inner loop is `vfwmaccbf16.vf`. 2. Numerics under `qemu-riscv64` A standalone harness calling the real tile functions matches an f32 reference exactly. A control run without the extension crashes, confirming `vfwmaccbf16` is genuinely executed. I used qemu 8.2 which exposes the extension as the experimental `-cpu max,x-zvfbfwma=true` property. 3. Tile selection under `qemu-riscv64`. Verified `iree_uk_mmt4d_select_tile_func_arch` (the `entry_point.c`/`tiles.inl` wiring) returns the correct bf16 tile for each M0 value when `cpu_data` has the `Zvfbfwma` bit, and does not select it when the bit is clear — i.e. the feature is only enabled for `Zvfbfwma`. 4. In-tree `mmt4d_test`. The added case exercises the same path in CI's RISC-V cross-build (falls back to the generic tile if the CI toolchain can't assemble `Zvfbfwma`, i think it's the same limitation as `Zvfh` in #22303). ### A note on CI coverage If I understand correctly, the RISC-V CI job runs under an older `qemu-riscv64` on a stock `rv64gcv` profile that doesn't expose `Zvfbfwma`, and `mmt4d_test` *skips* (does not fail) any case whose feature the runtime CPU lacks. Zvfbfwma is only assembled/executed on a toolchain + qemu new enough to support it (qemu ≥ 8.2 exposes it as the experimental `x-zvfbfwma` property). My point is: a green CI does not mean the bf16 kernel ran in CI — the functional validation of record is the local `qemu-riscv64` 8.2 numerics + tile-selection runs described above. Disclosure: this change was assisted by Claude Code, but the change was reviewed and thoroughly tested before being submitted. --------- Signed-off-by: Zmicier Prybysh <zprybysh@baylibre.com>
IREE (Intermediate Representation Execution Environment, pronounced as “eerie”) is an MLIR-based end-to-end compiler and runtime that lowers Machine Learning (ML) models to a unified IR that scales up to meet the needs of the datacenter and down to satisfy the constraints and special considerations of mobile and edge deployments.
See our website for project details, user guides, and instructions on building from source.
Releases notes are published on GitHub releases.
| Package | Release status |
|---|---|
| GitHub release (stable) | |
| GitHub release (nightly) | |
iree-base-compiler | |
iree-base-runtime |
For more details on the release process, see https://iree.dev/developers/general/release-management/.
| Operating system | Build status |
|---|---|
| Linux | |
| macOS | |
| macOS |
For the full list of workflows see https://iree.dev/developers/general/github-actions/.
See our website for more information.
Community meeting recordings: IREE YouTube channel
| Date | Title | Recording | Slides |
|---|---|---|---|
| 2025-06-10 | Data-Tiling in IREE: Achieving High Performance Through Compiler Design (AsiaLLVM) | recording | slides |
| 2025-05-17 | Introduction to GPU architecture and IREE's GPU CodeGen Pipeline | recording | slides |
| 2025-02-12 | The Long Tail of AI: SPIR-V in IREE and MLIR (Vulkanised) | recording | slides |
| 2024-10-01 | Unveiling the Inner Workings of IREE: An MLIR-Based Compiler for Diverse Hardware | recording | |
| 2021-06-09 | IREE Runtime Design Tech Talk | recording | slides |
| 2020-08-20 | IREE CodeGen (MLIR Open Design Meeting) | recording | slides |
| 2020-03-18 | Interactive HAL IR Walkthrough | recording | |
| 2020-01-31 | End-to-end MLIR Workflow in IREE (MLIR Open Design Meeting) | recording | slides |
IREE is licensed under the terms of the Apache 2.0 License with LLVM Exceptions. See LICENSE for more information.