[HAL] Fix inline execution models (initializer inlining and dispatch ABI) (#24827) The inline execution models were broken end-to-end in both the compiler and runtime. On the compiler side, the `HAL_Loader` and `HAL_Inline` dialects did not support inlining, preventing `CombineInitializers` pass from combining initializers containing their ops. This broke `inline-dynamic` compilation generally, and stateful programs using mutable globals under both `inline-dynamic` and `inline-static`. On the runtime side, the `hal_loader` dispatch shim incorrectly accounted for padding in the fixed portion of the VM cconv arguments by using the C struct size, causing `inline-dynamic` dispatches to fail with an argument/result signature mismatch. The runtime also leaked the inline HAL storage buffer wrapper on release. VM buffer destruction also assumed that every buffer was created by `iree_vm_buffer_create/clone` as a single aligned allocation containing both the `iree_vm_buffer_t` handle and its data. Inline HAL instead embeds the handle inside a wrapper struct that maps an `iree_hal_buffer_t`, initializing it in place with `iree_vm_buffer_initialize`. Such buffers must invoke the wrapper’s custom allocator to release the underlying HAL buffer and free the wrapper, rather than calling `iree_allocator_free_aligned` on the embedded handle. Since nothing else in the handle distinguishes the two cases at destroy time, create/clone now tag their buffers with `IREE_VM_BUFFER_ACCESS_COALLOCATED` (set internally), and destroy takes the free-aligned path only for tagged buffers; all others pass the handle to their allocator. This PR adds the required inliner interfaces, fixes the dispatch argument handling, distinguishes co-allocated and in-place-initialized VM buffers during destruction, and adds compiler/runtime coverage for the affected execution models and buffer lifetimes. Assisted-by: Claude Code --------- Signed-off-by: Pooja Hemashekar <hemashekar@roofline.ai>
IREE (Intermediate Representation Execution Environment, pronounced as “eerie”) is an MLIR-based end-to-end compiler and runtime that lowers Machine Learning (ML) models to a unified IR that scales up to meet the needs of the datacenter and down to satisfy the constraints and special considerations of mobile and edge deployments.
See our website for project details, user guides, and instructions on building from source.
Releases notes are published on GitHub releases.
| Package | Release status |
|---|---|
| GitHub release (stable) | |
| GitHub release (nightly) | |
iree-base-compiler | |
iree-base-runtime |
For more details on the release process, see https://iree.dev/developers/general/release-management/.
| Operating system | Build status |
|---|---|
| Linux | |
| macOS |
For the full list of workflows see https://iree.dev/developers/general/github-actions/.
See our website for more information.
Community meeting recordings: IREE YouTube channel
| Date | Title | Recording | Slides |
|---|---|---|---|
| 2025-06-10 | Data-Tiling in IREE: Achieving High Performance Through Compiler Design (AsiaLLVM) | recording | slides |
| 2025-05-17 | Introduction to GPU architecture and IREE's GPU CodeGen Pipeline | recording | slides |
| 2025-02-12 | The Long Tail of AI: SPIR-V in IREE and MLIR (Vulkanised) | recording | slides |
| 2024-10-01 | Unveiling the Inner Workings of IREE: An MLIR-Based Compiler for Diverse Hardware | recording | |
| 2021-06-09 | IREE Runtime Design Tech Talk | recording | slides |
| 2020-08-20 | IREE CodeGen (MLIR Open Design Meeting) | recording | slides |
| 2020-03-18 | Interactive HAL IR Walkthrough | recording | |
| 2020-01-31 | End-to-end MLIR Workflow in IREE (MLIR Open Design Meeting) | recording | slides |
IREE is licensed under the terms of the Apache 2.0 License with LLVM Exceptions. See LICENSE for more information.