[Metal] Fix copy_buffer_1byte dispatch grid (#24640)

copy_buffer_1byte silently dropped the trailing ~3/4 of every copy that
fell back to the compute-kernel path. The kernel copies one byte per
thread, but its dispatch grid was sized as ceil(length / (workgroup_size
* 4)) -- a quarter of the required ceil(length / workgroup_size)
workgroups -- so only the first min(ceil(length/128)*32, length) bytes
were written and the tail was left untouched. The bounds guard `if (id
>= spec.length) return` prevents over-run but cannot launch the missing
threads.

On macOS the compute path is taken whenever the source offset, target
offset, or length is not a multiple of 4 (the Metal blit fast-path
requires all three 4-byte aligned). It is also the path behind the
staging->device upload
(iree_hal_metal_command_buffer_prepare_update_buffer), so non-4-aligned
source uploads were truncated too -- a device copy then read a zeroed
source tail.

The *4 multiplier is correct for the fill kernels (fill_buffer_16byte /
fill_buffer_4byte / fill_buffer_1byte process 16/4/4 bytes per thread
respectively), which is where it was copied from; copy_buffer_1byte is a
flat 1-byte-per-thread loop with no such expansion.

Fix (dispatch only -- the copy_buffer_1byte.metal kernel is unchanged):
size the grid as ceil(length / workgroup_size) so the launched thread
count (ceil(length/32) * 32) covers every byte. The existing `if (id >=
spec.length) return` guard masks the surplus threads in the final
workgroup.

Verified:
CTS/CommandBufferCopyBufferTest.CopySizeAndAlignmentClasses/metal now
passes all alignment/size sub-cases that previously truncated
(offset/length not a multiple of 4; lengths 31/32/33/.../65536 across
all alignment classes). The aligned16_mib sub-case still fails,
independently, from a staging-buffer capacity issue (128 KiB staging
buffer vs. a 1 MiB single-command-buffer upload), not from this dispatch
sizing. Signed-off-by: Alex Vasile
<48962821+Alex-Vasile@users.noreply.github.com>

Signed-off-by: Alex Vasile <48962821+Alex-Vasile@users.noreply.github.com>
1 file changed
tree: 3ba41a8b34ca76dda2a4d3d9ff4874e3345a7d31
  1. .github/
  2. build_tools/
  3. compiler/
  4. docs/
  5. experimental/
  6. integrations/
  7. lib/
  8. llvm-external-projects/
  9. runtime/
  10. samples/
  11. tests/
  12. third_party/
  13. tools/
  14. .bazel_to_cmake.cfg.py
  15. .bazelignore
  16. .bazelrc
  17. .bazelversion
  18. .clang-format
  19. .git-blame-ignore-revs
  20. .gitattributes
  21. .gitignore
  22. .gitmodules
  23. .pre-commit-config.yaml
  24. .yamllint.yml
  25. AUTHORS
  26. BUILD.bazel
  27. CITATION.cff
  28. CMakeLists.txt
  29. configure_bazel.py
  30. CONTRIBUTING.md
  31. LICENSE
  32. MAINTAINERS.md
  33. MODULE.bazel
  34. README.md
  35. RELEASING.md
README.md

IREE: Intermediate Representation Execution Environment

IREE (Intermediate Representation Execution Environment, pronounced as “eerie”) is an MLIR-based end-to-end compiler and runtime that lowers Machine Learning (ML) models to a unified IR that scales up to meet the needs of the datacenter and down to satisfy the constraints and special considerations of mobile and edge deployments.

See our website for project details, user guides, and instructions on building from source.

IREE Discord Status pre-commit OpenSSF Best Practices

Project news

Project status

Release status

Releases notes are published on GitHub releases.

PackageRelease status
GitHub release (stable)GitHub Release
GitHub release (nightly)GitHub Release
iree-base-compilerPyPI version
iree-base-runtimePyPI version

For more details on the release process, see https://iree.dev/developers/general/release-management/.

Build status

CI PkgCI

Nightly build status

Operating systemBuild status
LinuxCI - Linux arm64 clang
macOSCI - macOS x64 clang
macOSCI - macOS arm64 clang

For the full list of workflows see https://iree.dev/developers/general/github-actions/.

Communication channels

Related project channels

  • MLIR topic within LLVM Discourse: IREE is enabled by and heavily relies on MLIR. IREE sometimes is referred to in certain MLIR discussions. Useful if you are also interested in MLIR evolution.

Architecture overview

IREE Architecture IREE Architecture

See our website for more information.

Presentations and talks

Community meeting recordings: IREE YouTube channel

DateTitleRecordingSlides
2025-06-10Data-Tiling in IREE: Achieving High Performance Through Compiler Design (AsiaLLVM)recordingslides
2025-05-17Introduction to GPU architecture and IREE's GPU CodeGen Pipelinerecordingslides
2025-02-12The Long Tail of AI: SPIR-V in IREE and MLIR (Vulkanised)recordingslides
2024-10-01Unveiling the Inner Workings of IREE: An MLIR-Based Compiler for Diverse Hardwarerecording
2021-06-09IREE Runtime Design Tech Talkrecordingslides
2020-08-20IREE CodeGen (MLIR Open Design Meeting)recordingslides
2020-03-18Interactive HAL IR Walkthroughrecording
2020-01-31End-to-end MLIR Workflow in IREE (MLIR Open Design Meeting)recordingslides

License

IREE is licensed under the terms of the Apache 2.0 License with LLVM Exceptions. See LICENSE for more information.