Port iree.runtime to nanobind. (#14214) I believe that this should be a no-op for users. There is one minor API change (the MappedMemory class no longer implements the buffer protocol, but I've seen no evidence that this was actually used since it was a less functional way to get a host ndarray). More adventurous use of nanobind is possible in the future (i.e. using `ndarray` and `dlpack` interop for sharing across frameworks), but using that will most likely necessitate API changes, which I was working to avoid. Aside from relatively mechanical differences from pybind11, the main issues were that the buffer protocol and array support was dropped in nanobind. This required some direct coding against the C API to achieve the same characteristics. I think this is actually an improvement as the pybind11 implementations of these features was neither efficient nor obvious what it was doing. A build time dependency on `nanobind` is added. When building Python wheels, this gets satisfied automatically. Otherwise, the docker images have been updated to pre-install the necessary Python package. In addition, there is now a build time dependency on NumPy headers, which should already be installed (pybind11 vendored stripped down copies of these headers in an effort to avoid this, but I opted to just do the normal thing). Nanobind's performance is [quite compelling](https://nanobind.readthedocs.io/en/latest/benchmark.html) and owes to a combination of favoring more efficient binding styles that would basically be a rewrite in pybind11 and exclusive use of the new Python 3.8+ vectorcall ABI. Since the runtime is performance critical and the cost of Python calls is already quite visibly adding overhead on traces, it makes sense to baseline on the most efficient implementation. In addition, the compile-time savings seem to be real and the build is noticeably faster (this was not a primary consideration, just a nice bonus).
IREE (Intermediate Representation Execution Environment, pronounced as “eerie”) is an MLIR-based end-to-end compiler and runtime that lowers Machine Learning (ML) models to a unified IR that scales up to meet the needs of the datacenter and down to satisfy the constraints and special considerations of mobile and edge deployments.
See our website for project details, user guides, and instructions on building from source.
IREE is still in its early phase. We have settled down on the overarching infrastructure and are actively improving various software components as well as project logistics. It is still quite far from ready for everyday use and is made available without any support at the moment. With that said, we welcome any kind of feedback on any communication channels!
See our website for more information.
IREE is licensed under the terms of the Apache 2.0 License with LLVM Exceptions. See LICENSE for more information.