MLX v0.32.3: Fixes for Running Models on Apple Silicon
MLX, the tool for running ML models, has released v0.32.3. The update fixes scan and sort operations for the zero-size axis, adds a correction parameter to std and var, eliminates a deadlock when calling mx.clear_streams(), and fixes an integer exponentiation bug with a negative exponent that zeroed out the entire SIMD vector. Details are in the changelog on GitHub.
What it means
For those running neural networks on their own machine via MLX, this release is primarily about stability. The deadlock fix in mx.clear_streams() removes a hang during stream cleanup, and the adaptive concurrent loading limit lets I/O scale with the machine — on more powerful systems, data will load faster.
Worth noting separately is the fix in sorted gather_qmm: an NAX row overflow above 32K has been eliminated, which could lead to incorrect results on large tensors. Validation of GGUF tensor dimensions reduces the risk of a crash when loading corrupted or incompatible model files.
The developer does not specify which models are affected by the changes, but the fixes touch core operations — from exponentiation to reductions with NaN. The update is worth installing for those who have encountered hangs or unstable operation when running models via MLX. You can follow news about tools for running neural networks locally in the CPU3D news feed.How the method works. The diagram was drawn based on this news note.