Nex-N2.5-mini-MLX-4bit achieves 133.6 tok/s on Apple M5 Max
A new benchmark entry on llm-bench.io records the performance of the Nex-N2.5-mini-MLX-4bit model running on Apple’s M5 Max hardware. Measured using the oMLX 0.31.3 framework on September 12, 2026, the model achieved a throughput of 133.6 tokens per second. This result highlights the growing efficiency of quantized, small-scale language models on modern Apple Silicon, which is increasingly…
Key points
- Nex-N2.5-mini-MLX-4bit achieved 133.6 tok/s on Apple M5 Max using oMLX 0.31.3.
- Benchmark dated September 12, 2026, highlights efficient local inference on Apple Silicon.
- Qualitative reviews praise logical planning and coding but note weaknesses in narrative completeness.
The benchmark report includes qualitative assessments of the model's capabilities across various tasks. Evaluations note strong logical structuring in planning tasks, effective implementation of code such as a Breakout game, and accurate analytical reasoning. However, the review also points out limitations, including incomplete narrative arcs in creative writing and uncertainty in scaling claims due to small, heterogeneous datasets. These findings suggest that while the model is robust for structured and coding tasks, it may require further refinement for complex creative or extrapolative workloads.
This data point is significant for developers looking to deploy AI locally without relying on cloud infrastructure. The combination of high speed and reasonable quality on consumer-grade hardware like the M5 Max indicates a shift toward more accessible, on-device AI solutions. As quantization techniques improve, models like Nex-N2.5-mini-MLX-4bit offer a balance between performance and resource efficiency, making them attractive for privacy-conscious users and edge computing applications.
Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io
llm-bench.io · 12 September 2026Benchmark result
Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s
Measured with oMLX 0.31.3 on September 12, 2026.
LLM / Model
LLM Quality Assessment
The plan is logically structured, uses exactly four steps, and cleanly maps to the report requirements. It decomposes the task well and includes a reasonable backup strategy. Tool choice is mostly appropriate, but read_file and execute_code are not strongly justified given the task can be completed largely via web sources and manual synthesis, and the response goes beyond planning by including detailed benchmark content.
A strong, mostly complete Breakout implementation with mobile controls, scoring, lives, restart flow, and canvas rendering. The main weaknesses are a few collision edge cases and a minor deliverable-format issue from the surrounding response text.
The response captures Aldwyn's haunted restraint, caution, and hidden grief very well, with vivid atmosphere and authentic voice. However, it only presents an opening beat rather than the full requested arc, so the narrative progression is incomplete.
Strong and mostly accurate analysis with clear task-by-task calculations, good comparative interpretation, and actionable recommendations. The main weaknesses are that the scaling claims rest on a very small, heterogeneous dataset and the extrapolation is necessarily uncertain.
This text was published by llm-bench.io . It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
1 source- Reddit discussion reddit.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Hardware & Compute
All →- Google signs record 396 MW geothermal deal with Fervo Energy · 1 src
- Scavenger Finds 12 RTX 3070 GPUs from Crypto Mining Era · 1 src
- ggml-cuda adds missing AMD GCN MMQ config for RDNA2 GPU support · 2 src
- Marvell positions itself as backbone of next‑gen AI data‑center infrastructure · 1 src
- ABF Substrates Under Pressure as AI Boom Strains Supply Chain · 1 src
Comments
via GitHub Discussions