July 30, 2026 · Tom's Hardware
M4 Max beats Nvidia GB10 and AMD Strix Halo in local LLM decode speed
Tom's Hardware benchmarked Apple's M4 Max Mac Studio against Nvidia's GB10 and AMD's Strix Halo platforms for local large language model (LLM) inference. Despite having less raw memory bandwidth (546GB/s) than some competitors, the M4 Max won on decode throughput, showing bandwidth alone doesn't determine local inference performance.
Why it matters: As more developers run LLMs locally for privacy or cost reasons, real-world benchmarks like this matter more than spec-sheet comparisons, suggesting memory architecture and software optimization can outweigh raw bandwidth numbers. This keeps Apple Silicon competitive as a local-inference platform even against dedicated AI hardware from Nvidia and AMD.