M5 Ultra Mac Studio 评测:运行本地 AI 智能体的梦幻 Mac
M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents

原始链接: https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/

全新的 M5 Ultra Mac Studio(256 GB 内存)已证明其作为运行本地 AI 智能体的首选设备,成为了云服务和高端游戏 PC 的强有力替代方案。 尽管 RTX 5090 在原始带宽上仍保持微弱优势,但 Mac Studio 的统一内存架构、散热效率和紧凑设计,使其在实际的长期运行 AI 工作流中更具优势。M5 Ultra 较 M3 Ultra 有了巨大飞跃,提示词处理速度提升了 150%,生成速度提高了 70%。这些改进实现了与 Qwen3.8-Flash-Next 等复杂多轮 AI 智能体的无缝、低延迟交互,且无需承担云端 API 的持续费用或隐私隐患。 对于作者而言,这台机器已成为执行复杂自动化研究任务的引擎,例如以零成本处理用于大型审查的数百份文档。M5 Ultra 将隐私、性能以及安静节能的机身设计融为一体,为本地 AI 树立了重要的里程碑,使“最前沿”的智能体工作流对于高阶用户和技术爱好者来说,既可行又极其可靠。

Hacker News 最新 | 往期 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 M5 Ultra Mac Studio 评测:运行本地 AI 智能体的梦想 Mac ( macstories.net ) 17 点 由 piotrgrabowski 39 分钟前 | 隐藏 | 往期 | 收藏 | 2 条评论 帮助 simonw 1 分钟前 | 下一条 [–] 我最感兴趣的数据隐藏在底部的图表中——Mac Studios 与 RTX 5090 的速度对比: Qwen3.8 27B 生成速度(token/秒) 提示词长度 8K 64K 128K 256K RTX 5090 PC 59 51 44 不适用 M5 Ultra 48 39 32 24 M3 Ultra 31 23.5 20 15 该部分的更多对比数据: https://www.macstories.net/stories/m5-ultra-mac-studio-revie... 回复 snarfy 1 分钟前 | 上一条 | 下一条 [–] 12,299 美元 回复 社区准则 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系 搜索:
相关文章

原文

For the past few days, I’ve been testing the (currently) top-of-the-line M5 Ultra Mac Studio with 256 GB of RAM.

I’ll cut to the chase: the M5 Ultra Mac Studio is a dream machine for local AI agents. This computer makes it possible to run personal assistants powered by local models with great performance and no additional cloud costs. If you’ve been skeptical of testing OpenClaw or Hermes Agent with local models because they’d never be even remotely near the intelligence and speed of cloud ones, this Mac will change your mind about that.

Since last Thursday, I’ve been comparing this Mac Studio to its predecessor, the M3 Ultra with 512 GB of RAM, as well as my own desktop gaming PC with an RTX 5090 inside. For its size, price, thermal performance – not to mention Apple’s approach to unified memory – the M5 Ultra Mac Studio has fundamentally changed how I think about models running locally and what they can enable now. A 5090, of course, still has an edge over the M5 Ultra thanks to its higher memory bandwidth. But considering the sheer size of my PC build, as well as its heat and noise, I would prefer an M5 Ultra Mac Studio any day. It also happens to be a Mac, with an operating system that looks nice and doesn’t suck, plus a vibrant app ecosystem. (Windows fans, I’m sorry, but Microsoft software will never get my sympathy.)

As I’ll explore in this article, running the latest Qwen3.8-Flash-Next model on the M5 Ultra Mac Studio has been so nice and fast, I’ve made it my default in both Open Minis for iOS and Hermes Agent. That’s right: the personal assistants I use the most – more than Siri AI, in fact – are now entirely powered by a model running locally on a Mac Studio. Furthermore, thanks to the M5 Ultra’s faster GPU and higher memory bandwidth, these agents start responding more quickly, stay fast at larger context windows, and can run long, multi-turn loops without slowing to a crawl as the session grows. Because of this, I’ve also been using local models in the Codex app on my Mac – either as main threads or subagents orchestrated by GPT-6 Astra – and I’ve had a great experience doing so.

I should note upfront that I’m not an AI developer by trade: I do not train or fine-tune models. I’m a tinkerer at heart, and I’ve been playing around with local AI models for over a year at this point. This summer, I went all-in on local AI usage for a big project I was working on, which I will explain in the following section.

My goal with this article is to provide you with a mix of two things: numbers and visualizations based on the (many) tests I’ve run over the course of four days, and an explanation of my practical use cases for local AI applied to my workflow and how I get things done for MacStories.

Let’s dive in.

Why Local AI?

Let’s address the elephant in the room first: why bother with local AI at all when cloud frontier models are better and often faster?

It’s a fair question. You need expensive hardware to run these models, and by the time you’ve repaid your investment, you could have used the most expensive Anthropic subscription for several years, still saved money, and got better performance in return.

Different people will have different answers to this question. Some might say they use local models because of privacy: they’d rather rely on local intelligence for sensitive data and documents than upload anything to an external cloud. Others might argue that it’s simply cool – and I do not disagree. For some, it’s a work-related task: if you’re an AI developer, it makes sense to have a great local setup for training your own adapters or fine-tuning models.

For me, the journey into local AI has been characterized by a mix of the “cool, why not?” factor of it all as well as considerations about privacy and costs.

As I will share later this week with Club MacStories members, my research and writing setup for the iOS and iPadOS 27 review this summer has been powered and made possible by local AI. Back in June, I created an internal app, called Desk, to organize hundreds of notes, sessions, PDF documents, and clipped webpages related to iOS and iPadOS 27, as well as chapters of the review. By the end of the process, the project consisted of 310 documents. In Desk, a team of agents – all based on DeepSeek V4 Flash, plus olmOCR for PDFs – ran 24/7, for 99 days, to perform the following tasks:

  • Transcribe my favorite WWDC sessions (using summarize plus LLM processing)
  • Extract features of iOS and iPadOS 27 from clipped webpages, sessions, PDF guides, and my own notes
  • Cross-reference features across different sources, and keep track of which features belonged to which chapter of the review
  • Extract features and bugs from screenshots I uploaded
  • Work with the Notion API to organize everything across multiple databases

When I started working with this setup in early June, I quickly realized that relying on the OpenAI or Anthropic APIs for this kind of always-on, persistent background task would be…cost-prohibitive, to say the least. So I pivoted to local AI, and the result is the iOS and iPadOS 27 review you can read on MacStories. It was all written by me, the old-fashioned human way. But the entire research stack, deep-linking between notes, and keeping track of new features and betas were all performed by my agents, running locally on the Mac Studio, for a total cost of $0.

If you don’t think that’s neat, or a powerful concept to explore, then this article probably isn’t for you – and I understand. Dealing with these models is fiddly, and it’s not something I would ever recommend to someone who (rightfully) just wants to pay $20 to use Claude Cowork. This kind of setup is, by definition, the bleeding edge of AI workflows at the moment.

If you fall on the other end of the spectrum, though, and if you think this kind of stuff is neat…let me tell you: the M5 Ultra Mac Studio is a massive leap in performance for local models powered by MLX, and I have a few examples to prove it.

A Leap for Prompt Processing and Generation

As you may have seen from the announcement and my initial coverage, the M5 Ultra Mac Studio looks identical to the M3 Ultra model it replaces, but it comes with an all-new Apple silicon architecture that uses UltraFusion to connect two dual-die M5 Max chips to form a quad-die architecture, which is a first for the Apple ecosystem. As far as local AI workloads are concerned, there are two areas we have to pay attention to (and which I have been following since my coverage of the M5 iPad Pro for local AI last year): GPU and memory bandwidth.

The M5 Ultra has a next-gen GPU with 80 cores, each with a Neural Accelerator that grants it up to 4.5× the peak GPU compute for AI compared to the M3 Ultra. As for memory, Apple’s unified memory architecture still tops out at 512 GB as before (although that model will come out in late October), but its bandwidth has jumped from 819 GB/s to 1.2 TB/s, or 50% higher than the M3 Ultra.

With these numbers in mind, I started testing the M5 Ultra against the M3 Ultra with 512 GB of RAM and my RTX 5090. I’ll share more details on testing below, but the short version is this: with the M5 Ultra, you spend considerably less time waiting for a model to read your prompt and begin generating a response; and when it does start answering, text appears much faster than it used to on the M3 Ultra. These two improvements alone make the machine viable for modern agentic loops that require fast iteration with a model and, as a result, larger context windows.

In my day-to-day experience with agents running on the M5 Ultra, these improvements to token prefill (or how quickly a prompt can be processed) and token generation are the changes I noticed immediately. When comparing a model running on the M3 Ultra and M5 Ultra side by side with Open Minis on iOS, the M5 Ultra was ~70% faster on average than the M3 Ultra at generating a response. As we’ll see later, having a model such as Qwen3.8-Flash-Next clear 100 tokens/second on short prompts and still write at 60 to 85 with 64K to 256K of context behind it is no joke, and it enables the kind of agentic back-and-forth between you and the model that feels great to use, particularly when tool calls are involved.

However, I was more impressed with the performance gains in token prefill. When you use agentic assistants such as Hermes or Codex, a model receives a whole block of instructions that include things like the system prompt, user personalization and session memories, skill and MCP descriptions, and more. Some agents are better than others at trimming the instructions they send, but, generally, whenever you use a modern agent, you’re not starting with an empty context window. Because of this, I’ve never been able to consistently use local models with this new wave of agents: they would work, but I’d stare at an empty screen and a loading indicator for a while before the model would start generating a response. And on every turn of the loop, performance would get worse (because of the larger context of the session), and I’d wait some more time.

In my tests, prompt processing is up 150% on average from the M3 Ultra – a ~2.5× improvement from my previous setup. This change alone makes local models solid choices in apps like Open Minis and Hermes Agent. When I ask Flash-Next on the M5 Ultra to get my tasks for the week with RemCTL, I don’t have to wait around for the agent to process my prompt and Open Minis’: in just a few seconds, it gets to work by reasoning, performing tool calls, and so forth. And when I’m working on a large project, such as the voxel Colosseum demo below, the model is able to process multi-turn loops quickly, dispatch and coordinate subagents, and do it all at 60 to 85 tokens per second as the thread grows longer.

I’m a big believer in assistants that can agentically perform tasks in addition to answering questions, but in order to feel nice to use, they have to be fast. Over the past few months, I’ve tested several “boutique” cloud providers with Open Minis: Inco, which serves Kimi K3 at over 300 TPS; Cerebras, with Qwen3.8-27B at a whopping 1,800 TPS; and the likes of Fireworks and Baseten, each breaking the 150 TPS barrier. All of those providers feel extremely good to use in Open Minis and Hermes, but they are expensive (I burned through $20 of Inco credits in literally 10 minutes last week), and, of course, all my data is going…somewhere when I use them. When I fire up Open Minis with Flash-Next and the collection of Apple CLIs I’m creating, everything stays local, inside a computer I can see and reboot whenever I want.

Most importantly: a model like Flash-Next can be “small” enough to run at higher quantizations on a 256 GB M5 Ultra (I can run 5-bit entirely in RAM; 6- and 8-bit can offload their n-gram tables to SSD with this new architecture) but also intelligent enough to sustain long threads and multiple agentic tool calls.

For my taste, 5-bit quantization hits the sweet spot on this version of the Ultra with a balance of intelligence, performance, and memory consumption. But I already know that, if I ever get to test a 512 GB M5 Ultra, I’d be really interested to measure performance of the 8-bit quant without SSD offloading.

I have not spent much time tinkering with offloading coding tasks for my various projects to a local model, but I’ve done a few interesting experiments. With this kind of performance, and especially given the ability to stack up to three concurrent Flash-Next sessions with subagents in oMLX with 256 GB of RAM (more later), I can now realistically consider handing off simpler coding tasks to a local model and have frontier cloud ones review their work. For instance, I was able to set up Qwen3.8-Flash-Next in Codex, which lets me use a local model with the Codex harness. This means that I can let a main GPT model orchestrate local subagents, have Flash-Next coordinate its own subagents, or even just use the model from my phone with Codex Remote on iOS.

I’m curious to read more on this topic from actual developers who are getting an M5 Ultra soon. With open-weights models now outperforming on consumer hardware what was considered “frontier” ~10 months ago, and with performance on an M5 Ultra now making agentic coding feasible, I think we’re going to see some fascinating experiments from the MLX community very soon.

M5 Ultra vs. RTX 5090

As you’ll see from the visualizations later in this article, NVIDIA’s RTX 5090 is still faster than Apple’s M5 Ultra despite its “meager” 32 GB of VRAM, for two different reasons.

Prompt processing speeds are dictated by compute: the model reads the whole prompt in one giant matrix multiplication, which is exactly the job NVIDIA’s Tensor Cores were built for. Apple’s new Neural Accelerators (one in each of the M5 Ultra’s 80 GPU cores) narrow the gap, but can’t close it. On a 6,000-token prompt, the M5 Ultra read at ~1,700 tok/s; the 5090 delivered a staggering ~3,000 with the Qwen model I tested in LM Studio. Token generation, on the other hand, is bandwidth: the model writes one token at a time and pulls the entire model back out of memory for each one, so the 5090’s 1.79 TB/s against the M5 Ultra’s 1.2 TB/s gives it a steady ~25% lead at every prompt size. What the 5090 doesn’t have is memory: at 256K, the 5090 only finishes with an 8-bit attention cache. 32 GB of VRAM only goes so far.

There are, however, two problems with this comparison. First, while the 5090 does still edge out the M5 Ultra with smaller models, its lack of a unified memory pool means that I’m limited to the 32 GB of VRAM in the GPU if I want to run a model at blazing-fast speeds. The moment I want to run anything exceeding 32 GB (such as the aforementioned higher Flash-Next quants), the 5090 must offload model layers over PCIe to (much slower) system RAM, and that’s no way to live.

Second, my gaming PC is massive compared to a Mac Studio that fits on my desk – and I have a compact build with a Lian-Li A3 case. Not to mention how loud and hot it gets when I’m running local models at high context windows: when I walked into my office after some benchmarks had run, it was uncomfortably warmer compared to the rest of my apartment. By contrast, the “diminutive” Mac Studio on my desk was warm to the touch, but it was also appreciably quieter than my 5090, the fans were not spinning as fast or loudly, and, most important, it allowed me to run larger models such as GLM-5.3-Flash locally with decent performance thanks to Apple silicon’s unified memory. In my day-to-day use, when I was running Flash-Next oQ4e all the time, I could never hear the fan of the Studio on my desk unless I placed my ear directly on top of the computer.

Judging by the progress Apple has made in recent years, I wouldn’t be surprised to see an M7 Ultra that outperforms the memory bandwidth of a 5090 in the near future. But that’s a story for another time.

A Note on Testing

Lastly, before we jump into raw numbers and charts: how did I test everything?

Automated tests were conducted with a testing harness I built with GPT-6 Astra, which coordinated multiple instances of Codex across my M3 Ultra and M5 Ultra Mac Studio, as well as my PC with the Codex app for Windows and Computer Use. On macOS, I chose oMLX (version 0.7.0.dev2) as the local backend for MLX models, and ran Qwen3.8-Flash-Next-oQ4e-mtp, GLM-5.3-Flash-MLX-mixed-4_8bit, and Qwen3.8-27B-oQ4e-mtp on macOS Golden Gate 27.0 for the majority of my tests. On Windows, I used LM Studio and Qwen3.8-27B-GGUF with CUDA 12 runtime and with all 66 layers offloaded to the GPU for the full-GPU tests, plus separate tests splitting the model between GPU and system RAM.

Alongside separate experiments with Open Minis’ native subagents, I used a custom testing harness to measure concurrent requests and workflows involving a lead model and multiple helpers, with oMLX serving the Mac models and LM Studio serving the Windows model.

Numbers were collected by Astra over the course of four days, and later visualized by Claude Fable 5.1 and Opus 5 using Anthropic’s upcoming Projects feature, which I was able to test early when working on this story. The interactive visualization was built with pure HTML and CSS based on MacStories’ style, and it includes comments and annotations by yours truly.

My goal with the following interactive widgets was not only to help you understand the numbers more clearly, but also to visualize what the stats mean in practice. I’m quite happy with the widgets that approximate what different tokens per second feel like, since that’s a metric that’s often tricky to visualize. I hope these animated charts will be more useful than regular “static” ones you’ve probably seen elsewhere (which are also included below).

Visualizing the M5 Ultra

The M5 Ultra for Local AI Agents

As should be clear at this point, the performance gains of the M5 Ultra are real, and they show how Apple’s investment in custom silicon and its unified memory architecture is paying dividends for tinkerers and developers.

Despite my tests, I feel like I’ve barely scratched the surface of what’s possible with the M5 Ultra and its 256 GB of RAM. As more developers and open-source maintainers get their hands (and agents) on the M5 Ultra, I’m sure we’ll see more optimizations in quantization to allow even larger models to run with superior performance on this computer. For instance, I didn’t even have time to test DwarfStar – a fascinating project (made in Italy!) that is making it possible to run local frontier models on all kinds of Mac configurations with even less memory; nor did I have time to check out Inco Splash, a new inference engine designed for Apple silicon and specific models. Likewise, I didn’t have time to test Exo, whose RDMA implementation should (in theory) allow me to split and distribute inference across M3 Ultra and M5 Ultra via Thunderbolt 5, all while running an OpenAI-compatible server in front of it to serve an API for local agents.

And, of course, I can’t even begin to imagine what the high-end M5 Ultra with 512 GB of RAM will allow in terms of scaling up models capable of running locally. I hope to be able to test it eventually, too.

At the end of this experiment, I have a simple, tangible result: the M5 Ultra lets me run local agents with incredible performance, with less time spent staring at a blank screen and everything happening on a single, compact, cool, and quiet machine on my desk.

This would have seemed impossible a couple of years ago. But here we are.

联系我们 contact @ memedata.com