{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-08-29T21:00:37.966Z","headline":"在本地运行 Qwen3.8 27B：来自我的 Mac Studio 的实际数据","description":"Qwen3.8 27B（27.3B 参数，混合注意力架构，262，144 token 上下文窗口，Apache 2.0 开源）在 Mac Studio M3 Ultra 上经 Ollama 以 Q4_K_M 量化（17GB）生成速度约 14 tokens/s。","url":"https://www.aioga.com/news/cmte242e701jdrog2vz9p2cy4/","mainEntityOfPage":"https://www.aioga.com/news/cmte242e701jdrog2vz9p2cy4/","datePublished":"2026-08-29T07:00:22.003Z","dateModified":"2026-08-29T07:00:22.003Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://terminalbytes.com/run-qwen-3-8-27b-locally","https://aihot.virxact.com/items/cmte242e701jdrog2vz9p2cy4"],"canonicalUrl":"https://www.aioga.com/news/cmte242e701jdrog2vz9p2cy4/","directAnswer":{"@type":"Answer","text":"文章记录了 Qwen3.8 27B 在 Mac Studio M3 Ultra 上的本地运行测试：采用 Ollama 的 Q4_K_M 量化后，模型文件约 17GB，摘要给出的生成速度约为每秒 14 tokens。","url":"https://www.aioga.com/news/cmte242e701jdrog2vz9p2cy4/","dateCreated":"2026-08-29T07:00:22.003Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"terminalbytes.com source article","url":"https://terminalbytes.com/run-qwen-3-8-27b-locally","datePublished":"2026-08-29T07:00:22.003Z","provider":{"@type":"Organization","name":"terminalbytes.com","url":"https://terminalbytes.com/run-qwen-3-8-27b-locally"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmte242e701jdrog2vz9p2cy4","datePublished":"2026-08-29T07:00:22.003Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmte242e701jdrog2vz9p2cy4"}}],"aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","originalPublisher":{"name":"terminalbytes.com","url":"https://terminalbytes.com/run-qwen-3-8-27b-locally"},"geoDeepAnswer":null,"article":{"id":"cmte242e701jdrog2vz9p2cy4","slug":"cmte242e701jdrog2vz9p2cy4","url":"https://www.aioga.com/news/cmte242e701jdrog2vz9p2cy4/","title":"在本地运行 Qwen3.8 27B：来自我的 Mac Studio 的实际数据","title_en":"","summary":"Qwen3.8 27B（27.3B 参数，混合注意力架构，262，144 token 上下文窗口，Apache 2.0 开源）在 Mac Studio M3 Ultra 上经 Ollama 以 Q4_K_M 量化（17GB）生成速度约 14 tokens/s。","source":"Hacker News 热门（buzzing.cc 中文翻译）","sourceUrl":"https://terminalbytes.com/run-qwen-3-8-27b-locally","aiHotUrl":"https://aihot.virxact.com/items/cmte242e701jdrog2vz9p2cy4","publishedAt":"2026-08-29T07:00:22.003Z","category":"技巧观点","score":75,"selected":true,"articleBody":["For the past 10 days, Qwen3.8 27B has been quietly running on my Mac Studio as a background assistant. It summarizes my RSS feeds into a morning digest, renames and files the PDFs I scan into something searchable, and handles whatever summarizing chore I throw at it. Mundane stuff. That’s the appeal: this is the first local model I’ve trusted enough to leave alone with mundane stuff.","Then last week the model was suddenly everywhere on r/LocalLLaMA, my feeds filled up with benchmark charts, and I realized I’d been sitting on the one thing most of those threads were missing: a machine that can actually run it properly, and time to measure it.","So I benchmarked it. Five timed runs per model, same prompts, same machine, plus a 1-bit experiment that surprised me twice. Here’s everything I measured, and what it means for the hardware you’d need to run this thing yourself.","Qwen3.8-27B is a 27.3B parameter dense model with a hybrid attention design (the architecture tag in the GGUF is qwen35 , which matters later). It’s multimodal, with image and video understanding built in, carries a 262,144-token native context window, and ships under Apache 2.0. The official model card：https://huggingface.co/Qwen/Qwen3.8-27B claims 61.7 on SWE-bench Pro and 89.2 on GPQA Diamond, numbers that would have been frontier-lab territory a year ago.","The community reaction skipped right past that benchmark table. What lit the threads up was what people did with the model in its first week: one team wired it into their coding pipeline as a drop-in for a paid API model and reported it held up, and OCR testers claimed quality above some commercial cloud tiers. The line from the most-upvoted thread that stuck with me: “this is the first local model that feels like more than a toy.”","My contribution is the one measurement most of those charts are missing: what this model actually does on Apple silicon you can buy today.","My daily machine is a Mac Studio M3 Ultra with 256GB of unified memory, the same box I used for the DeepSeek V4 Flash guide：/run-deepseek-v4-flash-at-home/. I ran five timed generations per model through ollama run --verbose , varied technical prompts, ~200-500 word answers, and averaged the stats. Both models are the default Ollama Q4_K_M quant, both almost exactly 17GB on disk.","The headline number first: the new model generates at half the speed of its predecessor. Same parameter count, same quant size, same machine. The hybrid attention architecture is new, and the Metal kernels in Ollama clearly haven’t caught up yet. I expect this gap to narrow as the runtimes mature; the same thing happened with other novel architectures.","It didn’t actually cost me time, though. Qwen3.8 answered the same prompts in roughly 1,000 tokens where 3.6 rambled through 2,000-3,300. The arithmetic: 2,058 tokens at 28.6 tok/s is 72 seconds, 955 tokens at 14.2 tok/s is 67 seconds. Slower per token, faster per answer.","While it generates, the CPU barely notices, because on Apple silicon the inference runs on the GPU through Metal. The cover image of this post is exactly that moment, captured with macmon：https://github.com/vladkens/macmon mid-generation: GPU pinned at 100% pulling 63.95W, CPU sipping 6W, answer streaming the whole time.","The Stats menu bar app：https://github.com/exelban/stats tells the same story from the GUI side: all 60 GPU cores at 100%, system power draw touching 291W:","The single most-upvoted Qwen3.8 thread of the week celebrated Unsloth’s 1-bit quant：https://huggingface.co/unsloth/Qwen3.8-27B-GGUF, a 6.7GB file the poster affectionately called the “brain damage quant”. A 27B model in the memory footprint of a 7B. I had to try it.","309 tok/s prompt processing, 27.2 tok/s generation. Nearly twice my Q4 speed, in under 8GB of RAM.","Then I asked it questions. Factual recall was genuinely fine: it knew Canberra is Australia’s capital and correctly explained the Sydney-Melbourne compromise behind it. But when I asked for a simple bash one-liner, it produced a working command and then couldn’t stop second-guessing itself, burning 400 tokens cycling through alternatives without ever committing to a final answer.","This matches what Unsloth themselves say: their quantization docs：https://unsloth.ai/docs/basics/dynamic-3.0-ggufs are blunt that 1-bit should not be used for agentic or tool-calling work, and their divergence testing shows accuracy on long tasks collapsing at 1-bit while general knowledge survives. Their stated minimum for tool calling is the Q2_K_XL quant at 9.8GB.","From my experience: the 1-bit quant is a party trick that teaches a real lesson. Quantization doesn’t degrade a model evenly. Facts survive, decisiveness dies. If your use case is “answer trivia fast on a potato,” it genuinely works. If it’s anything agentic, pay the extra 3GB for Q2.","Unsloth publishes the full GGUF ladder, so here’s the practical version. Budget the file size plus a few GB for context and the vision projector.","For the 32GB tier, boxes like the GEEKOM A6 with a Ryzen 7 6800H and 32GB：https://www.amazon.com/dp/B0H8YXYNPQ?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1 or the GMKtec M6 Ultra with DDR5：https://www.amazon.com/dp/B0FLJQW1RD?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1 run the Q4 quant the way my benchmarks above run it, just slower: think single digits of tokens per second on CPU inference rather than 14. That pace suits background jobs like mine; it would test your patience in an interactive chat.","：https://www.amazon.com/dp/B0H8YXYNPQ?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1","If you want the model at real speed without buying Apple, the community consensus target is AMD’s Strix Halo platform. The strix-halo-guide project：https://github.com/hogeheer499-commits/strix-halo-guide measured the official Q4_K_M at 20.4 tok/s generation and 292 tok/s prompt processing on a Ryzen AI Max+ 395, with raw CSVs to back it. The GMKtec EVO-X2 with 64GB：https://www.amazon.com/dp/B0F53QXNGH?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1 is the value entry into that platform at $1,999, and 128GB configs like the BOSGAME M5：https://www.amazon.com/dp/B0H94TVN8G?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1 open up the Q8 and BF16 rows of the table, plus much bigger models. I covered that whole platform decision in best mini PC for local LLMs：/best-mini-pc-for-local-llm-2026/.","：https://www.amazon.com/dp/B0F53QXNGH?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1","：https://www.amazon.com/dp/B0H94TVN8G?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1","One warning from the current market: RAM prices are still inflated. A 64GB DDR5 SODIMM kit：https://www.amazon.com/dp/B09S2QLBWC?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1 currently runs $750-870. If you’re buying a mini PC for local LLM work, buying it with the RAM already installed is currently cheaper than upgrading later, which is backwards from every instinct I’ve built over twenty years of buying computers.","：https://www.amazon.com/dp/B09S2QLBWC?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1","GPU owners scale differently: community reports put a dual RTX 3090 setup around 60 tok/s and an RTX 5090 at 75-140 tok/s depending on runtime, with 16GB cards running IQ4 quants with quantized KV cache.","Ollama is the short path. The model page is ollama.com/library/qwen3.8：https://ollama.com/library/qwen3.8:","Run it with --verbose and every answer ends with a stats block like this one:","You’ll need Ollama 0.32.12 or newer; the model metadata declares it as a minimum.","For llama.cpp, here’s the gotcha. My Homebrew llama.cpp was a few weeks old, and it flatly refused the file:","The hybrid architecture needs current kernels. brew update && brew upgrade llama.cpp fixed it, and the same vintage requirement applies to any llama.cpp-based frontend (LM Studio, Jan, koboldcpp): if Qwen3.8 fails to load, update the runtime before debugging anything else.","The benchmark numbers matter less to me than what the model has been doing since I pulled it: unglamorous background work that used to either not happen or leak to a cloud API.","The morning feed digest : a launchd job collects my RSS unread items overnight and has qwen3.8 compress them into one summary I read with coffee. The 262k context means a week of feeds fits in a single prompt.","Scan filing : paper mail gets scanned, and the model reads each PDF’s text and renames it into my YYYY-MM-vendor-what-it-is convention. The vision capability means it handles the scans OCR mangles.","And when a forum thread runs to 400 comments, it gets pasted in and summarized with positions attributed. This post’s research generated a few of those, which felt pleasantly circular.","None of this cares about tokens per second. The requirements are a model smart enough not to file the insurance letter as a takeout menu, hardware I already own, and nothing leaving the house. That’s the actual pitch for local models in 2026, and it’s the same argument I made in the self-hosting revolution：/self-hosting-revolution-mini-pcs/: even the small cloud dependencies are worth replacing.","Can I run Qwen3.8 27B on 16GB of RAM? Yes, at 1-bit or 2-bit quantization (6.7-9.8GB files). The 2-bit is the smallest quant Unsloth considers usable for tool calling. Q4 quality needs 32GB.","Is it better than Gemma 4? Different shapes. Gemma 4’s 26B-A4B is a sparse MoE that generates much faster on the same hardware (my Gemma 4 guide：/run-gemma-4-mini-pc-without-gpu/ has those numbers). Qwen3.8 27B is dense, slower per token, and the community currently rates it well ahead on coding and agentic work. For a background assistant I’d pick Qwen3.8; for interactive chat on modest hardware, Gemma 4 still makes sense.","Why is it slower than qwen3.6 on my machine too? The hybrid attention architecture is new and runtime kernels (Ollama Metal, llama.cpp Vulkan/CUDA) haven’t fully optimized for it yet. Expect the gap to narrow with updates. Partial consolation: it uses far fewer tokens per answer, so finished-answer latency is closer than the tok/s gap suggests.","Does vision work locally? Yes. The Ollama build ships the vision projector (about 460M parameters) and image input works out of the box. Video understanding support in local runtimes is still patchy.","What about Qwen3.8-Flash-Next? It shipped its weights while I was writing this: a 180B MoE, roughly 110GB at Q4. Different hardware class entirely: you need 128GB-class unified memory, which today means a $3,500+ Strix Halo box or a big Mac. If the “surprisingly local-friendly” architecture claims hold up, it’s a future post."],"articleImages":[{"sourceUrl":"https://terminalbytes.com/run-qwen-3-8-27b-locally/run-qwen-3-8-27b-locally.jpg","alt":"macmon on a Mac Studio M3 Ultra mid-generation: GPU pinned at 100 percent pulling 64W while the CPU draws 6W","afterParagraph":1,"url":"/media/articles/cmte242e701jdrog2vz9p2cy4/17d7fabebde13573.jpg"},{"sourceUrl":"https://terminalbytes.com/run-qwen-3-8-27b-locally/qwen36-vs-qwen38-ollama-comparison.jpg","alt":"Side by side terminal panes comparing qwen3.6 27B at 28 tokens per second with qwen3.8 27B at 13 tokens per second","afterParagraph":6,"url":"/media/articles/cmte242e701jdrog2vz9p2cy4/12a901b2c0633da4.jpg"},{"sourceUrl":"https://terminalbytes.com/run-qwen-3-8-27b-locally/qwen38-gpu-100-percent-stats-app.jpg","alt":"Stats app GPU panel showing Apple M3 Ultra with 60 cores at 100 percent utilization and 291W power draw during Qwen3.8 generation","afterParagraph":10,"url":"/media/articles/cmte242e701jdrog2vz9p2cy4/4f516a4cde253e4f.jpg"},{"sourceUrl":"https://terminalbytes.com/run-qwen-3-8-27b-locally/qwen38-1bit-llamabench.jpg","alt":"llama-bench results for the 1-bit Qwen3.8 27B quant showing 309 tokens per second prompt processing and 27 tokens per second generation","afterParagraph":11,"url":"/media/articles/cmte242e701jdrog2vz9p2cy4/703e5a1bfd07558d.jpg"},{"sourceUrl":"https://terminalbytes.com/run-qwen-3-8-27b-locally/geekom-a6-ryzen-7-6800h-32gb-terminalbytes.com.jpg","alt":"GEEKOM A6 mini PC with Ryzen 7 6800H and 32GB DDR5 RAM","afterParagraph":17,"url":"/media/articles/cmte242e701jdrog2vz9p2cy4/719f83a004b6f999.jpg"},{"sourceUrl":"https://terminalbytes.com/run-qwen-3-8-27b-locally/gmktec-evo-x2-ryzen-ai-max-395-terminalbytes.com.jpg","alt":"GMKtec EVO-X2 mini PC with Ryzen AI Max+ 395 and 64GB unified memory","afterParagraph":19,"url":"/media/articles/cmte242e701jdrog2vz9p2cy4/191d2a3964a7895c.jpg"}],"mediaStatus":"ok","articleBodyZh":["在过去的10天里，Qwen3.8 27B悄悄地在我的Mac Studio上作为后台助手运行。它会把我的RSS订阅整理成晨间摘要，将我扫描的PDF重命名并归档成可搜索的格式，并处理我交给它的各种总结任务。都是一些琐碎的事情。这就是它的吸引力所在：这是我第一次信任的本地模型，可以放心地让它处理琐碎的事务。","然后就在上周，这个模型突然在r/LocalLLaMA上无处不在，我的订阅里充满了基准测试图表，我意识到我手里拥有了那些主题帖大多数缺少的东西：一台可以真正运行它的机器，以及测量它的时间。","于是我对它进行了基准测试。每个模型进行了五次计时运行，相同的提示语，相同的机器外加一次让我惊讶两次的1-bit实验。以下是我测量的所有内容，以及这些数据对你自己运行这款模型所需硬件的意义。","Qwen3.8-27B是一个拥有273亿参数的密集模型，采用混合注意力设计（GGUF中的架构标签是qwen35，这在后面会很重要）。它是多模态的，内置图像和视频理解功能，拥有262,144标记的原生上下文窗口，并在Apache 2.0下发布。官方模型卡：https://huggingface.co/Qwen/Qwen3.8-27B 声称在SWE-bench Pro中得分61.7，在GPQA Diamond中得分89.2，这些数字在一年前还属于前沿实验室的领域。","社区的反应直接跳过了那张基准表。引爆讨论的，是人们在模型的第一周里做的事情：一个团队把它接入到他们的编码管道，作为付费API模型的替代品，并报告说它表现稳定；OCR测试者则声称其质量超过了一些商业云服务。最受点赞的帖子的那句话让我印象深刻：“这是第一个感觉不只是玩具的本地模型。”","我的贡献是那些图表大多缺少的一个测量：这款模型在你今天可以买到的苹果硅芯片上的实际表现。","我每天使用的机器是一台Mac Studio M3 Ultra，配备256GB统一内存，就是我用于DeepSeek V4 Flash指南的那台机器：/run-deepseek-v4-flash-at-home/。我通过ollama run --verbose运行了每个模型五次计时生成，变化了技术提示语，生成约200-500字的答案，并对统计数据取平均值。两个模型都是默认的Ollama Q4_K_M量化版本，磁盘占用几乎都正好是17GB。","先说头条数字：新模型的生成速度是前代的一半。参数数量相同，量化大小相同，使用的机器也相同。混合注意力架构是新的，而 Ollama 中的 Metal 内核显然还没有跟上。我预计随着运行时的发展，这一差距会缩小；其他新型架构也是如此。","不过，这实际上并没有花我多少时间。Qwen3.8 用大约 1,000 个 token 回答了同样的提示，而 3.6 则在 2,000-3,300 个 token 之间絮絮叨叨。计算方式：2,058 个 token 以 28.6 tok/s 处理需要 72 秒，955 个 token 以 14.2 tok/s 处理需要 67 秒。每个 token 更慢，但每次回答更快。","生成时，CPU 几乎没有注意，因为在 Apple Silicon 上，推理通过 GPU 和 Metal 运行。这篇帖子的封面图片正是这一时刻，用 macmon 捕捉：https://github.com/vladkens/macmon 生成中：GPU 固定在 100%，功耗 63.95W，CPU 仅消耗 6W，整个过程持续流式输出回答。","Stats 菜单栏应用：https://github.com/exelban/stats 从 GUI 侧讲述了同样的故事：所有 60 个 GPU 核心 100% 运行，系统功耗达到 291W：","本周单个最高赞的 Qwen3.8 讨论帖庆祝了 Unsloth 的 1 位量化：https://huggingface.co/unsloth/Qwen3.8-27B-GGUF，这是一个 6.7GB 的文件，贴主亲切地称为“脑损量化”。一个 27B 模型的内存占用只有 7B。我必须试试。","提示处理速度 309 tok/s，生成 27.2 tok/s。几乎是我的 Q4 速度的两倍，内存不足 8GB。","然后我问了一些问题。事实回忆确实很好：它知道堪培拉是澳大利亚首都，并正确解释了悉尼-墨尔本之间的妥协。但是当我要求一个简单的 bash 单行命令时，它生成了一个可用命令，却无法停止自我怀疑，浪费 400 个 token 在循环替代方案中，始终没有给出最终答案。","这与 Unsloth 自己的说法一致：他们的量化文档：https://unsloth.ai/docs/basics/dynamic-3.0-ggufs 直言 1 位量化不应用于具有自主决策或工具调用的任务，他们的发散测试表明在长任务中 1 位精度会崩溃，而一般知识仍然保留。他们规定工具调用的最低要求是 Q2_K_XL 量化，占用 9.8GB。","根据我的经验：1位量化是一种派对技巧，但它提供了一个真正的教训。量化不会均匀地减损模型。事实依旧存在，决定性消失。如果你的用例是“在低性能设备上快速回答冷知识”，它确实有效。如果涉及任何智能代理任务，请额外花3GB使用Q2。","Unsloth发布了完整的GGUF梯度，所以这里是实用版本。请预算文件大小加上几GB用于上下文和视觉投影器。","对于32GB级别，像GEEKOM A6搭载Ryzen 7 6800H和32GB内存：https://www.amazon.com/dp/B0H8YXYNPQ?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1 或者GMKtec M6 Ultra带DDR5：https://www.amazon.com/dp/B0FLJQW1RD?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1 可以运行Q4量化，就像我上面基准测试的那样，只是速度更慢：在CPU推理中每秒处理的token只有个位数，而不是14。这个速度适合像我这样的后台任务；在互动聊天中会考验你的耐心。","：https://www.amazon.com/dp/B0H8YXYNPQ?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1","如果你想在不买Apple的情况下以真实速度运行模型，社区共识的目标是AMD的Strix Halo平台。strix-halo-guide项目：https://github.com/hogeheer499-commits/strix-halo-guide 在Ryzen AI Max+ 395上测得官方Q4_K_M生成速度为20.4 tok/s，提示处理速度为292 tok/s，并提供了原始CSV支持。GMKtec EVO-X2 64GB：https://www.amazon.com/dp/B0F53QXNGH?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1 是该平台的性价比入门选择，价格为$1,999，而128GB配置如BOSGAME M5：https://www.amazon.com/dp/B0H94TVN8G?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1 可以开启Q8和BF16行，以及更大型模型。我在“用于本地LLM的最佳迷你PC”中全面介绍了该平台的选择：/best-mini-pc-for-local-llm-2026/。","：https://www.amazon.com/dp/B0F53QXNGH?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1","：https://www.amazon.com/dp/B0H94TVN8G?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1","当前市场的一个警告：RAM 价格仍然偏高。一套 64GB DDR5 SODIMM：https://www.amazon.com/dp/B09S2QLBWC?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1 目前售价为 750-870 美元。如果你购买迷你 PC 用于本地 LLM 工作，目前直接购买安装好 RAM 的版本，要比之后升级更便宜，这与我二十年来购买电脑的所有直觉相反。","：https://www.amazon.com/dp/B09S2QLBWC?tag=tdi7w27r-20&linkCode=osi&th=1&psc=1","GPU 用户的扩展方式不同：社区报告显示，双 RTX 3090 配置大约 60 tok/s，而 RTX 5090 则根据运行时间在 75-140 tok/s 之间，16GB 显卡运行 IQ4 quants 并使用量化 KV 缓存。","奥拉玛（Ollama）是最短的路径。模型页面为ollama.com/library/qwen3.8：https://ollama.com/library/qwen3.8：","使用 --verbose 运行时，每个答案都会以类似这样的统计块结尾：","你需要 Ollama 0.32.12 或更高版本；模型元数据声明这是最低要求。","对于 llama.cpp，需要注意这一点。我的 Homebrew 版本的 llama.cpp 是几周前的，它直接拒绝了该文件：","混合架构需要当前内核。运行 brew update && brew upgrade llama.cpp 解决了问题，对任何基于 llama.cpp 的前端（LM Studio、Jan、koboldcpp）同样适用：如果 Qwen3.8 无法加载，请先更新运行环境，再调试其他问题。","对我来说，基准数据不如模型自我下载以来所做的事情重要：那些过去不会发生或会泄露到云 API 的不显眼的后台工作。","早间信息摘要：一个 launchd 任务会在夜间收集我的 RSS 未读项，并让 qwen3.8 将它们压缩成一个摘要，我边喝咖啡边阅读。262k 上下文意味着一周的订阅内容可以放在一个提示中。","扫描归档：纸质邮件会被扫描，模型读取每个 PDF 的文本，并按照我的 YYYY-MM-供应商-内容 类型的命名规则重命名。视觉能力意味着它可以处理 OCR 损坏的扫描件。","当论坛帖子评论达到 400 条时，会将其粘贴进来并进行摘要，同时标注每个观点来源。这篇文章的研究生成了几个这样的摘要，感觉很令人愉快的循环。","这一切都不关心每秒令牌数量。要求是模型足够聪明，不会把保险信当成外卖菜单来归档，我已经拥有硬件，并且没有任何数据离开家。这就是2026年本地模型的实际卖点，也是我在自托管革命中提出的同样论点：/self-hosting-revolution-mini-pcs/：即使是小规模的云依赖，也值得被替换。","我可以在16GB内存上运行Qwen3.8 27B吗？可以，以1-bit或2-bit量化（6.7-9.8GB文件）。2-bit是Unsloth认为可用于工具调用的最小量化。Q4质量需要32GB。","它比Gemma 4好吗？形态不同。Gemma 4的26B-A4B是一个稀疏MoE，在相同硬件上生成速度更快（我的Gemma 4指南：/run-gemma-4-mini-pc-without-gpu/ 有这些数据）。Qwen3.8 27B是密集模型，每个令牌速度较慢，目前社区认为它在编码和代理任务上明显优于Gemma 4。作为后台助手，我会选择Qwen3.8；对于硬件有限的互动聊天，Gemma 4仍然合理。","为什么在我的机器上它也比qwen3.6慢？混合注意力架构是新的，运行时内核（Ollama Metal，llama.cpp Vulkan/CUDA）尚未完全优化。可以预期随着更新差距会缩小。部分安慰：它每个回答使用的令牌远少于qwen3.6，因此完成答案的延迟比每秒令牌差距显示的要接近。","视觉功能可以本地运行吗？可以。Ollama版本自带视觉投影器（约460M参数），图像输入开箱即可使用。本地运行时的视频理解支持仍不完善。","那Qwen3.8-Flash-Next呢？在我写这篇文章时它刚发布了权重：一个180B MoE，Q4约为110GB。完全不同的硬件级别：你需要128GB级别统一内存，现在意味着一台3500美元以上的Strix Halo或一台大型Mac。如果“意外地对本地友好”的架构声称成立，那就是未来的文章主题。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"文章记录了 Qwen3.8 27B 在 Mac Studio M3 Ultra 上的本地运行测试：采用 Ollama 的 Q4_K_M 量化后，模型文件约 17GB，摘要给出的生成速度约为每秒 14 tokens。","background":"材料称该模型拥有 273 亿参数、混合注意力架构、262144 token 原生上下文窗口，并以 Apache 2.0 发布。作者连续运行十天，将其用于 RSS 摘要、扫描 PDF 整理等日常任务。","viewpoint":"Aioga 判断，这篇内容的价值主要在于提供了具体 Apple silicon 设备上的实测记录，而非只引用基准分数。混合架构与 Ollama Metal 内核的适配表现，仍可能影响最终速度。","implications":"对希望在本地部署模型的读者而言，17GB 量化体积和约 14 tokens/s 的摘要数据，可作为评估硬件与运行方式的参考；但该结果来自单台 Mac Studio，不能直接代表所有设备。","nextStep":"值得关注后续运行时对新架构的适配变化，并在相同提示、量化和硬件条件下复测速度与任务质量。若评估实际用途，也应分别验证摘要、文件整理等具体工作流。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-08-29T15:22:34.597Z","sourceHash":"99af8cbdfce838e3","review":{"approved":true,"groundedness":96,"clarity":92,"duplicationRisk":28,"blockingIssues":[],"notes":["“Aioga 判断”属于候选编辑内容中的署名或编辑判断，来源材料未明确提及 Aioga；如无额外上下文，建议确认该署名是否准确。","“约 14 tokens/s”来自来源摘要，正文摘录未直接出现该数值，但与提供的来源材料一致。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["技巧观点","Hacker News 热门（buzzing.cc 中文翻译）"],"translations":{"zh-CN":{"title":"在本地运行 Qwen3.8 27B：来自我的 Mac Studio 的实际数据","summary":"Qwen3.8 27B（27.3B 参数，混合注意力架构，262，144 token 上下文窗口，Apache 2.0 开源）在 Mac Studio M3 Ultra 上经 Ollama 以 Q4_K_M 量化（17GB）生成速度约 14 tokens/s。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmte242e701jdrog2vz9p2cy4","category":"行业动态","source":"terminalbytes.com","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"在本地运行 Qwen3.8 27B：来自我的 Mac Studio 的实际数据 - Aioga AI资讯","description":"Qwen3.8 27B（27.3B 参数，混合注意力架构，262，144 token 上下文窗口，Apache 2.0 开源）在 Mac Studio M3 Ultra 上经 Ollama 以 Q4_K_M 量化（17GB）生成速度约 14 tokens/s。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/item...","url":"https://www.aioga.com/news/cmte242e701jdrog2vz9p2cy4/"},"en":{"title":"Running Qwen3.8 27B locally: Actual data from my Mac Studio","summary":"Qwen3.8 27B (27.3B parameters, mixed attention architecture, 262,144 token context window, Apache 2.0 open source) generates at a speed of about 14 tokens/s on Mac Studio M3 Ultra via Ollama with Q4_K_M quantization (17GB).","category":"Insights","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"Running Qwen3.8 27B locally: Actual data from my Mac Studio - Aioga AI News","description":"Qwen3.8 27B (27.3B parameters, mixed attention architecture, 262,144 token context window, Apache 2.0 open source) generates at a speed of about 14 tokens/s on Mac Studio M3 Ultra...","url":"https://www.aioga.com/en/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:20:02.822Z"},"ja":{"title":"ローカルで Qwen3.8 27B を実行：私の Mac Studio からの実際のデータ","summary":"Qwen3.8 27B（27.3B パラメータ、ハイブリッドアテンションアーキテクチャ、262,144 トークンコンテキストウィンドウ、Apache 2.0 オープンソース）は、Mac Studio M3 Ultra 上で Ollama によって Q4_K_M 量子化（17GB）され、生成速度は約 14 トークン/秒です。","category":"ヒントと視点","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"ローカルで Qwen3.8 27B を実行：私の Mac Studio からの実際のデータ - Aioga AIニュース","description":"Qwen3.8 27B（27.3B パラメータ、ハイブリッドアテンションアーキテクチャ、262,144 トークンコンテキストウィンドウ、Apache 2.0 オープンソース）は、Mac Studio M3 Ultra 上で Ollama によって Q4_K_M 量子化（17GB）され、生成速度は約 14 トークン/秒です。","url":"https://www.aioga.com/ja/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:20:18.987Z"},"ko":{"title":"로컬에서 Qwen3.8 27B 실행: 내 Mac Studio의 실제 데이터","summary":"Qwen3.8 27B(27.3B 파라미터, 혼합 주의 아키텍처, 262,144 토큰 컨텍스트 윈도우, Apache 2.0 오픈소스)는 Mac Studio M3 Ultra에서 Ollama를 통해 Q4_K_M 양자화(17GB)로 생성 속도가 약 14 tokens/s입니다.","category":"인사이트","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"로컬에서 Qwen3.8 27B 실행: 내 Mac Studio의 실제 데이터 - Aioga AI 뉴스","description":"Qwen3.8 27B(27.3B 파라미터, 혼합 주의 아키텍처, 262,144 토큰 컨텍스트 윈도우, Apache 2.0 오픈소스)는 Mac Studio M3 Ultra에서 Ollama를 통해 Q4_K_M 양자화(17GB)로 생성 속도가 약 14 tokens/s입니다.","url":"https://www.aioga.com/ko/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:21:10.403Z"},"es":{"title":"Ejecutando Qwen3.8 27B localmente: datos reales de mi Mac Studio","summary":"Qwen3.8 27B (27,3B parámetros, arquitectura de atención mixta, ventana de contexto de 262.144 tokens, código abierto Apache 2.0) en Mac Studio M3 Ultra a través de Ollama con cuantificación Q4_K_M (17GB) genera a una velocidad de aproximadamente 14 tokens/s.","category":"Ideas","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"Ejecutando Qwen3.8 27B localmente: datos reales de mi Mac Studio - Aioga Noticias de IA","description":"Qwen3.8 27B (27,3B parámetros, arquitectura de atención mixta, ventana de contexto de 262.144 tokens, código abierto Apache 2.0) en Mac Studio M3 Ultra a través de Ollama con cuant...","url":"https://www.aioga.com/es/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:21:07.428Z"},"fr":{"title":"Exécution locale de Qwen3.8 27B : données réelles de mon Mac Studio","summary":"Qwen3.8 27B (27,3 milliards de paramètres, architecture d'attention mixte, fenêtre de contexte de 262 144 tokens, open source Apache 2.0) génère à une vitesse d'environ 14 tokens/s sur Mac Studio M3 Ultra via Ollama avec quantification Q4_K_M (17 GB).","category":"Analyses","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"Exécution locale de Qwen3.8 27B : données réelles de mon Mac Studio - Aioga Actualités IA","description":"Qwen3.8 27B (27,3 milliards de paramètres, architecture d'attention mixte, fenêtre de contexte de 262 144 tokens, open source Apache 2.0) génère à une vitesse d'environ 14 tokens/s...","url":"https://www.aioga.com/fr/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:22:01.918Z"},"de":{"title":"Qwen3.8 27B lokal ausführen: Echtzeitdaten von meinem Mac Studio","summary":"Qwen3.8 27B (27,3B Parameter, Hybrid-Attention-Architektur, 262.144 Token Kontextfenster, Open Source unter Apache 2.0) wird auf dem Mac Studio M3 Ultra über Ollama mit Q4_K_M-Quantisierung (17 GB) mit einer Geschwindigkeit von ca. 14 Tokens/s generiert.","category":"技巧观点","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"Qwen3.8 27B lokal ausführen: Echtzeitdaten von meinem Mac Studio - Aioga KI-News","description":"Qwen3.8 27B (27,3B Parameter, Hybrid-Attention-Architektur, 262.144 Token Kontextfenster, Open Source unter Apache 2.0) wird auf dem Mac Studio M3 Ultra über Ollama mit Q4_K_M-Quan...","url":"https://www.aioga.com/de/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:21:59.315Z"},"pt-BR":{"title":"Executando Qwen3.8 27B localmente: dados reais do meu Mac Studio","summary":"Qwen3.8 27B (27,3B parâmetros, arquitetura de atenção mista, janela de contexto de 262.144 tokens, código aberto Apache 2.0) gera cerca de 14 tokens/s no Mac Studio M3 Ultra via Ollama com quantização Q4_K_M (17GB).","category":"技巧观点","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"Executando Qwen3.8 27B localmente: dados reais do meu Mac Studio - Aioga Notícias de IA","description":"Qwen3.8 27B (27,3B parâmetros, arquitetura de atenção mista, janela de contexto de 262.144 tokens, código aberto Apache 2.0) gera cerca de 14 tokens/s no Mac Studio M3 Ultra via Ol...","url":"https://www.aioga.com/pt-BR/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:22:52.867Z"},"ru":{"title":"Запуск Qwen3.8 27B локально: реальные данные с моего Mac Studio","summary":"Qwen3.8 27B (27,3B параметров, гибридная архитектура внимания, окно контекста 262 144 токена, открытый исходный код Apache 2.0) на Mac Studio M3 Ultra через Ollama с квантизацией Q4_K_M (17GB) скорость генерации примерно 14 токенов/с.","category":"技巧观点","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"Запуск Qwen3.8 27B локально: реальные данные с моего Mac Studio - Aioga Новости ИИ","description":"Qwen3.8 27B (27,3B параметров, гибридная архитектура внимания, окно контекста 262 144 токена, открытый исходный код Apache 2.0) на Mac Studio M3 Ultra через Ollama с квантизацией Q...","url":"https://www.aioga.com/ru/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:22:52.595Z"},"ar":{"title":"تشغيل Qwen3.8 27B محلياً: بيانات فعلية من جهاز Mac Studio الخاص بي","summary":"Qwen3.8 27B (27.3B معلمات، هيكل الانتباه المختلط، نافذة سياق 262,144 توكن، مفتوح المصدر برخصة Apache 2.0) على Mac Studio M3 Ultra من خلال Ollama باستخدام كمية Q4_K_M (17GB) سرعة التوليد حوالي 14 توكن/ثانية.","category":"技巧观点","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"تشغيل Qwen3.8 27B محلياً: بيانات فعلية من جهاز Mac Studio الخاص بي - Aioga أخبار الذكاء الاصطناعي","description":"Qwen3.8 27B (27.3B معلمات، هيكل الانتباه المختلط، نافذة سياق 262,144 توكن، مفتوح المصدر برخصة Apache 2.0) على Mac Studio M3 Ultra من خلال Ollama باستخدام كمية Q4_K_M (17GB) سرعة ال...","url":"https://www.aioga.com/ar/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:23:47.873Z"},"hi":{"title":"स्थानीय रूप से Qwen3.8 27B चलाना: मेरे Mac Studio से वास्तविक डेटा","summary":"Qwen3.8 27B（27.3B पैरामीटर्स, मिश्रित ध्यान संरचना, 262,144 टोकन संदर्भ विंडो, Apache 2.0 ओपन सोर्स）Mac Studio M3 Ultra पर Ollama के माध्यम से Q4_K_M क्वांटाइज़ेशन (17GB) में लगभग 14 टोकन/सेकंड की गति से उत्पन्न करता है।","category":"技巧观点","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"स्थानीय रूप से Qwen3.8 27B चलाना: मेरे Mac Studio से वास्तविक डेटा - Aioga AI समाचार","description":"Qwen3.8 27B（27.3B पैरामीटर्स, मिश्रित ध्यान संरचना, 262,144 टोकन संदर्भ विंडो, Apache 2.0 ओपन सोर्स）Mac Studio M3 Ultra पर Ollama के माध्यम से Q4_K_M क्वांटाइज़ेशन (17GB) में लगभग...","url":"https://www.aioga.com/hi/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:23:47.713Z"},"it":{"title":"Eseguire Qwen3.8 27B in locale: dati reali dal mio Mac Studio","summary":"Qwen3.8 27B (27,3B parametri, architettura a attenzione mista, finestra di contesto di 262.144 token, open source Apache 2.0) su Mac Studio M3 Ultra tramite Ollama con quantizzazione Q4_K_M (17GB) genera a una velocità di circa 14 token/s.","category":"技巧观点","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"Eseguire Qwen3.8 27B in locale: dati reali dal mio Mac Studio - Aioga Notizie IA","description":"Qwen3.8 27B (27,3B parametri, architettura a attenzione mista, finestra di contesto di 262.144 token, open source Apache 2.0) su Mac Studio M3 Ultra tramite Ollama con quantizzazio...","url":"https://www.aioga.com/it/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:24:42.180Z"},"nl":{"title":"Qwen3.8 27B lokaal draaien: echte gegevens van mijn Mac Studio","summary":"Qwen3.8 27B (27,3B parameters, hybride aandachtarchitectuur, contextvenster van 262.144 tokens, Apache 2.0 open source) genereert op Mac Studio M3 Ultra via Ollama met Q4_K_M-kwantisering (17GB) met een snelheid van ongeveer 14 tokens/s.","category":"技巧观点","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"Qwen3.8 27B lokaal draaien: echte gegevens van mijn Mac Studio - Aioga AI-nieuws","description":"Qwen3.8 27B (27,3B parameters, hybride aandachtarchitectuur, contextvenster van 262.144 tokens, Apache 2.0 open source) genereert op Mac Studio M3 Ultra via Ollama met Q4_K_M-kwant...","url":"https://www.aioga.com/nl/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:24:34.064Z"},"tr":{"title":"Qwen3.8 27B'yi Yerel Olarak Çalıştırmak: Mac Studio'mdan Gerçek Veriler","summary":"Qwen3.8 27B (27,3B parametre, hibrit dikkat mimarisi, 262.144 token bağlam penceresi, Apache 2.0 açık kaynak) Mac Studio M3 Ultra üzerinde Ollama ile Q4_K_M kübikleme (17GB) ile yaklaşık 14 token/s hızında üretiliyor.","category":"技巧观点","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"Qwen3.8 27B'yi Yerel Olarak Çalıştırmak: Mac Studio'mdan Gerçek Veriler - Aioga AI Haberleri","description":"Qwen3.8 27B (27,3B parametre, hibrit dikkat mimarisi, 262.144 token bağlam penceresi, Apache 2.0 açık kaynak) Mac Studio M3 Ultra üzerinde Ollama ile Q4_K_M kübikleme (17GB) ile ya...","url":"https://www.aioga.com/tr/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:25:34.988Z"},"vi":{"title":"Chạy Qwen3.8 27B cục bộ: Dữ liệu thực tế từ Mac Studio của tôi","summary":"Qwen3.8 27B (27,3 tỷ tham số, kiến trúc chú ý hỗn hợp, cửa sổ ngữ cảnh 262.144 token, mã nguồn mở Apache 2.0) trên Mac Studio M3 Ultra thông qua Ollama với lượng tử hóa Q4_K_M (17GB) có tốc độ sinh khoảng 14 token/s.","category":"技巧观点","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"Chạy Qwen3.8 27B cục bộ: Dữ liệu thực tế từ Mac Studio của tôi - Tin tức AI Aioga","description":"Qwen3.8 27B (27,3 tỷ tham số, kiến trúc chú ý hỗn hợp, cửa sổ ngữ cảnh 262.144 token, mã nguồn mở Apache 2.0) trên Mac Studio M3 Ultra thông qua Ollama với lượng tử hóa Q4_K_M (17G...","url":"https://www.aioga.com/vi/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:25:32.245Z"},"id":{"title":"Menjalankan Qwen3.8 27B secara lokal: Data nyata dari Mac Studio saya","summary":"Qwen3.8 27B (27,3B parameter, arsitektur perhatian campuran, jendela konteks 262.144 token, open source Apache 2.0) di Mac Studio M3 Ultra melalui Ollama dengan kuantisasi Q4_K_M (17GB) kecepatan generasinya sekitar 14 token/detik.","category":"技巧观点","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"Menjalankan Qwen3.8 27B secara lokal: Data nyata dari Mac Studio saya - Berita AI Aioga","description":"Qwen3.8 27B (27,3B parameter, arsitektur perhatian campuran, jendela konteks 262.144 token, open source Apache 2.0) di Mac Studio M3 Ultra melalui Ollama dengan kuantisasi Q4_K_M (...","url":"https://www.aioga.com/id/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:26:28.703Z"},"th":{"title":"เรียกใช้ Qwen3.8 27B ในเครื่อง: ข้อมูลจริงจาก Mac Studio ของฉัน","summary":"Qwen3.8 27B (27.3B พารามิเตอร์ สถาปัตยกรรมความสนใจแบบผสม หน้าต่างบริบท 262,144 token เปิดซอร์ส Apache 2.0) บน Mac Studio M3 Ultra ผ่าน Ollama โดยการทำควอนไทเซชัน Q4_K_M (17GB) มีความเร็วในการสร้างประมาณ 14 tokens/s","category":"技巧观点","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"เรียกใช้ Qwen3.8 27B ในเครื่อง: ข้อมูลจริงจาก Mac Studio ของฉัน - ข่าว AI Aioga","description":"Qwen3.8 27B (27.3B พารามิเตอร์ สถาปัตยกรรมความสนใจแบบผสม หน้าต่างบริบท 262,144 token เปิดซอร์ส Apache 2.0) บน Mac Studio M3 Ultra ผ่าน Ollama โดยการทำควอนไทเซชัน Q4_K_M (17GB) มีคว...","url":"https://www.aioga.com/th/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:26:32.069Z"},"pl":{"title":"Uruchamianie Qwen3.8 27B lokalnie: rzeczywiste dane z mojego Mac Studio","summary":"Qwen3.8 27B (27,3B parametrów, hybrydowa architektura uwagi, okno kontekstu 262 144 tokenów, open source na licencji Apache 2.0) w Mac Studio M3 Ultra za pośrednictwem Ollama w kwantyzacji Q4_K_M (17GB) generuje z prędkością około 14 tokenów/s.","category":"技巧观点","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"Uruchamianie Qwen3.8 27B lokalnie: rzeczywiste dane z mojego Mac Studio - Aioga Wiadomości AI","description":"Qwen3.8 27B (27,3B parametrów, hybrydowa architektura uwagi, okno kontekstu 262 144 tokenów, open source na licencji Apache 2.0) w Mac Studio M3 Ultra za pośrednictwem Ollama w kwa...","url":"https://www.aioga.com/pl/news/cmte242e701jdrog2vz9p2cy4/","contentTranslated":true,"sourceHash":"97c438a5747bdaee","translatedAt":"2026-08-29T17:27:32.129Z"}}}}