{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-09-21T02:00:36.969Z","headline":"Perplexity 开源 Rust + Metal 推理引擎 Lily，专为 Apple 硅上的 Qwen3.6-35B-A3B 打造","description":"Perplexity 开源了支撑其 Hybrid Compute 的本地推理引擎 Lily，用 Rust 驱动生成循环、手写 Metal kernel 执行模型，不依赖 PyTorch 或 MLX，仅针对 Qwen3.6-35B-A3B 这一个模型。","url":"https://www.aioga.com/news/cmtl6fytz0jl6roal8donbcuo/","mainEntityOfPage":"https://www.aioga.com/news/cmtl6fytz0jl6roal8donbcuo/","datePublished":"2026-09-03T06:57:07.000Z","dateModified":"2026-09-03T06:57:07.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.marktechpost.com/2026/09/02/perplexity-open-sources-lily-a-rust-metal-inference-engine-for-qwen3-6-35b-a3b-on-apple-silicon","https://aihot.virxact.com/items/cmtl6fytz0jl6roal8donbcuo"],"canonicalUrl":"https://www.aioga.com/news/cmtl6fytz0jl6roal8donbcuo/","directAnswer":{"@type":"Answer","text":"Perplexity 开源了本地推理引擎 Lily。该单进程运行时由 Rust 负责加载检查点和生成循环，手写 Metal kernel 执行模型，不使用 PyTorch 或 MLX，仅面向 Apple silicon 上的 Qwen3.6-35B-A3B。","url":"https://www.aioga.com/news/cmtl6fytz0jl6roal8donbcuo/","dateCreated":"2026-09-03T06:57:07.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"marktechpost.com source article","url":"https://www.marktechpost.com/2026/09/02/perplexity-open-sources-lily-a-rust-metal-inference-engine-for-qwen3-6-35b-a3b-on-apple-silicon","datePublished":"2026-09-03T06:57:07.000Z","provider":{"@type":"Organization","name":"marktechpost.com","url":"https://www.marktechpost.com/2026/09/02/perplexity-open-sources-lily-a-rust-metal-inference-engine-for-qwen3-6-35b-a3b-on-apple-silicon"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmtl6fytz0jl6roal8donbcuo","datePublished":"2026-09-03T06:57:07.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmtl6fytz0jl6roal8donbcuo"}}],"aggregationSource":"MarkTechPost（RSS）","originalPublisher":{"name":"marktechpost.com","url":"https://www.marktechpost.com/2026/09/02/perplexity-open-sources-lily-a-rust-metal-inference-engine-for-qwen3-6-35b-a3b-on-apple-silicon"},"geoDeepAnswer":null,"article":{"id":"cmtl6fytz0jl6roal8donbcuo","slug":"cmtl6fytz0jl6roal8donbcuo","url":"https://www.aioga.com/news/cmtl6fytz0jl6roal8donbcuo/","title":"Perplexity 开源 Rust + Metal 推理引擎 Lily，专为 Apple 硅上的 Qwen3.6-35B-A3B 打造","title_en":"","summary":"Perplexity 开源了支撑其 Hybrid Compute 的本地推理引擎 Lily，用 Rust 驱动生成循环、手写 Metal kernel 执行模型，不依赖 PyTorch 或 MLX，仅针对 Qwen3.6-35B-A3B 这一个模型。","source":"MarkTechPost（RSS）","sourceUrl":"https://www.marktechpost.com/2026/09/02/perplexity-open-sources-lily-a-rust-metal-inference-engine-for-qwen3-6-35b-a3b-on-apple-silicon","aiHotUrl":"https://aihot.virxact.com/items/cmtl6fytz0jl6roal8donbcuo","publishedAt":"2026-09-03T06:57:07.000Z","category":"行业动态","score":58,"selected":false,"articleBody":["Perplexity has open sourced Lily ,：https://github.com/perplexityai/pplx-garden/tree/main/lily the local inference engine behind Hybrid Compute in Perplexity Computer：https://www.perplexity.ai/hub/products/hybrid-compute. It is a single-process runtime: a Rust layer loads the checkpoint and drives the generation loop, an OpenAI-compatible chat-completions API streams tokens, and hand-written Metal kernels execute the model. Neither PyTorch nor MLX sits in the execution path. Lily is deliberately narrow with one model, Qwen3.6-35B-A3B：https://huggingface.co/Qwen/Qwen3.6-35B-A3B, on one hardware family and that narrowness is the performance argument.","Is it deployable? Yes. A standalone demo is public in the pplx-garden repository：https://github.com/perplexityai/pplx-garden/tree/main/lily. A Rust and Metal inference server offering greedy text generation through a minimal OpenAI-compatible HTTP API. The 4-bit checkpoint is 19.4 GB, so an Apple silicon Mac with 32 GB or more of unified memory is the realistic floor; Perplexity’s shipping Hybrid Compute product lists macOS 15+, 24 GB minimum and 32 GB for best results.","The default Mac stack is MLX：https://github.com/ml-explore/mlx plus MLX-LM：https://github.com/ml-explore/mlx-lm, which already ships a Qwen implementation：https://github.com/ml-explore/mlx-lm/blob/main/mlx_lm/models/qwen3_5.py with grouped expert work, a fused recurrent Metal kernel, and GQA-aware attention. But its operations must stay reusable across architectures. Lily gives that up and puts model structure, execution plans, and kernel selection in one runtime.","Qwen3.6-35B-A3B stores 35B parameters and activates roughly 3B per token. A router scores 256 experts and picks eight, alongside one shared expert that sees every token. It also mixes 10 full-attention layers using grouped-query attention：https://arxiv.org/pdf/2305.13245 (16 query heads, two KV heads) with 30 Gated DeltaNet：https://arxiv.org/abs/2412.06464 layers. That yields three patterns: uneven expert groups, attention over a growing KV cache, and a fixed-size recurrence.","The checkpoint uses groupwise affine 4-bit quantization, every group of 64 weights sharing a bfloat16 scale and bias, about 70 GB of bfloat16 weights compressed to 19.4 GB. Metal 4 tensor operations consume bfloat16, so weights must be reconstructed first. Lily does that one tile at a time inside the grouped GEMM, holding results in threadgroup memory and accumulating in FP32, so the expanded array never reaches unified memory. In Perplexity’s ablation that fusion raised end-to-end prefill 77.4% at a 512-token prompt.","Keeping the routing histogram, prefix scan, scatter and block map inside a single GPU command buffer added 89% at 512 tokens by removing CPU synchronization inside each MoE layer. Moving from 16-row to 32-row tiles with four simdgroups added 13.2% at 2K; a register-resident Gated DeltaNet scan added 5.6% . Expert GEMMs are roughly 90% of prefill time. Long prompts run in bounded chunks so temporary activations do not compete with weights and cache for memory.","Batch-1 decode has almost no weight reuse, so bandwidth sets the ceiling. One recorded step launched 795 kernels forming 555 sequential stages; Lily records real dependencies in a concurrent Metal pass so independent kernels overlap. The selected token is written straight into the next step’s GPU-resident input slot, removing a per-token CPU round trip, and four kernel chains are fused to keep intermediates in registers.","Coalesced cache reads lifted key bandwidth from 33.8 to 47.9 GB/s and value bandwidth from 42.0 to 61.8 GB/s. GQA packing, four query heads sharing one threadgroup so each KV row loads once, improved decode 23.8% at 32K. A fixed-block attention layout at 32K and above improved decode 7.7% at 32K, 27.4% at 64K, and 40.2% at 128K.","On one 40-core, 128 GB M5 Max at batch 1, loading identical 4-bit checkpoint bytes against MLX-LM’s fastest direct-generation path across ten lengths from 256 to 128K tokens, Lily averaged 4,156 prefill tokens/s versus 3,388 (1.23x) and 170.0 decode tokens/s versus 126.4 (1.35x) . At a 4K prompt and 4K context it reached 5,749.9 and 186.6 tokens/s against 4,737.5 and 140.9, and was faster at every recorded point: 1.12–1.42x prefill, 1.31–1.37x decode. A teacher-forced check across 192 positions put Lily’s perplexity 0.04% higher, with the same top-ranked token 96.35% of the time.","Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us ：https://forms.gle/wbash1wF6efRj8G58","Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences."],"articleImages":[{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2019/06/Screen-Shot-2021-09-14-at-9.02.24-AM-300x300.png","alt":"","afterParagraph":9,"url":"/media/articles/cmtl6fytz0jl6roal8donbcuo/787a6d54564e8e19.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/09/blog2222-6-100x70.png","alt":"Qwen Developers Open-Sources zg (zvec-grep): A Local-First Search Layer Unifying ripgrep, BM25, and Vector Search","afterParagraph":10,"url":"/media/articles/cmtl6fytz0jl6roal8donbcuo/3ba0fb4c7f0051a8.png"}],"mediaStatus":"ok","articleBodyZh":["Perplexity 已开源 Lily：https://github.com/perplexityai/pplx-garden/tree/main/lily，这是 Perplexity Computer 中 Hybrid Compute 背后的本地推理引擎：https://www.perplexity.ai/hub/products/hybrid-compute。它是一个单进程运行时：Rust 层加载检查点并驱动生成循环，一个兼容 OpenAI 的聊天完成 API 流式传输令牌，手写的 Metal 内核执行模型。执行路径中没有 PyTorch 或 MLX。Lily 是有意设计得很窄，仅支持一个模型 Qwen3.6-35B-A3B：https://huggingface.co/Qwen/Qwen3.6-35B-A3B，在一个硬件家族上运行，这种窄化正是其性能优势所在。","它可以部署吗？可以。pplx-garden 仓库中有一个独立演示公开：https://github.com/perplexityai/pplx-garden/tree/main/lily。一个使用 Rust 和 Metal 的推理服务器，通过最小化的兼容 OpenAI 的 HTTP API 提供贪婪文本生成。4-bit 检查点大小为 19.4 GB，因此一台具有 32 GB 或更多统一内存的苹果硅 Mac 是实际的最低要求；Perplexity 的 Hybrid Compute 商用产品列出了 macOS 15+，最低 24 GB，最佳效果为 32 GB。","默认的 Mac 堆栈是 MLX：https://github.com/ml-explore/mlx 以及 MLX-LM：https://github.com/ml-explore/mlx-lm，它已经提供 Qwen 的实现：https://github.com/ml-explore/mlx-lm/blob/main/mlx_lm/models/qwen3_5.py，包含分组专家工作、融合的循环 Metal 内核，以及 GQA 感知注意力。但其操作必须在各架构间保持可复用。Lily 放弃了这一点，将模型结构、执行计划和内核选择集中在一个运行时中。","Qwen3.6-35B-A3B 存储 350 亿个参数，每个令牌激活约 30 亿个。一个路由器对 256 个专家进行评分并选择八个，同时还有一个共享专家处理每个令牌。它还混合了 10 层全注意力层，使用分组查询注意力：https://arxiv.org/pdf/2305.13245（16 个查询头，两个 KV 头）与 30 层门控 DeltaNet：https://arxiv.org/abs/2412.06464。这产生三种模式：不均衡的专家组、对增长的 KV 缓存的注意力，以及固定大小的循环。","检查点使用分组仿射4位量化，每组64个权重共享bfloat16的尺度和偏置，约70 GB的bfloat16权重压缩为19.4 GB。金属4张量运算会消耗bfloat16，因此权重必须先重建。Lily在分组后的GEMM中逐格进行，按住会进入线程组内存，累积到FP32，因此扩展后的数组永远不会到达统一内存。在Perplexity的消融中，该融合在512标记提示下提升了端到端预填充率77.4%。","将路由直方图、前缀扫描、散射和块状映射保持在单一 GPU 命令缓冲区内，通过移除每个 MoE 层的 CPU 同步，在 512 个令牌中增加了 89%。从 16 行瓦片到 32 行并配备四个 simdgroup 在 2K 时增加了 13.2%;寄存器驻留的门控 DeltaNet 扫描增加了 5.6%。专家 GEMM 大约占预填充时间的 90%。长提示以有界块形式运行，因此临时激活不会与权重和缓存争夺内存。","第一批译码几乎没有权重重复用，因此带宽限制了上限。一个记录的步骤启动了795个内核，形成了555个顺序阶段;Lily在并发的Metal通道中记录真实依赖，因此独立内核会重叠。所选令牌直接写入下一步的GPU驻留输入槽，省去每个令牌CPU的往返，四条内核链融合以保持寄存器中的中间节点。","合并缓存读取将密钥带宽从33.8提升至47.9 GB/s，价值带宽从42.0提升至61.8 GB/s。GQA打包、四个查询头共用一个线程组，使每行KV加载一次，在32K时解码提升了23.8%。在32K及以上采用固定块注意力布局，解码提升了7.7%，64K时提升了27.4%，128K时提升了40.2%。","在一台40核、128GB M5 Max的第1批次中，加载相同的4位检查点字节，匹配MLX-LM最快的直接生成路径，跨越10个长度，介于256至128K令牌，Lily平均预填充令牌为4,156个，对比3,388个（1.23x），译码令牌为170.0个，译码令牌为126.4个（1.35倍）。在4K提示和4K上下文下，它达到了5,749.9和186.6个令牌/秒，而4,737.5和140.9个，且在所有记录点（1.12–1.42x预填充，1.31–1.37x译码）都更快。教师强制检查了192个位置，结果发现莉莉的困惑度提高了0.04%，而同样排名第一的代币有96.35%的比例。","需要与我们合作推广您的 GitHub 仓库或 Hugging Face 页面或产品发布或网络研讨会等吗？请联系我们：https://forms.gle/wbash1wF6efRj8G58","Asif Razzaq 是 Marktechpost Media Inc. 的首席执行官。作为一位富有远见的企业家和工程师，Asif 致力于利用人工智能的潜力造福社会。他最近的工作是推出人工智能媒体平台 Marktechpost，该平台以深入报道机器学习和深度学习新闻而著称，既技术性强又易于广大读者理解。该平台每月访问量超过 200 万次，显示出其在观众中的受欢迎程度。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Perplexity 开源了本地推理引擎 Lily。该单进程运行时由 Rust 负责加载检查点和生成循环，手写 Metal kernel 执行模型，不使用 PyTorch 或 MLX，仅面向 Apple silicon 上的 Qwen3.6-35B-A3B。","background":"Lily 的公开演示提供贪心文本生成和最小化 OpenAI 兼容 HTTP API。其 4-bit 检查点为 19.4 GB，材料称 32 GB 统一内存是较现实的起点；Hybrid Compute 产品列出 macOS 15+、24 GB 最低和 32 GB 最佳配置。","viewpoint":"Aioga 判断：Lily 的主要取舍是缩小支持范围，把模型结构、执行计划与 kernel 选择集中在一个运行时中。材料中的优化围绕单一模型和 Apple silicon 展开，因此不宜据此推断其适用于其他模型或硬件平台。","implications":"可能影响：Lily 可能为本地推理的专用化优化提供参考，但其适用范围需要限定。来源列出的性能变化对应特定提示长度、上下文规模或消融实验，不代表所有 Apple silicon 设备和工作负载都能复现，也不足以证明普遍优势。","nextStep":"后续观察：建议核对公开演示的运行配置、检查点加载方式和测试条件，并关注不同统一内存配置下的实际表现。还需要观察项目后续是否继续维持单模型定位，以及公开实现能否在材料所述场景之外稳定运行。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-09-03T07:47:00.088Z","sourceHash":"789537ce03c1ca66","review":{"approved":true,"groundedness":96,"clarity":91,"duplicationRisk":8,"blockingIssues":[],"notes":["“Apple silicon 上”与来源所述单一硬件家族一致；未将性能数据外推为普遍优势。","viewpoint 和 implications 明确标注为判断或可能影响，未将观点冒充事实。","可进一步在 evidenceRefs 中补充更具体的性能段落或仓库链接，但不影响审核通过。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":1,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","MarkTechPost（RSS）"],"translations":{"zh-CN":{"title":"Perplexity 开源 Rust + Metal 推理引擎 Lily，专为 Apple 硅上的 Qwen3.6-35B-A3B 打造","summary":"Perplexity 开源了支撑其 Hybrid Compute 的本地推理引擎 Lily，用 Rust 驱动生成循环、手写 Metal kernel 执行模型，不依赖 PyTorch 或 MLX，仅针对 Qwen3.6-35B-A3B 这一个模型。","category":"行业动态","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 开源 Rust + Metal 推理引擎 Lily，专为 Apple 硅上的 Qwen3.6-35B-A3B 打造 - Aioga AI资讯","description":"Perplexity 开源了支撑其 Hybrid Compute 的本地推理引擎 Lily，用 Rust 驱动生成循环、手写 Metal kernel 执行模型，不依赖 PyTorch 或 MLX，仅针对 Qwen3.6-35B-A3B 这一个模型。","url":"https://www.aioga.com/news/cmtl6fytz0jl6roal8donbcuo/","articleBody":["Perplexity 已开源 Lily：https://github.com/perplexityai/pplx-garden/tree/main/lily，这是 Perplexity Computer 中 Hybrid Compute 背后的本地推理引擎：https://www.perplexity.ai/hub/products/hybrid-compute。它是一个单进程运行时：Rust 层加载检查点并驱动生成循环，一个兼容 OpenAI 的聊天完成 API 流式传输令牌，手写的 Metal 内核执行模型。执行路径中没有 PyTorch 或 MLX。Lily 是有意设计得很窄，仅支持一个模型 Qwen3.6-35B-A3B：https://huggingface.co/Qwen/Qwen3.6-35B-A3B，在一个硬件家族上运行，这种窄化正是其性能优势所在。","它可以部署吗？可以。pplx-garden 仓库中有一个独立演示公开：https://github.com/perplexityai/pplx-garden/tree/main/lily。一个使用 Rust 和 Metal 的推理服务器，通过最小化的兼容 OpenAI 的 HTTP API 提供贪婪文本生成。4-bit 检查点大小为 19.4 GB，因此一台具有 32 GB 或更多统一内存的苹果硅 Mac 是实际的最低要求；Perplexity 的 Hybrid Compute 商用产品列出了 macOS 15+，最低 24 GB，最佳效果为 32 GB。","默认的 Mac 堆栈是 MLX：https://github.com/ml-explore/mlx 以及 MLX-LM：https://github.com/ml-explore/mlx-lm，它已经提供 Qwen 的实现：https://github.com/ml-explore/mlx-lm/blob/main/mlx_lm/models/qwen3_5.py，包含分组专家工作、融合的循环 Metal 内核，以及 GQA 感知注意力。但其操作必须在各架构间保持可复用。Lily 放弃了这一点，将模型结构、执行计划和内核选择集中在一个运行时中。","Qwen3.6-35B-A3B 存储 350 亿个参数，每个令牌激活约 30 亿个。一个路由器对 256 个专家进行评分并选择八个，同时还有一个共享专家处理每个令牌。它还混合了 10 层全注意力层，使用分组查询注意力：https://arxiv.org/pdf/2305.13245（16 个查询头，两个 KV 头）与 30 层门控 DeltaNet：https://arxiv.org/abs/2412.06464。这产生三种模式：不均衡的专家组、对增长的 KV 缓存的注意力，以及固定大小的循环。","检查点使用分组仿射4位量化，每组64个权重共享bfloat16的尺度和偏置，约70 GB的bfloat16权重压缩为19.4 GB。金属4张量运算会消耗bfloat16，因此权重必须先重建。Lily在分组后的GEMM中逐格进行，按住会进入线程组内存，累积到FP32，因此扩展后的数组永远不会到达统一内存。在Perplexity的消融中，该融合在512标记提示下提升了端到端预填充率77.4%。","将路由直方图、前缀扫描、散射和块状映射保持在单一 GPU 命令缓冲区内，通过移除每个 MoE 层的 CPU 同步，在 512 个令牌中增加了 89%。从 16 行瓦片到 32 行并配备四个 simdgroup 在 2K 时增加了 13.2%;寄存器驻留的门控 DeltaNet 扫描增加了 5.6%。专家 GEMM 大约占预填充时间的 90%。长提示以有界块形式运行，因此临时激活不会与权重和缓存争夺内存。","第一批译码几乎没有权重重复用，因此带宽限制了上限。一个记录的步骤启动了795个内核，形成了555个顺序阶段;Lily在并发的Metal通道中记录真实依赖，因此独立内核会重叠。所选令牌直接写入下一步的GPU驻留输入槽，省去每个令牌CPU的往返，四条内核链融合以保持寄存器中的中间节点。","合并缓存读取将密钥带宽从33.8提升至47.9 GB/s，价值带宽从42.0提升至61.8 GB/s。GQA打包、四个查询头共用一个线程组，使每行KV加载一次，在32K时解码提升了23.8%。在32K及以上采用固定块注意力布局，解码提升了7.7%，64K时提升了27.4%，128K时提升了40.2%。","在一台40核、128GB M5 Max的第1批次中，加载相同的4位检查点字节，匹配MLX-LM最快的直接生成路径，跨越10个长度，介于256至128K令牌，Lily平均预填充令牌为4,156个，对比3,388个（1.23x），译码令牌为170.0个，译码令牌为126.4个（1.35倍）。在4K提示和4K上下文下，它达到了5,749.9和186.6个令牌/秒，而4,737.5和140.9个，且在所有记录点（1.12–1.42x预填充，1.31–1.37x译码）都更快。教师强制检查了192个位置，结果发现莉莉的困惑度提高了0.04%，而同样排名第一的代币有96.35%的比例。","需要与我们合作推广您的 GitHub 仓库或 Hugging Face 页面或产品发布或网络研讨会等吗？请联系我们：https://forms.gle/wbash1wF6efRj8G58","Asif Razzaq 是 Marktechpost Media Inc. 的首席执行官。作为一位富有远见的企业家和工程师，Asif 致力于利用人工智能的潜力造福社会。他最近的工作是推出人工智能媒体平台 Marktechpost，该平台以深入报道机器学习和深度学习新闻而著称，既技术性强又易于广大读者理解。该平台每月访问量超过 200 万次，显示出其在观众中的受欢迎程度。"]},"en":{"title":"Perplexity Open-Sources Rust + Metal Inference Engine Lily, Specifically Built for Qwen3.6-35B-A3B on Apple Silicon","summary":"Perplexity has open-sourced its local inference engine Lily, which supports its Hybrid Compute, generating loops driven by Rust and executing hand-written Metal kernels to run the model, without relying on PyTorch or MLX, and is designed specifically for the Qwen3.6-35B-A3B model.","category":"Industry","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity Open-Sources Rust + Metal Inference Engine Lily, Specifically Built for Qwen3.6-35B-A3B on Apple Silicon - Aioga AI News","description":"Perplexity has open-sourced its local inference engine Lily, which supports its Hybrid Compute, generating loops driven by Rust and executing hand-written Metal kernels to run the...","url":"https://www.aioga.com/en/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:40:52.387Z"},"ja":{"title":"PerplexityがRust＋Metalベースの推論エンジンLilyをオープンソース化、Apple Silicon向けのQwen3.6-35B-A3B専用","summary":"Perplexityは、ハイブリッドコンピュートを支えるローカル推論エンジンLilyをオープンソース化しました。Rustで生成ループを駆動し、Metalカーネルを手書きでモデル実行に使う方式で、PyTorchやMLXには依存せず、Qwen3.6-35B-A3Bこの1モデルのみに特化しています。","category":"業界動向","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"PerplexityがRust＋Metalベースの推論エンジンLilyをオープンソース化、Apple Silicon向けのQwen3.6-35B-A3B専用 - Aioga AIニュース","description":"Perplexityは、ハイブリッドコンピュートを支えるローカル推論エンジンLilyをオープンソース化しました。Rustで生成ループを駆動し、Metalカーネルを手書きでモデル実行に使う方式で、PyTorchやMLXには依存せず、Qwen3.6-35B-A3Bこの1モデルのみに特化しています。","url":"https://www.aioga.com/ja/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:40:53.731Z"},"ko":{"title":"Perplexity가 Rust + Metal 추론 엔진 Lily를 오픈소스로 공개, Apple 실리콘용 Qwen3.6-35B-A3B 전용","summary":"Perplexity가 하이브리드 컴퓨트(Hybrid Compute)를 지원하는 로컬 추론 엔진 Lily를 오픈소스로 공개했습니다. Rust로 루프를 생성하고, Metal 커널을 직접 작성하여 모델을 실행하며, PyTorch나 MLX에 의존하지 않고, 오직 Qwen3.6-35B-A3B 모델만을 대상으로 합니다.","category":"업계 동향","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity가 Rust + Metal 추론 엔진 Lily를 오픈소스로 공개, Apple 실리콘용 Qwen3.6-35B-A3B 전용 - Aioga AI 뉴스","description":"Perplexity가 하이브리드 컴퓨트(Hybrid Compute)를 지원하는 로컬 추론 엔진 Lily를 오픈소스로 공개했습니다. Rust로 루프를 생성하고, Metal 커널을 직접 작성하여 모델을 실행하며, PyTorch나 MLX에 의존하지 않고, 오직 Qwen3.6-35B-A3B 모델만을 대상으로 합니다.","url":"https://www.aioga.com/ko/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:40:57.245Z"},"es":{"title":"Perplexity lanza el motor de inferencia de código abierto Lily con Rust + Metal, creado específicamente para Qwen3.6-35B-A3B en Apple Silicon","summary":"Perplexity ha abierto el motor de inferencia local Lily que soporta su Hybrid Compute, usando Rust para impulsar bucles generativos y kernels Metal escritos a mano para ejecutar modelos, sin depender de PyTorch o MLX, y diseñado únicamente para el modelo Qwen3.6-35B-A3B.","category":"Industria","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity lanza el motor de inferencia de código abierto Lily con Rust + Metal, creado específicamente para Qwen3.6-35B-A3B en Apple Silicon - Aioga Noticias de IA","description":"Perplexity ha abierto el motor de inferencia local Lily que soporta su Hybrid Compute, usando Rust para impulsar bucles generativos y kernels Metal escritos a mano para ejecutar mo...","url":"https://www.aioga.com/es/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:40:56.837Z"},"fr":{"title":"Perplexity ouvre le moteur d'inférence Rust + Metal Lily, spécialement conçu pour Qwen3.6-35B-A3B sur les puces Apple","summary":"Perplexity a publié le moteur d'inférence local Lily soutenant son Hybrid Compute, utilisant Rust pour générer des boucles et exécuter des kernels Metal écrits à la main pour le modèle, sans dépendre de PyTorch ou MLX, uniquement pour le modèle Qwen3.6-35B-A3B.","category":"Industrie","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity ouvre le moteur d'inférence Rust + Metal Lily, spécialement conçu pour Qwen3.6-35B-A3B sur les puces Apple - Aioga Actualités IA","description":"Perplexity a publié le moteur d'inférence local Lily soutenant son Hybrid Compute, utilisant Rust pour générer des boucles et exécuter des kernels Metal écrits à la main pour le mo...","url":"https://www.aioga.com/fr/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:41:01.318Z"},"de":{"title":"Perplexity veröffentlicht den Open-Source Rust + Metal Inferenz-Engine Lily, speziell für Qwen3.6-35B-A3B auf Apple Silicon entwickelt","summary":"Perplexity hat den lokalen Inferenz-Engine Lily veröffentlicht, der seine Hybrid-Compute unterstützt. Er wird mit Rust betrieben, führt Schleifen aus und führt handgeschriebene Metal-Kernels aus, um Modelle auszuführen. Er ist nicht von PyTorch oder MLX abhängig und ist ausschließlich für das Modell Qwen3.6-35B-A3B vorgesehen.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity veröffentlicht den Open-Source Rust + Metal Inferenz-Engine Lily, speziell für Qwen3.6-35B-A3B auf Apple Silicon entwickelt - Aioga KI-News","description":"Perplexity hat den lokalen Inferenz-Engine Lily veröffentlicht, der seine Hybrid-Compute unterstützt. Er wird mit Rust betrieben, führt Schleifen aus und führt handgeschriebene Met...","url":"https://www.aioga.com/de/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:41:01.088Z"},"pt-BR":{"title":"Perplexity lança engine de inferência Lily de código aberto em Rust + Metal, desenvolvido especialmente para Qwen3.6-35B-A3B em Apple Silicon","summary":"A Perplexity lançou o engine de inferência local Lily, que suporta seu Hybrid Compute, usando Rust para gerar loops e executar o modelo com kernels Metal escritos à mão, sem depender de PyTorch ou MLX, apenas para o modelo Qwen3.6-35B-A3B.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity lança engine de inferência Lily de código aberto em Rust + Metal, desenvolvido especialmente para Qwen3.6-35B-A3B em Apple Silicon - Aioga Notícias de IA","description":"A Perplexity lançou o engine de inferência local Lily, que suporta seu Hybrid Compute, usando Rust para gerar loops e executar o modelo com kernels Metal escritos à mão, sem depend...","url":"https://www.aioga.com/pt-BR/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:41:04.263Z"},"ru":{"title":"Perplexity открыла исходный код Rust + Metal движка вывода Lily, созданного специально для Qwen3.6-35B-A3B на чипах Apple","summary":"Perplexity открыла исходный код локального движка вывода Lily, поддерживающего их гибридные вычисления. Он использует Rust для генерации циклов и выполнения моделей через написанные вручную Metal kernel, не зависит от PyTorch или MLX и предназначен исключительно для модели Qwen3.6-35B-A3B.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity открыла исходный код Rust + Metal движка вывода Lily, созданного специально для Qwen3.6-35B-A3B на чипах Apple - Aioga Новости ИИ","description":"Perplexity открыла исходный код локального движка вывода Lily, поддерживающего их гибридные вычисления. Он использует Rust для генерации циклов и выполнения моделей через написанны...","url":"https://www.aioga.com/ru/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:41:04.440Z"},"ar":{"title":"Perplexity تطلق محرك الاستدلال المفتوح المصدر Lily باستخدام Rust وMetal، مصمّم خصيصاً ل Qwen3.6-35B-A3B على معالجات Apple Silicon","summary":"Perplexity أطلقت محرك الاستدلال المحلي Lily الذي يدعم حوسبتها الهجينة، باستخدام Rust لتشغيل التوليد التكراري وكتابة نواة Metal يدويًا لتنفيذ النموذج، دون الاعتماد على PyTorch أو MLX، ومخصص فقط لهذا النموذج Qwen3.6-35B-A3B.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity تطلق محرك الاستدلال المفتوح المصدر Lily باستخدام Rust وMetal، مصمّم خصيصاً ل Qwen3.6-35B-A3B على معالجات Apple Silicon - Aioga أخبار الذكاء الاصطناعي","description":"Perplexity أطلقت محرك الاستدلال المحلي Lily الذي يدعم حوسبتها الهجينة، باستخدام Rust لتشغيل التوليد التكراري وكتابة نواة Metal يدويًا لتنفيذ النموذج، دون الاعتماد على PyTorch أو ML...","url":"https://www.aioga.com/ar/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:41:07.306Z"},"hi":{"title":"Perplexity ने Rust + Metal इनफेरेंस इंजन Lily को ओपनसोर्स किया, जो विशेष रूप से Apple सिलिकॉन पर Qwen3.6-35B-A3B के लिए बनाया गया है","summary":"Perplexity ने अपना लोकल इनफेरेंस इंजन Lily ओपनसोर्स किया है, जो उनके हाइब्रिड कंप्यूट का समर्थन करता है, Rust का उपयोग करके वे मॉडल के लिए लूप उत्पन्न करता है और मेटल कर्नेल को हाथ से निष्पादित करता है, PyTorch या MLX पर निर्भर नहीं है, और केवल Qwen3.6-35B-A3B मॉडल के लिए ही बनाया गया है।","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity ने Rust + Metal इनफेरेंस इंजन Lily को ओपनसोर्स किया, जो विशेष रूप से Apple सिलिकॉन पर Qwen3.6-35B-A3B के लिए बनाया गया है - Aioga AI समाचार","description":"Perplexity ने अपना लोकल इनफेरेंस इंजन Lily ओपनसोर्स किया है, जो उनके हाइब्रिड कंप्यूट का समर्थन करता है, Rust का उपयोग करके वे मॉडल के लिए लूप उत्पन्न करता है और मेटल कर्नेल को हाथ...","url":"https://www.aioga.com/hi/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:41:07.828Z"},"it":{"title":"Perplexity rende open source il motore di inferenza Lily in Rust + Metal, progettato specificamente per Qwen3.6-35B-A3B su chip Apple","summary":"Perplexity ha reso open source il motore di inferenza locale Lily che supporta il suo Hybrid Compute, usando Rust per gestire cicli di generazione e kernel Metal scritti a mano per eseguire modelli, senza dipendere da PyTorch o MLX, e progettato esclusivamente per il modello Qwen3.6-35B-A3B.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity rende open source il motore di inferenza Lily in Rust + Metal, progettato specificamente per Qwen3.6-35B-A3B su chip Apple - Aioga Notizie IA","description":"Perplexity ha reso open source il motore di inferenza locale Lily che supporta il suo Hybrid Compute, usando Rust per gestire cicli di generazione e kernel Metal scritti a mano per...","url":"https://www.aioga.com/it/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:41:10.619Z"},"nl":{"title":"Perplexity heeft de open-source Rust + Metal-inferentiemotor Lily uitgebracht, speciaal ontworpen voor Qwen3.6-35B-A3B op Apple Silicon.","summary":"Perplexity heeft Lily uitgebracht, de lokale inferentiemotor die hun Hybrid Compute ondersteunt. Deze gebruikt Rust voor het aansturen van generatielussen en handgeschreven Metal-kernels om modellen uit te voeren, zonder afhankelijkheid van PyTorch of MLX, en is uitsluitend gericht op het model Qwen3.6-35B-A3B.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity heeft de open-source Rust + Metal-inferentiemotor Lily uitgebracht, speciaal ontworpen voor Qwen3.6-35B-A3B op Apple Silicon. - Aioga AI-nieuws","description":"Perplexity heeft Lily uitgebracht, de lokale inferentiemotor die hun Hybrid Compute ondersteunt. Deze gebruikt Rust voor het aansturen van generatielussen en handgeschreven Metal-k...","url":"https://www.aioga.com/nl/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:41:11.567Z"},"tr":{"title":"Perplexity, Apple çipi üzerinde Qwen3.6-35B-A3B için özel olarak tasarlanmış Rust + Metal çıkarım motoru Lily'yi açtı","summary":"Perplexity, Hybrid Compute desteği sağlayan yerel çıkarım motoru Lily'yi açtı; Rust ile döngü üretimi yapıyor ve Metal kernel kullanarak modeli çalıştırıyor, PyTorch veya MLX'e bağımlı değil, sadece Qwen3.6-35B-A3B modeli için özel olarak tasarlanmış.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity, Apple çipi üzerinde Qwen3.6-35B-A3B için özel olarak tasarlanmış Rust + Metal çıkarım motoru Lily'yi açtı - Aioga AI Haberleri","description":"Perplexity, Hybrid Compute desteği sağlayan yerel çıkarım motoru Lily'yi açtı; Rust ile döngü üretimi yapıyor ve Metal kernel kullanarak modeli çalıştırıyor, PyTorch veya MLX'e bağ...","url":"https://www.aioga.com/tr/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:41:15.710Z"},"vi":{"title":"Perplexity mã nguồn mở engine suy luận Rust + Metal Lily, được tạo riêng cho Qwen3.6-35B-A3B trên chip Apple","summary":"Perplexity đã mã nguồn mở engine suy luận cục bộ Lily, hỗ trợ Hybrid Compute, sử dụng Rust để điều khiển vòng lặp sinh và viết tay kernel Metal thực hiện mô hình, không phụ thuộc PyTorch hay MLX, chỉ dành riêng cho mô hình Qwen3.6-35B-A3B này.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity mã nguồn mở engine suy luận Rust + Metal Lily, được tạo riêng cho Qwen3.6-35B-A3B trên chip Apple - Tin tức AI Aioga","description":"Perplexity đã mã nguồn mở engine suy luận cục bộ Lily, hỗ trợ Hybrid Compute, sử dụng Rust để điều khiển vòng lặp sinh và viết tay kernel Metal thực hiện mô hình, không phụ thuộc P...","url":"https://www.aioga.com/vi/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:41:14.900Z"},"id":{"title":"Perplexity merilis Lily, mesin inferensi Rust + Metal open-source, dibuat khusus untuk Qwen3.6-35B-A3B di chip Apple","summary":"Perplexity merilis mesin inferensi lokal Lily yang mendukung Hybrid Compute mereka, menggunakan Rust untuk menggerakkan loop generasi dan menjalankan model dengan kernel Metal yang ditulis tangan, tidak bergantung pada PyTorch atau MLX, hanya ditujukan untuk model Qwen3.6-35B-A3B ini.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity merilis Lily, mesin inferensi Rust + Metal open-source, dibuat khusus untuk Qwen3.6-35B-A3B di chip Apple - Berita AI Aioga","description":"Perplexity merilis mesin inferensi lokal Lily yang mendukung Hybrid Compute mereka, menggunakan Rust untuk menggerakkan loop generasi dan menjalankan model dengan kernel Metal yang...","url":"https://www.aioga.com/id/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:41:18.379Z"},"th":{"title":"Perplexity เปิดโอเพ่นซอร์ส Rust + Metal inference engine Lily ออกมา ซึ่งสร้างขึ้นมาเฉพาะสำหรับ Qwen3.6-35B-A3B บนชิป Apple","summary":"Perplexity เปิดโอเพ่นซอร์ส inference engine Lily สำหรับรองรับ Hybrid Compute ในเครื่อง จำหน่ายโดยใช้ Rust เพื่อขับเคลื่อนการวนซ้ำการสร้างและ kernel Metal ที่เขียนด้วยมือเพื่อรันโมเดลโดยไม่พึ่งพา PyTorch หรือ MLX โดยทำขึ้นเฉพาะสำหรับโมเดล Qwen3.6-35B-A3B เท่านั้น","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity เปิดโอเพ่นซอร์ส Rust + Metal inference engine Lily ออกมา ซึ่งสร้างขึ้นมาเฉพาะสำหรับ Qwen3.6-35B-A3B บนชิป Apple - ข่าว AI Aioga","description":"Perplexity เปิดโอเพ่นซอร์ส inference engine Lily สำหรับรองรับ Hybrid Compute ในเครื่อง จำหน่ายโดยใช้ Rust เพื่อขับเคลื่อนการวนซ้ำการสร้างและ kernel Metal ที่เขียนด้วยมือเพื่อรันโมเ...","url":"https://www.aioga.com/th/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:41:19.342Z"},"pl":{"title":"Perplexity udostępnia open source silnik inferencyjny Lily w Rust + Metal, stworzony specjalnie dla Qwen3.6-35B-A3B na procesorach Apple Silicon","summary":"Perplexity udostępniło silnik inferencyjny Lily, wspierający ich Hybrydowe Obliczenia, napisany w Rust i wykorzystujący pętle generacyjne oraz ręcznie pisane jądra Metal do wykonywania modelu, bez zależności od PyTorch czy MLX, stworzony wyłącznie dla modelu Qwen3.6-35B-A3B.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity udostępnia open source silnik inferencyjny Lily w Rust + Metal, stworzony specjalnie dla Qwen3.6-35B-A3B na procesorach Apple Silicon - Aioga Wiadomości AI","description":"Perplexity udostępniło silnik inferencyjny Lily, wspierający ich Hybrydowe Obliczenia, napisany w Rust i wykorzystujący pętle generacyjne oraz ręcznie pisane jądra Metal do wykonyw...","url":"https://www.aioga.com/pl/news/cmtl6fytz0jl6roal8donbcuo/","contentTranslated":true,"sourceHash":"56f80aa17ec2d970","translatedAt":"2026-09-03T07:41:22.777Z"}},"evidenceTier":"verified-news","reviewStatus":"automated-ingest","indexable":true,"editorialCover":""}}