小红书引擎架构团队在OSDI 2026提出HELMSMAN,一个面向全闪存服务器的高性能向量近似最近邻搜索系统。
该系统通过聚类式索引、定制化存储栈和分层学习式搜索剪枝,用约40台全闪存服务器承载了过去约35,000 CPU Core和约350 TB DRAM的负载,硬件成本节省超过90%。
The Xiaohongshu Engine Architecture Team proposed HELMSMAN at OSDI 2026, a high-performance vector approximate nearest neighbor search system for all-flash...
The Xiaohongshu Engine Architecture Team proposed HELMSMAN at OSDI 2026, a high-performance vector
approximate nearest neighbor search system for all-flash servers. This system uses clustered indexing, a customized storage stack, and hierarchical learning-based search pruning to handle workloads that previously required about 35,000 CPU cores and approximately 350 TB of DRAM with around 40 all-flash servers, resulting in hardware cost savings of over 90%.
小红书引擎架构团队在OSDI 2026提出HELMSMAN,一个面向全闪存服务器的高性能向量近似最近邻搜索系统。
该系统通过聚类式索引、定制化存储栈和分层学习式搜索剪枝,用约40台全闪存服务器承载了过去约35,000 CPU Core和约350 TB DRAM的负载,硬件成本节省超过90%。
小红书引擎架构团队在OSDI 2026提出HELMSMAN。该系统面向全闪存服务器,以聚类式索引、定制化存储栈和分层学习式搜索剪枝实现向量近似最近邻搜索。
公开材料称,HELMSMAN用约40台全闪存服务器承载了过去约35000个CPU核心和约350 TB DRAM对应的负载,报告的硬件成本节省幅度超过90%。
Aioga判断,HELMSMAN值得关注之处在于将索引结构、存储栈与搜索剪枝协同设计,并以全闪存服务器承接原有大规模计算与内存负载。
Aioga判断,该结果可能为高性能向量检索的硬件配置提供新的工程参考;但现有材料未披露测试数据集、延迟、召回率及成本核算口径,不宜扩大结论。 值得关注后续论文或团队披露的完整实验设置,包括近似最近邻搜索质量、延迟与吞吐表现、闪存读写特征,以及超过90%硬件成本节省的计算边界。
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Ingestion channel: Summary aggregation · Source domain: mp.weixin.qq.com
Source: 公众号:小红书技术(dots.llm)
Original link: Open original source
Aioga archive: Open intelligence page
Content record: summary-fallback · Updated: 2026-07-23T04:00:00.000Z

统一接入主流 AI 模型 API,为开发、测试与生产环境提供稳定调用入口。
立即访问 api.w173.comAioga aggregates global AI updates and preserves source information for verification and citation.