{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-08-22T23:00:08.505Z","headline":"SGLang 推出 Weight Cache Daemon，实现亚秒级引擎重启","description":"SGLang 团队推出 Weight Cache Daemon，通过 CUDA IPC 零拷贝映射将模型权重加载从约 495 秒降至约 0.63 秒（约 785 倍加速），端到端启动时间减少 93.9%。该守护进程在 GPU 内存中持久化后量化权重，支持多实例共享和亚秒级主备切换，是 Fast Engine Recovery Framework 的第一阶段。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","url":"https://www.aioga.com/news/cmt393qow0kfiro6tpe87m4nu/","mainEntityOfPage":"https://www.aioga.com/news/cmt393qow0kfiro6tpe87m4nu/","datePublished":"2026-08-21T17:56:25.000Z","dateModified":"2026-08-21T17:56:25.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.lmsys.org/blog/2026-08-21-sglang-fast-recovery","https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu"],"canonicalUrl":"https://www.aioga.com/news/cmt393qow0kfiro6tpe87m4nu/","directAnswer":{"@type":"Answer","text":"SGLang 团队推出 Weight Cache Daemon，将后量化、张量并行切分后的模型权重持续保存在 GPU 内存中，并通过 CUDA IPC 零拷贝映射提供给新引擎实例。材料称，权重加载时间可由约 495 秒降至约 0.63 秒，端到端启动时间减少 93.9%。","url":"https://www.aioga.com/news/cmt393qow0kfiro6tpe87m4nu/","dateCreated":"2026-08-21T17:56:25.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"LMSYS：Blog（Chatbot Arena 团队 source article","url":"https://www.lmsys.org/blog/2026-08-21-sglang-fast-recovery","datePublished":"2026-08-21T17:56:25.000Z","provider":{"@type":"Organization","name":"LMSYS：Blog（Chatbot Arena 团队","url":"https://www.lmsys.org/blog/2026-08-21-sglang-fast-recovery"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","datePublished":"2026-08-21T17:56:25.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu"}}],"aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","originalPublisher":{"name":"LMSYS：Blog（Chatbot Arena 团队","url":"https://www.lmsys.org/blog/2026-08-21-sglang-fast-recovery"},"geoDeepAnswer":null,"article":{"id":"cmt393qow0kfiro6tpe87m4nu","slug":"cmt393qow0kfiro6tpe87m4nu","url":"https://www.aioga.com/news/cmt393qow0kfiro6tpe87m4nu/","title":"SGLang 推出 Weight Cache Daemon，实现亚秒级引擎重启","title_en":"","summary":"SGLang 团队推出 Weight Cache Daemon，通过 CUDA IPC 零拷贝映射将模型权重加载从约 495 秒降至约 0.63 秒（约 785 倍加速），端到端启动时间减少 93.9%。该守护进程在 GPU 内存中持久化后量化权重，支持多实例共享和亚秒级主备切换，是 Fast Engine Recovery Framework 的第一阶段。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","source":"LMSYS：Blog（Chatbot Arena 团队","sourceUrl":"https://www.lmsys.org/blog/2026-08-21-sglang-fast-recovery","aiHotUrl":"https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","publishedAt":"2026-08-21T17:56:25.000Z","category":"行业动态","score":72,"selected":true,"articleBody":["Nowadays, State-of-the-Art (SOTA) models are getting much bigger and reloading the model service after a crash is very expensive. Therefore, we are introducing the Weight Cache Daemon , a persistent GPU process that holds post-quantized model weights in GPU memory and serves them to new SGLang engine instances via CUDA IPC zero-copy mapping. This reduces weight loading from minutes to seconds.","The Weight Cache Daemon is the first phase of our Fast Engine Recovery Framework , which targets < 10 second cold restarts and < 1 second warm standby switches for production LLM serving.","As LLM models grow larger — Qwen3-235B, Ling-2.6-1T, and the newly released 2.8T Kimi K3 — the cold-start time of serving engines has become a critical bottleneck for production efficiency. A Ling-2.6-1T FP8 instance on 8×H20-3e GPUs takes ~8.52 minutes just to become ready to serve, weights stay in 3.5T NVME SSD. In production, this means:","Where does the time go? We profiled a complete SGLang engine startup for Ling-2.6-1T FP8:","The bottleneck is clear: weight loading from disk accounts for 93.2% of startup time . For Ling-2.6-1T FP8 model, each TP rank reads ~120GB of safetensors from disk, deserializes, applies TP sharding, and runs post-quantization transforms (FP8 quantization, weight repacking). This work is repeated identically on every restart , even though the resulting GPU tensors are deterministic and often already present in GPU memory.","Can we avoid reloading from disk every time? The answer is yes — by keeping weights in GPU memory across engine restarts.","The Weight Cache Daemon is a persistent GPU process that holds post-quantized, TP-sharded weights in GPU memory. On engine restart, the new engine process maps weights from the daemon via CUDA IPC zero-copy — no disk I/O, no deserialization, no quantization.","Each GPU runs one daemon process for its TP rank. The daemon:","The engine connects to the daemon, validates config compatibility, and maps weights directly into its address space — the engine and daemon share the same physical GPU memory via CUDA IPC.","The key to sub-second loading is zero-copy : the engine's param.data pointer is set directly to the IPC-mapped GPU tensor. No data is copied.","To achieve this, the engine initializes the model on the meta device (no GPU/CPU memory allocation), then replaces each parameter's data pointer with the IPC-mapped tensor.","Post-quantization parameters (e.g., weight_scale from FP8 quantization) that were created by process_weights_after_loading() are also cached by the daemon and mapped directly — no re-quantization needed.","Any mismatch between the engine's config and the daemon's cached config triggers a full disk reload , ensuring correctness:","The last two fields form an environment stamp : a daemon and a client that ran different post-processing branches (different compute capability or torch/kernel version) can produce weights that map cleanly over IPC yet serve garbage — stamping the environment into CacheConfig turns that into a clean mismatch.","This is critical for production safety: if an operator changes the model or quantization config, the engine will detect the mismatch and fall back to disk loading rather than mapping incompatible weights.","On top of config validation, quantization methods are gated by an IPC allowlist . CUDA IPC zero-copy exports only raw tensor data, so it is correct only when the entire effect of process_weights_after_loading() is captured by that data. Methods that stamp Python-side metadata or repack/transpose weights (per-tensor FP8, Marlin, AWQ/GPTQ) would silently serve wrong numerics — they raise a hard error instead. Currently verified: unquantized and block-wise FP8 ( weight_block_size set); more methods will be added after end-to-end verification.","In daemon mode, the engine spawns daemon processes during startup and waits for them to load weights from disk. The first start is still slow (daemons must load from disk), but subsequent restarts are instant.","In client mode, the engine connects to already-running daemons. This is the fast-restart path — the daemon was started earlier and already holds weights in GPU memory.","The Weight Cache Daemon is designed to be non-intrusive and safe :","The Weight Cache Daemon unlocks production patterns that are impractical with traditional disk-based loading:","A single daemon per GPU holds weights in memory; multiple engine instances (e.g., independent services) map to the same IPC handles via zero-copy. Weights are loaded from disk and quantized exactly once per GPU , regardless of how many instances consume them.","Run a high-priority online service and a low-priority batch job on the same GPU, backed by the same weight cache daemon. The low-priority instance can be evicted and re-spawned in sub-second time without reloading weights from disk — enabling flexible GPU time-sharing without the usual startup penalty.","Deploy a standby engine alongside the primary, both backed by the same weight cache daemon. The standby maps weights via zero-copy and stays warm. When the primary fails, the standby takes over in < 1 second — no weight loading, no disk I/O.","This achieves near-zero-downtime failover without dedicating a full set of GPUs to an idle replica , avoiding the expensive GPU resource waste of traditional hot-standby deployments.","One command launches all TP rank daemons:","Wait for daemons to become ready (they write a .ready file per rank):","In a multi-node deployment, each node runs its own daemon for its local TP ranks. All daemons join the same distributed group, so --nnodes , --node-rank , and --dist-init-method must be consistent across nodes, with $MASTER_ADDR pointing at node 0:","Once every node reports its daemons ready, start the engine clients. They use a separate rendezvous port ( 29600 ) from the daemons ( 29500 ):","The Weight Cache Daemon is Phase 1 of a broader Fast Recovery Framework targeting < 10s cold restarts and < 1s warm standby switches :","Support for more models is also on the way.","The Weight Cache Daemon is just the first step — there is still a lot to build, and we are excited about the road ahead. Phase 1 today covers TP + PP, single- and multi-node launch, per-GPU zero-copy CUDA IPC, and unquantized plus block-wise FP8. Beyond that, many high-impact directions remain open:","This is very much a community effort. The full plan is tracked publicly in sgl-project/sglang#33522：https://github.com/sgl-project/sglang/issues/33522 — contributions and feedback are very welcome , and there is plenty of impactful work to pick up.","Ant Ling Infra Team, Ant Group : Michael Qiu：https://github.com/QiuMike qiudayu.qdy@antgroup.com：mailto:qiudayu.qdy@antgroup.com","Alibaba : Siyu Liu：https://github.com/liusy58 liusy58@smail.nju.edu.cn：mailto:liusy58@smail.nju.edu.cn","SGLang Team : Alex Nails：https://github.com/alexnails"],"articleImages":[{"sourceUrl":"https://www.lmsys.org/images/blog/sglang-fast-recovery/architecture.svg","alt":"","afterParagraph":6,"url":"/media/articles/cmt393qow0kfiro6tpe87m4nu/79bc59537ecbdf20.jpg"}],"mediaStatus":"ok","articleBodyZh":["如今，最先进（SOTA）的模型越来越大，模型服务在崩溃后重新加载代价非常高。因此，我们引入了权重缓存守护进程（Weight Cache Daemon），这是一个持久的 GPU 进程，将量化后的模型权重保存在 GPU 内存中，并通过 CUDA IPC 零拷贝映射将其提供给新的 SGLang 引擎实例。这将权重加载时间从数分钟减少到几秒。","权重缓存守护进程是我们的快速引擎恢复框架（Fast Engine Recovery Framework）的第一阶段，其目标是在生产 LLM 服务中实现 <10 秒的冷重启和 <1 秒的热待机切换。","随着 LLM 模型越来越大——Qwen3-235B、Ling-2.6-1T，以及新发布的 2.8T Kimi K3——服务引擎的冷启动时间已成为生产效率的关键瓶颈。在 8×H20-3e GPU 上运行的 Ling-2.6-1T FP8 实例，仅准备就绪就需要约 8.52 分钟，权重存储在 3.5T NVME SSD 上。在生产环境中，这意味着：","时间都花到哪里去了？我们对 Ling-2.6-1T FP8 的完整 SGLang 引擎启动进行了性能分析：","瓶颈很明显：从磁盘加载权重占启动时间的 93.2%。对于 Ling-2.6-1T FP8 模型，每个 TP 计算节点需要从磁盘读取约 120GB 的 safetensors，反序列化，应用 TP 分片，并运行后量化转换（FP8 量化、权重重打包）。即使生成的 GPU 张量是确定性的且经常已存在于 GPU 内存中，这些工作在每次重启时都会被完全重复。","我们能否避免每次都从磁盘重新加载？答案是肯定的——通过在引擎重启时将权重保存在 GPU 内存中。","权重缓存守护进程是一个持久的 GPU 进程，将量化后的、TP 分片的权重保存在 GPU 内存中。在引擎重启时，新引擎进程通过 CUDA IPC 零拷贝从守护进程映射权重——无需磁盘 I/O，无需反序列化，无需量化。","每个 GPU 为其 TP 计算节点运行一个守护进程。该守护进程：","引擎连接到守护进程，验证配置兼容性，并将权重直接映射到其地址空间——引擎和守护进程通过 CUDA IPC 共享相同的物理 GPU 内存。","实现亚秒级加载的关键是零拷贝：引擎的 param.data 指针直接指向 IPC 映射的 GPU 张量。数据不会被复制。","为了实现这一点，引擎在元设备上初始化模型（不分配 GPU/CPU 内存），然后将每个参数的数据指针替换为 IPC 映射的张量。","后量化参数（例如，FP8 量化生成的 weight_scale）由 process_weights_after_loading() 创建后，也会被守护进程缓存并直接映射——无需重新量化。","引擎的配置与守护进程缓存的配置之间的任何不匹配都会触发完整的磁盘重新加载，以确保正确性：","最后两个字段形成环境标记：使用不同后处理分支（不同计算能力或 torch/kernel 版本）运行的守护进程和客户端可能生成可以通过 IPC 映射，但实际上无效的权重——将环境标记写入 CacheConfig 会将其转化为明显的不匹配。","这对生产安全至关重要：如果操作员更改了模型或量化配置，引擎将检测到不匹配，并回退到磁盘加载，而不是映射不兼容的权重。","在配置验证基础上，量化方法受 IPC 白名单限制。CUDA IPC 零拷贝仅导出原始张量数据，因此仅在 process_weights_after_loading() 的所有效果都被该数据捕获时才正确。那些在 Python 端打标元数据或重新打包/转置权重的方法（每张量 FP8、Marlin、AWQ/GPTQ）将会悄悄生成错误数值——因此改为抛出硬错误。目前验证：未量化和按块 FP8（设置了 weight_block_size）；更多方法将在端到端验证后添加。","在守护进程模式下，引擎在启动期间生成守护进程，并等待它们从磁盘加载权重。首次启动仍然较慢（守护进程必须从磁盘加载），但随后重启是瞬时完成的。","在客户端模式下，引擎连接到已在运行的守护进程。这是快速重启路径——守护进程已提前启动并已将权重保存在 GPU 内存中。","权重缓存守护进程的设计是非侵入且安全的：","权重缓存守护进程解锁了传统基于磁盘加载无法实现的生产模式：","每个 GPU 上只运行一个守护程序以在内存中保存权重；多个引擎实例（例如独立服务）通过零拷贝映射到相同的 IPC 句柄。权重从磁盘加载并且在每个 GPU 上仅量化一次，无论有多少实例使用它们。","在同一个 GPU 上运行高优先级的在线服务和低优先级的批处理作业，由同一个权重缓存守护程序支持。低优先级实例可以被驱逐并在不到一秒的时间内重新启动，而无需从磁盘重新加载权重——实现灵活的 GPU 时间共享，而无需通常的启动开销。","在主引擎旁部署一个备用引擎，两者均由同一个权重缓存守护程序支持。备用引擎通过零拷贝映射权重并保持热状态。当主引擎发生故障时，备用引擎在不到 1 秒内接管——无需加载权重，无需磁盘 I/O。","这实现了接近零停机时间的故障切换，而无需为空闲副本专门分配整套 GPU，从而避免了传统热备用部署中昂贵的 GPU 资源浪费。","一条命令即可启动所有 TP 等级守护程序：","等待守护程序准备就绪（每个等级写入一个 .ready 文件）：","在多节点部署中，每个节点为其本地 TP 等级运行自己的守护程序。所有守护程序加入同一个分布式组，因此 --nnodes、--node-rank 和 --dist-init-method 必须在各节点保持一致，$MASTER_ADDR 指向节点 0：","一旦每个节点报告其守护程序就绪，即可启动引擎客户端。它们使用不同的集合端口（29600），与守护程序端口（29500）分开：","权重缓存守护程序是更广泛的快速恢复框架的第一阶段，目标是实现<10 秒的冷启动和<1 秒的热备用切换：","对更多模型的支持也在路上。","权重缓存守护程序只是第一步——仍有大量工作需要构建，我们对未来的发展充满期待。今天的第一阶段涵盖 TP + PP、单节点和多节点启动、每 GPU 的零拷贝 CUDA IPC 以及未量化和块状 FP8。除此之外，还有许多高影响力的方向等待探索：","这非常依赖社区的努力。完整计划公开跟踪在 sgl-project/sglang#33522：https://github.com/sgl-project/sglang/issues/33522——非常欢迎贡献和反馈，还有大量有影响力的工作可供参与。","蚂蚁灵基础设施团队，蚂蚁集团：Michael Qiu：https://github.com/QiuMike qiudayu.qdy@antgroup.com：mailto:qiudayu.qdy@antgroup.com","阿里巴巴：Siyu Liu：https://github.com/liusy58 liusy58@smail.nju.edu.cn：mailto:liusy58@smail.nju.edu.cn","SGLang 团队：Alex Nails：https://github.com/alexnails"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"SGLang 团队推出 Weight Cache Daemon，将后量化、张量并行切分后的模型权重持续保存在 GPU 内存中，并通过 CUDA IPC 零拷贝映射提供给新引擎实例。材料称，权重加载时间可由约 495 秒降至约 0.63 秒，端到端启动时间减少 93.9%。","background":"该组件是 SGLang Fast Engine Recovery Framework 的第一阶段。原文以 Ling-2.6-1T FP8 为例：8 张 H20-3e GPU 的实例约需 8.52 分钟才能提供服务，其中磁盘权重加载占启动时间的 93.2%，包含读取、反序列化、张量并行切分和后量化处理。","viewpoint":"Aioga 判断，Weight Cache Daemon 将重启时反复执行的权重处理转化为 GPU 内存中的持久化缓存，直接触及大模型服务冷启动的主要耗时来源。其实际收益仍可能取决于 GPU 内存占用、故障范围与部署配置。","implications":"该方案可能降低引擎重启造成的服务空窗，并为多实例共享权重和更快的主备切换提供基础。值得关注的是，材料仅说明了目标指标与测试数据，尚未说明长期稳定性、缓存失效处理及不同模型和硬件配置下的表现。","nextStep":"值得关注后续阶段是否实现原文提出的低于 10 秒冷启动和低于 1 秒温备切换目标，并观察官方是否披露更多模型、硬件与生产环境验证数据。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-08-21T19:02:56.410Z","sourceHash":"8bb599c890ebde56","review":{"approved":true,"groundedness":94,"clarity":92,"duplicationRisk":12,"blockingIssues":[],"notes":["“其实际收益仍可能取决于 GPU 内存占用、故障范围与部署配置”属于合理的条件性判断，已使用“可能”限定，不构成事实错误。","“长期稳定性、缓存失效处理及不同模型和硬件配置下的表现”属于对后续验证重点的合理提示，但来源材料未明确展开这些方面。","“低于 10 秒冷启动和低于 1 秒温备切换”是原文提出的目标指标，不应与已完成的实测结果混淆；当前 nextStep 的表述已明确为后续观察目标。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","LMSYS：Blog（Chatbot Arena 团队）"],"translations":{"zh-CN":{"title":"SGLang 推出 Weight Cache Daemon，实现亚秒级引擎重启","summary":"SGLang 团队推出 Weight Cache Daemon，通过 CUDA IPC 零拷贝映射将模型权重加载从约 495 秒降至约 0.63 秒（约 785 倍加速），端到端启动时间减少 93.9%。该守护进程在 GPU 内存中持久化后量化权重，支持多实例共享和亚秒级主备切换，是 Fast Engine Recovery Framework 的第一阶段。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"行业动态","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang 推出 Weight Cache Daemon，实现亚秒级引擎重启 - Aioga AI资讯","description":"SGLang 团队推出 Weight Cache Daemon，通过 CUDA IPC 零拷贝映射将模型权重加载从约 495 秒降至约 0.63 秒（约 785 倍加速），端到端启动时间减少 93.9%。该守护进程在 GPU 内存中持久化后量化权重，支持多实例共享和亚秒级主备切换，是 Fast Engine Recovery Framework 的第一阶段。...","url":"https://www.aioga.com/news/cmt393qow0kfiro6tpe87m4nu/","articleBody":["如今，最先进（SOTA）的模型越来越大，模型服务在崩溃后重新加载代价非常高。因此，我们引入了权重缓存守护进程（Weight Cache Daemon），这是一个持久的 GPU 进程，将量化后的模型权重保存在 GPU 内存中，并通过 CUDA IPC 零拷贝映射将其提供给新的 SGLang 引擎实例。这将权重加载时间从数分钟减少到几秒。","权重缓存守护进程是我们的快速引擎恢复框架（Fast Engine Recovery Framework）的第一阶段，其目标是在生产 LLM 服务中实现 <10 秒的冷重启和 <1 秒的热待机切换。","随着 LLM 模型越来越大——Qwen3-235B、Ling-2.6-1T，以及新发布的 2.8T Kimi K3——服务引擎的冷启动时间已成为生产效率的关键瓶颈。在 8×H20-3e GPU 上运行的 Ling-2.6-1T FP8 实例，仅准备就绪就需要约 8.52 分钟，权重存储在 3.5T NVME SSD 上。在生产环境中，这意味着：","时间都花到哪里去了？我们对 Ling-2.6-1T FP8 的完整 SGLang 引擎启动进行了性能分析：","瓶颈很明显：从磁盘加载权重占启动时间的 93.2%。对于 Ling-2.6-1T FP8 模型，每个 TP 计算节点需要从磁盘读取约 120GB 的 safetensors，反序列化，应用 TP 分片，并运行后量化转换（FP8 量化、权重重打包）。即使生成的 GPU 张量是确定性的且经常已存在于 GPU 内存中，这些工作在每次重启时都会被完全重复。","我们能否避免每次都从磁盘重新加载？答案是肯定的——通过在引擎重启时将权重保存在 GPU 内存中。","权重缓存守护进程是一个持久的 GPU 进程，将量化后的、TP 分片的权重保存在 GPU 内存中。在引擎重启时，新引擎进程通过 CUDA IPC 零拷贝从守护进程映射权重——无需磁盘 I/O，无需反序列化，无需量化。","每个 GPU 为其 TP 计算节点运行一个守护进程。该守护进程：","引擎连接到守护进程，验证配置兼容性，并将权重直接映射到其地址空间——引擎和守护进程通过 CUDA IPC 共享相同的物理 GPU 内存。","实现亚秒级加载的关键是零拷贝：引擎的 param.data 指针直接指向 IPC 映射的 GPU 张量。数据不会被复制。","为了实现这一点，引擎在元设备上初始化模型（不分配 GPU/CPU 内存），然后将每个参数的数据指针替换为 IPC 映射的张量。","后量化参数（例如，FP8 量化生成的 weight_scale）由 process_weights_after_loading() 创建后，也会被守护进程缓存并直接映射——无需重新量化。","引擎的配置与守护进程缓存的配置之间的任何不匹配都会触发完整的磁盘重新加载，以确保正确性：","最后两个字段形成环境标记：使用不同后处理分支（不同计算能力或 torch/kernel 版本）运行的守护进程和客户端可能生成可以通过 IPC 映射，但实际上无效的权重——将环境标记写入 CacheConfig 会将其转化为明显的不匹配。","这对生产安全至关重要：如果操作员更改了模型或量化配置，引擎将检测到不匹配，并回退到磁盘加载，而不是映射不兼容的权重。","在配置验证基础上，量化方法受 IPC 白名单限制。CUDA IPC 零拷贝仅导出原始张量数据，因此仅在 process_weights_after_loading() 的所有效果都被该数据捕获时才正确。那些在 Python 端打标元数据或重新打包/转置权重的方法（每张量 FP8、Marlin、AWQ/GPTQ）将会悄悄生成错误数值——因此改为抛出硬错误。目前验证：未量化和按块 FP8（设置了 weight_block_size）；更多方法将在端到端验证后添加。","在守护进程模式下，引擎在启动期间生成守护进程，并等待它们从磁盘加载权重。首次启动仍然较慢（守护进程必须从磁盘加载），但随后重启是瞬时完成的。","在客户端模式下，引擎连接到已在运行的守护进程。这是快速重启路径——守护进程已提前启动并已将权重保存在 GPU 内存中。","权重缓存守护进程的设计是非侵入且安全的：","权重缓存守护进程解锁了传统基于磁盘加载无法实现的生产模式：","每个 GPU 上只运行一个守护程序以在内存中保存权重；多个引擎实例（例如独立服务）通过零拷贝映射到相同的 IPC 句柄。权重从磁盘加载并且在每个 GPU 上仅量化一次，无论有多少实例使用它们。","在同一个 GPU 上运行高优先级的在线服务和低优先级的批处理作业，由同一个权重缓存守护程序支持。低优先级实例可以被驱逐并在不到一秒的时间内重新启动，而无需从磁盘重新加载权重——实现灵活的 GPU 时间共享，而无需通常的启动开销。","在主引擎旁部署一个备用引擎，两者均由同一个权重缓存守护程序支持。备用引擎通过零拷贝映射权重并保持热状态。当主引擎发生故障时，备用引擎在不到 1 秒内接管——无需加载权重，无需磁盘 I/O。","这实现了接近零停机时间的故障切换，而无需为空闲副本专门分配整套 GPU，从而避免了传统热备用部署中昂贵的 GPU 资源浪费。","一条命令即可启动所有 TP 等级守护程序：","等待守护程序准备就绪（每个等级写入一个 .ready 文件）：","在多节点部署中，每个节点为其本地 TP 等级运行自己的守护程序。所有守护程序加入同一个分布式组，因此 --nnodes、--node-rank 和 --dist-init-method 必须在各节点保持一致，$MASTER_ADDR 指向节点 0：","一旦每个节点报告其守护程序就绪，即可启动引擎客户端。它们使用不同的集合端口（29600），与守护程序端口（29500）分开：","权重缓存守护程序是更广泛的快速恢复框架的第一阶段，目标是实现<10 秒的冷启动和<1 秒的热备用切换：","对更多模型的支持也在路上。","权重缓存守护程序只是第一步——仍有大量工作需要构建，我们对未来的发展充满期待。今天的第一阶段涵盖 TP + PP、单节点和多节点启动、每 GPU 的零拷贝 CUDA IPC 以及未量化和块状 FP8。除此之外，还有许多高影响力的方向等待探索：","这非常依赖社区的努力。完整计划公开跟踪在 sgl-project/sglang#33522：https://github.com/sgl-project/sglang/issues/33522——非常欢迎贡献和反馈，还有大量有影响力的工作可供参与。","蚂蚁灵基础设施团队，蚂蚁集团：Michael Qiu：https://github.com/QiuMike qiudayu.qdy@antgroup.com：mailto:qiudayu.qdy@antgroup.com","阿里巴巴：Siyu Liu：https://github.com/liusy58 liusy58@smail.nju.edu.cn：mailto:liusy58@smail.nju.edu.cn","SGLang 团队：Alex Nails：https://github.com/alexnails"]},"en":{"title":"SGLang launches Weight Cache Daemon, enabling a sub-second engine reboot","summary":"The SGLang team launched Weight Cache Daemon, which reduced model weight loading from about 495 seconds to about 0.63 seconds (approximately 785x acceleration) through CUDA IPC zero-copy mapping, reducing end-to-end startup time by 93.9%. This daemon is persisted in GPU memory and quantized for weights, supporting multi-instance sharing and sub-second master-backup switching, representing the first phase of the Fast Engine Recovery Framework. 🔗 Read the original article via AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"Industry","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang launches Weight Cache Daemon, enabling a sub-second engine reboot - Aioga AI News","description":"The SGLang team launched Weight Cache Daemon, which reduced model weight loading from about 495 seconds to about 0.63 seconds (approximately 785x acceleration) through CUDA IPC zer...","url":"https://www.aioga.com/en/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:23:55.730Z"},"ja":{"title":"SGLangはWeight Cache Daemonを起動し、1秒未満のエンジン再起動を可能にします","summary":"SGLangチームはWeight Cache Daemonを立ち上げ、CUDA IPCゼロコピーマッピングによりモデルのウェイトロード時間を約495秒から約0.63秒(約785倍の加速)に短縮し、エンドツーエンドの起動時間を93.9%短縮しました。 このデーモンはGPUメモリに永続化され、重みで量子化され、マルチインスタンス共有およびサブ秒間マスターバックアップスイッチングをサポートし、Fast Engine Recovery Frameworkの第一段階を表しています。 🔗 原文記事はAIHOTより読むことができます。 https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"業界動向","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLangはWeight Cache Daemonを起動し、1秒未満のエンジン再起動を可能にします - Aioga AIニュース","description":"SGLangチームはWeight Cache Daemonを立ち上げ、CUDA IPCゼロコピーマッピングによりモデルのウェイトロード時間を約495秒から約0.63秒(約785倍の加速)に短縮し、エンドツーエンドの起動時間を93.9%短縮しました。 このデーモンはGPUメモリに永続化され、重みで量子化され、マルチインスタンス共有およびサブ秒間マスターバックア...","url":"https://www.aioga.com/ja/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:23:57.371Z"},"ko":{"title":"SGLang이 무게 캐시 데몬을 실행하여 초도 이하의 엔진 재부팅을 가능하게 합니다","summary":"SGLang 팀은 CUDA IPC 제로 복사 매핑을 통해 모델 가중치 로딩을 약 495초에서 약 0.63초(약 785배 가속)로 줄여 시작 시간을 93.9% 단축하는 Weight Cache Daemon을 출시했습니다. 이 데몬은 GPU 메모리에 저장되며 가중치에 대해 양자화되어 다중 인스턴스 공유와 1초 미만의 마스터 백업 스위칭을 지원하며, 이는 Fast Engine Recovery Framework의 첫 단계를 나타냅니다. 🔗 원문 기사는 AIHOT를 통해 읽을 수 있습니다. https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"업계 동향","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang이 무게 캐시 데몬을 실행하여 초도 이하의 엔진 재부팅을 가능하게 합니다 - Aioga AI 뉴스","description":"SGLang 팀은 CUDA IPC 제로 복사 매핑을 통해 모델 가중치 로딩을 약 495초에서 약 0.63초(약 785배 가속)로 줄여 시작 시간을 93.9% 단축하는 Weight Cache Daemon을 출시했습니다. 이 데몬은 GPU 메모리에 저장되며 가중치에 대해 양자화되어 다중 인스턴스 공유와 1초 미만의 마스터 백...","url":"https://www.aioga.com/ko/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:06.100Z"},"es":{"title":"SGLang lanza Weight Cache Daemon, permitiendo un reinicio del motor en menos de un segundo","summary":"El equipo de SGLang lanzó Weight Cache Daemon, que redujo la carga de peso del modelo de unos 495 segundos a unos 0,63 segundos (aproximadamente 785 veces la aceleración) mediante mapeo de copia cero IPC de CUDA, reduciendo el tiempo de arranque de extremo a extremo en un 93,9%. Este daemon se mantiene en la memoria GPU y se cuantifica para pesos, soportando compartición multiinstancia y conmutación de copia de seguridad maestra en menos de un segundo, representando la primera fase del Marco de Recuperación del Motor Rápido. 🔗 Lee el artículo original a través de AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"Industria","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang lanza Weight Cache Daemon, permitiendo un reinicio del motor en menos de un segundo - Aioga Noticias de IA","description":"El equipo de SGLang lanzó Weight Cache Daemon, que redujo la carga de peso del modelo de unos 495 segundos a unos 0,63 segundos (aproximadamente 785 veces la aceleración) mediante...","url":"https://www.aioga.com/es/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:06.275Z"},"fr":{"title":"SGLang lance Weight Cache Daemon, permettant un redémarrage moteur en moins d’une seconde","summary":"L’équipe SGLang a lancé Weight Cache Daemon, qui a réduit la charge du poids du modèle d’environ 495 secondes à environ 0,63 seconde (environ 785 fois l’accélération) grâce à la cartographie IPC zéro copie CUDA, réduisant le temps de démarrage de bout en bout de 93,9 %. Ce démon est conservé en mémoire GPU et quantifié pour les poids, prenant en charge le partage multi-instances et la commutation maître-sauvegarde en moins de seconde, représentant la première phase du Fast Engine Recovery Framework. 🔗 Lisez l’article original via AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"Industrie","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang lance Weight Cache Daemon, permettant un redémarrage moteur en moins d’une seconde - Aioga Actualités IA","description":"L’équipe SGLang a lancé Weight Cache Daemon, qui a réduit la charge du poids du modèle d’environ 495 secondes à environ 0,63 seconde (environ 785 fois l’accélération) grâce à la ca...","url":"https://www.aioga.com/fr/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:14.019Z"},"de":{"title":"SGLang startet Weight Cache Daemon, was einen Neustart der Engine in weniger als einer Sekunde ermöglicht","summary":"Das SGLang-Team brachte Weight Cache Daemon auf den Markt, das das Modellgewicht von etwa 495 Sekunden auf etwa 0,63 Sekunden (etwa 785-fache Beschleunigung) durch CUDA IPC Zero-Copy Mapping reduzierte und die Startzeit von Ende zu Ende um 93,9 % verkürzte. Dieser Daemon wird im GPU-Speicher gespeichert und nach Gewichten quantisiert, unterstützt Multi-Instanz-Sharing und Master-Backup-Switching in Subsekunden und stellt die erste Phase des Fast Engine Recovery Frameworks dar. 🔗 Lesen Sie den Originalartikel über AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"行业动态","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang startet Weight Cache Daemon, was einen Neustart der Engine in weniger als einer Sekunde ermöglicht - Aioga KI-News","description":"Das SGLang-Team brachte Weight Cache Daemon auf den Markt, das das Modellgewicht von etwa 495 Sekunden auf etwa 0,63 Sekunden (etwa 785-fache Beschleunigung) durch CUDA IPC Zero-Co...","url":"https://www.aioga.com/de/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:14.255Z"},"pt-BR":{"title":"SGLang lança o Weight Cache Daemon, permitindo uma reinicialização do motor em menos de um segundo","summary":"A equipe SGLang lançou o Weight Cache Daemon, que reduziu o carregamento do peso do modelo de cerca de 495 segundos para cerca de 0,63 segundos (aproximadamente 785x de aceleração) por meio do mapeamento zero-copy IPC da CUDA, reduzindo o tempo de inicialização de ponta a ponta em 93,9%. Esse daemon é mantido na memória da GPU e quantizado para pesos, suportando compartilhamento multi-instância e comutação mestre-backup de menos de segundo, representando a primeira fase do Fast Engine Recovery Framework. 🔗 Leia o artigo original via AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"行业动态","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang lança o Weight Cache Daemon, permitindo uma reinicialização do motor em menos de um segundo - Aioga Notícias de IA","description":"A equipe SGLang lançou o Weight Cache Daemon, que reduziu o carregamento do peso do modelo de cerca de 495 segundos para cerca de 0,63 segundos (aproximadamente 785x de aceleração)...","url":"https://www.aioga.com/pt-BR/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:23.009Z"},"ru":{"title":"SGLang запускает Weight Cache Daemon, позволяющий перезагрузку движка менее чем за секунду","summary":"Команда SGLang запустила Weight Cache Daemon, который снизил загрузку веса модели с примерно 495 секунд до примерно 0,63 секунды (примерно 785-кратное ускорение) благодаря CUDA IPC zero-copy mapping, сократив время запуска от конца до конца на 93,9%. Этот демон сохраняется в памяти GPU и квантизируется для весов, поддерживая многоэкземплярное совместное использование и мастер-резервное переключение на секунду, представляя собой первый этап Fast Engine Recovery Framework. 🔗 Прочитайте оригинальную статью на сайте AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"行业动态","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang запускает Weight Cache Daemon, позволяющий перезагрузку движка менее чем за секунду - Aioga Новости ИИ","description":"Команда SGLang запустила Weight Cache Daemon, который снизил загрузку веса модели с примерно 495 секунд до примерно 0,63 секунды (примерно 785-кратное ускорение) благодаря CUDA IPC...","url":"https://www.aioga.com/ru/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:22.180Z"},"ar":{"title":"تطلق SGLang لعبة Weight Cache Daemon، مما يتيح إعادة تشغيل المحرك خلال أقل من الثانية","summary":"أطلق فريق SGLang جهاز Weight Cache Daemon، الذي خفض تحميل وزن النموذج من حوالي 495 ثانية إلى حوالي 0.63 ثانية (حوالي 785 ضعف التسارع) عبر رسم الخرائط بدون نسخ CUDA IPC، مما قلل من وقت بدء التشغيل من البداية إلى النهاية بنسبة 93.9٪. يتم الاحتفاظ بهذا الخادم في ذاكرة وحدة معالجة الرسومات ويتم تكميمها للأوزان، مما يدعم مشاركة النسخ المتعددة وتبديل النسخ الاحتياطي الرئيسي دون ثانية، مما يمثل المرحلة الأولى من إطار عمل استعادة المحرك السريع. 🔗 اقرأ المقال الأصلي عبر AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"行业动态","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"تطلق SGLang لعبة Weight Cache Daemon، مما يتيح إعادة تشغيل المحرك خلال أقل من الثانية - Aioga أخبار الذكاء الاصطناعي","description":"أطلق فريق SGLang جهاز Weight Cache Daemon، الذي خفض تحميل وزن النموذج من حوالي 495 ثانية إلى حوالي 0.63 ثانية (حوالي 785 ضعف التسارع) عبر رسم الخرائط بدون نسخ CUDA IPC، مما قلل من...","url":"https://www.aioga.com/ar/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:31.797Z"},"hi":{"title":"SGLang ने वेट कैश डेमॉन लॉन्च किया, जो सब-सेकंड इंजन रिबूट को सक्षम करता है","summary":"SGLang टीम ने वेट कैश डेमॉन लॉन्च किया, जिसने CUDA IPC जीरो-कॉपी मैपिंग के माध्यम से मॉडल वजन लोडिंग को लगभग 495 सेकंड से घटाकर लगभग 0.63 सेकंड (लगभग 785x त्वरण) कर दिया, जिससे एंड-टू-एंड स्टार्टअप समय 93.9% कम हो गया। यह डेमॉन GPU मेमोरी में बना रहता है और वजन के लिए क्वांटाइज़ किया जाता है, मल्टी-इंस्टेंस शेयरिंग और सब-सेकंड मास्टर-बैकअप स्विचिंग का समर्थन करता है, जो फास्ट इंजन रिकवरी फ्रेमवर्क के पहले चरण का प्रतिनिधित्व करता है। 🔗 AIHOT के माध्यम से मूल लेख पढ़ें · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"行业动态","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang ने वेट कैश डेमॉन लॉन्च किया, जो सब-सेकंड इंजन रिबूट को सक्षम करता है - Aioga AI समाचार","description":"SGLang टीम ने वेट कैश डेमॉन लॉन्च किया, जिसने CUDA IPC जीरो-कॉपी मैपिंग के माध्यम से मॉडल वजन लोडिंग को लगभग 495 सेकंड से घटाकर लगभग 0.63 सेकंड (लगभग 785x त्वरण) कर दिया, जिससे एंड...","url":"https://www.aioga.com/hi/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:31.276Z"},"it":{"title":"SGLang lancia Weight Cache Daemon, consentendo un riavvio del motore in meno di un secondo","summary":"Il team SGLang ha lanciato Weight Cache Daemon, che ha ridotto il carico del peso del modello da circa 495 secondi a circa 0,63 secondi (circa 785 volte l'accelerazione) tramite la mappatura zero-copy IPC CUDA, riducendo il tempo di avvio end-to-end del 93,9%. Questo daemon è conservato nella memoria GPU e quantizzato per i pesi, supportando la condivisione multi-istanze e la commutazione master-backup in meno di un secondo, rappresentando la prima fase del Fast Engine Recovery Framework. 🔗 Leggi l'articolo originale su AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"行业动态","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang lancia Weight Cache Daemon, consentendo un riavvio del motore in meno di un secondo - Aioga Notizie IA","description":"Il team SGLang ha lanciato Weight Cache Daemon, che ha ridotto il carico del peso del modello da circa 495 secondi a circa 0,63 secondi (circa 785 volte l'accelerazione) tramite la...","url":"https://www.aioga.com/it/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:40.516Z"},"nl":{"title":"SGLang start Weight Cache Daemon, waardoor een motorherstart binnen een seconde mogelijk wordt","summary":"Het SGLang-team lanceerde Weight Cache Daemon, waarmee het laadvermogen van het model werd verminderd van ongeveer 495 seconden tot ongeveer 0,63 seconden (ongeveer 785x versnelling) via CUDA IPC zero-copy mapping, waardoor de end-to-end opstarttijd met 93,9% werd verminderd. Deze daemon wordt opgeslagen in het GPU-geheugen en gekwantiseerd op gewichten, met ondersteuning voor multi-instance sharing en sub-seconde master-backup switching, wat de eerste fase van het Fast Engine Recovery Framework vertegenwoordigt. 🔗 Lees het originele artikel via AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"行业动态","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang start Weight Cache Daemon, waardoor een motorherstart binnen een seconde mogelijk wordt - Aioga AI-nieuws","description":"Het SGLang-team lanceerde Weight Cache Daemon, waarmee het laadvermogen van het model werd verminderd van ongeveer 495 seconden tot ongeveer 0,63 seconden (ongeveer 785x versnellin...","url":"https://www.aioga.com/nl/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:40.529Z"},"tr":{"title":"SGLang, Weight Cache Daemon'u başlatarak motorun saniyeden kısa sürede yeniden başlatılmasını mümkün kılarak","summary":"SGLang ekibi, model ağırlık yüklemesini yaklaşık 495 saniyeden yaklaşık 0,63 saniyeye (yaklaşık 785 kat hızlama) düşüren ve CUDA IPC sıfır kopya eşlemesi sayesinde uçtan uca başlatma süresini %93,9 oranında azaltan Weight Cache Daemon'u başlattı. Bu daemon, GPU belleğinde saklanır ve ağırlıklar için kuantize edilir; çoklu örnek paylaşımı ve saniye altı ana yedekleme anahtarlamasını destekler; bu da Hızlı Motor Kurtarma Çerçevesi'nin ilk aşamasını temsil eder. 🔗 Orijinal makaleyi AIHOT üzerinden okuyun · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"行业动态","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang, Weight Cache Daemon'u başlatarak motorun saniyeden kısa sürede yeniden başlatılmasını mümkün kılarak - Aioga AI Haberleri","description":"SGLang ekibi, model ağırlık yüklemesini yaklaşık 495 saniyeden yaklaşık 0,63 saniyeye (yaklaşık 785 kat hızlama) düşüren ve CUDA IPC sıfır kopya eşlemesi sayesinde uçtan uca başlat...","url":"https://www.aioga.com/tr/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:49.211Z"},"vi":{"title":"SGLang khởi chạy Weight Cache Daemon, cho phép khởi động lại engine dưới một giây","summary":"Nhóm SGLang đã ra mắt Weight Cache Daemon, giúp giảm thời gian tải trọng lượng mô hình từ khoảng 495 giây xuống còn khoảng 0,63 giây (tăng tốc khoảng 785 lần) thông qua ánh xạ zero-copy của CUDA IPC, giảm thời gian khởi động đầu cuối xuống 93,9%. Daemon này được lưu trữ trong bộ nhớ GPU và được lượng tử hóa theo trọng lượng, hỗ trợ chia sẻ đa phiên bản và chuyển mạch sao lưu chính dưới một giây, đại diện cho giai đoạn đầu tiên của Khung Phục hồi Động cơ Nhanh. 🔗 Đọc bài viết gốc qua AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"行业动态","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang khởi chạy Weight Cache Daemon, cho phép khởi động lại engine dưới một giây - Tin tức AI Aioga","description":"Nhóm SGLang đã ra mắt Weight Cache Daemon, giúp giảm thời gian tải trọng lượng mô hình từ khoảng 495 giây xuống còn khoảng 0,63 giây (tăng tốc khoảng 785 lần) thông qua ánh xạ zero...","url":"https://www.aioga.com/vi/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:49.170Z"},"id":{"title":"SGLang meluncurkan Weight Cache Daemon, memungkinkan reboot mesin dalam waktu kurang dari detik","summary":"Tim SGLang meluncurkan Weight Cache Daemon, yang mengurangi beban berat model dari sekitar 495 detik menjadi sekitar 0,63 detik (sekitar akselerasi 785x) melalui pemetaan zero-copy CUDA IPC, mengurangi waktu startup end-to-end sebesar 93,9%. Daemon ini disimpan di memori GPU dan dikuantisasi untuk bobot, mendukung berbagi multi-instance dan switching master-backup sub-detik, yang merupakan fase pertama dari Fast Engine Recovery Framework. 🔗 Baca artikel asli melalui AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"行业动态","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang meluncurkan Weight Cache Daemon, memungkinkan reboot mesin dalam waktu kurang dari detik - Berita AI Aioga","description":"Tim SGLang meluncurkan Weight Cache Daemon, yang mengurangi beban berat model dari sekitar 495 detik menjadi sekitar 0,63 detik (sekitar akselerasi 785x) melalui pemetaan zero-copy...","url":"https://www.aioga.com/id/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:57.249Z"},"th":{"title":"SGLang เปิดตัว Weight Cache Daemon ช่วยให้สามารถรีบูตเครื่องยนต์ได้ภายในวินาที","summary":"ทีม SGLang ได้เปิดตัว Weight Cache Daemon ซึ่งลดเวลาการโหลดน้ําหนักของโมเดลจากประมาณ 495 วินาทีเหลือประมาณ 0.63 วินาที (ประมาณ 785 เท่าของการเร่งความเร็ว) ผ่านการแมปแบบ zero-copy ของ CUDA IPC ลดเวลาการเริ่มต้นแบบครบวงจรลง 93.9% เดมอนนี้ถูกเก็บไว้ในหน่วยความจํา GPU และถูกควอนไทซ์เป็นน้ําหนัก รองรับการแชร์หลายอินสแตนซ์และการสลับข้อมูลสํารองหลักแบบต่ํากว่าวินาที ซึ่งเป็นเฟสแรกของกรอบการกู้คืนเครื่องยนต์ความเร็วสูง 🔗 อ่านบทความต้นฉบับผ่าน AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"行业动态","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang เปิดตัว Weight Cache Daemon ช่วยให้สามารถรีบูตเครื่องยนต์ได้ภายในวินาที - ข่าว AI Aioga","description":"ทีม SGLang ได้เปิดตัว Weight Cache Daemon ซึ่งลดเวลาการโหลดน้ําหนักของโมเดลจากประมาณ 495 วินาทีเหลือประมาณ 0.63 วินาที (ประมาณ 785 เท่าของการเร่งความเร็ว) ผ่านการแมปแบบ zero-copy ข...","url":"https://www.aioga.com/th/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:24:57.861Z"},"pl":{"title":"SGLang uruchamia Weight Cache Daemon, umożliwiając restart silnika w krótszym czasie","summary":"Zespół SGLang uruchomił Weight Cache Daemon, który zmniejszył obciążenie wagi modelu z około 495 sekund do około 0,63 sekundy (około 785x przyspieszenia) dzięki mapowaniu CUDA IPC zero-copy, skracając czas uruchomienia od początku do końca o 93,9%. Ten demon jest przechowywany w pamięci GPU i kwantyzowany pod względem wag, obsługując wieloinstancyjne udostępnianie i przełączanie kopii zapasowej w czasie poniżej sekundy, reprezentując pierwszą fazę Fast Engine Recovery Framework. 🔗 Przeczytaj oryginalny artykuł za pośrednictwem AIHOT · https://aihot.virxact.com/items/cmt393qow0kfiro6tpe87m4nu","category":"行业动态","source":"LMSYS：Blog（Chatbot Arena 团队","aggregationSource":"LMSYS：Blog（Chatbot Arena 团队","pageTitle":"SGLang uruchamia Weight Cache Daemon, umożliwiając restart silnika w krótszym czasie - Aioga Wiadomości AI","description":"Zespół SGLang uruchomił Weight Cache Daemon, który zmniejszył obciążenie wagi modelu z około 495 sekund do około 0,63 sekundy (około 785x przyspieszenia) dzięki mapowaniu CUDA IPC...","url":"https://www.aioga.com/pl/news/cmt393qow0kfiro6tpe87m4nu/","contentTranslated":true,"sourceHash":"5f54d525bd3149e9","translatedAt":"2026-08-21T18:25:05.925Z"}},"evidenceTier":"verified-news","reviewStatus":"editorial-selected","indexable":true,"editorialCover":""}}