{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-23T06:40:50.084Z","headline":"SkewAdam：为MoE训练提出分层优化器状态分配策略","description":"SkewAdam针对MoE模型提出分层优化器状态分配：为密集骨干网络保留float32动量与分解二阶矩，为专家层仅保留分解二阶矩，为路由器保留精确二阶矩。在6.78B参数MoE上，优化器状态仅占1.29 GB（AdamW的2.6%），峰值训练内存从81.4 GB降至31.3 GB，可装入40 GB加速器。","url":"https://www.aioga.com/news/cmrvs7s6600cobipzahp9o5lb/","mainEntityOfPage":"https://www.aioga.com/news/cmrvs7s6600cobipzahp9o5lb/","datePublished":"2026-07-21T00:00:00.000Z","dateModified":"2026-07-21T00:00:00.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://arxiv.org/abs/2607.19058","https://aihot.virxact.com/items/cmrvs7s6600cobipzahp9o5lb"],"canonicalUrl":"https://www.aioga.com/news/cmrvs7s6600cobipzahp9o5lb/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：SkewAdam针对MoE模型提出分层优化器状态分配：为密集骨干网络保留float32动量与分解二阶矩，为专家层仅保留分解二阶矩，为路由器保留精确二阶矩。 Aioga 将其归入「论文研究」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cmrvs7s6600cobipzahp9o5lb/","dateCreated":"2026-07-21T00:00:00.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"arXiv source article","url":"https://arxiv.org/abs/2607.19058","datePublished":"2026-07-21T00:00:00.000Z","provider":{"@type":"Organization","name":"arXiv","url":"https://arxiv.org/abs/2607.19058"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrvs7s6600cobipzahp9o5lb","datePublished":"2026-07-21T00:00:00.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrvs7s6600cobipzahp9o5lb"}}],"aggregationSource":"HuggingFace Daily Papers（社区热门论文）","originalPublisher":{"name":"arXiv","url":"https://arxiv.org/abs/2607.19058"},"article":{"id":"cmrvs7s6600cobipzahp9o5lb","slug":"cmrvs7s6600cobipzahp9o5lb","url":"https://www.aioga.com/news/cmrvs7s6600cobipzahp9o5lb/","title":"SkewAdam：为MoE训练提出分层优化器状态分配策略","title_en":"Where Should Optimizer State Live？ Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training","summary":"SkewAdam针对MoE模型提出分层优化器状态分配：为密集骨干网络保留float32动量与分解二阶矩，为专家层仅保留分解二阶矩，为路由器保留精确二阶矩。在6.78B参数MoE上，优化器状态仅占1.29 GB（AdamW的2.6%），峰值训练内存从81.4 GB降至31.3 GB，可装入40 GB加速器。","source":"HuggingFace Daily Papers（社区热门论文）","sourceUrl":"https://arxiv.org/abs/2607.19058","aiHotUrl":"https://aihot.virxact.com/items/cmrvs7s6600cobipzahp9o5lb","publishedAt":"2026-07-21T00:00:00.000Z","category":"论文研究","score":53,"selected":false,"articleBody":["arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.","Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs ：https://info.arxiv.org/labs/index.html."],"articleImages":[{"sourceUrl":"https://arxiv.org/static/base/1.0.1/images/funders/simons-foundation.png","alt":"Simons Foundation","afterParagraph":1,"url":"/media/articles/cmrvs7s6600cobipzahp9o5lb/e2d7f38d62f5ca91.png"},{"sourceUrl":"https://arxiv.org/static/base/1.0.1/images/funders/simons-foundation-international.png","alt":"Simons Foundation International","afterParagraph":1,"url":"/media/articles/cmrvs7s6600cobipzahp9o5lb/1d56e29c5557cbdc.png"},{"sourceUrl":"https://arxiv.org/static/base/1.0.1/images/funders/schmidt-sciences.png","alt":"Schmidt Sciences","afterParagraph":1,"url":"/media/articles/cmrvs7s6600cobipzahp9o5lb/8e18212b8219c104.png"}],"mediaStatus":"ok","articleBodyZh":["arXivLabs 是一个框架，允许合作者直接在我们的网站上开发和分享新的 arXiv 功能。","有一个可以为 arXiv 社区增加价值的项目想法吗？了解更多关于 arXivLabs 的信息：https://info.arxiv.org/labs/index.html。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：SkewAdam针对MoE模型提出分层优化器状态分配：为密集骨干网络保留float32动量与分解二阶矩，为专家层仅保留分解二阶矩，为路由器保留精确二阶矩。 Aioga 将其归入「论文研究」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：模型与研究类动态需要结合能力边界、开放方式、成本、可用性和真实任务表现判断，单项指标领先不等于已经形成稳定采用。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察官方文档、实际可用性、价格变化、开发者反馈和竞品回应。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-23T06:49:19.092Z","sourceHash":"20c1f8412772f659","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["论文研究","HuggingFace Daily Papers（社区热门论文）"],"translations":{"zh-CN":{"title":"SkewAdam：为MoE训练提出分层优化器状态分配策略","summary":"SkewAdam针对MoE模型提出分层优化器状态分配：为密集骨干网络保留float32动量与分解二阶矩，为专家层仅保留分解二阶矩，为路由器保留精确二阶矩。在6.78B参数MoE上，优化器状态仅占1.29 GB（AdamW的2.6%），峰值训练内存从81.4 GB降至31.3 GB，可装入40 GB加速器。","category":"论文研究","source":"arXiv","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam：为MoE训练提出分层优化器状态分配策略 - Aioga AI资讯","description":"SkewAdam针对MoE模型提出分层优化器状态分配：为密集骨干网络保留float32动量与分解二阶矩，为专家层仅保留分解二阶矩，为路由器保留精确二阶矩。在6.78B参数MoE上，优化器状态仅占1.29 GB（AdamW的2.6%），峰值训练内存从81.4 GB降至31.3 GB，可装入40 GB加速器。","url":"https://www.aioga.com/news/cmrvs7s6600cobipzahp9o5lb/"},"en":{"title":"SkewAdam: Proposing a Hierarchical Optimizer State Allocation Strategy for MoE Training","summary":"SkewAdam proposes a hierarchical optimizer state allocation for MoE models: retain float32 momentum and decomposed second-order moments for the dense backbone network, retain only decomposed second-order moments for the expert layers, and retain exact second-order moments for the router. On a 6.78B parameter MoE, the optimizer state only takes up 1.29 GB (2.6% of AdamW), peak training memory drops from 81.4 GB to 31.3 GB, making it possible to fit on a 40 GB accelerator.","category":"Research","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: Proposing a Hierarchical Optimizer State Allocation Strategy for MoE Training - Aioga AI News","description":"SkewAdam proposes a hierarchical optimizer state allocation for MoE models: retain float32 momentum and decomposed second-order moments for the dense backbone network, retain only...","url":"https://www.aioga.com/en/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:42:13.275Z"},"ja":{"title":"SkewAdam：MoEトレーニングのための階層的オプティマイザーステート割り当て戦略を提案","summary":"SkewAdamはMoEモデルに対して階層的なオプティマイザーステートの割り当てを提案する：密集バックボーンネットワークにはfloat32のモーメンタムと分解二次モーメントを保持し、エキスパート層には分解二次モーメントのみを保持し、ルーターには正確な二次モーメントを保持する。6.78BパラメータのMoEでは、オプティマイザーステートはわずか1.29 GB（AdamWの2.6%）を占め、トレーニング時のピークメモリは81.4 GBから31.3 GBに減少し、40 GBのアクセラレータに収めることができる。","category":"論文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam：MoEトレーニングのための階層的オプティマイザーステート割り当て戦略を提案 - Aioga AIニュース","description":"SkewAdamはMoEモデルに対して階層的なオプティマイザーステートの割り当てを提案する：密集バックボーンネットワークにはfloat32のモーメンタムと分解二次モーメントを保持し、エキスパート層には分解二次モーメントのみを保持し、ルーターには正確な二次モーメントを保持する。6.78BパラメータのMoEでは、オプティマイザーステートはわずか1.29 GB（A...","url":"https://www.aioga.com/ja/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:42:24.800Z"},"ko":{"title":"SkewAdam: MoE 훈련을 위한 계층별 옵티마이저 상태 할당 전략 제안","summary":"SkewAdam은 MoE 모델을 위해 계층별 옵티마이저 상태 할당을 제안합니다: 밀집 백본 네트워크에는 float32 모멘텀과 분해된 2차 모멘트를 유지하고, 전문가 층에는 분해된 2차 모멘트만 유지하며, 라우터에는 정확한 2차 모멘트를 유지합니다. 6.78B 파라미터 MoE에서 옵티마이저 상태는 단지 1.29 GB(AdamW의 2.6%)를 차지하고, 최고 훈련 메모리는 81.4 GB에서 31.3 GB로 감소하여 40 GB 가속기에 적재할 수 있습니다.","category":"연구","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: MoE 훈련을 위한 계층별 옵티마이저 상태 할당 전략 제안 - Aioga AI 뉴스","description":"SkewAdam은 MoE 모델을 위해 계층별 옵티마이저 상태 할당을 제안합니다: 밀집 백본 네트워크에는 float32 모멘텀과 분해된 2차 모멘트를 유지하고, 전문가 층에는 분해된 2차 모멘트만 유지하며, 라우터에는 정확한 2차 모멘트를 유지합니다. 6.78B 파라미터 MoE에서 옵티마이저 상태는 단지 1.29 GB(Ad...","url":"https://www.aioga.com/ko/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:43:18.123Z"},"es":{"title":"SkewAdam: Propuesta de una estrategia de asignación de estados de optimizador jerárquico para el entrenamiento MoE","summary":"SkewAdam propone una asignación jerárquica del estado del optimizador para modelos MoE: conserva el momentum en float32 y los momentos de segundo orden descompuestos para la red troncal densa, conserva solo los momentos de segundo orden descompuestos para las capas de expertos y conserva momentos de segundo orden precisos para el enrutador. En un MoE de 6.78B parámetros, el estado del optimizador ocupa solo 1.29 GB (el 2.6% de AdamW), y la memoria máxima de entrenamiento se reduce de 81.4 GB a 31.3 GB, pudiéndose cargar en un acelerador de 40 GB.","category":"Investigación","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: Propuesta de una estrategia de asignación de estados de optimizador jerárquico para el entrenamiento MoE - Aioga Noticias de IA","description":"SkewAdam propone una asignación jerárquica del estado del optimizador para modelos MoE: conserva el momentum en float32 y los momentos de segundo orden descompuestos para la red tr...","url":"https://www.aioga.com/es/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:43:06.609Z"},"fr":{"title":"SkewAdam : Proposer une stratégie d'allocation des états de l'optimiseur hiérarchique pour l'entraînement MoE","summary":"SkewAdam propose une allocation hiérarchique des états de l'optimiseur pour les modèles MoE : conserver la quantité de mouvement en float32 et la décomposition des secondes matrices pour le réseau central dense, conserver uniquement la décomposition des secondes matrices pour les couches d'experts, et conserver les secondes matrices exactes pour le routeur. Sur un MoE de 6,78 milliards de paramètres, l'état de l'optimiseur ne représente que 1,29 Go (2,6 % de AdamW), la mémoire de formation maximale passe de 81,4 Go à 31,3 Go, ce qui permet de la charger sur un accélérateur de 40 Go.","category":"Recherche","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam : Proposer une stratégie d'allocation des états de l'optimiseur hiérarchique pour l'entraînement MoE - Aioga Actualités IA","description":"SkewAdam propose une allocation hiérarchique des états de l'optimiseur pour les modèles MoE : conserver la quantité de mouvement en float32 et la décomposition des secondes matrice...","url":"https://www.aioga.com/fr/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:44:03.135Z"},"de":{"title":"SkewAdam: Ein hierarchisches Optimierer-Zustandszuweisungsstrategie für das MoE-Training vorgeschlagen","summary":"SkewAdam schlägt für MoE-Modelle eine hierarchische Optimiererzustandsaufteilung vor: Für das dichte Backbonenetzwerk werden Float32-Momente und zerlegte zweite Momente beibehalten, für die Expertenebene nur zerlegte zweite Momente, für den Router werden exakte zweite Momente beibehalten. Bei einem 6,78-Milliarden-Parameter-MoE nimmt der Optimiererzustand nur 1,29 GB ein (2,6 % von AdamW), der Spitzenwert des Trainingsspeichers sinkt von 81,4 GB auf 31,3 GB und kann auf einen 40-GB-Beschleuniger geladen werden.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: Ein hierarchisches Optimierer-Zustandszuweisungsstrategie für das MoE-Training vorgeschlagen - Aioga KI-News","description":"SkewAdam schlägt für MoE-Modelle eine hierarchische Optimiererzustandsaufteilung vor: Für das dichte Backbonenetzwerk werden Float32-Momente und zerlegte zweite Momente beibehalten...","url":"https://www.aioga.com/de/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:44:01.391Z"},"pt-BR":{"title":"SkewAdam: Propõe uma estratégia de alocação de estado do otimizador em camadas para o treinamento MoE","summary":"SkewAdam propôs uma alocação hierárquica do estado do otimizador para modelos MoE: mantém momentum em float32 e momentos de segunda ordem decompostos para a rede densa principal, mantém apenas momentos de segunda ordem decompostos para a camada de especialistas, e mantém momentos de segunda ordem precisos para o roteador. Em um MoE de 6,78 bilhões de parâmetros, o estado do otimizador ocupa apenas 1,29 GB (2,6% do AdamW), e a memória máxima de treinamento caiu de 81,4 GB para 31,3 GB, podendo ser carregada em um acelerador de 40 GB.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: Propõe uma estratégia de alocação de estado do otimizador em camadas para o treinamento MoE - Aioga Notícias de IA","description":"SkewAdam propôs uma alocação hierárquica do estado do otimizador para modelos MoE: mantém momentum em float32 e momentos de segunda ordem decompostos para a rede densa principal, m...","url":"https://www.aioga.com/pt-BR/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:44:50.262Z"},"ru":{"title":"SkewAdam: предложение стратегии распределения состояния оптимизатора для иерархической тренировки MoE","summary":"SkewAdam предложил распределение состояния оптимизатора по уровням для модели MoE: для плотной основной сети сохраняются моменты float32 и разложенные вторые моменты, для экспертного слоя сохраняются только разложенные вторые моменты, а для маршрутизатора сохраняются точные вторые моменты. На модели MoE с 6,78 миллиарда параметров состояние оптимизатора занимает всего 1,29 ГБ (2,6% от AdamW), пик использования памяти при обучении снизился с 81,4 ГБ до 31,3 ГБ, что позволяет загрузить модель на 40 ГБ ускоритель.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: предложение стратегии распределения состояния оптимизатора для иерархической тренировки MoE - Aioga Новости ИИ","description":"SkewAdam предложил распределение состояния оптимизатора по уровням для модели MoE: для плотной основной сети сохраняются моменты float32 и разложенные вторые моменты, для экспертно...","url":"https://www.aioga.com/ru/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:44:53.078Z"},"ar":{"title":"SkewAdam: اقتراح استراتيجية توزيع حالة المحسّن الهرمي لتدريب MoE","summary":"قدم SkewAdam توزيع حالات المحسن على شكل طبقات لنماذج MoE: احتفظ بزخم float32 والمصفوفة المربعة الثانية المفككة للشبكة الكثيفة الأساسية، واحتفظ فقط بالمصفوفة المربعة الثانية المفككة لطبقة الخبراء، واحتفظ بالمصفوفة المربعة الثانية الدقيقة للموجه. على نموذج MoE بمقدار 6.78 مليار معلمة، يحتل حالة المحسن مساحة 1.29 غيغابايت فقط (2.6% من AdamW)، وانخفضت ذروة ذاكرة التدريب من 81.4 غيغابايت إلى 31.3 غيغابايت، ويمكن تحميلها على مسرّع بسعة 40 غيغابايت.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: اقتراح استراتيجية توزيع حالة المحسّن الهرمي لتدريب MoE - Aioga أخبار الذكاء الاصطناعي","description":"قدم SkewAdam توزيع حالات المحسن على شكل طبقات لنماذج MoE: احتفظ بزخم float32 والمصفوفة المربعة الثانية المفككة للشبكة الكثيفة الأساسية، واحتفظ فقط بالمصفوفة المربعة الثانية المفككة...","url":"https://www.aioga.com/ar/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:45:38.736Z"},"hi":{"title":"SkewAdam: MoE प्रशिक्षण के लिए स्तरित ऑप्टिमाइज़र स्थिति आवंटन रणनीति प्रस्तुत की","summary":"SkewAdam ने MoE मॉडल के लिए स्तरित ऑप्टिमाइज़र स्टेट असाइनमेंट प्रस्तावित किया: डेंस बैकबोन नेटवर्क के लिए float32 मोमेंटम और फैक्टराइज्ड सेकंड-ऑर्डर मैट्रिक्स सुरक्षित रखें, विशेषज्ञ स्तर के लिए केवल फैक्टराइज्ड सेकंड-ऑर्डर मैट्रिक्स और राउटर के लिए सटीक सेकंड-ऑर्डर मैट्रिक्स सुरक्षित रखें। 6.78B पैरामीटर MoE पर, ऑप्टिमाइज़र स्टेट केवल 1.29 GB रखता है (AdamW का 2.6%), पीक ट्रेनिंग मेमोरी 81.4 GB से घटकर 31.3 GB हो गई, और इसे 40 GB एक्सेलेरेटर में डाल सकते हैं।","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: MoE प्रशिक्षण के लिए स्तरित ऑप्टिमाइज़र स्थिति आवंटन रणनीति प्रस्तुत की - Aioga AI समाचार","description":"SkewAdam ने MoE मॉडल के लिए स्तरित ऑप्टिमाइज़र स्टेट असाइनमेंट प्रस्तावित किया: डेंस बैकबोन नेटवर्क के लिए float32 मोमेंटम और फैक्टराइज्ड सेकंड-ऑर्डर मैट्रिक्स सुरक्षित रखें, विशेष...","url":"https://www.aioga.com/hi/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:45:43.747Z"},"it":{"title":"SkewAdam: Proposta di una strategia di assegnazione dello stato dell'ottimizzatore gerarchico per l'addestramento MoE","summary":"SkewAdam propone una distribuzione gerarchica dello stato dell'ottimizzatore per i modelli MoE: riserva il momentum in float32 e i secondi momenti fattorizzati per la rete backbone densa, conserva solo i secondi momenti fattorizzati per il livello degli esperti e mantiene i secondi momenti esatti per il router. Su un MoE da 6,78 miliardi di parametri, lo stato dell'ottimizzatore occupa solo 1,29 GB (il 2,6% di AdamW), la memoria di picco dell'addestramento scende da 81,4 GB a 31,3 GB, consentendo di caricarlo in un acceleratore da 40 GB.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: Proposta di una strategia di assegnazione dello stato dell'ottimizzatore gerarchico per l'addestramento MoE - Aioga Notizie IA","description":"SkewAdam propone una distribuzione gerarchica dello stato dell'ottimizzatore per i modelli MoE: riserva il momentum in float32 e i secondi momenti fattorizzati per la rete backbone...","url":"https://www.aioga.com/it/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:46:26.489Z"},"nl":{"title":"SkewAdam: Een gelaagde optimizer-statustoewijzingsstrategie voorstellen voor MoE-training","summary":"SkewAdam stelt een hiërarchische optimizer-statusallocatie voor MoE-modellen voor: behoud float32-momentum en gedecodeerde tweede-orde momenten voor het dense backbone-netwerk, behoud alleen gedecodeerde tweede-orde momenten voor de expertlaag, en behoud exacte tweede-orde momenten voor de router. Voor een 6,78 miljard parameter MoE neemt de optimizer-status slechts 1,29 GB in beslag (2,6% van AdamW), en het maximale trainingsgeheugen daalt van 81,4 GB naar 31,3 GB, waardoor het in een 40 GB accelerator kan worden geladen.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: Een gelaagde optimizer-statustoewijzingsstrategie voorstellen voor MoE-training - Aioga AI-nieuws","description":"SkewAdam stelt een hiërarchische optimizer-statusallocatie voor MoE-modellen voor: behoud float32-momentum en gedecodeerde tweede-orde momenten voor het dense backbone-netwerk, beh...","url":"https://www.aioga.com/nl/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:46:32.048Z"},"tr":{"title":"SkewAdam: MoE eğitimi için katmanlı optimize edici durum dağıtım stratejisi önerdi","summary":"SkewAdam, MoE modelleri için katmanlı bir optimize edici durumu tahsisi öneriyor: Yoğun omurga ağı için float32 momentum ve ayrıştırılmış ikinci momentleri koruyor, uzman katmanı için sadece ayrıştırılmış ikinci momentleri tutuyor, yönlendirici için ise hassas ikinci momentleri saklıyor. 6.78 milyar parametreli MoE üzerinde, optimize edici durumu yalnızca 1.29 GB (AdamW'nin %2.6'sı) yer kaplıyor, maksimum eğitim belleği 81.4 GB'den 31.3 GB'ye düşüyor ve 40 GB hızlandırıcıya sığabiliyor.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: MoE eğitimi için katmanlı optimize edici durum dağıtım stratejisi önerdi - Aioga AI Haberleri","description":"SkewAdam, MoE modelleri için katmanlı bir optimize edici durumu tahsisi öneriyor: Yoğun omurga ağı için float32 momentum ve ayrıştırılmış ikinci momentleri koruyor, uzman katmanı i...","url":"https://www.aioga.com/tr/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:47:20.064Z"},"vi":{"title":"SkewAdam: Đề xuất chiến lược phân bổ trạng thái bộ tối ưu phân cấp cho đào tạo MoE","summary":"SkewAdam đề xuất phân bổ trạng thái bộ tối ưu theo lớp cho mô hình MoE: giữ động lượng float32 và ma trận hạng hai phân rã cho mạng xương sống dày đặc, chỉ giữ ma trận hạng hai phân rã cho lớp chuyên gia, và giữ ma trận hạng hai chính xác cho bộ định tuyến. Trên MoE 6,78 tỷ tham số, trạng thái bộ tối ưu chỉ chiếm 1,29 GB (2,6% của AdamW), bộ nhớ đỉnh trong huấn luyện giảm từ 81,4 GB xuống còn 31,3 GB, có thể vừa với bộ tăng tốc 40 GB.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: Đề xuất chiến lược phân bổ trạng thái bộ tối ưu phân cấp cho đào tạo MoE - Tin tức AI Aioga","description":"SkewAdam đề xuất phân bổ trạng thái bộ tối ưu theo lớp cho mô hình MoE: giữ động lượng float32 và ma trận hạng hai phân rã cho mạng xương sống dày đặc, chỉ giữ ma trận hạng hai phâ...","url":"https://www.aioga.com/vi/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:47:23.217Z"},"id":{"title":"SkewAdam: Mengusulkan strategi distribusi status optimizer bertingkat untuk pelatihan MoE","summary":"SkewAdam mengusulkan distribusi status optimizer bertingkat untuk model MoE: mempertahankan momentum float32 dan momen kuadrat terdekomposisi untuk jaringan inti padat, hanya mempertahankan momen kuadrat terdekomposisi untuk lapisan ahli, dan mempertahankan momen kuadrat presisi untuk pengarah. Pada MoE dengan 6,78 miliar parameter, status optimizer hanya memakan 1,29 GB (2,6% dari AdamW), puncak memori pelatihan turun dari 81,4 GB menjadi 31,3 GB, sehingga bisa dimuat ke akselerator 40 GB.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: Mengusulkan strategi distribusi status optimizer bertingkat untuk pelatihan MoE - Berita AI Aioga","description":"SkewAdam mengusulkan distribusi status optimizer bertingkat untuk model MoE: mempertahankan momentum float32 dan momen kuadrat terdekomposisi untuk jaringan inti padat, hanya mempe...","url":"https://www.aioga.com/id/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:48:03.647Z"},"th":{"title":"SkewAdam: เสนอแนวทางการจัดสรรสถานะตัวปรับแต่งแบบหลายชั้นสำหรับการฝึก MoE","summary":"SkewAdam เสนอการจัดสรรสถานะตัวปรับแต่งแบบชั้นสำหรับโมเดล MoE: เก็บความเฉื่อย float32 และเมทริกซ์อันดับสองแยกของเครือข่ายหลักหนาแน่น สำหรับชั้นผู้เชี่ยวชาญเก็บเพียงเมทริกซ์อันดับสองแยก สำหรับตัวกำหนดเส้นทางเก็บเมทริกซ์อันดับสองที่แม่นยำ ใน MoE ที่มีพารามิเตอร์ 6.78B สถานะตัวปรับแต่งใช้เพียง 1.29 GB (2.6% ของ AdamW) หน่วยความจำการฝึกสูงสุดลดจาก 81.4 GB เหลือ 31.3 GB สามารถโหลดเข้าอุปกรณ์เร่งความเร็ว 40 GB ได้","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: เสนอแนวทางการจัดสรรสถานะตัวปรับแต่งแบบหลายชั้นสำหรับการฝึก MoE - ข่าว AI Aioga","description":"SkewAdam เสนอการจัดสรรสถานะตัวปรับแต่งแบบชั้นสำหรับโมเดล MoE: เก็บความเฉื่อย float32 และเมทริกซ์อันดับสองแยกของเครือข่ายหลักหนาแน่น สำหรับชั้นผู้เชี่ยวชาญเก็บเพียงเมทริกซ์อันดับสอง...","url":"https://www.aioga.com/th/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:48:16.560Z"},"pl":{"title":"SkewAdam: Propozycja strategii przydzielania stanów optymalizatora warstwowego dla treningu MoE","summary":"SkewAdam zaproponował hierarchiczne przydzielanie stanów optymalizatora dla modelu MoE: dla gęstej sieci szkieletowej zachowuje momenty float32 i rozłożone drugie momenty, dla warstw ekspertów zachowuje tylko rozłożone drugie momenty, a dla routera zachowuje dokładne drugie momenty. W modelu MoE o 6,78 mld parametrów stan optymalizatora zajmuje tylko 1,29 GB (2,6% w porównaniu z AdamW), a szczytowa pamięć treningowa spadła z 81,4 GB do 31,3 GB, co pozwala na załadowanie na akcelerator 40 GB.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SkewAdam: Propozycja strategii przydzielania stanów optymalizatora warstwowego dla treningu MoE - Aioga Wiadomości AI","description":"SkewAdam zaproponował hierarchiczne przydzielanie stanów optymalizatora dla modelu MoE: dla gęstej sieci szkieletowej zachowuje momenty float32 i rozłożone drugie momenty, dla wars...","url":"https://www.aioga.com/pl/news/cmrvs7s6600cobipzahp9o5lb/","contentTranslated":true,"sourceHash":"8b287da15f961afb","translatedAt":"2026-07-23T03:49:03.646Z"}}}}