{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-23T08:01:28.298Z","headline":"SLPO：通过代理策略实现潜在推理的规模化扩展","description":"四川大学团队提出 Surrogate Latent Policy Optimization （SLPO），将结果奖励强化学习引入自回归潜在推理模型。SLPO 通过定义潜在过渡的代理策略密度实现轨迹级信用分配，并利用正确性监督的停止头使推理步长可学习。实验表明，SLPO 在连续和软思考设置下均能提升 Pass@ 指标，并为更难的问题分配更长的潜在计算。","url":"https://www.aioga.com/news/cmrwwy3d8038arobhj9nbyrbj/","mainEntityOfPage":"https://www.aioga.com/news/cmrwwy3d8038arobhj9nbyrbj/","datePublished":"2026-07-22T00:00:00.000Z","dateModified":"2026-07-22T00:00:00.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://arxiv.org/abs/2607.19691","https://aihot.virxact.com/items/cmrwwy3d8038arobhj9nbyrbj"],"canonicalUrl":"https://www.aioga.com/news/cmrwwy3d8038arobhj9nbyrbj/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：四川大学团队提出 Surrogate Latent Policy Optimization （SLPO），将结果奖励强化学习引入自回归潜在推理模型。 Aioga 将其归入「论文研究」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cmrwwy3d8038arobhj9nbyrbj/","dateCreated":"2026-07-22T00:00:00.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"arXiv source article","url":"https://arxiv.org/abs/2607.19691","datePublished":"2026-07-22T00:00:00.000Z","provider":{"@type":"Organization","name":"arXiv","url":"https://arxiv.org/abs/2607.19691"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrwwy3d8038arobhj9nbyrbj","datePublished":"2026-07-22T00:00:00.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrwwy3d8038arobhj9nbyrbj"}}],"aggregationSource":"HuggingFace Daily Papers（社区热门论文）","originalPublisher":{"name":"arXiv","url":"https://arxiv.org/abs/2607.19691"},"article":{"id":"cmrwwy3d8038arobhj9nbyrbj","slug":"cmrwwy3d8038arobhj9nbyrbj","url":"https://www.aioga.com/news/cmrwwy3d8038arobhj9nbyrbj/","title":"SLPO：通过代理策略实现潜在推理的规模化扩展","title_en":"SLPO： Scaling Latent Reasoning via a Surrogate Policy","summary":"四川大学团队提出 Surrogate Latent Policy Optimization （SLPO），将结果奖励强化学习引入自回归潜在推理模型。SLPO 通过定义潜在过渡的代理策略密度实现轨迹级信用分配，并利用正确性监督的停止头使推理步长可学习。实验表明，SLPO 在连续和软思考设置下均能提升 Pass@ 指标，并为更难的问题分配更长的潜在计算。","source":"HuggingFace Daily Papers（社区热门论文）","sourceUrl":"https://arxiv.org/abs/2607.19691","aiHotUrl":"https://aihot.virxact.com/items/cmrwwy3d8038arobhj9nbyrbj","publishedAt":"2026-07-22T00:00:00.000Z","category":"论文研究","score":51,"selected":false,"articleBody":["arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.","Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs ：https://info.arxiv.org/labs/index.html."],"articleImages":[{"sourceUrl":"https://arxiv.org/static/base/1.0.1/images/funders/simons-foundation.png","alt":"Simons Foundation","afterParagraph":1,"url":"/media/articles/cmrwwy3d8038arobhj9nbyrbj/e2d7f38d62f5ca91.png"},{"sourceUrl":"https://arxiv.org/static/base/1.0.1/images/funders/simons-foundation-international.png","alt":"Simons Foundation International","afterParagraph":1,"url":"/media/articles/cmrwwy3d8038arobhj9nbyrbj/1d56e29c5557cbdc.png"}],"mediaStatus":"ok","articleBodyZh":["arXivLabs 是一个框架，允许合作者直接在我们的网站上开发和分享新的 arXiv 功能。","有一个可以为 arXiv 社区增加价值的项目想法吗？了解更多关于 arXivLabs 的信息：https://info.arxiv.org/labs/index.html。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：四川大学团队提出 Surrogate Latent Policy Optimization （SLPO），将结果奖励强化学习引入自回归潜在推理模型。 Aioga 将其归入「论文研究」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：模型与研究类动态需要结合能力边界、开放方式、成本、可用性和真实任务表现判断，单项指标领先不等于已经形成稳定采用。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察官方文档、实际可用性、价格变化、开发者反馈和竞品回应。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-23T08:10:16.706Z","sourceHash":"8364c747ba1e1749","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["论文研究","HuggingFace Daily Papers（社区热门论文）"],"translations":{"zh-CN":{"title":"SLPO：通过代理策略实现潜在推理的规模化扩展","summary":"四川大学团队提出 Surrogate Latent Policy Optimization （SLPO），将结果奖励强化学习引入自回归潜在推理模型。SLPO 通过定义潜在过渡的代理策略密度实现轨迹级信用分配，并利用正确性监督的停止头使推理步长可学习。实验表明，SLPO 在连续和软思考设置下均能提升 Pass@ 指标，并为更难的问题分配更长的潜在计算。","category":"论文研究","source":"arXiv","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO：通过代理策略实现潜在推理的规模化扩展 - Aioga AI资讯","description":"四川大学团队提出 Surrogate Latent Policy Optimization （SLPO），将结果奖励强化学习引入自回归潜在推理模型。SLPO 通过定义潜在过渡的代理策略密度实现轨迹级信用分配，并利用正确性监督的停止头使推理步长可学习。实验表明，SLPO 在连续和软思考设置下均能提升 Pass@ 指标，并为更难的问题分配更长的潜在计算。","url":"https://www.aioga.com/news/cmrwwy3d8038arobhj9nbyrbj/"},"en":{"title":"SLPO: Scaling Up Latent Reasoning Through Proxy Policies","summary":"The team from Sichuan University proposed Surrogate Latent Policy Optimization (SLPO), introducing result-based reinforcement learning into autoregressive latent reasoning models. SLPO achieves trajectory-level credit assignment by defining a surrogate policy density over latent transitions and uses a correctness-supervised stopping head to make inference steps learnable. Experiments show that SLPO can improve the Pass@ metric in both continuous and soft-thinking settings, and allocates longer latent computation to more difficult problems.","category":"Research","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: Scaling Up Latent Reasoning Through Proxy Policies - Aioga AI News","description":"The team from Sichuan University proposed Surrogate Latent Policy Optimization (SLPO), introducing result-based reinforcement learning into autoregressive latent reasoning models....","url":"https://www.aioga.com/en/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:03:06.730Z"},"ja":{"title":"SLPO：エージェント戦略によって潜在推論の規模拡張を実現","summary":"四川大学のチームは Surrogate Latent Policy Optimization（SLPO）を提案し、結果報酬の強化学習を自己回帰潜在推論モデルに導入しました。SLPO は潜在遷移の代理方策密度を定義することで軌跡レベルのクレジット配分を実現し、正確性監督の停止ヘッドを利用して推論ステップ長を学習可能にします。実験により、SLPO は連続およびソフト思考設定の下で Pass@ 指標を向上させ、より難しい問題にはより長い潜在計算を割り当てることが示されました。","category":"論文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO：エージェント戦略によって潜在推論の規模拡張を実現 - Aioga AIニュース","description":"四川大学のチームは Surrogate Latent Policy Optimization（SLPO）を提案し、結果報酬の強化学習を自己回帰潜在推論モデルに導入しました。SLPO は潜在遷移の代理方策密度を定義することで軌跡レベルのクレジット配分を実現し、正確性監督の停止ヘッドを利用して推論ステップ長を学習可能にします。実験により、SLPO は連続およびソ...","url":"https://www.aioga.com/ja/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:03:22.589Z"},"ko":{"title":"SLPO: 에이전트 전략을 통해 잠재 추론의 규모화 확장을 실현","summary":"쓰촨대학교 팀은 Surrogate Latent Policy Optimization(SLPO)를 제안했으며, 결과 보상 강화 학습을 자기회귀 잠재 추론 모델에 도입했습니다. SLPO는 잠재 전이의 대리 정책 밀도를 정의하여 궤적 수준의 신용 할당을 실현하고, 정답 감독을 통한 중지 헤드를 활용하여 추론 단계 길이를 학습 가능하게 합니다. 실험 결과, SLPO는 연속 및 부드러운 사고 설정 모두에서 Pass@ 지표를 향상시키며, 더 어려운 문제에 대해 더 긴 잠재 계산을 할당합니다.","category":"연구","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: 에이전트 전략을 통해 잠재 추론의 규모화 확장을 실현 - Aioga AI 뉴스","description":"쓰촨대학교 팀은 Surrogate Latent Policy Optimization(SLPO)를 제안했으며, 결과 보상 강화 학습을 자기회귀 잠재 추론 모델에 도입했습니다. SLPO는 잠재 전이의 대리 정책 밀도를 정의하여 궤적 수준의 신용 할당을 실현하고, 정답 감독을 통한 중지 헤드를 활용하여 추론 단계 길이를 학습 가...","url":"https://www.aioga.com/ko/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:04:14.255Z"},"es":{"title":"SLPO: Escalando la inferencia latente a gran escala mediante estrategias de agencia","summary":"El equipo de la Universidad de Sichuan propuso la Optimización de Políticas Latentes Sustitutas (SLPO), que introduce el aprendizaje por refuerzo basado en recompensas en modelos de inferencia latente autorregresivos. SLPO logra la asignación de crédito a nivel de trayectoria definiendo la densidad de política sustituta de la transición latente, y utiliza una cabeza de detención supervisada por corrección para hacer que los pasos de inferencia sean aprendibles. Los experimentos muestran que SLPO mejora el indicador Pass@ tanto en configuraciones continuas como de pensamiento blando, y asigna más cómputo latente a problemas más difíciles.","category":"Investigación","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: Escalando la inferencia latente a gran escala mediante estrategias de agencia - Aioga Noticias de IA","description":"El equipo de la Universidad de Sichuan propuso la Optimización de Políticas Latentes Sustitutas (SLPO), que introduce el aprendizaje por refuerzo basado en recompensas en modelos d...","url":"https://www.aioga.com/es/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:04:09.352Z"},"fr":{"title":"SLPO : Extension à grande échelle du raisonnement latent grâce à la stratégie d'agent","summary":"L'équipe de l'Université du Sichuan a proposé l'Optimisation de Politique Latente de Substitut (SLPO), introduisant l'apprentissage par renforcement basé sur la récompense de résultat dans les modèles de raisonnement latent autoregressif. SLPO réalise l'allocation de crédit au niveau des trajectoires en définissant la densité de politique de l'agent pour les transitions latentes, et utilise une tête d'arrêt supervisée par la correction pour rendre la longueur des étapes de raisonnement apprenable. Les expériences montrent que SLPO améliore l'indicateur Pass@ dans des configurations continues et de réflexion douce, et attribue un calcul latent plus long pour des problèmes plus difficiles.","category":"Recherche","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO : Extension à grande échelle du raisonnement latent grâce à la stratégie d'agent - Aioga Actualités IA","description":"L'équipe de l'Université du Sichuan a proposé l'Optimisation de Politique Latente de Substitut (SLPO), introduisant l'apprentissage par renforcement basé sur la récompense de résul...","url":"https://www.aioga.com/fr/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:04:58.658Z"},"de":{"title":"SLPO: Skalierbare Erweiterung des latenten Schließens durch Agentenstrategien","summary":"Ein Team der Sichuan-Universität schlug das Surrogate Latent Policy Optimization (SLPO) vor, das Ergebnisbelohnungsverstärkendes Lernen in autoregressive latente Inferenzmodelle einführt. SLPO realisiert eine Trajektorien-Ebenen-Kreditvergabe durch die Definition der Proxy-Politikdichte latenter Übergänge und nutzt einen Correctness-Supervised-Stop-Head, um die Inferenzschritte lernbar zu machen. Experimente zeigen, dass SLPO sowohl in kontinuierlichen als auch in Soft-Think-Einstellungen den Pass@-Indikator verbessern kann und für schwierigere Aufgaben längere latente Berechnungen zuweist.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: Skalierbare Erweiterung des latenten Schließens durch Agentenstrategien - Aioga KI-News","description":"Ein Team der Sichuan-Universität schlug das Surrogate Latent Policy Optimization (SLPO) vor, das Ergebnisbelohnungsverstärkendes Lernen in autoregressive latente Inferenzmodelle ei...","url":"https://www.aioga.com/de/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:05:04.133Z"},"pt-BR":{"title":"SLPO: Expansão em escala do raciocínio potencial através de estratégias de agente","summary":"A equipe da Universidade de Sichuan propôs o Surrogate Latent Policy Optimization (SLPO), que introduz reforço de recompensa de resultados em modelos de inferência latent autoregressivos. O SLPO realiza a atribuição de crédito em nível de trajetória definindo a densidade da política substituta para transições latentes e utiliza uma cabeça de parada supervisionada pela correção para tornar os passos de inferência aprendíveis. Experimentos mostram que o SLPO melhora o indicador Pass@ tanto em configurações contínuas quanto em pensamento suave, e aloca cálculos latentes mais longos para problemas mais difíceis.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: Expansão em escala do raciocínio potencial através de estratégias de agente - Aioga Notícias de IA","description":"A equipe da Universidade de Sichuan propôs o Surrogate Latent Policy Optimization (SLPO), que introduz reforço de recompensa de resultados em modelos de inferência latent autoregre...","url":"https://www.aioga.com/pt-BR/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:05:46.679Z"},"ru":{"title":"SLPO: Масштабирование потенциального вывода с помощью стратегии агентов","summary":"Команда Сычуаньского университета предложила метод Surrogate Latent Policy Optimization (SLPO), вводящий усиленное обучение с результативной наградой в авторегрессионные модели латентного вывода. SLPO реализует распределение кредитов на уровне траектории через определение плотности прокси-политики для латентных переходов и использует контрольный механизм остановки с корректностью для обучения числа шагов вывода. Эксперименты показывают, что SLPO повышает показатель Pass@ как в непрерывных, так и в мягких режимах размышлений, распределяя более длительные латентные вычисления для более сложных задач.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: Масштабирование потенциального вывода с помощью стратегии агентов - Aioga Новости ИИ","description":"Команда Сычуаньского университета предложила метод Surrogate Latent Policy Optimization (SLPO), вводящий усиленное обучение с результативной наградой в авторегрессионные модели лат...","url":"https://www.aioga.com/ru/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:05:56.156Z"},"ar":{"title":"SLPO: تحقيق التوسع الكمي للتفكير المحتمل من خلال سياسة الوكيل","summary":"اقترحت فريق جامعة سيتشوان أسلوب تحسين السياسة الكامنة البديلة (SLPO)، حيث يتم إدخال تعلم التعزيز القائم على مكافآت النتائج في نموذج الاستدلال الكامن الذاتي التراجعي. يقوم SLPO بتوزيع المصداقية على مستوى المسار من خلال تعريف كثافة سياسة وكيلة للانتقالات الكامنة، ويستفيد من رأس التوقف المشرف على الصحة لجعل طول خطوات الاستدلال قابلاً للتعلم. تظهر التجارب أن SLPO يحسن مؤشر Pass@ في كل من الإعدادات المستمرة وذات التفكير الناعم، ويوزع حسابًا كامنًا أطول على المشكلات الأكثر صعوبة.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: تحقيق التوسع الكمي للتفكير المحتمل من خلال سياسة الوكيل - Aioga أخبار الذكاء الاصطناعي","description":"اقترحت فريق جامعة سيتشوان أسلوب تحسين السياسة الكامنة البديلة (SLPO)، حيث يتم إدخال تعلم التعزيز القائم على مكافآت النتائج في نموذج الاستدلال الكامن الذاتي التراجعي. يقوم SLPO بتوز...","url":"https://www.aioga.com/ar/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:06:39.711Z"},"hi":{"title":"SLPO：एजेंट नीति के माध्यम से संभाव्य तर्क के पैमाने पर विस्तार को प्राप्त करना","summary":"सिचुआन विश्वविद्यालय की टीम ने Surrogate Latent Policy Optimization (SLPO) पेश किया, जो परिणाम-इनाम सुदृढ़न सीखने को स्व-प्रतिगामी गुप्त तर्क मॉडल में शामिल करता है। SLPO ट्रैजेक्टरी-स्तरीय क्रेडिट आवंटन को लागू करने के लिए गुप्त संक्रमण के एजेंट नीति घनत्व को परिभाषित करता है, और सहीपन निगरानी वाले स्टॉप हेड का उपयोग करके तर्क चरणों को सीखने योग्य बनाता है। प्रयोगों से पता चलता है कि SLPO निरंतर और सॉफ़्ट थिंकिंग सेटिंग्स दोनों में Pass@ सूचकांक बढ़ा सकता है, और कठिन समस्याओं के लिए लंबी गुप्त गणना आवंटित करता है।","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO：एजेंट नीति के माध्यम से संभाव्य तर्क के पैमाने पर विस्तार को प्राप्त करना - Aioga AI समाचार","description":"सिचुआन विश्वविद्यालय की टीम ने Surrogate Latent Policy Optimization (SLPO) पेश किया, जो परिणाम-इनाम सुदृढ़न सीखने को स्व-प्रतिगामी गुप्त तर्क मॉडल में शामिल करता है। SLPO ट्रैजेक्ट...","url":"https://www.aioga.com/hi/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:06:53.042Z"},"it":{"title":"SLPO: Espansione su larga scala del ragionamento potenziale attraverso strategie di delega","summary":"Il team dell'Università del Sichuan ha proposto il Surrogate Latent Policy Optimization (SLPO), introducendo l'apprendimento per rinforzo basato sulla ricompensa dei risultati nei modelli di ragionamento latente autoregressivi. SLPO realizza l'assegnazione del credito a livello di traiettoria definendo la densità della politica sostitutiva per le transizioni latenti e utilizza una testa di interruzione supervisionata per la correttezza, rendendo apprese le lunghezze dei passaggi di ragionamento. Esperimenti mostrano che SLPO può migliorare l'indice Pass@ sia in contesti continui che di soft thinking, assegnando calcoli latenti più lunghi a problemi più difficili.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: Espansione su larga scala del ragionamento potenziale attraverso strategie di delega - Aioga Notizie IA","description":"Il team dell'Università del Sichuan ha proposto il Surrogate Latent Policy Optimization (SLPO), introducendo l'apprendimento per rinforzo basato sulla ricompensa dei risultati nei...","url":"https://www.aioga.com/it/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:07:41.986Z"},"nl":{"title":"SLPO: Schaalbare uitbreiding van latente redenering via beleid代理","summary":"Het team van de Sichuan Universiteit stelde Surrogate Latent Policy Optimization (SLPO) voor, waarmee resultaatbeloning in versterkend leren wordt geïntroduceerd in autoregressieve latente inferentiemodellen. SLPO realiseert trajectniveau krediettoewijzing door de agentstrategie-dichtheid van latente transities te definiëren en maakt de inferentiestappen leerbaar met behulp van een correctheidsgebaseerd stopmechanisme. Experimenten tonen aan dat SLPO de Pass@-score kan verbeteren in zowel continue als zachte gedachte-instellingen, en langere latente berekening toekent voor moeilijkere problemen.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: Schaalbare uitbreiding van latente redenering via beleid代理 - Aioga AI-nieuws","description":"Het team van de Sichuan Universiteit stelde Surrogate Latent Policy Optimization (SLPO) voor, waarmee resultaatbeloning in versterkend leren wordt geïntroduceerd in autoregressieve...","url":"https://www.aioga.com/nl/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:07:40.886Z"},"tr":{"title":"SLPO: Temsilci stratejisiyle potansiyel akıl yürütmenin ölçeklenebilir genişletilmesi","summary":"Sichuan Üniversitesi ekibi, sonuç ödül takviyeli öğrenmeyi kendi kendine geribildirimli gizli akıl yürütme modeline dahil eden Surrogate Latent Policy Optimization (SLPO) yöntemini önerdi. SLPO, iz düzeyinde kredi dağılımını gerçekleştirmek için gizli geçişlerin vekil politika yoğunluğunu tanımlar ve akıl yürütme adım boyutunu öğrenilebilir kılmak için doğruluk denetimli durdurma başlığını kullanır. Deneyler, SLPO'nun hem sürekli hem de yumuşak düşünme ayarlarında Pass@ metriğini artırabildiğini ve daha zor problemlere daha uzun gizli hesaplama atadığını göstermektedir.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: Temsilci stratejisiyle potansiyel akıl yürütmenin ölçeklenebilir genişletilmesi - Aioga AI Haberleri","description":"Sichuan Üniversitesi ekibi, sonuç ödül takviyeli öğrenmeyi kendi kendine geribildirimli gizli akıl yürütme modeline dahil eden Surrogate Latent Policy Optimization (SLPO) yöntemini...","url":"https://www.aioga.com/tr/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:08:28.744Z"},"vi":{"title":"SLPO: Mở rộng quy mô suy luận tiềm năng thông qua chiến lược đại lý","summary":"Nhóm nghiên cứu của Đại học Tứ Xuyên đã đề xuất Surrogate Latent Policy Optimization (SLPO), đưa phần thưởng kết quả trong học tăng cường vào mô hình suy luận tiềm ẩn tự hồi quy. SLPO thực hiện phân bổ tín nhiệm ở mức quỹ đạo thông qua việc định nghĩa mật độ chính sách đại diện cho các chuyển đổi tiềm ẩn, và sử dụng đầu dừng với giám sát tính đúng đắn để làm cho bước suy luận có thể học được. Thí nghiệm cho thấy, SLPO có thể cải thiện chỉ số Pass@ cả trong cài đặt liên tục và suy nghĩ mềm, đồng thời phân bổ tính toán tiềm ẩn dài hơn cho các vấn đề khó hơn.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: Mở rộng quy mô suy luận tiềm năng thông qua chiến lược đại lý - Tin tức AI Aioga","description":"Nhóm nghiên cứu của Đại học Tứ Xuyên đã đề xuất Surrogate Latent Policy Optimization (SLPO), đưa phần thưởng kết quả trong học tăng cường vào mô hình suy luận tiềm ẩn tự hồi quy. S...","url":"https://www.aioga.com/vi/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:08:30.900Z"},"id":{"title":"SLPO: Mewujudkan ekspansi skala inferensi potensial melalui strategi agen","summary":"Tim dari Universitas Sichuan mengajukan Surrogate Latent Policy Optimization (SLPO), yang memperkenalkan pembelajaran penguatan berbasis hasil ke dalam model penalaran laten autoregresif. SLPO mencapai distribusi kredit tingkat trajektori dengan mendefinisikan densitas kebijakan agen dari transisi laten, dan menggunakan kepala penghentian yang diawasi oleh kebenaran untuk membuat langkah penalaran dapat dipelajari. Eksperimen menunjukkan bahwa SLPO dapat meningkatkan metrik Pass@ baik dalam pengaturan kontinu maupun pemikiran lunak, serta mengalokasikan perhitungan laten yang lebih lama untuk masalah yang lebih sulit.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: Mewujudkan ekspansi skala inferensi potensial melalui strategi agen - Berita AI Aioga","description":"Tim dari Universitas Sichuan mengajukan Surrogate Latent Policy Optimization (SLPO), yang memperkenalkan pembelajaran penguatan berbasis hasil ke dalam model penalaran laten autore...","url":"https://www.aioga.com/id/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:09:17.975Z"},"th":{"title":"SLPO: การขยายการอนุมานเชิงลึกอย่างเป็นระบบโดยใช้กลยุทธ์ตัวแทน","summary":"ทีมงานของมหาวิทยาลัยเสฉวนได้เสนอ Surrogate Latent Policy Optimization (SLPO) ซึ่งนำรางวัลการเรียนรู้แบบเสริมเข้ามาใช้กับโมเดลการอนุมานแฝงแบบออโต้รีเกรสซีฟ SLPO ทำการจัดสรรเครดิตระดับเส้นทางโดยการกำหนดความหนาแน่นของนโยบายตัวแทนของการเปลี่ยนแปลงแฝง และใช้หัวหยุดที่ควบคุมความถูกต้องเพื่อให้ขั้นตอนการอนุมานสามารถเรียนรู้ได้ การทดลองแสดงให้เห็นว่า SLPO สามารถปรับปรุงตัวชี้วัด Pass@ ทั้งในการตั้งค่าแบบต่อเนื่องและการคิดแบบนุ่มนวล และจัดสรรการคำนวณแฝงที่ยาวขึ้นสำหรับปัญหาที่ยากกว่า","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: การขยายการอนุมานเชิงลึกอย่างเป็นระบบโดยใช้กลยุทธ์ตัวแทน - ข่าว AI Aioga","description":"ทีมงานของมหาวิทยาลัยเสฉวนได้เสนอ Surrogate Latent Policy Optimization (SLPO) ซึ่งนำรางวัลการเรียนรู้แบบเสริมเข้ามาใช้กับโมเดลการอนุมานแฝงแบบออโต้รีเกรสซีฟ SLPO ทำการจัดสรรเครดิตระด...","url":"https://www.aioga.com/th/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:09:23.639Z"},"pl":{"title":"SLPO: Skala rozszerzenia potencjalnego wnioskowania poprzez strategię reprezentowania","summary":"Zespół Uniwersytetu Syczuanu zaproponował Surrogate Latent Policy Optimization (SLPO), wprowadzając wzmocnione uczenie nagród wyników do autoregresyjnego modelu wnioskowania latentnego. SLPO realizuje przydzielanie kredytu na poziomie trajektorii, definiując gęstość polityki zastępczej dla przejść latentnych, oraz korzysta z nadzorowanej głowy zatrzymania poprawności, aby umożliwić uczenie kroków wnioskowania. Eksperymenty wykazały, że SLPO może zwiększyć wskaźnik Pass@ zarówno w ustawieniach ciągłych, jak i w miękkim myśleniu, a także przydziela dłuższe obliczenia latentne dla trudniejszych problemów.","category":"论文研究","source":"HuggingFace Daily Papers（社区热门论文）","aggregationSource":"HuggingFace Daily Papers（社区热门论文）","pageTitle":"SLPO: Skala rozszerzenia potencjalnego wnioskowania poprzez strategię reprezentowania - Aioga Wiadomości AI","description":"Zespół Uniwersytetu Syczuanu zaproponował Surrogate Latent Policy Optimization (SLPO), wprowadzając wzmocnione uczenie nagród wyników do autoregresyjnego modelu wnioskowania latent...","url":"https://www.aioga.com/pl/news/cmrwwy3d8038arobhj9nbyrbj/","contentTranslated":true,"sourceHash":"e6676fb14b96fc13","translatedAt":"2026-07-23T08:10:12.599Z"}}}}