SurviAGI
← 全部更新
Tencent · 2026-09-24 · 研究成果

Tencent Hunyuan research extends critical-batch-size theory to online LLM RL原文

Tencent Hunyuan published new research revisiting classical critical-batch-size theory and extending it to online LLM reinforcement learning.

这条更新涉及 1 条赛道的 1 项工作,最高采信到 L1 辅助。

最高采信
L1辅助
工作
1
赛道
1
浏览
1.44 万

它涉及的工作

级别高的在前。点开一项工作,查看推动它的其他更新。

发布于