EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning Paper • 2608.06197 • Published 7 days ago • 43
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Paper • 2608.05987 • Published 7 days ago • 92
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Paper • 2608.01837 • Published 10 days ago • 38
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search Paper • 2607.24280 • Published 17 days ago • 83
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 28 days ago • 106
TACO: Tool-Augmented Credit Optimization for Agentic Tool Use Paper • 2606.30251 • Published Jun 29 • 22
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning Paper • 2606.26790 • Published Jun 25 • 57
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 158
RobotEQ: Transitioning from Passive Intelligence to Active Intelligence in Embodied AI Paper • 2605.06234 • Published May 7 • 4
Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation Paper • 2606.09131 • Published Jun 8 • 3
Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles Paper • 2605.22177 • Published May 21 • 21
From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills Paper • 2604.24026 • Published Apr 27 • 22
KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation Paper • 2604.08455 • Published Apr 9 • 48
SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization Paper • 2604.02268 • Published Apr 2 • 103
HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household Scenarios Paper • 2603.11975 • Published Mar 12 • 12
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Execution-Free Reward Models Paper • 2602.17684 • Published Feb 4 • 23
Query as Anchor: Scenario-Adaptive User Representation via Large Language Model Paper • 2602.14492 • Published Feb 16 • 18