Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 3 days ago • 326
nguyenthilaitrieulong/DeepSeek-R1-Distill-Llama-70B-abliterated Text Generation • 71B • Updated 9 days ago • 441 • • 3
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 14 days ago • 302
baohao/OPD_Search_Qwen3-8B_SFT-RL_to_Search_Qwen3-4B-Instruct-2507_SFT 4B • Updated 20 days ago • 21 • 1
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 28 days ago • 212
Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES Paper • 2607.05691 • Published Jul 6 • 4
TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning Paper • 2606.32017 • Published Jun 30 • 12
ENPIRE: Agentic Robot Policy Self-Improvement in the Real World Paper • 2606.19980 • Published Jun 18 • 15