Dongsheng Ma, Jiayu Li, Zhengren Wang, Yijie Wang, Jiahao Kong
Multimodal Large Language Models (MLLMs) have significantly advanced document understanding, yet current Doc-VQA evaluations score only the final answer and leave the supporting evidence unchecked. This answer-only…
huggingface_papers
Taewon Yun, Jisu Shin, Jeonghwan Choi, Seunghwan Bang, Hwanjun Song
Distilling large reasoning models is essential for making Long-CoT reasoning practical, as full-scale inference remains computationally prohibitive. Existing curation-based approaches select complete reasoning traces…
huggingface_papers
Yuchen Cai, Ding Cao, Liang Lin, Chunxi Luo, Xin Xu
On-policy distillation (OPD) has emerged as an efficient post-training paradigm for large language models. However, existing studies largely attribute this advantage to denser and more stable supervision, while the…
huggingface_papers
Hanxun Yu, Xuan Qu, Yuxin Wang, Jianke Zhu, Lei ke
Vision-Language Models (VLMs) excel at 2D tasks such as grounding and captioning, yet remain limited in 3D understanding. A key limitation is their text-only supervision paradigm, which under-constrains fine-grained…
huggingface_papers
Quanjian Song, Yefeng Shen, Mengting Chen, Hao Sun, Jinsong Lan
Human-centric video customization, particularly at the garment level, has shown significant commercial value. However, existing approaches cannot support low-latency and interactive garment control, which is crucial for…
huggingface_papers
Mengjie Ren, Jie Lou, Boxi Cao, Xueru Wen, Hongyu Lin
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an effective paradigm for improving the reasoning capabilities of large language models. However, RLVR training is often hindered by sparse binary…
huggingface_papers
Chanuk Lee, Sangwoo Park, Minki Kang, Sung Ju Hwang
Reinforcement learning with verifiable rewards (RLVR) has emerged as a scalable paradigm for improving the reasoning capabilities of large language models. However, its effectiveness is fundamentally limited by…
huggingface_papers
Jingxuan Wei, Xi Bai, Shan Liu, Caijun Jia, Zheng Sun
Large vision-language models have significantly advanced GUI agents, enabling executable interaction across web, mobile, and desktop interfaces. Yet these gains largely rely on a forgiving region-tolerant paradigm,…
huggingface_papers
Han Li, Jinyu Tian, Rili Feng, Yuqiao Du, Chong Zheng
Large language models (LLMs) still struggle with the rigorous reasoning demands of hard competitive programming. While recent multi-agent frameworks attempt to bridge this reliability gap, they remain fundamentally…
huggingface_papers
Alberto Pepe, Chien-Yu Lin, Despoina Magka, Bilge Acun, Yannan Nellie Wu
Toward recursive self-improvement, we investigate LLM agents autonomously designing foundation models beyond standard Transformers. We introduce a dual-framework approach: AIRA-Compose for high-level architecture…
huggingface_papers
Xiaoxuan He, Siming Fu, Zeyue Xue, Weijie Wang, Ruizhe He
Group Relative Policy Optimization has emerged as essential for aligning video diffusion models with human preferences, but faces a critical computational bottleneck: training a 14B parametered model typically demands…
huggingface_papers
Yang Yue, Fangyun Wei, Tianyu He, Jinjing Zhao, Zanlin Ni
Text and faces are among the most perceptually salient and practically important patterns in visual generation, yet they remain challenging for autoregressive generators built on discrete tokenization. A central…
huggingface_papers
Ziang Ye, Wentao Shi, Yuxin Liu, Yu Wang, Zhengzhou Cai
Large language model based agents often fail in unfamiliar environments due to premature exploitation: a tendency to act on prior knowledge before acquiring sufficient environment-specific information. We identify…
huggingface_papers
Devin Yasith De Silva, Dhaval Patel, Christodoulos Constantinides, Shuxin Lin, Nianjun Zhou
Monitoring complex industrial assets relies on engineer-authored symbolic rules that trigger based on sensor conditions and prompt technicians to perform corrective actions. The bottleneck is not detection but response:…
huggingface_papers