Yuqiao Tan, Minzheng Wang, Bo Liu, Zichen Liu, Tian Liang, Shizhu He, Jun Zhao, Kang Liu
While reinforcement learning with verifiable rewards (RLVR) significantly enhances LLM reasoning by optimizing the conditional distribution P(y|x), its potential is fundamentally bounded by the base…
cs.LGcs.AIcs.CL
Sumeet Ramesh Motwani, Daniel Nichols, Charles London et al.
As language models are increasingly deployed for complex autonomous tasks, their ability to reason accurately over longer horizons becomes critical.
cs.LGcs.AI
Itay Itzhak, Eliya Habba, Gabriel Stanovsky, Yonatan Belinkov
Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness.
cs.CLcs.AIcs.LG
Louie Hong Yao, Vishesh Anand, Yuan Zhuang, Tianyu Jiang
Rhetorical questions are asked not to seek information but to persuade or signal stance. How large language models internally represent them remains unclear.
cs.CLcs.AIcs.LG
Tianshuo Yang, Guanyu Chen, Yutian Chen et al.
While end-to-end Vision-Language-Action (VLA) models offer a promising paradigm for robotic manipulation, fine-tuning them on narrow control data often compromises the profound reasoning capabilities…
cs.CVcs.AIcs.RO
Fei Tang, Bofan Chen, Zhengxi Lu et al.
GUI grounding, which localizes interface elements from screenshots given natural language queries, remains challenging for small icons and dense layouts.
cs.CVcs.AIcs.CL
Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, Alborz Geramifard
On-policy knowledge distillation (OPD) trains a student on its own rollouts under token-level supervision from a teacher.
cs.LGcs.AI
Kavya Gupta, Nektarios Kalampalikis, Christoph Heitz, Isabel Valera
Fairness in algorithmic decision-making is often defined in the predictive space, where predictive performance - used as a proxy for decision-maker (DM) utility - is traded off against…
cs.LGcs.AI
Zheyu Zhang, Ziqi Pang, Shixing Chen, Xiang Hao, Vimal Bhat, Yu-Xiong Wang
Long video understanding is inherently challenging for vision-language models (VLMs) because of the extensive number of frames.
cs.CV
Song Tang, Guangquan Jie, Henghui Ding, Yu-Gang Jiang
Existing segmentation models based on multimodal large language models (MLLMs), such as LISA, often struggle with novel or emerging entities due to their inability to incorporate up-to-date…
cs.CV
Dinging Li, Yingxiu Zhao, Xinrui Cheng et al.
Spatial reasoning over three-dimensional scenes is a core capability for embodied intelligence, yet continuous model improvement remains bottlenecked by the cost of geometric annotation.
cs.CVcs.CL
Erjia Yan, Chaoqun Ni
Generative AI systems such as ChatGPT are increasingly used in scientific writing, yet their broader implications for the organization of scientific knowledge remain unclear.
cs.DL
Zipeng Ling, Shuliang Liu, Shenghong Fu, Yuehao Tang, Seonil Son, Yao Wan, Xuming Hu
LLM reasoning traces suffer from complex flaws -- *Step Internal Flaws* (logical errors, hallucinations, etc.) and *Step-wise Flaws* (overthinking, underthinking), which vary by sample.
cs.CL