Jian Xie, Tianhe Lin, Zilu Wang, Yuting Ning, Yuekun Yao
QUEST is a family of open deep research agents (2B to 35B) trained on 8K synthesized tasks using unified rubric trees, approaching or surpassing frontier closed-source agents across eight deep…
cs.AIcs.CL
Hoang Phan, Quang H. Nguyen, Hung T. Q. Le, Xiusi Chen, Heng Ji
Large Reasoning Models use a 'hidden critique ability' — an internal mechanism that detects errors and triggers self-correction even when the chain-of-thought doesn't verbalize corrections — and steering latent representations…
cs.CLcs.AI
Bang Liu, Yongfeng Gu, Jiayi Zhang, Zhaoyang Yu, Sirui Hong
Foundation Protocol (FP) is a graph-first coordination layer for an emerging human-AI society that unifies agents, tools, resources, humans, and organizations, providing economic primitives for metering and settlement while treating…
cs.AIcs.MA
Fancy Kong, Congjie Zheng, Murphy Zhuang, Rio Yang, Sueky Zhang
Macaron-A2UI generates natural language together with lightweight executable UI actions for personal agents, reaching 75.6 on A2UI-Bench and surpassing full-schema frontier baselines, moving beyond text-only chat interaction.
cs.AIcs.HC
Yoav Gur-Arieh, Ana Marasović, Mor Geva
BonaFide, a benchmark of 3,066 ground-truth labeled CoTs, reveals that most faithfulness metrics for chain-of-thought reasoning perform near chance, with the best reaching only 0.70 AUROC at CoT level —…
cs.AIcs.CL
Yifan Lan, Yuanpu Cao, Hanyu Wang, Lu Lin, Jinghui Chen
Zero-CoT Probe (ZCP) truncates chain-of-thought to expose latent shortcut mappings and detect evasive data contamination in LLMs, introducing Contamination Confidence to quantify both likelihood and severity of contamination beyond binary…
cs.CLcs.AI
Yusong Lin, Xinyuan Liang, Haiyang Wang, Qipeng Gu, Siqi Cheng
Claw-Anything benchmarks always-on personal assistants with months of simulated user activity, finding GPT-5.5 achieves only 34.5% pass@1, substantially below prior benchmarks and revealing a significant gap between current agent capabilities…
cs.AIcs.HC
Haoyi Hu, Qirong Lyu, Xianghan Kong, Weiwen Liu, Jianghao Lin
ProAct is a proactive agent architecture that leverages idle-time compute between user interactions to anticipate and prepare for likely upcoming needs, reducing task completion turns by 14.8% and hallucination rates…
cs.AIcs.CL
Yihao Hu, Zhihao Wen, Xiujin Liu, Pan Wang, Xin Zhang
SEAL co-evolves both the agent policy and its training environment in a closed loop, using turn-level failure diagnoses as a shared signal for both environment adaptation and policy optimization, yielding…
cs.AIcs.ML
Guohong Liu, Jialei Ye, Pengzhi Gao, Wei Liu, Jian Luan
SimuWoB is a fully synthetic benchmark for mobile GUI agents with 120 challenging tasks in a virtual environment, revealing that state-of-the-art agents achieve only 27.92% average success rate, dropping to…
cs.AIcs.HC
Han Chen, Zining Zhang, Wenqi Pei, Bingsheng He, Ming Wu
MemForest reformulates agent memory as a temporal data management problem using parallel chunk extraction and a hierarchical temporal index (MemTree), achieving 79.8% pass@1 accuracy on LongMemEval-S with 6x higher memory…
cs.AIcs.ML
Zuhao Yang, Kaichen Zhang, Sudong Wang, Keming Wu, Zhongyu Yang
ParaVT introduces parallel video tool calling via multi-agent RL (dispatching multiple time-window crops in one turn), resolving the Tool Prior Paradox through PARA-GRPO which applies targeted format rewards and frame-budget…
cs.AIcs.LG
Yingtie Lei, Zhongwei Wan, Jiankun Zhang, Samiul Alam, Zixuan Zhong
SkillEvolBench diagnoses whether LLM agents can distill episodic task experience into reusable procedural skills, finding that current agents often adapt locally but rarely form robust reusable skills, and that raw-trajectory…
cs.AIcs.ML
Benhao Huang, Zhengyang Geng, Zico Kolter
Equilibrium Reasoners (EqR) learn task-conditioned attractors — latent dynamical systems whose fixed points correspond to valid solutions — enabling test-time scaling via depth (more iterations) and breadth (multiple initializations), boosting…
cs.AIcs.LG
Guochao Jiang, Jingyi Song, Guofeng Quan, Chuzhan Hao, Guohua Liu
DVAO dynamically adjusts combination weights based on empirical reward variance of each objective within a rollout group, up-weighting objectives with stronger learning signals while suppressing noisy ones, achieving superior multi-objective…
cs.LGcs.AI
Siyu An, Junru Lu, Junnan Dong, Qiufeng Wang, Yinghui Li
This paper formalizes a roadmap for native multimodal modeling (NMM), defining architectural nativity and organizing existing native models into three categories (Multi-to-Text, Multi-to-Target, Multi-to-Multi), providing an industrial-grade investigation from architectural…
cs.CVcs.AI
Yuqian Yuan, Wentong Li, Zhaocheng Li, Yutong Lin, Juncheng Li
InstructSAM bridges a vision-language model with SAM3 through learnable instance queries, enabling multi-instance segmentation under arbitrary natural language instructions without modifying SAM3's core architecture.
cs.CVcs.AI
Zhuoqun Li, Boxi Cao, Guiping Jiang, Fangrui Lv, Ruotong Pan
MetaphorVU-Bench is the first benchmark for metaphorical video understanding, revealing that current MLLMs lag far behind human level primarily due to defective cross-domain mapping, and MetaphorBoost improves performance through metaphor…
cs.CVcs.AI
Guiyao Tie, Jiawen Shi, Dingjie Song, Yixiao Huang, Ziji Sheng
This survey examines AI-powered scientific workflow automation (AutoResearch), analyzing how research systems redistribute control, evidence, execution, validation, and accountability across workflows.
cs.AIcs.SE
Guijin Son, Jehyun Park, Seyeon Park, Sunghee Ahn, Youngjae Yu
This work shows that GPT-5.5 and Claude Code agents achieve at best ~20% first-attempt success on engineering CAD tasks validated via FEA, and introduces blueprint schema and 21-view image rendering…
cs.AIcs.CV
Chao Tang, Jianzong Wu, Qingyu Shi, Ye Tian, Aixi Zhang
UniCharacter enables customized multimodal role-play by jointly customizing a character's persona, dialogue style, and visual identity from only 10 images, using unified SFT and character-specific GRPO.
cs.AIcs.CV
Kaining Ying, Hengrui Hu, Siyu Ren, Jiamu Li, Fengjiao Chen
WBench evaluates interactive world models across 5 dimensions using 289 test cases and 1,058 interaction turns, finding that no single model performs strongly across all dimensions.
cs.CVcs.AI