Sy-Tuyen Ho, Minghui Liu, Huy Nghiem, Furong Huang
Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review. However …
HuggingFacedaily curated papers2026-05-28
Stine Lyngsø Beltoft, William Brach, Federico Torrielli, Jacob Nielsen, Annemette Brok Pirchert
Monitoring autonomous language model agents currently relies mostly on surface behavior. But what happens when agent populations invent new languages with the goal of avoiding human oversight. …
HuggingFacedaily curated papers2026-05-29
Killian Steunou, Anas Filali Razzouki, Khalil Guetari, Mounîm A. El-Yacoubi, Yannis Tevissen
Video-language models can process only a limited number of frames, making frame selection a key bottleneck for efficient video captioning. Most captioning pipelines still rely on uniform sampling …
HuggingFacedaily curated papers2026-05-29
Chang-Bin Zhang, Yujie Zhong, Qiang Zhang, Kai Han
While visually grounded Chain-of-Thought (CoT) has emerged as a promising paradigm to enhance fine-grained perception in multimodal large language models (MLLMs), its efficacy during the inference phase remains underexplored …
HuggingFacedaily curated papers2026-05-29
Bill Psomas, Dionysis Christopoulos, Thanasis Petropoulos, Nikos Efthymiadis, Ioannis Kakogeorgiou
Remote sensing composed image retrieval (RSCIR) enables search in large satellite image archives using composed queries that combine a reference image with a textual modifier …
HuggingFacedaily curated papers2026-05-23
Jian Mu, Tianyi Lin, Chengwei Qin, Zhongxiang Dai, Yao Shu
Large language models are increasingly deployed in multi-turn interactive settings where users or environments can iteratively provide lightweight feedback. Unfortunately …
HuggingFacedaily curated papers2026-05-29
Tianpeng Bu, Xin Liu, Qihua Chen, Hao Jiang, Shurui Li
While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge this gap at both the evaluation and data levels …
HuggingFacedaily curated papers2026-05-28
Yue Zhang, Zun Wang, Han Lin, Yonatan Bitton, Idan Szpektor
Spatial reasoning is a fundamental capability for vision-language models (VLMs) deployed in real-world environments. However, visual observations are inherently limited representations of a 3D world: occlusion can render objects invisible, …
HuggingFacedaily curated papers2026-05-28
Wai-Chung Kwan, Aryo Pradipta Gema, Joshua Ong Jun Leang, Pasquale Minervini
Self-play can train language models without external supervision. However, existing methods require rule-checkable answers, leaving open-ended tasks dependent on curated prompts or frontier-model judges. …
HuggingFacedaily curated papers2026-05-29
Xiaobo Wang, Tong Wu, Min Tang, Jiaqi Li, Qi Liu
Building strong reward models (RMs) for language model alignment is bottlenecked by the cost and difficulty of acquiring diverse and reliable preference data from human annotation or judge models …
HuggingFacedaily curated papers2026-05-29
Mengqi Lei, Shuokun Cheng, Wei Bao, Shaoyi Du, Jun-Hai Yong
Object counting remains fragmented across domain-specific datasets and task formulations, despite rapid progress in generalist vision models. Existing counting models are often tailored to scenarios such as crowds, vehicles, cells …
HuggingFacedaily curated papers2026-05-29
Yiheng Li, Zhuo Li, Ruibing Hou, Yingjie Chen, Hong Chang
Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configurations and task-specific architectures …
HuggingFacedaily curated papers2026-05-28
Shu Wan, Abhinav Gorantla, Huan Liu, K. Selçuk Candan
Under standard graphical assumptions, the Markov boundary of a target variable is the smallest set of features that renders every other feature redundant. Once the boundary is observed …
HuggingFacedaily curated papers2026-05-28
Junlin Wang
Learning visuomotor policies via behavior cloning typically involves mimicking expert demonstrations collected by human operators. However, natural human demonstrations inherently contain high-frequency noise, such as intermittent jerks …
HuggingFacedaily curated papers2026-05-27
Changhao Pan, Rui Yang, Han Wang, Zhuan Zhou, Xuming He
Recent advances in speech generation have enabled high-fidelity synthesis, yet systematic evaluation of models under long-context conditions remains largely underexplored …
HuggingFacedaily curated papers2026-05-27
Ke Lei, Yu Zhang, Changhao Pan, Xueyi Pu, Wenxiang Guo
Real-time and accurate spatial audio generation is pivotal for delivering an immersive experience. However, existing spatial audio synthesis technologies are often encumbered by a tradeoff between generation quality and high inference late …
HuggingFacedaily curated papers2026-05-29
Ruiqi Li, Yu Zhang, Changhao Pan, Ke Lei, Xiang Yin
Zero-shot text-to-speech (TTS) has improved substantially for single-speaker synthesis, yet expressive long-form multi-speaker dialogue remains difficult …
HuggingFacedaily curated papers2026-05-29
Dan Jacobellis, Neeraja J. Yadwadkar
Media compression standards have reached a plateau in terms of the rate-distortion-complexity trade-off, limiting the ability to offload expensive AI perception to the cloud in applications like robotics, wearables, and remote sensing …
HuggingFacedaily curated papers2026-05-27
Jiahao Ying, Boxian Ai, Wei Tang, Siyuan Liu, Yixin Cao
Skills, i.e., structured workflow instructions distilled for large language models (LLMs), are becoming an increasingly important mechanism for improving agent performance on real-world downstream tasks. However …
HuggingFacedaily curated papers2026-05-28
Tianyi Zhou, Dongrui Liu, Leitao Yuan, Jing Shao, Xia Hu
LLM agents are increasingly expected not only to complete isolated tasks, but also to carry bounded representations of human expertise, judgment, and interaction style …
HuggingFacedaily curated papers2026-05-29
Tao Zou, Yichen He, Tian Qiu, Yuan Lin, Hang Li
Long-term memory is essential for multimodal agents to build coherent experience, accumulate world knowledge, and achieve continual learning. However …
HuggingFacedaily curated papers2026-05-29
Yuyang Zhao, Yicheng Pan, Qiyuan He, Jincheng Yu, Junsong Chen
Real-time streaming video-to-video editing (V2V) is critical for interactive applications such as live broadcasting and gaming, yet it remains a formidable challenge due to the stringent requirements for temporal consistency and inference …
HuggingFacedaily curated papers2026-05-28
Sicheng Feng, Zigeng Chen, Gongfan Fang, Xinyin Ma, Xinchao Wang
Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive models, offering competitive performance while naturally supporting parallel decoding. However …
HuggingFacedaily curated papers2026-05-29
Suryash Yagnik, Shubham Gaur, Saksham Thakur, Vinija Jain, Aman Chadha
Machine unlearning evaluation is structurally skewed: Why-type questions, which probe causal and relational knowledge, comprise less than 0.06% of CounterFact, 0.6% of ZSRE, and less than 1.3% of TOFU, MUSE, and WMDP-Cyber …
HuggingFacedaily curated papers2026-05-28
Yuanyi Wang, Su Lu, Yanggan Gu, Pengkai Wang, Yifan Yang
On-policy distillation (OPD) trains a student on its own rollouts with token-level teacher supervision. Recent selective OPD methods exploit the non-uniformity of OPD signals by prioritizing high-entropy or high-disagreement tokens. …
HuggingFacedaily curated papers2026-05-26
Xiangtao Kong, Jixin Zhao, Lingchen Sun, Rongyuan Wu, Lei Zhang
Real-world image restoration (IR) is bottlenecked by the scarcity of high-quality paired training data. Synthetic datasets are abundant but often fail to model real-world degradations …
HuggingFacedaily curated papers2026-05-29
Cristobal Eyzaguirre, Jiajun Wu, Juan Carlos Niebles
Video vision-language models (VLMs) are increasingly used in long-horizon and streaming settings, yet most video encoders still rely on spatiotemporal self-attention …
HuggingFacedaily curated papers2026-05-29
Jiacheng Lu, Haoyi Zhu, Sipei Yi, Enze Xie, Yu Li
Interactive video world models generate video chunk by chunk in response to user-controlled camera movements, enabling applications such as real-time game simulation, virtual scene navigation, and embodied AI training. However …
HuggingFacedaily curated papers2026-05-29
Jiazheng Xing, Hangjie Yuan, Lingling Cai, Xinyu Liu, Yujie Wei
Connector-based video unified models have demonstrated strong capability in instruction-grounded video synthesis, but integrating a large high-fidelity generator into the unified training loop is computationally prohibitive …
HuggingFacedaily curated papers2026-05-29
Jiejun Tan, Zhicheng Dou, Xinyu Yang, Yuyang Hu, Yiruo Cheng
LLM agents are evolving from conversational chatbots to operational tools in real-world workspaces. In local agentic harnesses, an LLM can read and write files, call tools, and reuse workspace state across sessions. …
HuggingFacedaily curated papers2026-05-29
Nianyi Lin, Jiajie Zhang, Lei Hou, Juanzi Li
Long-context reasoning remains a central challenge for large language models, which often fail to locate and integrate key information in extensive distracting content …
HuggingFacedaily curated papers2026-05-29
Yuqing Wang, Zhijie Lin, Ceyuan Yang, Yang Zhao, Fei Xiao
Unified multimodal models (UMMs) aim to handle perception and generation in a single model. Yet existing UMMs still rely on a frozen, separately pretrained VAE for image generation, imposing a structural bottleneck. …
HuggingFacedaily curated papers2026-05-29
Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu, Vikas Chandra
Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semantic understanding. …
HuggingFacedaily curated papers2026-05-28
Yujie Luo, Xiangyuan Ru, Jingsheng Zheng, Jingjing Wang, Yuqi Zhu
Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without high-quality domain-specific data …
HuggingFacedaily curated papers2026-05-28
Ruiqi Wang, Qimin Chen, Daniel Ritchie, Angel X. Chang, Manolis Savva
Most text-driven 3D indoor scene synthesis methods generate rooms from object-centric prompts, asking what furniture should be placed rather than how the space is used. Yet in real interior design …
HuggingFacedaily curated papers2026-05-29
Seongheon Park, Wendi Li, Changdae Oh, Samuel Yeh, Zsolt Kira
Vision-Language-Action (VLA) models enable robots to follow natural language instructions and generalize across diverse tasks, but they remain vulnerable to execution failures that compromise reliability in real-world deployment …
HuggingFacedaily curated papers2026-05-29
Kewei Xu, Xiaoben Lu, Shuofei Qiao, Zihan Ding, Haoming Xu
Real-world data analysis is inherently iterative, yet existing benchmarks mostly evaluate isolated or short interactive tasks, leaving agents' ability to track evolving analytical context over long horizons untested. We introduce LongDS …
HuggingFacedaily curated papers2026-05-28