Farima Fatahi Bayat, Moin Aminnaseri, Pouya Pezeshkpour, Estevam Hruschka
Large language models (LLMs) have become increasingly capable of following instructions and complex reasoning, making prompting a flexible interface for adapting models without parameter updates …
HuggingFacedaily curated papers2026-05-20
Cheolhong Min, Jaeyun Jung, Daeun Lee, Hyeonseong Jeon, Yu Su
Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D understanding or reliance on statistical shortcuts in natural images …
HuggingFacedaily curated papers2026-05-28
Yubo Li, Yidi Miao
Long-horizon LLM inference turns the key--value (KV) cache into the dominant GPU memory consumer and makes per-token attention increasingly expensive. Many common eviction policies use static recency windows or historical attention …
HuggingFacedaily curated papers2026-05-24
Yubo Li, Yidi Miao, Yuntian Shen, Yuxin Liu
Recent advances in multimodal web agents often rely on increased inference-time computation, including rollout search, verifier passes, offline skill discovery, and specialist model stacks …
HuggingFacedaily curated papers2026-05-26
Miria Feng, William Tan, Mert Pilanci
Globalization and multiculturalism continue to produce increasingly diverse speech varieties. Yet current spoken dialogue systems frequently fail on under-represented dialects and accents …
HuggingFacedaily curated papers2026-05-22
Jusuk Lee, Seungjae Lee, Jonghun Shin, Hoseong Jung, Sungha Kim
Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recognition or vision-language alignment …
HuggingFacedaily curated papers2026-05-28
Xiaona Zhou, Muntasir Wahed, Tianjiao Yu, Constantin Brif, Ismini Lourentzou
Recent advances in Vision-Language Models (VLMs) have achieved impressive performance across many tasks, yet prior studies report unsatisfactory performance when applying large language or multimodal models to finding abnormal patterns in …
HuggingFacedaily curated papers2026-05-28
Long Phan, Devin Kim, Alexander Pan, Alice Blair, Adam Khoja
Large language models (LLMs) exhibit systematic political bias across a variety of sensitive contexts. We find that LLMs handle counterpart topics from opposing political sides asymmetrically. …
HuggingFacedaily curated papers2026-05-28
Aviral Chharia, Fernando De la Torre
High-fidelity 3D Gaussian head avatar generation is critical for applications such as AR/VR, telepresence, and digital humans. Existing methods depend on multi-view datasets, 3D captures, or intermediate 2D view synthesis. …
HuggingFacedaily curated papers2026-05-24
Parsa Mazaheri
One-shot Program-of-Thought (PoT) emits a Python program that prints a primitive-action plan; a single invalid action silently invalidates the trajectory …
HuggingFacedaily curated papers2026-05-28
Zhixin Cai, Jun Bai, Yang Liu, Jiaqi Li, Yichi Zhang
Explaining why dense retrievers assign high relevance scores remains challenging because retrieval decisions are made through opaque high-dimensional embeddings. Existing explanations often focus on surface signals …
HuggingFacedaily curated papers2026-05-28
Vaishali Senthil, Ashutosh Hathidara, Sebastian Schreiber
Tool retrieval over large API catalogs is a core bottleneck for LLM agents: user queries arrive in colloquial, often underspecified language, while the catalog uses technical API vocabulary that no fixed encoder can bridge on its own …
HuggingFacedaily curated papers2026-05-28
Hesong Wang, Xin Jin, Lu Lu, Chenhaowen Li, Jian Chen
Video large language models (Video-LLMs) have demonstrated strong capabilities in video understanding tasks. However, their practical deployment is still hindered by the inefficiency introduced by processing massive amounts of visual token …
HuggingFacedaily curated papers2026-05-28
Samson Gourevitch, Yazid Janati, Dario Shariatian, Umut Simsekli, Eric Moulines
Discrete diffusion models are often trained through clean-data prediction, but the prediction can be used in different ways to define the reverse dynamics. In Masked Diffusion Models (MDM) these choices largely coincide …
HuggingFacedaily curated papers2026-05-21
Ngoc Phan Phuoc Loc, Toan Huynh La Viet, Thanh Tran Khanh, Duy A Nguyen, Tuan Anh Nguyen Pham
The rapid growth in submissions to machine learning venues has strained the scientific peer-review system and intensified interest in LLM-based automated peer reviewers. However, how good these systems are actually …
HuggingFacedaily curated papers2026-05-27
Zhu Yu, Jingnan Gao, Runmin Zhang, Lingteng Qiu, Zhengyi Zhao
This work presents ViGeo, a feed-forward foundation model for recovering spatially dense and temporally consistent geometry from video sequences …
HuggingFacedaily curated papers2026-05-28
Dongyoon Hahm, Dylan Hadfield-Menell, Kimin Lee
Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we introduce alignment tampering …
HuggingFacedaily curated papers2026-05-26
Corrado Rainone, Davide Belli, Bence Major, Arash Behboodi
The design space of agentic AI inference spans two extremes: frontier large language models (LLMs), typically hosted in the cloud and offering strong performance across a wide range of tasks at substantially high cost …
HuggingFacedaily curated papers2026-05-28
Artur Jesslen, Olaf Dünkel, Adam Kortylewski
Foundation features from self-supervised vision models and text-to-image diffusion models have proven effective for semantic correspondence estimation. However, because these features are learned primarily from 2D image objectives …
HuggingFacedaily curated papers2026-05-28
Ngoc Trinh Hung Nguyen, Alonso Silva, Laith Zumot, Liubov Tupikina, Armen Aghasaryan
Natural generation allows Large Language Models (LLMs) to produce free-form responses with rich reasoning, yet the lack of structure makes outputs difficult to verify. Conversely …
HuggingFacedaily curated papers2026-05-28
Omer Benishu, Gal Fiebelman, Sagie Benaim
We address the task of generating physically accurate and visually faithful 4D Human-Object Interaction (HOI). Given a static 3D human and target object represented as 3D Gaussian Splats (3DGS) …
HuggingFacedaily curated papers2026-05-28
Yingdong Shi, Ruiming Zhang, Changming Li, Zhiyu Yang, Kaixing Zhang
Activation-based control steers large language models (LLMs) by intervening on their internal representations during inference, and has emerged as an effective paradigm for controlling behaviors such as persona and style. However …
HuggingFacedaily curated papers2026-05-28
Wenbo Gao, Songbai Tan, Zhongan Wang, Fei Shen, Gang Xu
Smartphone scams are increasingly prevalent and typically manifest as multi-stage, cross-application processes with gradually emerging intent. Effective intervention thus requires anticipating scams before the intent becomes explicit. …
HuggingFacedaily curated papers2026-05-09
Haoxiang Jiang, Zihan Dong, Tianci Liu, Wanying Wang, Ran Xu
Pointwise reward modeling offers critical signals for LLM post-training, yet struggles with absolute scoring in subjective, non-verifiable settings. Rubric-based methods address this by decomposing evaluation into explicit criteria …
HuggingFacedaily curated papers2026-05-27
Víctor Gallego
We study two-level autoresearch for cooperation: an outer-loop AI agent autonomously redesigns the inner-loop pipeline of an LLM policy-synthesis system for multi-agent Sequential Social Dilemmas (SSDs) …
HuggingFacedaily curated papers2026-05-28
Shicheng Fan, Haochang Hao, Dehai Min, Weihao Liu, Philip S. Yu
Applying reinforcement learning to improve factual accuracy in knowledge-intensive question answering faces a reward design dilemma. Response-level rewards provide only coarse supervision and cannot distinguish correct from incorrect state …
HuggingFacedaily curated papers2026-05-28
Chenghao Zhang, Guanting Dong, Yufan Liu, Tong Zhao, Zhicheng Dou
Large Language Models (LLMs) have advanced autonomous agents from deep search, which retrieves concise factual answers, to deep research, which synthesizes scattered evidence into long-form reports. However …
HuggingFacedaily curated papers2026-05-28
Tiantian Feng, Anfeng Xu, Xuan Shi, Aditya Kommineni, Shakhrul Iman Siam
We present ChildVox, a novel benchmark for characterizing the diverse acoustic signals through which children communicate. Specifically, ChildVox follows the full developmental trajectory from birth through school age …
HuggingFacedaily curated papers2026-05-28
Hadar Davidson, Noam Issachar, Sagie Benaim
Diffusion models achieve state-of-the-art image synthesis, with their generative trajectories fundamentally exhibiting a spectral bias, resolving low-frequency global structures early and high-frequency fine details later …
HuggingFacedaily curated papers2026-05-28
Jiapeng Zhu, Jianxiang Yu, Yibo Zhao, Chengcheng Han, Qi Gu
Equipping large language models with explicit skills has emerged as a promising paradigm for enabling autonomous agents to solve complex tasks …
HuggingFacedaily curated papers2026-05-27
Daegon Yu, SeungYoon Han, Woomyoung Park
Dense retrievers exhibit positional bias, favoring documents whose query-relevant information appears near the beginning and degrading retrieval performance when the information appears later …
HuggingFacedaily curated papers2026-05-26
Junlin Yang, Dylan Zhang, Xiangchen Song, Qirun Dai, Xiao Liu
We introduce CausaLab, a scalable environment for evaluating interactive causal discovery by LLM agents. Unlike prior evaluations, CausaLab evaluates both whether an agent can solve a problem using causal evidence and whether its answer is …
HuggingFacedaily curated papers2026-05-28
Travis Lelle
We show that LoRA adapters, the dominant distribution format for fine-tuned LLMs, can be reliably backdoored through training data poisoning while preserving baseline task performance. On a Qwen 2.5 1.5B prompt-injection classifier …
HuggingFacedaily curated papers2026-05-28
Fangtai Wu, Hailong Guo, Shijie Huang, Jiayi Song, Yubo Huang
Customized image editing aims to equip pre-trained diffusion models with specific visual effects using limited paired data, typically via Low-Rank Adaptation (LoRA). As the number of desired effects grows …
HuggingFacedaily curated papers2026-05-25
You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee, Kaipeng Zhang, Yu-Lun Liu
As video diffusion models (VDMs) advance toward world models, a key question arises: do they truly understand causality, or merely overfit to statistical temporal patterns? Existing benchmarks mostly rely on synthetic data …
HuggingFacedaily curated papers2026-05-28
Zhengyang Tang, Yuxuan Liu, Xin Lai, Junyi Li, Pengyuan Lyu
A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Existing mobile-agent benchmarks have made important progress on evaluation …
HuggingFacedaily curated papers2026-05-28
Jing Huang, Daniel Wurgaft, Rachit Bansal, Laura Ruis, Naomi Saphra
Larger models learn tasks smaller models do not. What drives this phenomenon? …
HuggingFacedaily curated papers2026-05-28
Chengzhi Liu, Yuzhe Yang, Sophia Xiao Pu, Yepeng Liu, Lin Long
Multimodal large language models are increasingly deployed as long-horizon agents, where memory must do more than recall: it must track an evolving world, revise what has gone stale, and surface the right evidence at decision time …
HuggingFacedaily curated papers2026-05-28
Jie Jia, Yaofeng Su, Zeyu Bao, Yun Hong, Bingzhao Gao
Occlusion-aware prediction remains a critical challenge in autonomous driving due to the inherent uncertainty of unobserved regions. Existing approaches either overestimate risk based on reachable states or struggle to predict accurate tra …
HuggingFacedaily curated papers2026-05-21
Feng Han, Zhixiong Zhang, Zheming Liang, Yibin Wang, Jiaqi Wang
Vision-Language Models (VLMs) have achieved substantial progress across a wide range of understanding and reasoning tasks, driven by large-scale image-text training aimed at multimodal fusion. Ideally …
HuggingFacedaily curated papers2026-05-28
Junyan Ye, Jun He, Zilong Huang, Dongzhi Jiang, Xuan Yang
Image generation models have evolved from text-conditioned pixel synthesis toward multimodal agents endowed with visual comprehension and tool invocation capabilities. Yet …
HuggingFacedaily curated papers2026-05-28
Xudong Lu, Xueying Li, Annan Wang, Yang Bo, Jinpeng Chen
We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-visual streams. Unlike offline video understanding or text-prompted streaming QA …
HuggingFacedaily curated papers2026-05-26
Min Zhao, Hongzhou Zhu, Bokai Yan, Zihan Zhou, Yimin Chen
Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive video world models remains challenging …
HuggingFacedaily curated papers2026-05-28
Chen Geng, Guangzhao He, Yue Gao, Yunzhi Zhang, Shangzhe Wu
Data-driven approaches have revolutionized 3D vision, enabling transformers to effectively reconstruct and generate static 3D objects. However …
HuggingFacedaily curated papers2026-05-28
Jinheon Baek, Soyeong Jeong, Sangwoo Park, Woongyeong Yeo, Minki Kang
Real-world information needs require access to structurally diverse knowledge sources, from unstructured text and relational tables to knowledge graphs and property graphs. Existing retrievers, however …
HuggingFacedaily curated papers2026-05-28
Dongrui Liu, Yu Li, Zhonghao Yang, Peng Wang, Guanxu Chen
Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI models drastically lower attack barriers …
HuggingFacedaily curated papers2026-05-28
Chun-Hsiao Yeh, Shengyi Qian, Manchen Wang, Yi Ma, Joseph Tighe
Vision-Language Models (VLMs) often struggle with robust 3D spatial reasoning. Prevailing methods that rely on fine-tuning with 3D visual question-answering (VQA) datasets may overfit dataset-specific biases …
HuggingFacedaily curated papers2026-05-28
Yusuf Dalva, Pinar Yanardag
Autoregressive video diffusion models generate streaming video by producing frames sequentially, conditioning each chunk on previously generated content …
HuggingFacedaily curated papers2026-05-28
Yifei Zuo, Dhruv Pai, Zhichen Zeng, Alec Dewulf, Shuming Hu
Large Language Models (LLMs) have become the central paradigm in artificial intelligence, yet the core computational primitive of attention has remained structurally unchanged …
HuggingFacedaily curated papers2026-05-27
Yuxiang Chai, Han Xiao, Xinyu Fu, Jinpeng Chen, Rui Liu
Recent advances in mobile GUI agents have shown strong potential for automating mobile tasks, but most effective systems still depend on large vision-language models for screenshot understanding and long-horizon planning …
HuggingFacedaily curated papers2026-05-28