Rui Yang, Qianhui Wu, Yuxi Chen, Hao Bai, Wenlin Yao
Building capable visual web agents requires long-horizon reasoning, precise grounding, and robust interaction with dynamic real-world websites. Despite rapid progress, the strongest systems remain largely proprietary …
HuggingFacedaily curated papers2026-06-01
Pengcheng Jiang, Zhiyi Shi, Kelly Hong, Xueqiang Xu, Jiashuo Sun
Search agents are often trained as policies over growing transcripts: the model must decide how to search while also remembering what it has seen, which evidence is useful, which constraints remain open …
HuggingFacedaily curated papers2026-06-01
Ning Lu, Baijiong Lin, Shengcai Liu, Jiahao Wu, Haoze Lv
Reinforcement learning (RL) improves large language model (LLM) agents by teaching them which actions lead to high rewards, but provides little supervision on what those actions do to the environment …
HuggingFacedaily curated papers2026-06-01
Abdelaziz M. A. Ibrahim, Zihao Li, Jörg Tiedemann, Shaoxiong Ji
Large-scale multilingual bitext often contains two distinct problems: non-parallel sentence pairs and low-quality translations. We decompose model-based assessment for such data into two independent components: parallelism assessment with …
HuggingFacedaily curated papers2026-05-29
Yubo Li, Rema Padman, Ramayya Krishnan
A retrieval-augmented generation (RAG) system deployed over a multi-author institutional corpus can give a different answer to the same question depending on which source it retrieves -- a failure mode the dominant single-gold-answer parad …
HuggingFacedaily curated papers2026-05-27
Yubo Li, Ramayya Krishnan, Rema Padman
Reasoning models are evaluated on single-turn benchmarks but deployed in multi-turn dialogue, where users push back on correct answers …
HuggingFacedaily curated papers2026-05-27
Ethan Caballero, Priyank Jaini, David Krueger, Irina Rish
We present a functional form (that we refer to as a Unified Neural Scaling Law (UNSL)) that accurately models and extrapolates the scaling behaviors of deep neural networks as multiple dimensions all vary simultaneously (i.e …
HuggingFacedaily curated papers2026-05-25
Pengfei Zhou, Shengcong Chen, Di Chen, Jiaxu Wang, Rongjun Jin
Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present τ_0-World Model (τ_0-WM) …
HuggingFacedaily curated papers2026-05-31
Zhihao Wu, Gracia Gong, Qinglin Zhu, Yudong Chen, Runcong Zhao
Watermarking embeds statistical signatures in AI-generated text for detection and attribution. We reveal a fundamental vulnerability: when users access multiple models (today's reality), watermarks trivially fail. …
HuggingFacedaily curated papers2026-05-28
Jianuo Huang, Yaojie Zhang, Qituan Zhang, Hao Lin, Hanlin Xu
Speculative decoding accelerates LLM inference by drafting multiple tokens and verifying them in parallel with the target model. However …
HuggingFacedaily curated papers2026-05-28
Ahan Chatterjee, Matthias Schöffel, Matthias Aßenmacher, Marinus Wiedner, Esteban Garces Arias
The diachronic evolution from Latin to the Romance languages involved a restructuring of the grammatical gender system from a tripartite configuration (masculine, feminine, neuter) to a bipartite one (masculine …
HuggingFacedaily curated papers2026-05-26
Gábor Recski, Szilveszter Tóth, Nadia Verdha, István Boros, Ádám Kovács
Academic researchers need efficient and reliable methods for collecting high-quality information from trusted sources, but modern tools for AI-assisted research still suffer from the tendency of Large Language Models (LLMs) to produce fact …
HuggingFacedaily curated papers2026-05-20
Binxiao Xu, Ruichuan An, Bocheng Zou, Hang Hua
Reusable skills are a key mechanism for extending agent capabilities, allowing agents to accumulate experience and solve increasingly complex tasks. Yet most existing skill-learning methods store reusable experience as text-only assets …
HuggingFacedaily curated papers2026-05-31
Hans Ole Hatzel, Sebastian Steindl, Jan Strich
LLM-generated reviews for scientific papers are gaining considerable traction and are even being officially piloted by major conferences. We have to assume that not only reviewers are using LLM-assistance …
HuggingFacedaily curated papers2026-05-27
Olaf Dünkel, Basavaraj Sunagad, Haoran Wang, David T. Hoffmann, Christian Theobalt
Measuring structured object understanding in vision foundation models remains challenging due to inconsistent evaluation protocols and limited part-level supervision …
HuggingFacedaily curated papers2026-05-29
Maharshi Gor, Yoo Yeon Sung, Yu Hou, Eve Fleisig, Irene Ying
AI systems are fallible, and humans can make mistakes in deciding whether to trust AI over their own judgment. Thus, improving human-AI collaboration requires understanding when, why, and how humans decide to rely on AI. …
HuggingFacedaily curated papers2026-05-27
Tianjie Ju, Yueqing Sun, Zheng Wu, Wei Zhang, Yaqi Huo
Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to sustain exploration in dynamic open worlds remains unclear. …
HuggingFacedaily curated papers2026-05-29
Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli, Yanran Chen, Christian Greisinger
Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who produced the annotations and how the annotation process was controlled …
HuggingFacedaily curated papers2026-06-01
Selim Kuzucu, Alessio Tonioni, Vasile Lup, Bernt Schiele, Federico Tombari
Large Vision-Language Models (LVLMs) map visual inputs into dense token sequences, imposing a quadratic computational bottleneck for inference …
HuggingFacedaily curated papers2026-05-28
Jaeung Lee, Dohyun Kim, Jaemin Jo
Large language model (LLM) unlearning has emerged as a crucial post-hoc mechanism for privacy protection and AI safety, yet auditing whether target knowledge is truly erased remains challenging …
HuggingFacedaily curated papers2026-05-23
Raphael Coelho
We describe a library of mathematical finance built in the Lean 4 proof assistant, on top of Mathlib and the BrownianMotion package. It is broad: more than two hundred sorry-free theorems across eleven areas …
HuggingFacedaily curated papers2026-05-31
Jose Marie Antonio Miñoza, Erika Fille T. Legara, Christopher P. Monterola
In this paper, training a neural network is identified, exactly, as a search through Hamilton--Jacobi initial-value problems: each gradient step selects the initial data of a viscous Hamilton--Jacobi equation whose Hopf--Cole propagator be …
HuggingFacedaily curated papers2026-05-27
Shashi Kumar, Yacouba Kaloga, Petr Motlicek, Ina Kodrasi, Andrea Cavallaro
Large language models solve complex problems by generating lengthy chains of explicit reasoning tokens. While effective, this makes reasoning expensive, length-sensitive, and constrained to (discrete) natural language. …
HuggingFacedaily curated papers2026-06-01