An Luo, Jie Ding
As generative models enable rapid creation of high-fidelity images, societal concerns about misinformation and authenticity have intensified. A promising remedy is multi-bit image watermarking, which embeds a multi-bit message into an image so that a verifier…
🤗 HuggingFacedaily curated papers2026-04-13
Krzysztof Ociepa, Łukasz Flis, Remigiusz Kinas, Krzysztof Wróbel, et al.
The development of the Bielik v3 PL series, encompassing both the 7B and 11B parameter variants, represents a significant milestone in the field of language-specific large language model (LLM) optimization. While general-purpose models often demonstrate…
🤗 HuggingFacedaily curated papers2026-04-12
Zunhai Su, Hengyuan Zhang, Wei Wu, Yifan Zhang, et al.
As the foundational architecture of modern machine learning, Transformers have driven remarkable progress across diverse AI domains. Despite their transformative impact, a persistent challenge across various Transformers is Attention Sink (AS), in which a…
🤗 HuggingFacedaily curated papers2026-04-11
Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, et al.
We present Audio Flamingo Next (AF-Next), the next-generation and most capable large audio-language model in the Audio Flamingo series, designed to advance understanding and reasoning over speech, environmental sounds and music. Compared to Audio Flamingo 3…
🤗 HuggingFacedaily curated papers2026-04-13
Matteo Spanio, Ilay Guler, Antonio Rodà
Symbolic music research has relied almost exclusively on MIDI-based datasets; text-based engraving formats such as LilyPond remain unexplored for music understanding. We present BMdataset, a musicologically curated dataset of 393 LilyPond scores (2,646…
🤗 HuggingFacedaily curated papers2026-04-12
CocoaBench Team, Shibo Hao, Zhining Zhang, Zhiqi Liang, et al.
LLM agents now perform strongly in software engineering, deep research, GUI automation, and various other applications, while recent agent scaffolds and models are increasingly integrating these capabilities into unified systems. Yet, most evaluations still…
🤗 HuggingFacedaily curated papers2026-04-13
Han Li, Yifan Yao, Letian Zhu, Rili Feng, et al.
Code agents are advancing rapidly, but debugging them is becoming increasingly difficult. As frameworks orchestrate parallel tool calls and multi-stage workflows over complex tasks, making the agent's state transitions and error propagation hard to observe.
🤗 HuggingFacedaily curated papers2026-04-13
Shanchuan Lin, Ceyuan Yang, Zhijie Lin, Hao Chen, et al.
We propose continuous adversarial flow models, a type of continuous-time flow model trained with an adversarial objective. Unlike flow matching, which uses a fixed mean-squared-error criterion, our approach introduces a learned discriminator to guide training.
🤗 HuggingFacedaily curated papers2026-04-13
Song Jin, Juntian Zhang, Xun Zhang, Zeying Tian, et al.
Recent advancements in Vision-Language Models (VLMs) have revolutionized general visual understanding. However, their application in the food domain remains constrained by benchmarks that rely on coarse-grained categories, single-view imagery, and inaccurate…
🤗 HuggingFacedaily curated papers2026-04-12
Haolin Li, Shuyang Jiang, Ruipeng Zhang, Jiangchao Yao, et al.
While large language models hold promise for complex medical applications, their development is hindered by the scarcity of high-quality reasoning data. To address this issue, existing approaches typically distill chain-of-thought reasoning traces from large…
🤗 HuggingFacedaily curated papers2026-04-13
Chenchen Zhang
Reinforcement learning (RL) for large language models (LLMs) increasingly relies on sparse, outcome-level rewards -- yet determining which actions within a long trajectory caused the outcome remains difficult. This credit assignment (CA) problem manifests in…
🤗 HuggingFacedaily curated papers2026-04-13
Yifan Yu, Yuqing Jian, Junxiong Wang, Zhongzhu Zhou, et al.
Diffusion language models promise parallel generation, yet still lag behind autoregressive (AR) models in quality. We stem this gap to a failure of introspective consistency: AR models agree with their own generations, while DLMs often do not.
🤗 HuggingFacedaily curated papers2026-04-13
Zhipeng Chen, Tao Qian, Wayne Xin Zhao, Ji-Rong Wen
Recently, scaling reinforcement learning with verifiable rewards (RLVR) for large language models (LLMs) has emerged as an effective training paradigm for significantly improving model capabilities, which requires guiding the model to perform extensive…
🤗 HuggingFacedaily curated papers2026-04-13
Ivan Sedykh, Nikita Sorokin, Valentin Malykh
Recent advances in masked diffusion language models (MDLMs) narrow the quality gap to autoregressive LMs, but their sampling remains expensive because generation requires many full-sequence denoising passes with a large Transformer and, unlike autoregressive…
🤗 HuggingFacedaily curated papers2026-04-11
Hanqi Xiao, Vaidehi Patil, Zaid Khan, Hyunji Lee, et al.
As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dialogue partners (i.e., form and use a theory-of-mind, or ToM) becomes increasingly critical for safe…
🤗 HuggingFacedaily curated papers2026-04-13
Lester James V. Miranda, Ivan Vulić, Anna Korhonen
Synthesizing supervised finetuning (SFT) data from language models (LMs) to teach smaller models multilingual tasks has become increasingly common. However, teacher model selection is often ad hoc, typically defaulting to the largest available option, even…
🤗 HuggingFacedaily curated papers2026-04-13
Gordon Chen, Ziqi Huang, Ziwei Liu
Video diffusion models have achieved remarkable progress in generating high-quality videos. However, these models struggle to represent the temporal succession of multiple events in real-world videos and lack explicit mechanisms to control when semantic…
🤗 HuggingFacedaily curated papers2026-04-11
Songlin Yang, Xianghao Kong, Anyi Rao
Unified multimodal models (UMMs) were designed to combine the reasoning ability of large language models (LLMs) with the generation capability of vision models. In practice, however, this synergy remains elusive: UMMs fail to transfer LLM-like reasoning to…
🤗 HuggingFacedaily curated papers2026-04-13
Ali Slim, Haydar Hamieh, Jawad Kotaich, Yehya Ghosn, et al.
Large Language Models (LLMs) are increasingly used for code generation, yet quantum code generation is still evaluated mostly within single frameworks, making it difficult to separate quantum reasoning from framework familiarity. We introduce QuanBench+, a…
🤗 HuggingFacedaily curated papers2026-03-25
Shahar Levy, Eliya Habba, Reshef Mintz, Barak Raveh, et al.
Many disciplines pose natural-language research questions over large document collections whose answers typically require structured evidence, traditionally obtained by manually designing an annotation schema and exhaustively labeling the corpus, a slow and…
🤗 HuggingFacedaily curated papers2026-04-10
Udari Madhushani Sehwag, Elaine Lau, Haniyeh Ehsani Oskouie, Shayan Shabihi, et al.
Accelerating scientific discovery requires the identification of which experiments would yield the best outcomes before committing resources to costly physical validation. While existing benchmarks evaluate LLMs on scientific knowledge and reasoning, their…
🤗 HuggingFacedaily curated papers2026-04-12
Han Luo, Guy Laban
Large language models are increasingly deployed in multi-turn settings such as tutoring, support, and counseling, where reliability depends on preserving consistent roles, personas, and goals across long horizons. This requirement becomes critical when LLMs…
🤗 HuggingFacedaily curated papers2026-04-10
Talor Abramovich, Maor Ashkenazi, Carl, Putterman, et al.
Speculative Decoding (SD) has emerged as a critical technique for accelerating Large Language Model (LLM) inference. Unlike deterministic system optimizations, SD performance is inherently data-dependent, meaning that diverse and representative workloads are…
🤗 HuggingFacedaily curated papers2026-02-10
Rui Xu, Dafei Qin, Kaichun Qiao, Qiujie Dong, et al.
Recent advancements in autoregressive transformers have demonstrated remarkable potential for generating artist-quality meshes. However, the token ordering strategies employed by existing methods typically fail to meet professional artist standards, where…
🤗 HuggingFacedaily curated papers2026-04-10
Shuquan Lian, Juncheng Liu, Yazhe Chen, Yuhong Chen, et al.
Prior representative ReAct-style approaches in autonomous Software Engineering (SWE) typically lack the explicit System-2 reasoning required for deep analysis and handling complex edge cases. While recent reasoning models demonstrate the potential of extended…
🤗 HuggingFacedaily curated papers2026-04-13
Ao Li, Yonggen Ling, Yiyang Lin, Yuji Wang, et al.
Accurate 3D human keypoints localization is a critical technology enabling robots to achieve natural and safe physical interaction with users. Conventional 3D human keypoints estimation methods primarily focus on the whole-body reconstruction quality relative…
🤗 HuggingFacedaily curated papers2026-04-10
Yinyi Luo, Wenwen Wang, Hayes Bai, Hongyu Zhu, et al.
Recent advances in unified multimodal models (UMMs) have led to a proliferation of architectures capable of understanding, generating, and editing across visual and textual modalities. However, developing a unified framework for UMMs remains challenging due…
🤗 HuggingFacedaily curated papers2026-04-12
Yu Li, Xiaoran Shang, Qizhi Pei, Yun Zhu, et al.
Post-training data plays a pivotal role in shaping the capabilities of Large Language Models (LLMs), yet datasets are often treated as isolated artifacts, overlooking the systemic connections that underlie their evolution. To disentangle these complex…
🤗 HuggingFacedaily curated papers2026-04-12
Luozheng Qin, Jia Gong, Qian Qiao, Tianjiao Li, et al.
Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher computational costs than understanding, particularly for video. This imbalance motivates us to invert the…
🤗 HuggingFacedaily curated papers2026-04-09
Khai Loong Aw, Klemen Kotar, Wanhee Lee, Seungwoo Kim, et al.
Young children demonstrate early abilities to understand their physical world, estimating depth, motion, object coherence, interactions, and many other aspects of physical scene understanding. Children are both data-efficient and flexible cognitive systems…
🤗 HuggingFacedaily curated papers2026-04-11