Shangding Gu
This paper studies the next major bottleneck in agentic AI as system scaling, not only model scaling: the design of audi…
HuggingFace
Jeongeun Park, Janghyeok Han, Geonung Kim, Hyun-Seung Lee, Kyuha Choi
Video outpainting generates plausible visual content beyond the original spatial extent of a video, playing a key role i…
HuggingFace
Josef Chen
Physical AI systems, including robots, autonomous vehicles, embodied agents and edge copilots, often run a different inf…
HuggingFace
Jianlin Ye, Christos Kyrkou, Panayiotis Kolios
The integration of Unmanned Aerial Vehicles(UAVs) into Intelligent Transportation Systems (ITS) offers synoptic visibili…
HuggingFace
Marko Kojic, Ivan Bondyrev, Aral de Moor, Joseph Shtok, Petr Borovlev
We present Mellum 2, an open-weight 12B-parameter Mixture-of-Experts (MoE) language model with 2.5B active parameters pe…
HuggingFace
Mingjian Gao, Wenqiao Zhang, Yuqian Yuan, Yang Dai, Binhe Yu
Recent work has begun to equip vision-language-action (VLA) policies with explicit intermediate reasoning. In embodied c…
HuggingFace
Sanghyun Jo, Seo Jin Lee, Seohyung Hong, Yoorim Gang, Hyeongsub Kim
Cell instance segmentation models trained on cell-specific datasets suffer severe performance drops on out-of-distributi…
HuggingFace
Arnas Uselis, Darina Koishigarina, Seong Joon Oh
Humans easily determine which color belongs to which shape in multi-object scenes, an ability known as concept binding. …
HuggingFace
Alireza Salemi, Chang Zeng, Atharva Nijasure, Jui-Hui Chung, Razieh Rahimi
Large Language Model (LLM) search agents have shown strong promise for knowledge-intensive language tasks through multip…
HuggingFace
Jungwon Park, Jimyeong Kim, Jungmin Ko, Nojun Kwak, Wonjong Rhee
Diffusion language models decode text by iteratively denoising masked token sequences, making the choice of which positi…
HuggingFace
Xiaohang Tang, Keyue Jiang, Che Liu, Qifang Zhao, Xiaoxiao Xu
Reinforcement learning (RL) can be used to improve the policy (denoiser) of diffusion large language models (dLLMs), whi…
HuggingFace
Daniil Plyusov, Alexey Gorbatovski, Alexey Malakhov, Nikita Balagansky, Boris Shaposhnikov
On-policy distillation (OPD) trains a student on prefixes sampled from its own policy while matching a stronger teacher.…
HuggingFace
Shuang Liang, Chaochuan Hou, Xu Yao, Shiping Wang, Hailiang Huang
While previous research in multivariate time series forecasting has focused on developing complex holistic models, this …
HuggingFace
Aditya Chetan, Eric Cai, Peeyush Kushwaha, Bharath Raj Nagoor Kani, Utkarsh Mall
The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. Howev…
HuggingFace
Yunbo Tang, Chengyi Yang, Shiyu Liu, Zhishang Xiang, Zerui Chen
Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search. Despit…
HuggingFace
Zhenhao Yang, Xiaoshi Wu, Zhengyao Lv, Xiaoyu Shi, Xintao Wang
Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaini…
HuggingFace