Chengshuai Shi, Wenzhe Li, Xinran Liang, Yizhou Lu, Wenjia Yang
Given the rapidly growing capabilities of vision-language models (VLMs), extending them to interactive decision-making tasks such as video games has emerged as a promising…
HuggingFacearXiv:2605.00347
Laki Iinbor, Zhiyang Dou, Wojciech Matusik
We introduce Soft Anisotropic Diagrams (SAD), an explicit and differentiable image representation parameterized by a set of adaptive sites in the image plane. In SAD, each site…
HuggingFacearXiv:2604.21984
Jona te Lintelo, Lichao Wu, Marina Krček, Sengim Karayalçin, Stjepan Picek
Mixture-of-Experts (MoE) architectures in Large Language Models (LLMs) have significantly reduced inference costs through sparse activation. However, this sparse activation…
HuggingFacearXiv:2604.27818
Zhen Ye, Xu Tan, Aoxiong Yin, Hongzhan Lin, Guangyan Zhang
Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities…
HuggingFacearXiv:2604.23586