Yinyi Luo, Wenwen Wang, Hayes Bai, Marios Savvides, Jindong Wang
Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhibit functional inconsistency between these two capabilities. We observe that this issue d...
HuggingFace
Shuofei Qiao, Yunxiang Wei, Jiazheng Fan, Bin Wu, Busheng Zhang
The exponential growth of global academic output has confronted researchers and AI agents with an unprecedented ``information explosion,'' where fragmented and unstructured knowledge organization impedes deep interdisciplinary integration. ...
HuggingFace
Juncheng Wu, Hardy Chen, Haoqin Tu, Xianfeng Tang, Freda Shi
Recent advances in vision-language models (VLMs) emphasize long chain-of-thought reasoning; yet, we find that their performance on visual tasks is primarily limited by a lack of visual perception as opposed to reasoning itself. In this work...
HuggingFace
Shuhong Zheng, Michael Oechsle, Erik Sandström, Marie-Julie Rakotosaona, Federico Tombari
Visual geometry transformers have become powerful architectures for multi-view 3D reconstruction, enabling joint prediction of multiple 3D attributes in a feed-forward manner. However, their computational cost grows quadratically with the i...
HuggingFace
Jiarui Guo, Haojia Wei, Yiming Zhang, Yifei Liu, Yuning Gong
Virtual photography asks an agent to enter a prepared 3D scene with no preselected camera pose or reference image, infer a suitable shot from scene information and a language intent, choose executable camera parameters, and render the final...
HuggingFace
Karan Goyal
The rapid proliferation of Vision-Language Models (VLMs) is often framed as enabling unified multimodal knowledge discovery but rests on an under-examined assumption: that current VLMs faithfully synthesise multimodal data. We argue they of...
HuggingFace
Boyuan Sun, Bowen Yin, Yuanming Li, Xihan Wei, Qibin Hou
We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine-grained object understanding solely from textual prompts. Unlike existing approaches that require explicit visual pr...
HuggingFace
Zizun Li, Haoyu Guo, Runzhe Teng, Chunhua Shen, Tong He
Camera-controlled video generation has achieved remarkable progress in recent years. However, existing video-to-video re-rendering methods primarily rely on Supervised Fine-Tuning using synthetic datasets. At present, there is an extreme sc...
HuggingFace
Chao Xu, Maohua Li, Qirui Li, Yixuan Xu, Yanke Zhou
Diffusion Transformers (DiTs) have become a de facto backbone of modern visual generation, and nearly every major axis of their design -- tokenization, attention, conditioning, objectives, and latent autoencoders -- has been extensively rev...
HuggingFace
Siyong Jian, Siyuan Li, Luyuan Zhang, Zedong Wang, Xin Jin
Discrete autoregressive (AR) text-to-image (T2I) models pair a VQ tokenizer with an AR policy, and current post-training pipelines optimize only the policy while keeping the VQ decoder frozen. Recent diffusion T2I work, exemplified by REPA-...
HuggingFace
Zizhao Tong, Hongfeng Lai, Zeqing Wang, Zhaohu Xing, Kexu Cheng
Interactive world models for first-person shooter (FPS) games must resolve high-frequency overlapping control signals at every frame without disrupting unaffected regions. Existing methods inject actions globally and train on single titles,...
HuggingFace
Jinho Park, Youbin Kim, Hogun Park, Eunbyung Park
Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential challenge. However, existing spatio-temporal reasoning benchma...
HuggingFace
Bin Lin, Bo Zhao, Boyong Wu, Chao Yan, Chen Wu
Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks. However, existing unified foundations often struggle to mat...
HuggingFace
Woongyeng Yeo, Yumin Choi, Taekyung Ki, Sung Ju Hwang
Training long-horizon LLM agents with reinforcement learning is challenging because sparse outcome rewards reveal whether a task succeeds, but not which intermediate actions caused the outcome or how they should be corrected. Recent methods...
HuggingFace
Chongyu Fan, Gaowen Liu, Mingyi Hong, Ramana Rao Kompella, Sijia Liu
Muon is a matrix-aware optimizer that leverages Newton-Schulz (NS) iterations to enforce spectral gradient orthogonalization by driving all singular values of the momentum matrix toward 1. While this uniform spectral whitening enhances expl...
HuggingFace
Dong Chen, Fangyun Wei, Ziyu Wan, Dongdong Chen, Jiawei Zhang
We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with more than 6B parameters across various benchmarks, while requiring significantly less tr...
HuggingFace
Beichen Zhang, Yuhong Liu, Jinsong Li, Yuhang Zang, Jiaqi Wang
Multimodal Large Language Models have advanced visual reasoning, yet a purely textual chain of thought remains a bottleneck for questions that require fine-grained focus or view transformations. The ''think with images'' paradigm narrows th...
HuggingFace
Katharina Schmid, Nicolas von Lützow, Jozef Hladký, Angela Dai, Matthias Nießner
We introduce a new approach to high-fidelity 3D scene reconstruction from multi-view RGB images that tightly couples reconstruction with a strong generative 3D prior. We cast scene reconstruction as conditional 3D generation over a set of s...
HuggingFace
Xu Ouyang, Deyi Liu, Yuhang Cai, Jing Liu, Yuan Yang
Existing scaling laws for Large Language Models (LLMs), predominantly monotonic power laws, fail to explain emerging non-monotonic phenomena such as catastrophic overtraining and quantization-induced degradation, where performance deteriora...
HuggingFace
Zisu Huang, Jingwen Xu, Yifan Yang, Ziyang Gong, Qihao Yang
Language agents increasingly improve by reusing skills -- structured procedural artifacts distilled from past experience. In particular, domain-level and model-generated skills are especially promising. They offer fast adaptation within a d...
HuggingFace
Yifan Yang, Ziyang Gong, Weiquan Huang, Qihao Yang, Ziwei Zhou
Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning optimizer for the skill, and none of which reliably improves over its starting point un...
HuggingFace
Yifan Lu, Qi Wu, Jay Zhangjie Wu, Zian Wang, Huan Ling
Most practical high-resolution text-to-image systems, including latent diffusion and autoregressive models, perform generation in a compact latent space, and a decoder maps the generated latents back to pixels. Yet the latent-to-pixel decod...
HuggingFace
Pablo Marcos-Manchón, Rishi Jha, Lluís Fuentemilla
The Strong Platonic Representation Hypothesis suggests that representational convergence in artificial neural networks can be harnessed constructively: embeddings can be translated across models through a universal latent space without pair...
HuggingFace
Víctor Yeste, Paolo Rosso
Detecting Schwartz values in political text is difficult because implicit cues often depend on surrounding arguments and fine-grained distinctions between neighboring values. We study when context and explicit moral knowledge help sentence-...
HuggingFace
Yifan Dai, Zhenhua Wu, Bohan Zeng, Daili Hua, Jialing Liu
Joint audio-visual reasoning is essential for omnimodal understanding, yet current multimodal large language models (MLLMs) still struggle when reasoning requires fine-grained evidence from both modalities. A central limitation is that expl...
HuggingFace