Hyeonwoo Kim, Jeonghwan Kim, Kyungwon Cho, Hanbyul Joo
DeVI leverages text-conditioned synthetic videos for physically plausible dexterous robot control using hybrid 3D/2D tracking rewards, enabling zero-shot generalization across diverse objects and interaction types.
cs.CV
Yupeng Zheng, Xiang Li, Songen Gu, Yuhang Zheng, Shuai Tian, Weize Li, Linbo Wang, Senyu Fei, Pengfei Li, Yinfeng Gao, Zebin Xing, Yilun Chen, Qichao Zhang, Haoran Li, Wenchao Ding
PokeVLA is a lightweight foundation model for embodied manipulation that infuses vision-language understanding into action learning via a two-stage training paradigm, achieving strong performance on pocket-sized hardware.
cs.RO
Sina Gholami, Abdulmoneam Ali, Tania Haghighi, Ahmed Arafa, Minhaj Nur Alam
FedSIR is a multi-stage framework for robust federated learning under noisy labels that leverages spectral structure of client feature representations to identify mislabeled clients and relabel them, improving performance without…
cs.LG cs.AI cs.CV cs.DC eess.SP
Ana Sanchez-Fernandez, Thomas Pinetz, Werner Zellinger, Günter Klambauer
Batch effects systematically undermine deep learning on new experimental batches in biomedical imaging. Control-Stabilized Adaptive (CSA) uses in-context control samples to close the domain gap across heterogeneous lab conditions.
cs.LG q-bio.QM
Thorsten Hoeser, Felix Bachofer, Claudia Kuenzer
A global Sentinel-1 SAR time series dataset enables high-temporal-resolution monitoring of offshore wind infrastructure deployment and operation at global scale, filling a critical gap in temporally dense, semantically fine-grained offshore…
cs.CV cs.LG
Yiming Bian, Joshua M. Akey
CQS Divide, derived from cyclic quorum sets theory, decomposes attention into composable operations that remove the assumption that full QKV tensors fit in device memory, enabling long-context LLMs without OOM…
cs.LG cs.DC
Deqing Fu, Tianyi Zhou, Mikhail Belkin, Vatsal Sharan, Robin Jia
Language models trained on natural text learn periodic number features with dominant periods at T=2, 5, 10. This paper identifies a two-tiered hierarchy: only some models learn geometrically separable features…
cs.CL cs.AI cs.LG
Shelly Golan, Michael Finkelson, Ariel Bereslavsky, Yotam Nitzan, Or Patashnik
ParetoSlider enables inference-time control over inherently conflicting goals during diffusion model post-training, replacing fixed scalarization with a Pareto-based approach that maintains multiple trade-off points.
cs.LG cs.CV
Qiguang Chen, Chengyu Luan, Jiajun Wu, Qiming Yu, Yi Yang, Yizhuo Li, Jingqi Tong, Xiachong Feng, Libo Qin, Wanxiang Che
OMIBench evaluates Olympiad-level reasoning when required evidence is distributed over multiple images, spanning biology, chemistry, mathematics, and physics. It reveals that current LVLMs fail to exploit cross-image contextual information that…
cs.CV cs.AI cs.CL
Dimitrije Antić, Alvaro Budria, George Paschalidis, Sai Kumar Dwivedi, Dimitrios Tzionas
InterFields encodes dense, continuous proximity across body and object surfaces to reconstruct 3D Human-Object Interaction from a single RGB image, going beyond sparse binary contact cues to model subtle physical…
cs.CV cs.LG
Pranava Madhyastha, Dagmar Adamcova
Human-like working memory constraints integrated into Transformers via fixed-width windows and temporal decay attention variants enable modified GPT-2 models trained on developmentally plausible datasets to align with human reading time…
cs.CL cs.AI cs.LG
Zesheng Liu, Maryam Rahnemoonfar
Graph-based models for ice stratigraphy extended to handle incomplete radar layer traces and entirely missing layers using physics-conditioned synthesis, enabling prediction of deeper-layer thickness from partially observed shallow layers.
cs.LG
Chao Wang, Luca Nepote, Giulio Franzese, Pietro Michiardi
A general framework for estimating KL divergence between probability distributions in function space, with applications to trajectory inference in single-cell genomics where path-space laws are non-identifiable from finitely many marginals.
cs.LG