2605.14438 | Juntong Wu, Jialiang Cheng, Qishen Yin, Yue Dai et al.
BEAM introduces trainable binary masks for token-adaptive expert selection in Mixture-of-Experts models, achieving 85% MoE FLOP reduction while retaining over 98% of model performance via a custom…
HUGGINGFACE_PAPERS
2605.14454 | Minbeom Kim, Lesly Miculicich, Bhavana Dalvi Mishra, Mihir Parmar et al.
LiSA is a conservative policy induction framework for lifelong guardrail adaptation in AI agents, converting sparse user-reported failures into reusable policy abstractions with conflict-aware local…
HUGGINGFACE_PAPERS
2605.13852 | Ido Sobol, Kihyuk Sohn, Yoav Blum, Egor Zakharov et al.
Realiz3D decouples control signals from visual domain in diffusion model training via co-variate residual adapters, preventing models from learning spurious associations between controls and…
HUGGINGFACE_PAPERS
2605.14445 | Runyuan He, Qiuyang Mang, Shang Zhou, Kaiyuan Liu et al.
FrontierSmith automatically evolves open-ended coding problems from closed-ended competitive programming tasks using idea divergence metrics to select problems that elicit diverse solutions.
HUGGINGFACE_PAPERS
2605.14068 | Amirreza Mohseni, Mona Mohammadi, Morteza Saghafian, Naser Talebizadeh Saradari | 1
CurveBench tests hierarchical topological reasoning from visual input; even Gemini 3.1 Pro achieves only 71.1% on easy and 19.1% on hard configurations, revealing topology-aware reasoning remains…
HUGGINGFACE_PAPERS
2605.11458 | Zihao Han, Tiangang Zhang, Huaibin Wang, Yilun Sun | 3
ATESD addresses teacher-side exposure mismatch in LLM self-distillation by treating teacher exposure as a learnable Beta-policy controller conditioned on training-state statistics.
HUGGINGFACE_PAPERS
2605.08703 | Yuxuan Zhang, Penghui Du, Bo Li, Cong Wei et al.
RewardHarness reframes reward modeling as context evolution, iteratively refining a library of tools and skills from as few as 100 preference demonstrations to align image-edit evaluation with human…
HUGGINGFACE_PAPERS
2605.11459 | Yanyan Zhang, Chaoda Song, Vikash Singh, Xinpeng Li et al.
Pace-and-Path Correction is a training-free inference-time operator for Vision-Language-Action models that decomposes motion correction into pace and path channels, improving success rates by up to…
HUGGINGFACE_PAPERS
2605.13169 | Changpeng Wang, Xin Lin, Junhan Liu, Yuheng Liu et al.
PanoWorld introduces pano-native understanding for MLLMs, requiring reasoning over equirectangular panorama as continuous observer-centered space, with Spherical Spatial Cross-Attention and…
HUGGINGFACE_PAPERS
2605.13027 | Zihang Xu, Xiaoyang Liu, Zheng Chen, Yulun Zhang et al.
PRISM addresses unreliable text conditions and ambiguous stroke boundaries in text-image super-resolution through Flow-Matching Prior Rectification and Structure-guided Uncertainty-aware Residual…
HUGGINGFACE_PAPERS
2605.14876 | Hanbo Cheng, Limin Lin, Ruo Zhang, Yicheng Pan et al.
CLVR couples visual-language planning with pixel-level diffusion generation via automated step-level visual verification and Proxy Prompt Reinforcement Learning to resolve long-context optimization…
HUGGINGFACE_PAPERS
2605.00180 | Jingjun Xu, Hongji Pu, Tao Feng, Haozhen Zhang et al.
RouteProfile characterizes LLM profiling along four dimensions, showing structured query-level profiles outperform flat ones for routing.
HUGGINGFACE_PAPERS
2605.09681 | Yicheng Ji, Zhizhou Zhong, Jun Zhang, Qin Yang et al.
Forcing-KV classifies attention heads into static and dynamic categories and applies hybrid KV cache pruning, achieving 29 fps on H200 with 30% cache memory reduction.
HUGGINGFACE_PAPERS
2605.14323 | Fangyuan Yu, Xin Su, Amir Abdullah | 1
DLR jointly learns discrete latent codes, routing policies, and model parameters through dynamic search in a single training stage, matching or outperforming SFT with a mean gain of 6.6 percentage…
HUGGINGFACE_PAPERS
2605.06527 | Hanxiang Chao, Yihan Bai, Rui Sheng, Tianle Li et al.
STALE identifies Implicit Conflict as a failure mode where later observations invalidate earlier memories without explicit negation, introducing a 400-scenario benchmark and CUPMem baseline for…
HUGGINGFACE_PAPERS
2605.15141 | Min Zhao, Hongzhou Zhu, Kaiwen Zheng, Zihan Zhou et al.
Causal Forcing++ uses causal consistency distillation for few-step AR video generation initialization, surpassing prior 4-step chunk-wise methods while reducing first-frame latency by 50% and…
HUGGINGFACE_PAPERS
2605.10912 | Shuangrui Ding, Xuanlang Dai, Long Xing, Shengyuan Ding et al.
WildClawBench evaluates agents in native Docker CLI runtimes with 60 bilingual tasks averaging 8 minutes and 20 tool calls; Claude Opus 4.7 reaches only 62.2%, showing long-horizon agent evaluation…
HUGGINGFACE_PAPERS
2605.14712 | Shijie Lian, Bin Yu, Xiaopeng Lin, Zhaolong Shen et al.
IntentVLA encodes recent visual observations into a short-horizon intent representation to condition chunk generation, reducing inter-chunk conflict in aliased robot manipulation.
HUGGINGFACE_PAPERS
2605.14906 | Xiyu Ren, Zhaowei Wang, Yiming Du, Zhongwei Xie et al.
MemLens benchmarks 27 LVLMs and 7 memory-augmented agents across 789 questions, finding long-context models degrade with conversation length while memory agents lose visual fidelity.
HUGGINGFACE_PAPERS
2605.15055 | Quanhao Li, Junqiu Yu, Kaixun Jiang, Yujie Wei et al.
DiffusionOPD trains task-specific teachers independently then distills them into a unified student along the student's own rollout trajectories, unifying single-task and multi-task RL for diffusion…
HUGGINGFACE_PAPERS
2605.15167 | Kam Man Wu, Haolin Yang, Qingyu Chen, Yihu Tang et al.
Pure synthetic layered graphic design data is shown to be a scalable and effective substitute for scarce proprietary layered assets in training layer decomposition models.
HUGGINGFACE_PAPERS
2605.15178 | Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye et al.
SANA-WM is a 2.6B open-source world model generating minute-long 720p video with precise camera control, trained on 213K public clips in 15 days on 64 H100s, deployable on a single RTX 5090.
HUGGINGFACE_PAPERS
2605.15186 | Kaixin Zhu, Yiwen Tang, Yifan Yang, Renrui Zhang et al.
VGGT-Edit performs text-conditioned native 3D scene editing in a single forward pass using depth-synchronized text injection and residual 3D geometric displacement prediction.
HUGGINGFACE_PAPERS
2605.14386 | Taebong Kim, Youngsik Hong, Minsik Kim, Sunyoung Choi et al.
Darwin-27B-Opus achieves 86.9% on GPQA Diamond (#6 of 1,252 models) with no gradient training, using MRI-Trust-weighted evolutionary merging across 14-dimensional adaptive merge genomes.
HUGGINGFACE_PAPERS
2605.13880 | Yumin Choi, Sangwoo Park, Minki Kang, Jinheon Baek et al.
Preping uses proposer-guided synthetic practice (Proposer-Solver-Validator) to construct agent memory before any target-environment tasks, achieving competitive performance with 2.99x lower…
HUGGINGFACE_PAPERS
2605.15128 | Minghao Guo, Qingyue Jiao, Zeru Shi, Yihao Quan et al.
MemEye evaluates visual evidence preservation in agent memory across 8 life-scenario tasks, showing current architectures struggle with fine-grained visual details and state change reasoning.
HUGGINGFACE_PAPERS
2605.15040 | Baolin Peng, Wenlin Yao, Qianhui Wu, Hao Cheng et al.
Orchard provides open-source scalable agentic modeling with Orchard-SWE achieving 67.5% on SWE-bench (new SOTA for open models) and Orchard-GUI training a 4B computer-use agent from 2.2K open-ended…
HUGGINGFACE_PAPERS
2605.13941 | Jiaqi Liu, Xinyu Ye, Peng Xia, Zeyu Zheng et al.
EvolveMem autonomously diagnoses retrieval failures and adjusts its retrieval configuration through LLM-powered meta-analysis, converging on effective strategies including novel configuration…
HUGGINGFACE_PAPERS
2605.13301 | Yafu Li, Runzhe Zhan, Haoran Zhang, Shunkai Zhang et al.
SU-01 applies reverse-perplexity curriculum SFT followed by two-stage RL and test-time scaling to a 30B-A3B backbone, reaching IMO gold-medal level with 100K+ token trajectories.
HUGGINGFACE_PAPERS
2605.13834 | Dongzhe Zheng, Tao Zhong, Christine Allen-Blanchette | 2
Hodge Spectral Duality (HSD) isolates topological degrees of freedom from learnable geometric dynamics in neural operators on geometric meshes, achieving superior accuracy and fidelity to physical…
HUGGINGFACE_PAPERS
2605.15182 | Yifan Wang, Tong He | 29
Warp-as-History is a training-free interface enabling a frozen video model to follow camera trajectories by constructing camera-warped pseudo-history with positional alignment.
HUGGINGFACE_PAPERS
2605.14389 | Sarkar Snigdha Sarathi Das, Palash Goyal, Mihir Parmar, Nanyun Peng et al.
Nexus decomposes time-series forecasting into specialized macro/micro temporal stages with contextual integration, outperforming state-of-the-art TSFMs on real estate and stock market data strictly…
HUGGINGFACE_PAPERS
2605.14392 | Yucheng Shi, Zhenwen Liang, Kishan Panaganti, Dian Yu et al.
EvoEnv synthesizes Python environments from seed tasks with staged validation and solver-relative difficulty calibration, reframing self-improvement as environment construction with reusable…
HUGGINGFACE_PAPERS
2605.15155 | Zhengxi Lu, Zhiyuan Yao, Zhuowen Han, Zi-Han Wang et al.
SDAR treats On-Policy Self-Distillation as a gated auxiliary objective with sigmoid gating for token-level positive-gap reinforcement, improving over GRPO by 9.4% on ALFWorld and 10.2% on WebShop.
HUGGINGFACE_PAPERS
2605.15185 | Jiaxin Wu, Yihao Pi, Yinling Zhang, Yuheng Li et al.
PDI-Bench introduces projective-geometry residuals (scale-depth alignment, 3D motion consistency, structural rigidity) for quantitative auditing of geometric coherence in generated videos.
HUGGINGFACE_PAPERS
2605.14354 | Sinclair Schneider, Florian Steuber, Gabi Dreo Rodosek | 0
An unsupervised framework combining few-shot filtering, UMAP dimensionality reduction, and HDBSCAN clustering identifies 41 distinct manipulative narrative clusters from 1.2 million social media…
HUGGINGFACE_PAPERS
2605.15198 | Ziyu Guo, Rain Liu, Xinyan Chen, Pheng-Ann Heng | 15
ATLAS uses a single discrete functional token as both an agentic operation and latent visual reasoning unit, with Latent-Anchored GRPO stabilizing RL training without architectural modifications.
HUGGINGFACE_PAPERS
2605.14892 | Shihao Qi, Jie Ma, Rui Xing, Wei Guo et al.
A unified survey of multi-agent LLM systems organized around the LIFE progression: Lay capability foundation, Integrate agents, Find faults, Evolve, identifying open challenges at stage boundaries.
HUGGINGFACE_PAPERS
2605.15188 | Shashwat Goel, Nikhil Chandak, Arvindh Arun, Ameya Prabhu et al.
FutureSim replays real-world news chronologically to test whether agents can predict world events beyond their knowledge cutoff; the best agent achieves only 25% accuracy.
HUGGINGFACE_PAPERS
2605.14269 | Yidong Huang, Zun Wang, Han Lin, Dong-Ki Kim et al.
PhyMotion recovers SMPL body meshes from generated videos, simulates them in MuJoCo, and scores motion along kinematic plausibility, contact balance, and dynamic feasibility dimensions.
HUGGINGFACE_PAPERS
2605.14051 | Yusuke Ozaki, Dhaval Patel | 1
SPIN combines validated DAG planning with prefix-based execution control for industrial LLM agents, reducing executed tasks by 41% and tool calls by 42% on AssetOpsBench.
HUGGINGFACE_PAPERS
2605.14169 | Letian Peng, Ziche Liu, Yiming Huang, Longfei Yun et al.
BOOKMARKS uses search-based bookmark management for role-playing agents, actively initializing and synchronizing task-relevant bookmarks at storyline points.
HUGGINGFACE_PAPERS
2605.15190 | Yanzuo Lu, Ronglai Zuo, Jiankang Deng | 4
RAVEN repacks self-rollouts into interleaved clean historical endpoints and noisy states for better training-inference alignment in causal AR video diffusion.
HUGGINGFACE_PAPERS
2605.14352 | Sinclair Schneider, Florian Steuber, Joao A. G. Schneider, Gabi Dreo Rodosek | 1
A transformer model projecting German political texts onto a continuous left-right spectrum achieves F1=0.844 in-domain and MAE=0.172 on newspaper out-of-domain tests.
HUGGINGFACE_PAPERS
2605.06607 | Nithin Somasekharan, Rabi Pathak, Manushri Dhanakoti, Tingwen Zhang et al.
AI CFD Scientist is the first AI scientist for computational fluid dynamics spanning literature-grounded ideation, OpenFOAM execution, vision-based physics verification, and manuscript writing.
HUGGINGFACE_PAPERS
2605.07865 | Minjae Oh, Sangjun Song, Gyubin Choi, Yunho Choi et al.
vOPD stabilizes On-Policy Distillation by casting it as policy-gradient RL with a control variate baseline derived from per-token reverse KL divergence, consistently outperforming vanilla OPD.
HUGGINGFACE_PAPERS
2605.14306 | Yuwen Du, Tian Jin, Jing Kang, Xianghe Pang et al.
PaSaMaster is a self-evolving agentic retrieval system using iterative intent analysis and ranking, improving F1 by 15.6X over keyword search across 38 scientific disciplines.
HUGGINGFACE_PAPERS