Marius Dragoi, Ioana Pintilie, Alexandra Dragomir, Antonio Barbalau, Florin Brad
Uses singular bases of pre-trained weights as a fixed reference frame with spectral penalty to route fine-grained adaptation into long-tail singular values, protecting principal components and reducing catastrophic forgetting.
cs.LG
Qintong Xie, Edward Koh, Xavier Cadet, Peter Chin
A solver-in-the-loop equilibrium supervision framework for training agents in multi-turn simultaneous games under partial observability, achieving significantly better performance than prior game-theoretic learning methods.
cs.GTcs.LG
August Y. Chen, Ahmed El Alaoui
Establishes a large deviation principle for how rare well-performing interpolating linear classifiers are under proportional scaling between sample size and dimension, with implications for neural network generalization theory.
math.STcs.LGmath.PR
Mykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp Hochreiter
Segment-level reward redistribution enables more granular credit assignment for intermediate reasoning steps, accelerating RL convergence and improving final answer quality on math and logic benchmarks.
cs.LGcs.AI
Paul Jünger, Justin Lovelace, Linxi Zhao, Dongyoung Go, Kilian Q. Weinberger
Discarded tokens in diffusion language models serve as a lookahead signal for retrieval — even low-confidence tokens surface salient entities early, enabling stronger evidence retrieval and reducing hallucinations.
cs.CLcs.AIcs.LG
Shangheng Du, Xiangchao Yan, Jinxin Shi, Zongsheng Cao, et al.
An LLM-based self-evolving multi-agent framework for end-to-end machine learning algorithm discovery with hierarchical control, persistent memory, and structured exploration — achieving new NASBench SOTA and discovering human-competitive architectures.
cs.AIcs.CL
Senmiao Wang, Tiantian Fang, Haoran Zhang, Yushun Zhang, et al.
Polynomial preconditioning reshapes weight singular-value spectra during training, ensuring stable conditioning throughout LLM pre-training with no inference overhead after merging — demonstrated on Llama-1B.
cs.LGcs.AI
Akarsh Kumar, Phillip Isola
Supervised Memory Training sidesteps recurrent credit propagation entirely, enabling highly parallel RNN training with better long-range association learning than BPTT — potentially bridging RNNs and transformers.
cs.LGcs.AI
Mingyang Liu, Asuman Ozdaglar, Tiancheng Yu, Kaiqing Zhang
Introduces Repeated Policy Regret (RP-Regret) for learning against adaptive opponents who respond to your historical play — fundamentally different from standard i.i.d. opponent models, necessary for auctions and security games.
cs.LGcs.AIcs.GT
Christie Djidjev, Nicholas Kaminski
An event detection framework for learning parameter-to-KPI dependencies in AI-RAN architectures, enabling proactive identification of harmful AI function interactions before service degradation occurs.
cs.LG