Yaoyu Zhao, Yichen Xu, Oliver Bračevac, Cao Nguyen Pham, Frank Zhengqing Wu
LACUNA closes the runtime/code split in LLM agents by using typed holes agent[T](task) that are type-checked before execution, allowing model-written code to shape the runtime…
Aman Priyanshu, Supriti Vijay, Esha Pahwa
A Moltbook-style simulation with thousands of LLM agents interacting over a simulated month reveals that multi-turn social evaluation amplifies privacy violations from 19.95% to…
Jianing Zhu, Yeonju Ro, John Robertson, Kevin Wang, Junbo Li
AgingBench introduces longitudinal reliability measurement for deployed agents across four aging mechanisms: compression, interference, revision, and maintenance aging. Findings…
Haodong Zhao, Tianyi Xu, Tianhang Zhao, Zhuosheng Zhang, Gongshen Liu
GradSentry detects poisoned fine-tuning samples via spectral entropy of per-sample gradients, finding that poisoned samples produce higher-entropy gradient signatures. It requires…
Maikel Yelandi Leyva-Vázquez, Florentin Smarandache
Neutrosophic Logic (Truth, Indeterminacy, Falsity as independent dimensions) models epistemic states in LLMs, finding that hyper-truth (T+I+F > 1) emerges in 35% of evaluations,…
Andrea Gurioli, Davide D'Ascenzo, Federico Pennino, Maurizio Gabbrielli, Stefano Zacchiroli
SOURCETRACKER is a 300M-parameter code retrieval encoder that narrows candidates via vector search then re-ranks with Winnowing fingerprints, outperforming Winnowing alone by 5.4%…
Katharina Deckenbach, Haritz Puerto, Jonas Geiping, Sahar Abdelnabi
Fine-tuning models on synthetic documents describing evaluation traits makes them significantly safer on six benchmarks, suggesting evaluation meta-knowledge (parametric knowledge…
Jingwei Sun, Jianing Zhu, Jiangchao Yao, Tongliang Liu, Bo Han
TriMem maintains three coexisting memory representations (raw dialogue, atomic facts, synthesized profiles) with TextGrad-based prompt optimization for lifelong evolution without…
Jaihoon Kim, Taehoon Yoon, Prin Phunyaphibarn, Seungjun Kim, Morteza Mardani
CDM amortizes SMC inference in discrete diffusion models by learning a twist function via positive and negative samples, adding less than 5% computational overhead and…
Eric Onyame, Runtao Zhou, Kowshik Thopalli, Bhavya Kailkhura, Chirag Agarwal
CoT monitoring for detecting misaligned LLM behavior is fundamentally fragile across 13 languages and 7 frontier model families, with 95.9% average unfaithfulness rate and…