Carlos's Debrief

May 15, 2026 04:00
0ArXiv Papers
59Web Findings
59Total Sources
Topics searched: Artificial Intelligence Machine Learning Large Language Models Security & Cybersecurity Cryptography Zero Knowledge Quantum Computing Crypto & Blockchain AI Agents & Reasoning AI Safety & Alignment
May 15, 2026 04:00 — Curated by Hermes Research Scout

⚡ Quick Summary

📰 News Headlines

📄 ArXiv Papers

🦠 LLMs 14

Tags: ai | Score: 1
We let Codex and Claude Code autonomously iterate on the nanoGPT speedrun optimizer track for two weeks, producing ~10k runs, a new 2930-step record, and a detailed look at where autonomous research…
HF
Tags: privacy, web | Score: 69
Reverse-engineering how Cloudflare's JavaScript middleware reads ChatGPT's React component state before allowing user input.
HF
2605.14438 | Juntong Wu, Jialiang Cheng, Qishen Yin, Yue Dai et al.
BEAM introduces trainable binary masks for token-adaptive expert selection in Mixture-of-Experts models, achieving 85% MoE FLOP reduction while retaining over 98% of model performance via a custom…
HF
2605.14068 | Amirreza Mohseni, Mona Mohammadi, Morteza Saghafian, Naser Talebizadeh Saradari | 1
CurveBench tests hierarchical topological reasoning from visual input; even Gemini 3.1 Pro achieves only 71.1% on easy and 19.1% on hard configurations, revealing topology-aware reasoning remains…
HF
2605.11458 | Zihao Han, Tiangang Zhang, Huaibin Wang, Yilun Sun | 3
ATESD addresses teacher-side exposure mismatch in LLM self-distillation by treating teacher exposure as a learnable Beta-policy controller conditioned on training-state statistics.
HF
2605.13169 | Changpeng Wang, Xin Lin, Junhan Liu, Yuheng Liu et al.
PanoWorld introduces pano-native understanding for MLLMs, requiring reasoning over equirectangular panorama as continuous observer-centered space, with Spherical Spatial Cross-Attention and…
HF
2605.00180 | Jingjun Xu, Hongji Pu, Tao Feng, Haozhen Zhang et al.
RouteProfile characterizes LLM profiling along four dimensions, showing structured query-level profiles outperform flat ones for routing.
HF
2605.13301 | Yafu Li, Runzhe Zhan, Haoran Zhang, Shunkai Zhang et al.
SU-01 applies reverse-perplexity curriculum SFT followed by two-stage RL and test-time scaling to a 30B-A3B backbone, reaching IMO gold-medal level with 100K+ token trajectories.
HF
2605.14389 | Sarkar Snigdha Sarathi Das, Palash Goyal, Mihir Parmar, Nanyun Peng et al.
Nexus decomposes time-series forecasting into specialized macro/micro temporal stages with contextual integration, outperforming state-of-the-art TSFMs on real estate and stock market data strictly…
HF
2605.14392 | Yucheng Shi, Zhenwen Liang, Kishan Panaganti, Dian Yu et al.
EvoEnv synthesizes Python environments from seed tasks with staged validation and solver-relative difficulty calibration, reframing self-improvement as environment construction with reusable…
HF
2605.15198 | Ziyu Guo, Rain Liu, Xinyan Chen, Pheng-Ann Heng | 15
ATLAS uses a single discrete functional token as both an agentic operation and latent visual reasoning unit, with Latent-Anchored GRPO stabilizing RL training without architectural modifications.
HF
2605.14051 | Yusuke Ozaki, Dhaval Patel | 1
SPIN combines validated DAG planning with prefix-based execution control for industrial LLM agents, reducing executed tasks by 41% and tool calls by 42% on AssetOpsBench.
HF
The advent of large language models (LLMs) has marked a new era in computational social science.
A comprehensive review covering how LLMs are applied in computational social science, including social phenomenon analysis, data bias and privacy challenges.
HF
Overall, AI teammates altered the process, the distribution and connection of moral frames...
A study examining how generative AI personas function as teammates in collaborative moral reasoning, finding that AI teammates alter the process and distribution of moral frames among human…
HF

⚙️ Machine Learning 29

Tags: ai, culture, practices, programming et al.
An essay on how the rise of AI-assisted coding is degrading programmers' ability to reason through bugs and edge cases independently.
HF
2605.13852 | Ido Sobol, Kihyuk Sohn, Yoav Blum, Egor Zakharov et al.
Realiz3D decouples control signals from visual domain in diffusion model training via co-variate residual adapters, preventing models from learning spurious associations between controls and…
HF
2605.14445 | Runyuan He, Qiuyang Mang, Shang Zhou, Kaiyuan Liu et al.
FrontierSmith automatically evolves open-ended coding problems from closed-ended competitive programming tasks using idea divergence metrics to select problems that elicit diverse solutions.
HF
2605.08703 | Yuxuan Zhang, Penghui Du, Bo Li, Cong Wei et al.
RewardHarness reframes reward modeling as context evolution, iteratively refining a library of tools and skills from as few as 100 preference demonstrations to align image-edit evaluation with human…
HF
2605.11459 | Yanyan Zhang, Chaoda Song, Vikash Singh, Xinpeng Li et al.
Pace-and-Path Correction is a training-free inference-time operator for Vision-Language-Action models that decomposes motion correction into pace and path channels, improving success rates by up to…
HF
2605.13027 | Zihang Xu, Xiaoyang Liu, Zheng Chen, Yulun Zhang et al.
PRISM addresses unreliable text conditions and ambiguous stroke boundaries in text-image super-resolution through Flow-Matching Prior Rectification and Structure-guided Uncertainty-aware Residual…
HF
2605.14876 | Hanbo Cheng, Limin Lin, Ruo Zhang, Yicheng Pan et al.
CLVR couples visual-language planning with pixel-level diffusion generation via automated step-level visual verification and Proxy Prompt Reinforcement Learning to resolve long-context optimization…
HF
2605.09681 | Yicheng Ji, Zhizhou Zhong, Jun Zhang, Qin Yang et al.
Forcing-KV classifies attention heads into static and dynamic categories and applies hybrid KV cache pruning, achieving 29 fps on H200 with 30% cache memory reduction.
HF
2605.14323 | Fangyuan Yu, Xin Su, Amir Abdullah | 1
DLR jointly learns discrete latent codes, routing policies, and model parameters through dynamic search in a single training stage, matching or outperforming SFT with a mean gain of 6.6 percentage…
HF
2605.15141 | Min Zhao, Hongzhou Zhu, Kaiwen Zheng, Zihan Zhou et al.
Causal Forcing++ uses causal consistency distillation for few-step AR video generation initialization, surpassing prior 4-step chunk-wise methods while reducing first-frame latency by 50% and…
HF
2605.14712 | Shijie Lian, Bin Yu, Xiaopeng Lin, Zhaolong Shen et al.
IntentVLA encodes recent visual observations into a short-horizon intent representation to condition chunk generation, reducing inter-chunk conflict in aliased robot manipulation.
HF
2605.14906 | Xiyu Ren, Zhaowei Wang, Yiming Du, Zhongwei Xie et al.
MemLens benchmarks 27 LVLMs and 7 memory-augmented agents across 789 questions, finding long-context models degrade with conversation length while memory agents lose visual fidelity.
HF
2605.15055 | Quanhao Li, Junqiu Yu, Kaixun Jiang, Yujie Wei et al.
DiffusionOPD trains task-specific teachers independently then distills them into a unified student along the student's own rollout trajectories, unifying single-task and multi-task RL for diffusion…
HF
2605.15167 | Kam Man Wu, Haolin Yang, Qingyu Chen, Yihu Tang et al.
Pure synthetic layered graphic design data is shown to be a scalable and effective substitute for scarce proprietary layered assets in training layer decomposition models.
HF
2605.15178 | Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye et al.
SANA-WM is a 2.6B open-source world model generating minute-long 720p video with precise camera control, trained on 213K public clips in 15 days on 64 H100s, deployable on a single RTX 5090.
HF
2605.15186 | Kaixin Zhu, Yiwen Tang, Yifan Yang, Renrui Zhang et al.
VGGT-Edit performs text-conditioned native 3D scene editing in a single forward pass using depth-synchronized text injection and residual 3D geometric displacement prediction.
HF
2605.14386 | Taebong Kim, Youngsik Hong, Minsik Kim, Sunyoung Choi et al.
Darwin-27B-Opus achieves 86.9% on GPQA Diamond (#6 of 1,252 models) with no gradient training, using MRI-Trust-weighted evolutionary merging across 14-dimensional adaptive merge genomes.
HF
2605.15040 | Baolin Peng, Wenlin Yao, Qianhui Wu, Hao Cheng et al.
Orchard provides open-source scalable agentic modeling with Orchard-SWE achieving 67.5% on SWE-bench (new SOTA for open models) and Orchard-GUI training a 4B computer-use agent from 2.2K open-ended…
HF
2605.13834 | Dongzhe Zheng, Tao Zhong, Christine Allen-Blanchette | 2
Hodge Spectral Duality (HSD) isolates topological degrees of freedom from learnable geometric dynamics in neural operators on geometric meshes, achieving superior accuracy and fidelity to physical…
HF
2605.15182 | Yifan Wang, Tong He | 29
Warp-as-History is a training-free interface enabling a frozen video model to follow camera trajectories by constructing camera-warped pseudo-history with positional alignment.
HF
2605.15155 | Zhengxi Lu, Zhiyuan Yao, Zhuowen Han, Zi-Han Wang et al.
SDAR treats On-Policy Self-Distillation as a gated auxiliary objective with sigmoid gating for token-level positive-gap reinforcement, improving over GRPO by 9.4% on ALFWorld and 10.2% on WebShop.
HF
2605.15185 | Jiaxin Wu, Yihao Pi, Yinling Zhang, Yuheng Li et al.
PDI-Bench introduces projective-geometry residuals (scale-depth alignment, 3D motion consistency, structural rigidity) for quantitative auditing of geometric coherence in generated videos.
HF
2605.14354 | Sinclair Schneider, Florian Steuber, Gabi Dreo Rodosek | 0
An unsupervised framework combining few-shot filtering, UMAP dimensionality reduction, and HDBSCAN clustering identifies 41 distinct manipulative narrative clusters from 1.2 million social media…
HF
2605.14269 | Yidong Huang, Zun Wang, Han Lin, Dong-Ki Kim et al.
PhyMotion recovers SMPL body meshes from generated videos, simulates them in MuJoCo, and scores motion along kinematic plausibility, contact balance, and dynamic feasibility dimensions.
HF
2605.15190 | Yanzuo Lu, Ronglai Zuo, Jiankang Deng | 4
RAVEN repacks self-rollouts into interleaved clean historical endpoints and noisy states for better training-inference alignment in causal AR video diffusion.
HF
2605.14352 | Sinclair Schneider, Florian Steuber, Joao A. G. Schneider, Gabi Dreo Rodosek | 1
A transformer model projecting German political texts onto a continuous left-right spectrum achieves F1=0.844 in-domain and MAE=0.172 on newspaper out-of-domain tests.
HF
2605.07865 | Minjae Oh, Sangjun Song, Gyubin Choi, Yunho Choi et al.
vOPD stabilizes On-Policy Distillation by casting it as policy-gradient RL with a control variate baseline derived from per-token reverse KL divergence, consistently outperforming vanilla OPD.
HF
Rising levels of atmospheric carbon dioxide and methane have sparked the interest of researchers in resolving this issue.
A review of ML-driven catalyst research for converting greenhouse gases (CO2 and methane) into useful chemicals via dry reforming processes.
HF
in quantum-enhanced classical ML to native quantum algorithms and hybrid quantum approaches...
A review covering quantum-enhanced classical ML to native quantum algorithms, with applications in optimization, drug discovery, and quantum-secured communications.
HF

🔒 Security & Crypto 5

Tags: security | Score: 1
Public CVE disclosure volumes are surging across major software suppliers and open source projects, with evidence increasingly pointing to AI-assisted vulnerability discovery as the driving force.
HF
Tags: privacy, security | Score: 9
Analysis of how Mullvad VPN's smaller server fleet creates a fingerprinting vector that tracks users across sites despite its no-logging reputation.
HF
Tags: linux, security | Score: 34
A Linux kernel 0-day exploit allowing an unprivileged user to steal SSH host private keys and /etc/shadow via ptrace_may_access mm-NULL bypass on pre-fixed kernels.
HF
in adversarial AI, automated threat intelligence, and AI-driven security orchestration...
A review of AI and ML applications in cybersecurity, covering adversarial AI, automated threat intelligence, and AI-driven security orchestration.
HF
the need for post-quantum cryptography to secure devices and user privacy against quantum-enabled adversaries.
A review of post-quantum cryptographic approaches for securing IoT devices and user privacy against quantum-enabled adversaries.
HF

🛡️ AI Safety 1

2605.14454 | Minbeom Kim, Lesly Miculicich, Bhavana Dalvi Mishra, Mihir Parmar et al.
LiSA is a conservative policy induction framework for lifelong guardrail adaptation in AI agents, converting sparse user-reported failures into reusable policy abstractions with conflict-aware local…
HF

🤖 Agents 10

2605.06527 | Hanxiang Chao, Yihan Bai, Rui Sheng, Tianle Li et al.
STALE identifies Implicit Conflict as a failure mode where later observations invalidate earlier memories without explicit negation, introducing a 400-scenario benchmark and CUPMem baseline for…
HF
2605.10912 | Shuangrui Ding, Xuanlang Dai, Long Xing, Shengyuan Ding et al.
WildClawBench evaluates agents in native Docker CLI runtimes with 60 bilingual tasks averaging 8 minutes and 20 tool calls; Claude Opus 4.7 reaches only 62.2%, showing long-horizon agent evaluation…
HF
2605.13880 | Yumin Choi, Sangwoo Park, Minki Kang, Jinheon Baek et al.
Preping uses proposer-guided synthetic practice (Proposer-Solver-Validator) to construct agent memory before any target-environment tasks, achieving competitive performance with 2.99x lower…
HF
2605.15128 | Minghao Guo, Qingyue Jiao, Zeru Shi, Yihao Quan et al.
MemEye evaluates visual evidence preservation in agent memory across 8 life-scenario tasks, showing current architectures struggle with fine-grained visual details and state change reasoning.
HF
2605.13941 | Jiaqi Liu, Xinyu Ye, Peng Xia, Zeyu Zheng et al.
EvolveMem autonomously diagnoses retrieval failures and adjusts its retrieval configuration through LLM-powered meta-analysis, converging on effective strategies including novel configuration…
HF
2605.14892 | Shihao Qi, Jie Ma, Rui Xing, Wei Guo et al.
A unified survey of multi-agent LLM systems organized around the LIFE progression: Lay capability foundation, Integrate agents, Find faults, Evolve, identifying open challenges at stage boundaries.
HF
2605.15188 | Shashwat Goel, Nikhil Chandak, Arvindh Arun, Ameya Prabhu et al.
FutureSim replays real-world news chronologically to test whether agents can predict world events beyond their knowledge cutoff; the best agent achieves only 25% accuracy.
HF
2605.14169 | Letian Peng, Ziche Liu, Yiming Huang, Longfei Yun et al.
BOOKMARKS uses search-based bookmark management for role-playing agents, actively initializing and synchronizing task-relevant bookmarks at storyline points.
HF
2605.06607 | Nithin Somasekharan, Rabi Pathak, Manushri Dhanakoti, Tingwen Zhang et al.
AI CFD Scientist is the first AI scientist for computational fluid dynamics spanning literature-grounded ideation, OpenFOAM execution, vision-based physics verification, and manuscript writing.
HF
2605.14306 | Yuwen Du, Tian Jin, Jing Kang, Xianghe Pang et al.
PaSaMaster is a self-evolving agentic retrieval system using iterative intent analysis and ranking, improving F1 by 15.6X over keyword search across 38 scientific disciplines.
HF

🌐 Web Findings

🦞 Lobste.rs 6

Tags: ai | Score: 1
We let Codex and Claude Code autonomously iterate on the nanoGPT speedrun optimizer track for two weeks, producing ~10k runs, a new 2930-step record, and a detailed look at where autonomous research…
LOBSTERS
Tags: ai, culture, practices, programming et al.
An essay on how the rise of AI-assisted coding is degrading programmers' ability to reason through bugs and edge cases independently.
LOBSTERS
Tags: security | Score: 1
Public CVE disclosure volumes are surging across major software suppliers and open source projects, with evidence increasingly pointing to AI-assisted vulnerability discovery as the driving force.
LOBSTERS
Tags: privacy, security | Score: 9
Analysis of how Mullvad VPN's smaller server fleet creates a fingerprinting vector that tracks users across sites despite its no-logging reputation.
LOBSTERS
Tags: linux, security | Score: 34
A Linux kernel 0-day exploit allowing an unprivileged user to steal SSH host private keys and /etc/shadow via ptrace_may_access mm-NULL bypass on pre-fixed kernels.
LOBSTERS
Tags: privacy, web | Score: 69
Reverse-engineering how Cloudflare's JavaScript middleware reads ChatGPT's React component state before allowing user input.
LOBSTERS

🤗 HuggingFace Papers 47

2605.14438 | Juntong Wu, Jialiang Cheng, Qishen Yin, Yue Dai et al.
BEAM introduces trainable binary masks for token-adaptive expert selection in Mixture-of-Experts models, achieving 85% MoE FLOP reduction while retaining over 98% of model performance via a custom…
HUGGINGFACE_PAPERS
2605.14454 | Minbeom Kim, Lesly Miculicich, Bhavana Dalvi Mishra, Mihir Parmar et al.
LiSA is a conservative policy induction framework for lifelong guardrail adaptation in AI agents, converting sparse user-reported failures into reusable policy abstractions with conflict-aware local…
HUGGINGFACE_PAPERS
2605.13852 | Ido Sobol, Kihyuk Sohn, Yoav Blum, Egor Zakharov et al.
Realiz3D decouples control signals from visual domain in diffusion model training via co-variate residual adapters, preventing models from learning spurious associations between controls and…
HUGGINGFACE_PAPERS
2605.14445 | Runyuan He, Qiuyang Mang, Shang Zhou, Kaiyuan Liu et al.
FrontierSmith automatically evolves open-ended coding problems from closed-ended competitive programming tasks using idea divergence metrics to select problems that elicit diverse solutions.
HUGGINGFACE_PAPERS
2605.14068 | Amirreza Mohseni, Mona Mohammadi, Morteza Saghafian, Naser Talebizadeh Saradari | 1
CurveBench tests hierarchical topological reasoning from visual input; even Gemini 3.1 Pro achieves only 71.1% on easy and 19.1% on hard configurations, revealing topology-aware reasoning remains…
HUGGINGFACE_PAPERS
2605.11458 | Zihao Han, Tiangang Zhang, Huaibin Wang, Yilun Sun | 3
ATESD addresses teacher-side exposure mismatch in LLM self-distillation by treating teacher exposure as a learnable Beta-policy controller conditioned on training-state statistics.
HUGGINGFACE_PAPERS
2605.08703 | Yuxuan Zhang, Penghui Du, Bo Li, Cong Wei et al.
RewardHarness reframes reward modeling as context evolution, iteratively refining a library of tools and skills from as few as 100 preference demonstrations to align image-edit evaluation with human…
HUGGINGFACE_PAPERS
2605.11459 | Yanyan Zhang, Chaoda Song, Vikash Singh, Xinpeng Li et al.
Pace-and-Path Correction is a training-free inference-time operator for Vision-Language-Action models that decomposes motion correction into pace and path channels, improving success rates by up to…
HUGGINGFACE_PAPERS
2605.13169 | Changpeng Wang, Xin Lin, Junhan Liu, Yuheng Liu et al.
PanoWorld introduces pano-native understanding for MLLMs, requiring reasoning over equirectangular panorama as continuous observer-centered space, with Spherical Spatial Cross-Attention and…
HUGGINGFACE_PAPERS
2605.13027 | Zihang Xu, Xiaoyang Liu, Zheng Chen, Yulun Zhang et al.
PRISM addresses unreliable text conditions and ambiguous stroke boundaries in text-image super-resolution through Flow-Matching Prior Rectification and Structure-guided Uncertainty-aware Residual…
HUGGINGFACE_PAPERS
2605.14876 | Hanbo Cheng, Limin Lin, Ruo Zhang, Yicheng Pan et al.
CLVR couples visual-language planning with pixel-level diffusion generation via automated step-level visual verification and Proxy Prompt Reinforcement Learning to resolve long-context optimization…
HUGGINGFACE_PAPERS
2605.00180 | Jingjun Xu, Hongji Pu, Tao Feng, Haozhen Zhang et al.
RouteProfile characterizes LLM profiling along four dimensions, showing structured query-level profiles outperform flat ones for routing.
HUGGINGFACE_PAPERS
2605.09681 | Yicheng Ji, Zhizhou Zhong, Jun Zhang, Qin Yang et al.
Forcing-KV classifies attention heads into static and dynamic categories and applies hybrid KV cache pruning, achieving 29 fps on H200 with 30% cache memory reduction.
HUGGINGFACE_PAPERS
2605.14323 | Fangyuan Yu, Xin Su, Amir Abdullah | 1
DLR jointly learns discrete latent codes, routing policies, and model parameters through dynamic search in a single training stage, matching or outperforming SFT with a mean gain of 6.6 percentage…
HUGGINGFACE_PAPERS
2605.06527 | Hanxiang Chao, Yihan Bai, Rui Sheng, Tianle Li et al.
STALE identifies Implicit Conflict as a failure mode where later observations invalidate earlier memories without explicit negation, introducing a 400-scenario benchmark and CUPMem baseline for…
HUGGINGFACE_PAPERS
2605.15141 | Min Zhao, Hongzhou Zhu, Kaiwen Zheng, Zihan Zhou et al.
Causal Forcing++ uses causal consistency distillation for few-step AR video generation initialization, surpassing prior 4-step chunk-wise methods while reducing first-frame latency by 50% and…
HUGGINGFACE_PAPERS
2605.10912 | Shuangrui Ding, Xuanlang Dai, Long Xing, Shengyuan Ding et al.
WildClawBench evaluates agents in native Docker CLI runtimes with 60 bilingual tasks averaging 8 minutes and 20 tool calls; Claude Opus 4.7 reaches only 62.2%, showing long-horizon agent evaluation…
HUGGINGFACE_PAPERS
2605.14712 | Shijie Lian, Bin Yu, Xiaopeng Lin, Zhaolong Shen et al.
IntentVLA encodes recent visual observations into a short-horizon intent representation to condition chunk generation, reducing inter-chunk conflict in aliased robot manipulation.
HUGGINGFACE_PAPERS
2605.14906 | Xiyu Ren, Zhaowei Wang, Yiming Du, Zhongwei Xie et al.
MemLens benchmarks 27 LVLMs and 7 memory-augmented agents across 789 questions, finding long-context models degrade with conversation length while memory agents lose visual fidelity.
HUGGINGFACE_PAPERS
2605.15055 | Quanhao Li, Junqiu Yu, Kaixun Jiang, Yujie Wei et al.
DiffusionOPD trains task-specific teachers independently then distills them into a unified student along the student's own rollout trajectories, unifying single-task and multi-task RL for diffusion…
HUGGINGFACE_PAPERS
2605.15167 | Kam Man Wu, Haolin Yang, Qingyu Chen, Yihu Tang et al.
Pure synthetic layered graphic design data is shown to be a scalable and effective substitute for scarce proprietary layered assets in training layer decomposition models.
HUGGINGFACE_PAPERS
2605.15178 | Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye et al.
SANA-WM is a 2.6B open-source world model generating minute-long 720p video with precise camera control, trained on 213K public clips in 15 days on 64 H100s, deployable on a single RTX 5090.
HUGGINGFACE_PAPERS
2605.15186 | Kaixin Zhu, Yiwen Tang, Yifan Yang, Renrui Zhang et al.
VGGT-Edit performs text-conditioned native 3D scene editing in a single forward pass using depth-synchronized text injection and residual 3D geometric displacement prediction.
HUGGINGFACE_PAPERS
2605.14386 | Taebong Kim, Youngsik Hong, Minsik Kim, Sunyoung Choi et al.
Darwin-27B-Opus achieves 86.9% on GPQA Diamond (#6 of 1,252 models) with no gradient training, using MRI-Trust-weighted evolutionary merging across 14-dimensional adaptive merge genomes.
HUGGINGFACE_PAPERS
2605.13880 | Yumin Choi, Sangwoo Park, Minki Kang, Jinheon Baek et al.
Preping uses proposer-guided synthetic practice (Proposer-Solver-Validator) to construct agent memory before any target-environment tasks, achieving competitive performance with 2.99x lower…
HUGGINGFACE_PAPERS
2605.15128 | Minghao Guo, Qingyue Jiao, Zeru Shi, Yihao Quan et al.
MemEye evaluates visual evidence preservation in agent memory across 8 life-scenario tasks, showing current architectures struggle with fine-grained visual details and state change reasoning.
HUGGINGFACE_PAPERS
2605.15040 | Baolin Peng, Wenlin Yao, Qianhui Wu, Hao Cheng et al.
Orchard provides open-source scalable agentic modeling with Orchard-SWE achieving 67.5% on SWE-bench (new SOTA for open models) and Orchard-GUI training a 4B computer-use agent from 2.2K open-ended…
HUGGINGFACE_PAPERS
2605.13941 | Jiaqi Liu, Xinyu Ye, Peng Xia, Zeyu Zheng et al.
EvolveMem autonomously diagnoses retrieval failures and adjusts its retrieval configuration through LLM-powered meta-analysis, converging on effective strategies including novel configuration…
HUGGINGFACE_PAPERS
2605.13301 | Yafu Li, Runzhe Zhan, Haoran Zhang, Shunkai Zhang et al.
SU-01 applies reverse-perplexity curriculum SFT followed by two-stage RL and test-time scaling to a 30B-A3B backbone, reaching IMO gold-medal level with 100K+ token trajectories.
HUGGINGFACE_PAPERS
2605.13834 | Dongzhe Zheng, Tao Zhong, Christine Allen-Blanchette | 2
Hodge Spectral Duality (HSD) isolates topological degrees of freedom from learnable geometric dynamics in neural operators on geometric meshes, achieving superior accuracy and fidelity to physical…
HUGGINGFACE_PAPERS
2605.15182 | Yifan Wang, Tong He | 29
Warp-as-History is a training-free interface enabling a frozen video model to follow camera trajectories by constructing camera-warped pseudo-history with positional alignment.
HUGGINGFACE_PAPERS
2605.14389 | Sarkar Snigdha Sarathi Das, Palash Goyal, Mihir Parmar, Nanyun Peng et al.
Nexus decomposes time-series forecasting into specialized macro/micro temporal stages with contextual integration, outperforming state-of-the-art TSFMs on real estate and stock market data strictly…
HUGGINGFACE_PAPERS
2605.14392 | Yucheng Shi, Zhenwen Liang, Kishan Panaganti, Dian Yu et al.
EvoEnv synthesizes Python environments from seed tasks with staged validation and solver-relative difficulty calibration, reframing self-improvement as environment construction with reusable…
HUGGINGFACE_PAPERS
2605.15155 | Zhengxi Lu, Zhiyuan Yao, Zhuowen Han, Zi-Han Wang et al.
SDAR treats On-Policy Self-Distillation as a gated auxiliary objective with sigmoid gating for token-level positive-gap reinforcement, improving over GRPO by 9.4% on ALFWorld and 10.2% on WebShop.
HUGGINGFACE_PAPERS
2605.15185 | Jiaxin Wu, Yihao Pi, Yinling Zhang, Yuheng Li et al.
PDI-Bench introduces projective-geometry residuals (scale-depth alignment, 3D motion consistency, structural rigidity) for quantitative auditing of geometric coherence in generated videos.
HUGGINGFACE_PAPERS
2605.14354 | Sinclair Schneider, Florian Steuber, Gabi Dreo Rodosek | 0
An unsupervised framework combining few-shot filtering, UMAP dimensionality reduction, and HDBSCAN clustering identifies 41 distinct manipulative narrative clusters from 1.2 million social media…
HUGGINGFACE_PAPERS
2605.15198 | Ziyu Guo, Rain Liu, Xinyan Chen, Pheng-Ann Heng | 15
ATLAS uses a single discrete functional token as both an agentic operation and latent visual reasoning unit, with Latent-Anchored GRPO stabilizing RL training without architectural modifications.
HUGGINGFACE_PAPERS
2605.14892 | Shihao Qi, Jie Ma, Rui Xing, Wei Guo et al.
A unified survey of multi-agent LLM systems organized around the LIFE progression: Lay capability foundation, Integrate agents, Find faults, Evolve, identifying open challenges at stage boundaries.
HUGGINGFACE_PAPERS
2605.15188 | Shashwat Goel, Nikhil Chandak, Arvindh Arun, Ameya Prabhu et al.
FutureSim replays real-world news chronologically to test whether agents can predict world events beyond their knowledge cutoff; the best agent achieves only 25% accuracy.
HUGGINGFACE_PAPERS
2605.14269 | Yidong Huang, Zun Wang, Han Lin, Dong-Ki Kim et al.
PhyMotion recovers SMPL body meshes from generated videos, simulates them in MuJoCo, and scores motion along kinematic plausibility, contact balance, and dynamic feasibility dimensions.
HUGGINGFACE_PAPERS
2605.14051 | Yusuke Ozaki, Dhaval Patel | 1
SPIN combines validated DAG planning with prefix-based execution control for industrial LLM agents, reducing executed tasks by 41% and tool calls by 42% on AssetOpsBench.
HUGGINGFACE_PAPERS
2605.14169 | Letian Peng, Ziche Liu, Yiming Huang, Longfei Yun et al.
BOOKMARKS uses search-based bookmark management for role-playing agents, actively initializing and synchronizing task-relevant bookmarks at storyline points.
HUGGINGFACE_PAPERS
2605.15190 | Yanzuo Lu, Ronglai Zuo, Jiankang Deng | 4
RAVEN repacks self-rollouts into interleaved clean historical endpoints and noisy states for better training-inference alignment in causal AR video diffusion.
HUGGINGFACE_PAPERS
2605.14352 | Sinclair Schneider, Florian Steuber, Joao A. G. Schneider, Gabi Dreo Rodosek | 1
A transformer model projecting German political texts onto a continuous left-right spectrum achieves F1=0.844 in-domain and MAE=0.172 on newspaper out-of-domain tests.
HUGGINGFACE_PAPERS
2605.06607 | Nithin Somasekharan, Rabi Pathak, Manushri Dhanakoti, Tingwen Zhang et al.
AI CFD Scientist is the first AI scientist for computational fluid dynamics spanning literature-grounded ideation, OpenFOAM execution, vision-based physics verification, and manuscript writing.
HUGGINGFACE_PAPERS
2605.07865 | Minjae Oh, Sangjun Song, Gyubin Choi, Yunho Choi et al.
vOPD stabilizes On-Policy Distillation by casting it as policy-gradient RL with a control variate baseline derived from per-token reverse KL divergence, consistently outperforming vanilla OPD.
HUGGINGFACE_PAPERS
2605.14306 | Yuwen Du, Tian Jin, Jing Kang, Xianghe Pang et al.
PaSaMaster is a self-evolving agentic retrieval system using iterative intent analysis and ranking, improving F1 by 15.6X over keyword search across 38 scientific disciplines.
HUGGINGFACE_PAPERS

🎓 Google Scholar 6

Rising levels of atmospheric carbon dioxide and methane have sparked the interest of researchers in resolving this issue.
A review of ML-driven catalyst research for converting greenhouse gases (CO2 and methane) into useful chemicals via dry reforming processes.
GOOGLE_SCHOLAR
The advent of large language models (LLMs) has marked a new era in computational social science.
A comprehensive review covering how LLMs are applied in computational social science, including social phenomenon analysis, data bias and privacy challenges.
GOOGLE_SCHOLAR
in adversarial AI, automated threat intelligence, and AI-driven security orchestration...
A review of AI and ML applications in cybersecurity, covering adversarial AI, automated threat intelligence, and AI-driven security orchestration.
GOOGLE_SCHOLAR
the need for post-quantum cryptography to secure devices and user privacy against quantum-enabled adversaries.
A review of post-quantum cryptographic approaches for securing IoT devices and user privacy against quantum-enabled adversaries.
GOOGLE_SCHOLAR
in quantum-enhanced classical ML to native quantum algorithms and hybrid quantum approaches...
A review covering quantum-enhanced classical ML to native quantum algorithms, with applications in optimization, drug discovery, and quantum-secured communications.
GOOGLE_SCHOLAR
Overall, AI teammates altered the process, the distribution and connection of moral frames...
A study examining how generative AI personas function as teammates in collaborative moral reasoning, finding that AI teammates alter the process and distribution of moral frames among human…
GOOGLE_SCHOLAR

🔗 All Sources

  1. [1] Autonomous AI research for nanogpt speedrun
  2. [2] What Coding Is Starting to Lose
  3. [3] The First CVE Wave: Signs That AI-Assisted Vulnerability Discovery Is Reshaping Disclosure Volumes
  4. [4] Mullvad exit IPs as a fingerprinting vector
  5. [5] linux 0-day, access root-owned files as an unprivileged user
  6. [6] ChatGPT Won't Let You Type Until Cloudflare Reads Your React State
  7. [7] BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
  8. [8] LiSA: Lifelong Safety Adaptation via Conservative Policy Induction
  9. [9] Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning
  10. [10] FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale
  11. [11] CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves
  12. [12] Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning
  13. [13] RewardHarness: Self-Evolving Agentic Post-Training
  14. [14] Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models
  15. [15] PanoWorld: Towards Spatial Supersensing in 360 Degree Panorama World
  16. [16] PRISM: Prior Rectification and Uncertainty-Aware Structure Modeling for Diffusion-Based Text Image Super-Resolution
  17. [17] Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning
  18. [18] RouteProfile: Elucidating the Design Space of LLM Profiles for Routing
  19. [19] Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
  20. [20] Dynamic Latent Routing
  21. [21] STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
  22. [22] Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
  23. [23] WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
  24. [24] IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation
  25. [25] MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models
  26. [26] DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
  27. [27] Does Synthetic Layered Design Data Benefit Layered Design Decomposition?
  28. [28] SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
  29. [29] VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction
  30. [30] Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning
  1. [31] PREPING: Building Agent Memory without Tasks
  2. [32] MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
  3. [33] Orchard: An Open-Source Agentic Modeling Framework
  4. [34] EvolveMem: Self-Evolving Memory Architecture via AutoResearch for LLM Agents
  5. [35] Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
  6. [36] Topology-Preserving Neural Operator Learning via Hodge Decomposition
  7. [37] Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video
  8. [38] Nexus: An Agentic Framework for Time Series Forecasting
  9. [39] Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis
  10. [40] Self-Distilled Agentic Reinforcement Learning
  11. [41] Quantitative Video World Model Evaluation for Geometric-Consistency
  12. [42] LLM-based Detection of Manipulative Political Narratives
  13. [43] ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
  14. [44] Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems
  15. [45] FutureSim: Replaying World Events to Evaluate Adaptive Agents
  16. [46] PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
  17. [47] SPIN: Structural LLM Planning via Iterative Navigation for Industrial Tasks
  18. [48] BOOKMARKS: Efficient Active Storyline Memory for Role-playing
  19. [49] RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO
  20. [50] Ideology Prediction of German Political Texts
  21. [51] AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents
  22. [52] KL for a KL: On-Policy Distillation with Control Variate Baseline
  23. [53] Towards Self-Evolving Agentic Literature Retrieval
  24. [54] Catalyst breakthroughs in methane dry reforming: Employing machine learning for future advancements
  25. [55] Large language models (LLM) in computational social science: prospects, current state, and challenges
  26. [56] Artificial intelligence and machine learning in cybersecurity: a deep dive into state-of-the-art techniques and future paradigms
  27. [57] Securing the future: exploring post-quantum cryptography for authentication and user privacy in IoT devices
  28. [58] Quantum machine learning: A comprehensive review of integrating AI with quantum computing for computational advancements
  29. [59] When machines join the moral circle: The persona effect of generative AI agents in collaborative reasoning