Mingkai Deng, Jinyu Hou, Lara Sá Neves, Varad Pimpalkhute, Taylor W. Killian
How should an agent decide when and how to plan? A dominant approach builds agents as reactive policies with adaptive computation (e.g., chain-of-thought), trained end-to-end expecting planning to emerge implicitly. Without control over the presence, structure, or horizon of planning, these systems dramatically…
arXiv:2605.22138
Jialin Lu, Soonho Kong, Rodrigo Stehling, Kaiyu Yang, Zhangyang Wang
We present Lean Refactor, a plug-and-play retrieval-augmented agentic framework for multi-objective, controllable, and version-robust refactoring of Lean proofs. LLM-generated proofs are notoriously correct-but-verbose and brittle across library versions, yet existing refactoring works overlook three practical…
arXiv:2605.20244
Sixiang Chen, Zhaohu Xing, Tian Ye, Xinyu Geng, Yunlong Lin
Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal generative ability with external resources. As requests become more diverse and demanding, we aim to develop a general image-generation agent that can self-evolve…
arXiv:2605.21605
Juncheng Wu, Letian Zhang, Yuhan Wang, Haoqin Tu, Hardy Chen
Large language models (LLMs) and agentic systems have shown promise for clinical decision support, but existing works largely assume that evidence has already been curated and handed to the model. Real-world clinical workflows instead require agents to actively seek, iteratively plan, and synthesize multimodal…
arXiv:2605.20176
Yufei Shi, Weilong Yan, Naixuan Huang, Yucheng Chen, Chenyu Zhang
Existing approaches for digital short-drama production typically rely on one-shot LLM generated scripts and loosely coupled pipelines, which fail to satisfy three key requirements of short-drama generation: (1) narrative pacing, resulting in weak hooks, insufficient escalation, and unattractive endings; (2) spatial…
arXiv:2605.22144
Jinyang Wu, Guocheng Zhai, Ruihan Jin, Yuhao Shen, Zhengxi Lu
The proliferation of large language models (LLMs) and modular skills has endowed autonomous agents with increasingly powerful capabilities. Existing frameworks typically rely on monolithic LLMs and fixed logic to interface with these skills. This gives rise to a critical bottleneck: different LLMs offer distinct…
arXiv:2605.22177
Jiahao Wang, Bo Sun, Yijing Bai, Vincent Casser, Songyou Peng
Robust training and validation of Autonomous Driving Systems (ADS) require massive, diverse datasets. Proprietary data collected by Autonomous Vehicle (AV) fleets, while high-fidelity, are limited in scale, diversity of sensor configurations, as well as geographic and long-tail-behavioral coverage. In contrast,…
arXiv:2605.22809
Haoran Zhang, Luxin Xu, Zhilin Wang, Runquan Gui, Shunkai Zhang
The rise of personal assistant agents, e.g., OpenClaw, highlights the growing potential of large language models to support users across everyday life and work. A core challenge in these settings is proactive assistance, since users often begin with underspecified requests and leave important needs, constraints, or…
arXiv:2605.14678
Zhaoyang Chu, Jiarui Hu, Xingyu Jiang, Pengyu Zou, Han Li
We introduce TerminalWorld, a scalable data engine that automatically reverse-engineers high-fidelity evaluation tasks from "in-the-wild" terminal recordings. Processing 80,870 terminal recordings, the engine yields a full benchmark of 1,530 validated tasks, spanning 18 real-world categories, ranging from short…
arXiv:2605.22535
Qisheng Su, Zhen Fang, Shiting Huang, Yu Zeng, Yiming Zhao
Recent development of agents has renewed demand for long-context reasoning capacity of LLMs. However, training LLMs for this capacity requires costly long-document curation or heuristic context synthesis. We observe that agents produce massive trajectories when solving problems, invoking tools and receiving…
arXiv:2605.21850
Qingnan Ren, Shun Zou, Shiting Huang, Ziao Zhang, Kou Shi
As autonomous coding agents become capable of handling increasingly long-horizon tasks, they have gradually demonstrated the potential to complete end-to-end software development. Although existing benchmarks have recently evolved from localized code editing to from-scratch project generation, they remain confined to…
arXiv:2605.17526
Introduction Modern ML workloads depend heavily on custom GPU kernels. Even when a model is expressed as clean tensor operations, the performance almost a...
Lobste.rs
Hello All
I am the author of this post.
A couple of weeks ago, my previous piece ('The 90 day disclosure policy is dead')
ended up here (https://lobste.rs/s/qxkdgl/90_day_disclosure_policy_is_dead). The thread had some good comments specifically the critique that 'saying the model is broken is a complaint, not…
Lobste.rs
Friends Matthew Garrett Jonathan McDowell Jo McIntyre Martin Michlmayr Andrew Mobbs Mike Pitt Daniel Silverstone Andy Simpkins Neil Williams
Lobste.rs
Megalodon: Mass GitHub Repo Backdooring via CI Workflows
Lobste.rs
Deleting a Google API key doesn't revoke it immediately. Our testing found successful authentications up to 23 minutes after deletion, and Google has declined to fix it.
Lobste.rs
FTC to Require Cox Media Group to Pay Nearly $1million to Settle Charges They Deceived Customers About “Active Listening” AI-Powered Marketing Service
Lobste.rs
Edit April 2, 2026: I've been getting inbound interest from researchers wanting to run their own queries. The MCP integration I use for my own research lets you analyze live mobile telemetry continuously collected from real devices in the wild, directly from Claude. To access it reach out at buchodi@
Lobste.rs
Launch HN: Superset (YC P26) – IDE for the agents era
Hacker News
Code Editor for the AI Agents Era - Run an army of Claude Code, Codex, etc. on your machine - superset-sh/superset
Hacker News
Master open-source intelligence through immersive training, hands-on challenges, and browser-based CTF competitions. Built by Stratir.
Hacker News
Show HN: How to analyze your LLM output – A behavioural health monitor for LLMs
Hacker News
U.S. to Award Nine Quantum-Computing Firms $2B and Take Equity Stakes
Hacker News
(https://www.wsj.com/tech/u-s-to-award-quantum-computing-firms-2-billion-and-take-equity-stakes-7382e6be)
Hacker News
US Government takes $2B equity stake in nine quantum computing firms
Hacker News
Beneficiaries include startup backed by firm with links to the Trump family.
Hacker News
Scott Aaronson – The Truth About Quantum Computing [video]
Hacker News