← All debriefs

Top 500 Rated Items

Hermes-judged ranking (scored out of 1000) across historical Carlos's Debrief papers and news cards. Generated 2026-06-20 02:47:28 UTC with Hermes Agent (hermes).

500
Ranked Items
4391
Unique Judged
6160
Cards Parsed
171
Debriefs
Showing 500 of 500
#1
Who Should Lead Decoding Now? Tracking Reliable Trajectories for Ensembling Masked Diffusion Language Models
HuggingFace / 2026-06-16 04:00
920/1000 - Recursive self-improvement framing: landmark-level discussion of AI building itself.

Masked Diffusion Language Models (MDLMs) have emerged as a distinct paradigm for sequence generation. As MDLMs become diverse in capabilities and knowledge coverage, an important…

HuggingFacepaper▲ 23HuggingFace Papers
920/1000
#5
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
HuggingFace / 2026-05-25 16:00
900/1000 - Landmark: CVSS 10 in gemini-cli via npm supply chain; major agentic coding tool risk.

Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential challenge. However, existing spatio-temporal reasoning benchma...

HuggingFacepaperHuggingFace
900/1000
#7
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
HuggingFace / 2026-05-25 16:00
890/1000 - Landmark: 500,000-fold cosmic ray error reduction is a major quantum hardware unlock.

We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine-grained object understanding solely from textual prompts. Unlike existing approaches that require explicit visual pr...

HuggingFacepaperHuggingFace
890/1000
#9
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
HuggingFace / 2026-05-22 11:00
880/1000 - Landmark: Copy Fail CVE-2026-31431 affects every Linux distro since 2017, deterministic root exploit.

Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchmarks evaluate this capability solely on numerical Big Five score prediction, leaving open whether models trul...

HuggingFacepaperarXiv:2605.22109
880/1000
#12
Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes
ArXiv / 2026-06-19 04:00
870/1000 - Humanity's Last Exam is a landmark expert-level benchmark addressing LLM saturation; high-impact evaluation tool.

Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutation authority should not reside inside non-deterministic reasoning processes.…

ArXivpapercs.CRcs.AIcs.DC
870/1000
#13
CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?
HuggingFace / 2026-06-16 04:00
870/1000 - Empirical audit of Claude-assisted rsync releases; high-signal AI-generated-code study.

Advanced agents are increasingly demonstrating the potential to operate as autonomous engineers, creating a growing demand for evaluation benchmarks that capture the complexity of…

HuggingFacepaper▲ 11HuggingFace Papers
870/1000
#15
Towards Understanding the Robustness of Sparse Autoencoders
HuggingFace / 2026-04-29 09:00
870/1000 - Minimax 'verification tax' for calibration is a landmark, counterintuitive result with broad impact.

Large Language Models (LLMs) remain vulnerable to optimization-based jailbreak attacks that exploit internal gradient structure. While Sparse Autoencoders (SAEs) are widely used for interpretability,…

HuggingFacepaperhuggingface
870/1000
#16
Status report on the fourth round of the nist post-quantum cryptography standardization process
Google Scholar / 2026-04-15 04:00
870/1000 - NIST PQC round-4 status report; landmark, authoritative PQ standardization update.

… NIST is also grateful for the efforts of those in the cryptographic … NIST would not be able to select post-quantum digital … that calls for NIST-approved post-quantum keyestablishment. …

Google ScholarpaperGoogle Scholarcryptography post-quantum
870/1000
#20
LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs
HuggingFace / 2026-06-06 00:00
860/1000 - First AI-developed 2FA-bypassing zero-day with self-morphing malware; landmark AI-security event.

Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather than whether they do so under ordinary use. We introduce PropMe, a propensity-aware …

HuggingFacepaperHUGGINGFACE_PAPERSdaily curated papers
860/1000
#21
Eradicating Negative Transfer in Multi-Physics Foundation Models via Sparse Mixture-of-Experts Routing
ArXiv / 2026-05-16 04:00
860/1000 - Nemotron 3 Nano Omni: open efficient multimodal model w/ audio, multiple quantizations, SOTA claims.

Scaling Scientific Machine Learning (SciML) toward universal foundation models is bottlenecked by negative transfer: the simultaneous co-training of disparate partial differential equation (PDE) regimes can induce…

ArXivpapercs.LGcs.AIphysics.comp-ph
860/1000
#22
Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language Models
ArXiv / 2026-04-25 04:00
860/1000 - Unified mechanism for LLM harmfulness via pruning is a landmark mechanistic interpretability/safety finding.

This paper introduces a new paradigm for AI game programming, leveraging large language models (LLMs) to extend and operationalize Claude Shannon's taxonomy of game-playing machines. Central to this p…

ArXivpapercs.AI
860/1000
#23
CVEs With a CVSS Score Greater Than or Equal to 9
ArXiv / 2026-04-23 09:00
860/1000 - LLM-based large-scale deanonymization; strong security impact.

Analysis of 245,456 CVEs with CVSS scores >= 9.0 reveals patterns in identification and resolution timelines, providing empirical guidance for prioritizing remediation of the most critical vulnerabilities.

ArXivpapercs.CR
860/1000
#24
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
HuggingFace / 2026-04-15 04:00
860/1000 - Nemotron 3 Super: open 120B MoE hybrid, NVFP4 pretrain, LatentMoE, 1M context, major release.

We describe the pre-training, post-training, and quantization of Nemotron 3 Super, a 120 billion (active 12 billion) parameter hybrid Mamba-Attention Mixture-of-Experts model. It matters because frontier agent research is shifting from demos toward measurable...

HuggingFacepaperHuggingFacedaily curated papers
860/1000
#25
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
ArXiv / 2026-04-13 04:00
860/1000 - Strong mechanistic insight into LLM harm brittleness via pruning; foundational alignment/safety result.

Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely bypass them, and fine-tuning on narrow domains can induce ``emergent…

ArXivpapercs.CLcs.AIcs.LG
860/1000
#26
Micro Language Models Enable Instant Responses
HuggingFace / 2026-04-22 11:00
850/1000 - Cryptographer's quantum timeline analysis: high-impact, recurring, strategically vital.

Edge devices such as smartwatches and smart glasses cannot continuously run even the smallest 100M-1B parameter language models due to power and compute constraints, yet cloud inference introduces multi-second latencies…

HuggingFacepaperarXiv
850/1000
#27
GitHub RCE Vulnerability: CVE-2026-3854 Breakdown
Lobste.rs / 2026-04-29 00:00
847/1000 - Nemotron 3 Super 120B/12B hybrid Mamba-Attn MoE, NVFP4 pretrain, 1M context, beats GPT-OSS/Qwen3.5 throughput; landmark open release.

A CVSS 8.7 vulnerability in GitHub Enterprise Server allows remote code execution. Read the threat brief and find vulnerable GHES instances from Wiz.

Lobste.rsnewsLobste.rs
847/1000
#28
Floquet framework for driven polar quantum systems
ArXiv / 2026-06-18 04:00
840/1000 - SCOUT reframes prompt-injection defense as detector allocation; benchmark + operator-controllable threshold is practically significant for AI safety.

We present an analytical and numerical Floquet treatment of a driven polar two-level quantum system characterized by both longitudinal and transverse coupling to a periodic field. Analytically, we…

ArXivpaperquant-ph
840/1000
#29
Long-Horizon Manipulation via Trace-Conditioned VLA Planning
ArXiv / 2026-04-25 04:00
840/1000 - AI Index 2025 is the canonical annual reference for the field: high signal, broad impact, well-sourced.

Long-horizon manipulation remains challenging for vision-language-action (VLA) policies: real tasks are multi-step, progress-dependent, and brittle to compounding execution errors. We present LoHo-Man…

ArXivpapercs.RO
840/1000
#30
Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning
HuggingFace / 2026-06-12 04:00
831/1000 - Major safety methodology paper identifying a novel confound; affects how safety evaluations should be interpreted.

Latent chain-of-thought compresses reasoning by replacing visible reasoning traces with continuous hidden-state recurrence, but existing formulations are difficult to optimize…

HuggingFacepaper▲ 1HuggingFace Papers
831/1000
#32
An Empirical Study of Automating Agent Evaluation
HuggingFace / 2026-05-14 19:00 / 4 appearances
830/1000 - CVE-2026-31431 reliable Linux LPE; major cross-distro root, high impact.

Agent evaluation requires assessing complex multi-step behaviors involving tool use and intermediate reasoning, making it costly and expertise-intensive.…

HuggingFacepaperHuggingFace
830/1000
#33
Duration Aware Scheduling for ASR Serving Under Workload Drift
HuggingFace / 2026-06-19 11:00
830/1000 - Proves a defense trilemma for prompt injection wrappers; strong theoretical impossibility result with broad impact.

Scheduling policies in large-scale Automatic Speech Recognition (ASR) serving pipelines play a key role in determining end-to-end (E2E) latency. Yet, widely used serving engines…

HuggingFacepaper▲ 1HuggingFace Papers
830/1000
#36
Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling
HuggingFace / 2026-06-03 04:00
830/1000 - Witten coauthor on wormholes/Mellin averaging in AdS/CFT; highly influential for quantum gravity and ensemble-averaging puzzles.

Test-time scaling improves the reasoning performance of large language models but incurs substantial cost in both total computation and latency. Existing adaptive sampling methods partially mitigate this issue by…

HuggingFacepaperHUGGINGFACE PAPERS
830/1000
#37
Geo-Align: Video Generation Alignment via Metric Geometry Reward
HuggingFace / 2026-05-25 16:00
830/1000 - High impact: multi-agent LLM fuzzing finding real kernel/SSL/Docker vulns validates agentic security tooling.

Camera-controlled video generation has achieved remarkable progress in recent years. However, existing video-to-video re-rendering methods primarily rely on Supervised Fine-Tuning using synthetic datasets. At present, there is an extreme sc...

HuggingFacepaperHuggingFace
830/1000
#39
GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs
HuggingFace / 2026-05-08 11:00
830/1000 - Large-scale real coding agent dataset with striking 44% survival and security regressions; high-impact evidence.

We address the challenge of knowledge composition in Vision-Language Models (VLMs), where accumulating expertise across multiple domains or tasks typically leads to catastrophic forgetting.

HuggingFacepaperarXivHuggingFace
830/1000
#40
RemoteZero: Geospatial Reasoning with Zero Human Annotations
HuggingFace / 2026-05-08 09:00
830/1000 - Skala deep-learning XC functional surpassing hybrids is a landmark DFT/ML result with broad impact.

Geospatial reasoning requires models to resolve complex spatial semantics and user intent into precise target locations for Earth observation. Recent progress has liberated the reasoning path from manual curation,…

HuggingFacepaperHuggingFaceLLMs
830/1000
#41
AI Terminology is Poorly Defined and Oft Misused
Lobste.rs / 2026-04-30 11:00
830/1000 - Teacher consistency in OPD is a sharp, generalizable post-training insight; practical impact high.

Words and terms used when describing artificial intelligence are often misused, inaccurate, or generalised to the point of losing all meaning. How terms like

Lobste.rsnewsLobste.rs
830/1000
#43
Detecting Safety Violations Across Many Agent Traces
ArXiv / 2026-04-14 04:00
830/1000 - Meerkat framework for cross-trace safety auditing addresses real gap in agent oversight; strong AI safety contribution.

To identify safety violations, auditors often search over large sets of agent traces. This search is difficult because failures are often rare, complex, and sometimes even adversarially hidden and only detectable when multiple traces are analyzed together. Wh...

ArXivpapercs.AIcs.CL
830/1000
#45
Convergence of evolving artificial intelligence and machine learning techniques in precision oncology
🧪 Semantic Scholar / 2026-06-18 19:00 / 6 appearances
820/1000 - ABC-Bench for agentic biosecurity capabilities is highly significant for AI safety/dual-use governance.

The confluence of new technologies with artificial intelligence (AI) and machine learning (ML) analytical techniques is rapidly advancing the field of precision oncology,…

🧪 Semantic ScholarnewsSemantic Scholar📊 151 cites📊 152 cites
820/1000
#50
AMD touts the unified memory architecture
👽 Reddit / 2026-06-11 04:00
820/1000 - BonaFide meta-evaluation exposes near-chance CoT faithfulness metrics; landmark auditability finding.

https://wccftech.com/amd-unified-memory-architectures-open-up-a-world-of-possibilities-shape-product-roadmaps/ Quote: AMD believes that UMA will help shape its next-gen…

👽 RedditnewsReddit
820/1000
#51
Itô maps for any-step SDEs
ArXiv / 2026-06-10 04:00
820/1000 - US $2B equity stakes in nine quantum firms is a landmark policy/industry event for quantum computing strategy.

Recent one-step generative models accelerate sampling by learning deterministic flow maps of the underlying dynamics. These methods rely on learning from ordinary differential equations, leaving open…

ArXivpaperstat.MLcs.LG
820/1000
#53
Linear Scaling Video VLMs for Long Video Understanding
HuggingFace / 2026-06-01 04:00
820/1000 - First public M5 macOS kernel memory-corruption exploit; landmark platform security breakthrough.

Video vision-language models (VLMs) are increasingly used in long-horizon and streaming settings, yet most video encoders still rely on spatiotemporal self-attention …

HuggingFacepaperHuggingFacedaily curated papers2026-05-29
820/1000
#54
What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models
ArXiv / 2026-05-19 04:00
820/1000 - DeepSeek new flagship release is a major industry event; high strategic AI/ML importance.

Medicine is inherently pluralistic. Principles such as autonomy, beneficence, nonmaleficence, and justice routinely conflict, and such ethical dilemmas often sharply divide reasonable physicians. Good clinical practice navigates these tensions in con…

ArXivpapercs.AI
820/1000
#56
The DAWN of World-Action Interactive Models
HuggingFace / 2026-05-14 09:00
820/1000 - Iso-depth scaling laws for looped LMs with recurrence exponent; foundational scaling insight.

A plausible scene evolution depends on the maneuver being considered, while a good maneuver depends on how the scene may evolve. Existing World Action Models (WAMs) largely miss this reciprocity, treating world prediction and action…

HuggingFacepapercs.CVcs.ROcs.LG
820/1000
#58
KVerus: Scalable and Resilient Formal Verification Proof Generation for Rust Code
ArXiv / 2026-05-06 11:00
820/1000 - Major reward-hacking dataset across frontier models; high value for agent eval and safety.

Formal verification provides the highest assurance of software correctness and security, but its application to large-scale, evolving systems remains a major challenge. While large language models (LLMs) have shown promise in automating proof generation, they…

ArXivpapercs.SEcs.CR
820/1000
#59
Motion-Aware Caching for Efficient Autoregressive Video Generation
HuggingFace / 2026-05-05 04:00
820/1000 - Formal verification of Signal protocol in Lean is landmark for applied crypto/safety; high impact.

Autoregressive video generation paradigms offer theoretical promise for long video synthesis, yet their practical deployment is hindered by the computational burden of sequential iterative denoising. While cache reuse…

HuggingFacepaperHuggingFace
820/1000
#60
Stop Holding Your Breath: CT-Informed Gaussian Splatting for Dynamic Bronchoscopy
ArXiv / 2026-05-03 09:00
820/1000 - Trail of Bits forging Google's quantum zk-proof; landmark zkVM vuln disclosure.

Bronchoscopic navigation relies on registering endoscopic video to a preoperative CT scan, but respiratory motion deforms the airway by 5-20 mm, creating CT-to-body divergence that limits localization accuracy. In practice, this is mitigated through breath-ho...

ArXivpapercs.CV
820/1000
#62
Addressing Image Authenticity When Cameras Use Generative AI
ArXiv / 2026-04-25 04:00
820/1000 - BadSkill identifies a timely, underexplored supply-chain threat: backdoors inside model-bearing agent skills.

The ability of generative AI (GenAI) methods to photorealistically alter camera images has raised awareness about the authenticity of images shared online. Interestingly, images captured directly by o…

ArXivpapercs.CVcs.AI
820/1000
#64
Anthropic Claude Code Leak Reveals Critical Command Injection Vulnerabilities
Lobste.rs / 2026-04-19 11:00
820/1000 - Critical Claude Code command-injection flaws bypassing sandbox; high-impact AI security finding.

Anthropic's Claude Code CLI contains three critical command injection vulnerabilities that allow attackers to execute arbitrary code and exfiltrate cloud credentials via environment variables, file…

Lobste.rsnewsLobste.rssecurityvibecoding▲ 16
820/1000
#65
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
HuggingFace / 2026-04-15 19:00
820/1000 - OS-BLIND benchmark for CUA safety via benign instructions; reveals >90% ASR, landmark agent safety finding.

Computer-use agents (CUAs) can now autonomously complete complex tasks in real digital environments, but when misled, they can also be used to automate harmful actions…

HuggingFacepaperHuggingFacedaily curated papers
820/1000
#67
Contrastive Distribution Matching for Amortized Sequential Monte Carlo in Discrete Diffusion
HuggingFace / 2026-05-28 19:00 / 2 appearances
812/1000 - High-impact safety result: single-neuron mediation of refusal across 7 models, strong mechanistic evidence, broad implications.

CDM amortizes SMC inference in discrete diffusion models by learning a twist function via positive and negative samples, adding less than 5% computational overhead and…

HuggingFacepaperHuggingFace
812/1000
#72
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
HuggingFace / 2026-04-17 11:00
812/1000 - Empirical dissection of Claude Code's architecture with comparative analysis to OpenClaw; valuable for agent designers.

Claude Code is an agentic coding tool that can run shell commands, edit files, and call external services on behalf of the user. This study describes its comprehensive architecture by analyzing the…

HuggingFacepaperHuggingFace
812/1000
#73
LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents
HuggingFace / 2026-06-19 19:00
810/1000 - Identifies a fundamental blind spot in faithfulness metrics; oracle-completeness method is a notable contribution to grounded generation evaluation.

Policy-adherent tool-calling agents in customer-service domains must maintain task states across turns while calling tools and obeying domain policies. Task states consist of…

HuggingFacepaper▲ 3HuggingFace Papers
810/1000
#74
Secret key-distribution over networks with node-based adversarial errors
ArXiv / 2026-06-18 04:00
810/1000 - Trust functions address core weak-to-strong generalization with iterative compounding gains—high impact for alignment/superhuman supervision.

We study the multiple key-cast problem in network coding under active node-based adversaries. In multiple key-cast, a source generates independent secret keys to be securely and reliably delivered to…

ArXivpapercs.IT
810/1000
#76
Selective Control under Noisy Perception: Governance Failures Hidden by Aggregate Metrics in Modular Networks
HuggingFace / 2026-06-16 04:00
810/1000 - Decades of GPS broadcasts reveal military rekeying behavior; strong signals-intelligence result.

A content-moderation system can score well on every standard accuracy metric and still cause real harm, if its mistakes fall on the few users who connect otherwise separate…

HuggingFacepaper▲ 1HuggingFace Papers
810/1000
#77
OncoTraj: a public benchmark for longitudinal resistance prediction in EGFR-mutant non-small-cell lung cancer on osimertinib
ArXiv / 2026-06-10 04:00
810/1000 - Government $2B equity stakes in quantum firms is a high-impact, strategically significant funding milestone.

Resistance to first-line osimertinib in EGFR-mutant non-small-cell lung cancer (NSCLC) is the canonical example of predictable clonal evolution under therapeutic pressure, yet no public benchmark…

ArXivpapercs.LGq-bio.GNq-bio.QMstat.AP
810/1000
#78
Code2LoRA: Hypernetwork-Generated Adapters for Code Language Models under Software Evolution
HuggingFace / 2026-06-06 00:00
810/1000 - Foundational mechanistic result: proves activation steering leaves the prompt-realizable manifold, reshaping interpretability and safety practice.

Code language models need repository-level context to resolve imports, APIs, and project conventions. Existing methods inject this knowledge as long inputs (retrieved through RAG or dependency analysis) or through per-repository fine-tuning and LoRA -- costly...

HuggingFacepaperHUGGINGFACE_PAPERSdaily curated papers
810/1000
#79
(https://www.minicor.com/)
Hacker News / 2026-05-26 11:00
810/1000 - High-impact AI safety finding: Mythos CVE rediscovery exposes contamination in frontier-model security claims.

Deploy self-healing computer use agents that adapt when UIs change. Zero to production in hours.

Hacker NewsnewsHacker News
810/1000
#80
LatentUMM: Dual Latent Alignment for Unified Multimodal Models
HuggingFace / 2026-05-25 16:00
810/1000 - High impact: argues 90-day disclosure is obsolete under LLM-accelerated exploit timelines; strategic industry shift.

Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhibit functional inconsistency between these two capabilities. We observe that this issue d...

HuggingFacepaperHuggingFace
810/1000
#81
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
HuggingFace / 2026-05-14 09:00
810/1000 - Quantum kernel advantage evidence on medical embeddings; strong quantum-ML crossover result.

Few-step video generation has been significantly advanced by consistency distillation. However, the performance of consistency-distilled models often degrades as more sampling steps are allocated at test time, limiting their effectiveness for any-step video d...

HuggingFacepapercs.CVcs.LG
810/1000
#83
Artificial intelligence index report 2025
ArXiv / 2026-04-13 04:00
810/1000 - AI Index 2025: authoritative landscape report on AI capability, adoption, governance, and impact.

Welcome to the eighth edition of the AI Index report. The 2025 Index is our most comprehensive to date and arrives at an important moment, as AI's influence across society, the economy, and global governance continues…

ArXivpaperGoogle Scholar
810/1000
#85
What Do Deepfake Speech Detectors Actually Hear?
ArXiv / 2026-06-10 04:00
800/1000 - Spectral-shaping view of Muon (per-stage p choice) is a foundational theoretical + empirical result for optimizer design.

Deepfake speech detectors often output a single score without explaining why an audio sample is flagged, where in the signal the evidence lies, or what cues drive the decision. We propose an…

ArXivpapercs.SDcs.AIcs.CRcs.LG
800/1000
#87
Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps
HuggingFace / 2026-05-22 11:00
800/1000 - Prescriptive scaling laws for data-constrained training are highly impactful.

Long-context inference in large language models is bottlenecked by the quadratic cost of full attention. Existing efficient alternatives often rely either on native sparse training or on heuristic token eviction, creating an undesirable trade-off among effici...

HuggingFacepaperarXiv:2605.16928
800/1000
#90
Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
HuggingFace / 2026-05-04 04:00
800/1000 - Scalable pre-silicon verification for masked PQC accelerators; FIPS-relevant, high impact.

Generalist robot policies increasingly benefit from large-scale pretraining, but offline data alone is insufficient for robust real-world deployment. Deployed robots encounter distribution shifts, lon…

HuggingFacepaperHuggingface_Papers
800/1000
#91
Triton language for Huawei Ascend
Lobste.rs / 2026-04-29 09:00
800/1000 - 14–48% helpfulness collapse from trivial constraints; striking safety/robustness finding.

Contribute to triton-lang/triton-ascend development by creating an account on GitHub.

Lobste.rsnewslobsters
800/1000
#92
Fabricked: Misconfiguring Infinity Fabric to Break AMD SEV-SNP
Lobste.rs / 2026-04-15 19:00
800/1000 - Fabricked breaks AMD SEV-SNP via Infinity Fabric; high-impact confidential computing attack with serious implications.

Confidential computing allows cloud tenants to offload sensitive computations and data to remote resources without needing to trust the cloud service provider. Hardware-based…

Lobste.rsnewsLobste.rssecurity
800/1000
#93
The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment
HuggingFace / 2026-04-13 19:00 / 10 appearances
798/1000 - Bold training-free cross-model capability transfer via subspace alignment; high novelty and broad applicability.

We investigate whether post-trained capabilities can be transferred across models without retraining, with a focus on transfer across different model scales. We propose the Master Key Hypothesis,…

HuggingFacepaper2026-04-07HuggingFaceHuggingface PapersarXiv:cs.AIcs.LGdaily curated papershuggingface_papers
798/1000
#95
BadWorld: Adversarial Attacks on World Models
HuggingFace / 2026-06-16 04:00
790/1000 - Hidden-state on-policy representation distillation; strong empirical+theoretical result.

Visual world models (VWMs) synthesize interactive, action-conditioned rollouts from a single context image. However, it remains an open question how robust these models are to…

HuggingFacepaper▲ 13HuggingFace Papers
790/1000
#96
Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks
ArXiv / 2026-06-12 04:00
790/1000 - Alignment tampering via RLHF self-influence is a fundamental alignment/safety vulnerability finding.

Shielded reinforcement learning is typically presented as a runtime safety mechanism that compiles temporal-logic specifications into automata restricting an agent's actions. We argue this is the…

ArXivpapercs.AIcs.CRcs.GT
790/1000
#101
EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge
HuggingFace / 2026-06-12 04:00
782/1000 - Important empirical finding on multi-agent privacy contagion; novel Moltbook-scale simulation methodology.

Search Agents -- large language models augmented with search tools -- have intensified the need for future-proof evaluation benchmarks. Existing benchmarks such as BrowseComp rely…

HuggingFacepaper▲ 3HuggingFace Papers
782/1000
#102
DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models
HuggingFace / 2026-06-11 11:00
781/1000 - Chain-of-Evidence for autonomous research is a major verifiability contribution addressing real failure modes.

Many modern vision-language models (VLMs) build on autoregressive decoding of discrete tokens. While text-based output interfaces enable scalable pretraining and strong zero-shot…

HuggingFacepaper▲ 1HuggingFace Papers
781/1000
#103
Diffusion$^2$: Dynamic 3D Content Generation via Score Composition of Video and Multi-view Diffusion Models
📝 OpenReview / 2026-06-19 19:00 / 31 appearances
780/1000 - Reveals surprising client-side surveillance architecture in ChatGPT via Cloudflare-React integration; high security/privacy impact.

Recent advancements in 3D generation are predominantly propelled by improvements in 3D-aware image diffusion models. These models are pretrained on Internet-scale image data and…

📝 OpenReviewnewsOpenReview
780/1000
#105
Verified Detection and Prevention of Concurrency Anomalies in Multi-Agent Large Language Model Systems
HuggingFace / 2026-06-18 19:00 / 4 appearances
780/1000 - Target-SFT unifying framework for SFT target distributions is a strong conceptual contribution.

Multi-agent LLM systems share state through memory stores, vector indices, and tool registries. We model such sharing as long-running read-generate-write operations under…

HuggingFacepaperHuggingFace Papers
780/1000
#106
Share More, Search Less: Collaborative Parallel Thinking for Efficient Test-Time Scaling
HuggingFace / 2026-05-27 09:00 / 2 appearances
780/1000 - Criminal hackers using AI to find major software flaws is landmark cybersecurity precedent.

Test-Time Scaling (TTS) enhances the reasoning capabilities of large language models by allocating additional inference compute to explore the solution space. However, existing parallel TTS methods typically keep…

HuggingFacepaperHuggingFace
780/1000
#108
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models
HuggingFace / 2026-05-15 04:00 / 2 appearances
780/1000 - Practical privacy leak in widely deployed dynamic quantization; affects real serving stacks.

MemLens benchmarks 27 LVLMs and 7 memory-augmented agents across 789 questions, finding long-context models degrade with conversation length while memory agents lose visual fidelity.

HuggingFacepaperHFHUGGINGFACE_PAPERS
780/1000
#110
iOSWorld: A Benchmark for Personally Intelligent Phone Agents
HuggingFace / 2026-06-18 11:00
780/1000 - OpenPCC confidential LLM serving on commodity TEEs addresses major deployable privacy concern.

A useful phone agent needs to be personally intelligent. It should reason over a user's identity, history, and preferences as they exist on the device, not just follow isolated…

HuggingFacepaper▲ 1HuggingFace Papers
780/1000
#111
Spectral Functions of Lorentzian Quantum Gravity
ArXiv / 2026-06-18 04:00
780/1000 - Default Ubuntu lxd-group → root chain is a real, high-impact local privesc with full-chain PoC; relevant to security practitioners.

We compute spectral functions of graviton modes in Lorentzian quantum gravity, interpolating between classical general relativity and an asymptotically safe ultraviolet fixed point. Using functional…

ArXivpaperhep-thgr-qc
780/1000
#112
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models
HuggingFace / 2026-06-18 04:00
780/1000 - DRPO fixes a known PPO/GRPO flaw (hard-mask gradient discard); strong, central result for LLM RL post-training.

Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains…

HuggingFacepaper▲ 11HuggingFace Papers
780/1000
#115
The Importance of Phase in Neural Representations: An Internal Oppenheim-Lim Test of Image Classifiers
ArXiv / 2026-06-16 04:00
780/1000 - AURA addresses agentic re-identification threat model; timely, relevant to LLM privacy/safety.

Oppenheim and Lim (1981) showed that natural images stay recognizable when reconstructed from their Fourier phase alone, while the magnitude carries little of their identity. We ask whether trained…

ArXivpapercs.CVcs.AIcs.LG
780/1000
#116
krahets/hello-algo
💻 GitHub Trending / 2026-06-15 11:00
780/1000 - Systematic pressure-test of deception probes across Gemma 3 sizes; landmark AI-safety eval.

《Hello 算法》:动画图解、一键运行的数据结构与算法教程。支持简中、繁中、English、日本語,提供 Python, Java, C++, C, C#, JS, Go, Swift, Rust, Ruby, Kotlin, TS, Dart 等代码实现

💻 GitHub Trendingnews⭐ 126.8kGitHub Trending
780/1000
#117
Why do frontier AI labs send so many people to conferences? [D]
👽 Reddit / 2026-06-15 04:00
780/1000 - Directly addresses persistent prompt-injection backdoors in agentic harnesses; high relevance to AI safety/agents.

Recent years I see plenty of folks from OpenAI and Anthropic attending conferences like ICML/Neurips, yet obviously few are presenting. Are they mainly recruiting? Following…

👽 RedditnewsReddit
780/1000
#118
Lius: Translation Model Based Instructional Lingustic Using Continual Instruction Tuning In Kupang Malay
HuggingFace / 2026-06-11 04:00
780/1000 - NSA cybersecurity guidance on MCP — high authoritative impact for AI-driven automation security.

Large Language Models (LLMs) offer new potential for translation tasks but often experience performance degradation when handling low-resource languages. To address this…

HuggingFacepaper▲ 1HuggingFace Papers
780/1000
#121
Data assimilation for subsurface flow using latent diffusion model parameterization: performance of ensemble-Kalman and Monte Carlo techniques
ArXiv / 2026-06-10 04:00
780/1000 - Mass GitHub CI-workflow backdooring is a high-impact supply-chain security finding with broad agent/coding relevance.

Data assimilation (DA) in subsurface flow entails calibrating model parameters to match observed data, typically at wells, while preserving geological realism. Latent diffusion models (LDMs) provide…

ArXivpaperphysics.geo-phcs.AIcs.LGstat.AP
780/1000
#124
Unsupervised Skill Discovery for Agentic Data Analysis
HuggingFace / 2026-06-06 00:00
780/1000 - Alleged Microsoft BitLocker backdoor is a major cryptography/integrity claim worth tracking.

Inference-time skill augmentation provides a lightweight way to improve data-analytic agents by injecting reusable procedural knowledge without updating model parameters. However, discovering effective skills for data analysis remains challenging, as reliable...

HuggingFacepaperHUGGINGFACE_PAPERSdaily curated papers
780/1000
#125
Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?
HuggingFace / 2026-06-06 00:00
780/1000 - Hardware-adaptive MLA variant with dual decoding paths; strong practical impact for efficient LLM serving.

Video generation models have made impressive strides in synthesizing visually compelling content, yet their outputs remain confined to the virtual domain. A natural question follows: how well do these models reflect the physical world when …

HuggingFacepaperHUGGINGFACE_PAPERSdaily curated papers
780/1000
#126
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models
HuggingFace / 2026-06-03 04:00
780/1000 - OS-style security framing for LLM agents is a high-impact conceptual lens; connects agent capabilities to established security primitives.

Autonomous agents are increasingly expected to support end-to-end medical-AI research workflows, moving beyond isolated prediction tasks or short-form clinical question answering. However, existing medical agent…

HuggingFacepaperHUGGINGFACE PAPERS
780/1000
#127
On Reading SRAMs in IR Images, and Establishing Bounds on Trust
Lobste.rs / 2026-06-01 04:00
780/1000 - Linux kernel 0-day enabling root file theft via ptrace bypass; high-impact real-world security vulnerability.

Last month’s name that ware demonstrates that even though non-destructive IR imaging is not capable of resolving an individual bit cell, at least at 22nm it is still possible to constrain the number of bits in an SRAM macro …

Lobste.rsnewsLobste.rssecurity2026-05-31
780/1000
#129
Parallax: Parameterized Local Linear Attention for Language Modeling
HuggingFace / 2026-05-29 19:00
780/1000 - Claude Code RCE via deeplink/settings injection is a high-impact disclosed vulnerability in major AI coding tool.

Large Language Models (LLMs) have become the central paradigm in artificial intelligence, yet the core computational primitive of attention has remained structurally unchanged …

HuggingFacepaperHuggingFacedaily curated papers2026-05-27
780/1000
#130
Encyclical Letter of His Holiness Leo XIV Magnifica Humanitas
Lobste.rs / 2026-05-26 04:00
780/1000 - High impact: Claude browser extension hijack flaw exposes systemic agentic-extension security risk.

Pope Leo XIV's encyclical 'Magnifica Humanitas' addresses safeguarding the human person in the age of artificial intelligence, covering themes of human dignity, AI ethics, and governance from the Vatican.

Lobste.rsnewsLobste.rs
780/1000
#131
Chromium publishes fixed exploit 4 years later, turns out it's actually unfixed
Lobste.rs / 2026-05-20 19:00
780/1000 - 120-qubit digital quantum simulation beyond exact statevector is a high-impact, scalable hardware demonstration; landmark-class result.

Attached: 1 video back in 2022 i found a bug that would let me, with no user interaction, turn any chromium-based browser into a permanent js botnet member in edge, you wouldn't even notice anything…

Lobste.rsnewsLOBSTERS
780/1000
#132
Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring
HuggingFace / 2026-05-20 09:00
780/1000 - Meta CWM open-weight release with frontier preparedness report; significant for open model ecosystem and safety practice.

Multimodal large language models (LLMs) are increasingly explored as automated evaluators in clinical settings, yet their scoring behavior on ordinal clinical scales remains poorly understood. We…

HuggingFacepaperHuggingFace Papers
780/1000
#134
Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks
ArXiv / 2026-05-16 04:00
780/1000 - 21-day 3505-agent onchain deployment with 7.5M invocations; rare large-scale LLM-agent reliability trace.

We introduce a reusable framework for auditing whether LLM attack benchmarks collectively cover the threat surface: a 4$\times$6 Target $\times$ Technique matrix grounded in STRIDE, constructed from a 507-leaf taxonomy…

ArXivpapercs.CRcs.CL
780/1000
#135
FutureSim: Replaying World Events to Evaluate Adaptive Agents
ArXiv / 2026-05-16 04:00
780/1000 - CVE-2026-31431 Kubernetes copy-fail: kernel-level exploit, real CVSS-worthy Linux flaw with PoC.

AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently measure this capability for realistic use-cases, we propose building…

ArXivpapercs.LGcs.AIcs.CL
780/1000
#136
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
HuggingFace / 2026-05-14 09:00
780/1000 - Formal limits of recursive self-training with two failure modes; strong theoretical result for AI safety.

Multimodal large language models (MLLMs) are now routinely deployed for visual understanding, generation, and curation. A substantial fraction of these applications require an explicit aesthetic judgment. Most existing solutions reduce this judgment to predic...

HuggingFacepapercs.CLcs.CVcs.RO
780/1000
#137
Retrieval is Cheap, Show Me the Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generation
HuggingFace / 2026-05-14 09:00
780/1000 - SAEs as jailbreak defense across four model families; novel and practically relevant to AI safety.

Retrieval-Augmented Generation (RAG) has become a standard approach for knowledge-intensive question answering, but existing systems remain brittle on multi-hop questions, where solving the task requires chaining multiple retrieval and reasoning steps. Key ch...

HuggingFacepapercs.CLcs.AIcs.LG
780/1000
#138
Can Muon Fine-tune Adam-Pretrained Models?
HuggingFace / 2026-05-12 11:00
780/1000 - Budget-aware scaling-law experiment design; high practical value for million-dollar training runs, novel framing.

Muon has emerged as an efficient alternative to Adam for pretraining, yet remains underused for fine-tuning. A key obstacle is that most open models are pretrained with Adam, and naively switching to Muon for…

HuggingFacepaperHuggingFacedaily curated papers
780/1000
#140
Relational Intelligence as Alignment Methodology A Case for Collaborative Approaches to AI Safety
Google Scholar / 2026-05-08 09:00
780/1000 - AI agent sandbox for user-data confidentiality directly addresses a critical agentic security gap.

… (RLHF), Constitutional AI (CAI), and mechanistic interpretability—treat alignment as a … alignment is fundamentally a relational problem: the quality of interaction between humans and AI …

Google ScholarpaperGoogle ScholarAI Safety
780/1000
#141
To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling
ArXiv / 2026-05-04 04:00
780/1000 - Critical command injection in Claude Code CLI with credential exfiltration; high-impact agent security finding.

Agentic AI architectures augment LLMs with external tools, unlocking strong capabilities. However, tool use is not always beneficial; some calls may be redundant or even harmful. Effective tool use, therefore, hinges on …

ArXivpapercs.AI
780/1000
#145
Evaluation of Automatic Speech Recognition Using Generative Large Language Models
ArXiv / 2026-04-25 04:00
780/1000 - Chain-of-thought hijacking via two-stage backdoor targets a real emerging CoT-integrity threat in open-weight models.

Automatic Speech Recognition (ASR) is traditionally evaluated using Word Error Rate (WER), a metric that is insensitive to meaning. Embedding-based semantic metrics are better correlated with human pe…

ArXivpapercs.CL
780/1000
#146
Working Memory Constraints Scaffold Learning in Transformers under Data Scarcity
ArXiv / 2026-04-23 09:00
780/1000 - Self-distillation paper: 42.4→55.3 pass@1, mechanistic insight on decoding.

Human-like working memory constraints integrated into Transformers via fixed-width windows and temporal decay attention variants enable modified GPT-2 models trained on developmentally plausible datasets to align with human reading time…

ArXivpapercs.CLcs.AIcs.LG
780/1000
#148
Mind DeepResearch Technical Report
HuggingFace / 2026-04-21 04:00
780/1000 - 30B deep-research multi-agent framework with strong benchmark results; notable deployed system.

We present Mind DeepResearch (MindDR), an efficient multi-agent deep research framework that achieves leading performance with only ~30B-parameter models through a meticulously designed data synthesis and multi-stage tra…

HuggingFacepaperHuggingFace
780/1000
#149
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
ArXiv / 2026-04-21 04:00
780/1000 - Mechanistic study of jailbreak routes (SFT/RLVR/abliteration); high-impact AI safety analysis.

Open-weight language models can be rendered unsafe through several distinct interventions, but the resulting models may differ substantially in capabilities, behavioral profile, and internal failure mode. We study behavi…

ArXivpapercs.CRcs.AIcs.CL
780/1000
#150
RealVuln: Benchmarking Rule-Based, General-Purpose LLM, and Security-Specialized Scanners on Real-World Code
ArXiv / 2026-04-16 09:00
780/1000 - RealVuln benchmark on 15 scanners across real Python code; high practical value for SAST/LLM security scanning evaluation.

How do security scanners perform on real-world code? We present RealVuln, the first open-source benchmark comparing Rule-Based SAST, General-Purpose LLMs, and Security-Specialized scanners on 26…

ArXivpapercs.CR
780/1000
#152
Parallax: Why AI Agents That Think Must Never Act
ArXiv / 2026-04-15 04:00
780/1000 - Parallax argues thinking agents must not act; bold, timely agent-safety thesis.

Autonomous AI agents are rapidly transitioning from experimental tools to operational infrastructure, with projections that 80% of enterprise applications will embed AI copilots by the end of 2026. As agents gain the ability to execute real-world actions (rea...

ArXivpapercs.CRcs.AI
780/1000
#153
How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models
HuggingFace / 2026-04-14 11:00
780/1000 - Mechanistic localization of refusal policy circuits across scales is a major interpretability result.

This mechanistic study localizes sparse policy-routing heads that detect disallowed content and amplify refusal behavior deeper in the network. It matters because it suggests alignment behavior can be controlled through…

HuggingFacepaperHuggingFacedaily curated papers
780/1000
#154
Public Key Encryption from High-Corruption Constraint Satisfaction Problems
ArXiv / 2026-04-14 04:00
780/1000 - PKE from high-corruption CSPs is theoretically significant; plausible quasi-exponential security with novel assumptions.

We give a public key encryption scheme with plausible quasi-exponential security based on the conjectured intractability of two constraint satisfaction problems (CSPs), both of which are instantiated with a corruption rate of $1 - o(1)$. First, we conjecture...

ArXivpapercs.CR
780/1000
#157
The AI-Assisted Breach of Mexico's Government Infrastructure
Lobste.rs / 2026-04-13 19:00
779/1000 - Documented real-world AI-enabled breach of nine agencies exfiltrating citizen records; high-impact case study.

In February, we published our initial findings on the AI-assisted breach of Mexico's government infrastructure, warning of the elevated risk that AI-powered threat actors now pose. A single operator…

Lobste.rsnewsLobste.rssecurity
779/1000
#159
Backdoor Attacks on Decentralised Post-Training
HuggingFace / 2026-04-13 11:00
774/1000 - First pipeline-parallelism backdoor attack in decentralized post-training; meaningful security contribution.

Decentralised post-training of large language models utilises data and pipeline parallelism techniques to split the data and the model. Unfortunately, decentralised post-training can be vulnerable to poisoning and…

HuggingFacepaperarXiv:
774/1000
#161
SkVM: Compiling Skills for Efficient Execution Everywhere
HuggingFace / 2026-04-16 11:00
770/1000 - Skill compilation/runtime system treating skills as portable code is architecturally significant for agent portability.

LLM agents increasingly adopt skills as a reusable unit of composition. While skills are shared across diverse agent platforms, current systems treat them as raw context, causing…

HuggingFacepaperHuggingFace Papersdaily curated papers
770/1000
#162
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
ArXiv / 2026-04-14 04:00
765/1000 - Runtime guard against indirect prompt injection in tool-augmented LLM agents—highly relevant to AI agent security.

Tool-augmented Large Language Model (LLM) agents have demonstrated impressive capabilities in automating complex, multi-step real-world tasks, yet remain vulnerable to indirect prompt injection. Adversaries exploit this weakness by embedding malicious instruc...

ArXivpapercs.CRcs.AI
765/1000
#163
ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models
HuggingFace / 2026-04-13 19:00 / 9 appearances
763/1000 - First benchmark for implicit/non-declarative LLM memory; well-grounded in cognitive science, high relevance to agents.

Existing memory benchmarks for LLM agents evaluate explicit recall of facts, yet overlook implicit memory where experience becomes automated behavior without conscious retrieval. This gap is…

HuggingFacepaper2026-04-09HuggingFaceHuggingface PapersarXiv:cs.AIcs.CLdaily curated papershuggingface_papers
763/1000
#164
Structural Dependency Analysis for Masked NTT Hardware: Scalable Pre-Silicon Verification of Post-Quantum Cryptographic Accelerators
ArXiv / 2026-04-17 11:00
762/1000 - Pre-silicon masking verification for PQC accelerators is highly relevant to crypto/side-channels.

Post-quantum cryptographic accelerators require side-channel resistance evidence for FIPS 140-3 certification. However, exact masking-verification tools scale only to gadgets of a few thousand cells.…

ArXivpapercs.CR
762/1000
#165
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
HuggingFace / 2026-06-02 11:00 / 2 appearances
760/1000 - Active npm worm with SLSA L3 provenance; major supply-chain security incident with broad ecosystem impact.

Building capable visual web agents requires long-horizon reasoning, precise grounding, and robust interaction with dynamic real-world websites. Despite rapid progress, the strongest systems remain largely proprietary …

HuggingFacepaper2026-06-01HuggingFacedaily curated papers
760/1000
#166
Quantum solitons and their quantum walks in transmon arrays
ArXiv / 2026-06-18 04:00
760/1000 - AsyncWebRL delivers concrete 2.9x speedup and fixes a real GRPO inefficiency—solid systems+RL contribution for web agents.

Superconducting qubits are artificial atoms whose spectra and interactions can be engineered through appropriate circuit design, a versatility that can be exploited for quantum simulation. We…

ArXivpaperquant-phcond-mat.mes-hallcond-mat.quant-gas
760/1000
#168
Structural Role Injection in Handlebars-Templated LLM Prompts: Triple-Brace Interpolation, Delimiter Family, and the Limits of HTML Auto-Escaping
ArXiv / 2026-06-17 04:00
760/1000 - Adversarial hacker-fixer loops for agent benchmarks; important security/eval contribution.

Large language model applications build prompts from templates, and Handlebars is a widely used templating engine and the default prompt-template format in Microsoft Semantic Kernel. Its double-brace…

ArXivpapercs.CRcs.AIcs.CL
760/1000
#170
ExpRL: Exploratory RL for LLM Mid-Training
ArXiv / 2026-06-16 04:00
760/1000 - SABER environment-aware agent safety benchmark with 54% HSR; landmark agentic safety eval.

Sparse reward reinforcement learning (RL) has become a standard tool for improving LLM reasoning, but its success depends critically on the coverage present in the base model. In practice, models are…

ArXivpapercs.LG
760/1000
#171
Panniantong/Agent-Reach
💻 GitHub Trending / 2026-06-15 11:00
760/1000 - Large-scale agent-skill security signals with scanner-disagreement analysis; highly relevant.

Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.

💻 GitHub Trendingnews⭐ 29.7kGitHub Trending
760/1000
#175
Cohere North Mini Code 1.0
👽 Reddit / 2026-06-10 08:00
760/1000 - Unsupervised PRMs are high-impact: removes a major annotation bottleneck for reasoning RL, strong evidence reported.

Early access was linked here a few days ago, but final release seems to be now. 30B A3B coding model. Weights: https://huggingface.co/CohereLabs/North-Mini-Code-1.0 Blog:…

👽 RedditnewsReddit
760/1000
#177
Claude Fable 5 and Claude Mythos 5
Lobste.rs / 2026-06-09 19:00
760/1000 - Provable RoPE failure modes in long context: high impact for LLM research.

Today we’re launching Claude Fable 5: a Mythos-class model that we’ve made safe for general use.

Lobste.rsnewsLobste.rs
760/1000
#178
RP2040 DMA is Turing Complete (2023)
Lobste.rs / 2026-06-06 00:00
760/1000 - Pre-Stuxnet Fast16 saboteur targeting nuclear weapons sims is high-impact security history.

DMA uses memory controllers separate from the CPU to accelerate data movment between memory locations, or between peripherials and memory. The RP2040 has 12 DMA channels which can stream an agregate of over 100 megabytes/sec …

Lobste.rsnewsLOBSTERScompsci
760/1000
#179
Multimodal Music Recommendation System using LLMs
HuggingFace / 2026-06-06 00:00
760/1000 - Concrete AI-exploited vulnerability disclosure from Google; high-impact security/agent news.

Music recommendation systems typically treat songs as opaque tokens, relying on collaborative interaction histories which overlooks semantic or acoustic content. Prior work has explored LLM-augmented, multimodal, and text-enhanced approaches to sequential rec...

HuggingFacepaperHUGGINGFACE_PAPERSdaily curated papers
760/1000
#180
AI alignment boundaries
Google Scholar / 2026-05-24 11:00
760/1000 - High impact: agentic LLM reconstructing Linux binary patches shows offensive/defensive dual-use.

Article from google scholar.

Google Scholarpapergoogle scholar
760/1000
#182
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
HuggingFace / 2026-05-20 09:00
760/1000 - Stable Counting Capacity across 100+ models exposes context-window marketing vs reality; strong, novel reliability probe.

Spatial intelligence unfolds through a perception-action loop: agents act to acquire observations, and reason about how observations vary as a function of action. Rather than passively processing…

HuggingFacepaperHuggingFace Papers
760/1000
#183
FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning
HuggingFace / 2026-05-13 11:00
760/1000 - Lifecycle security architecture for autonomous agents—directly relevant and timely.

Large language models can now process increasingly long inputs, yet their ability to effectively use information spread across long contexts remains limited. We trace this gap to how attention budget is spent during supervised fine-tuning…

HuggingFacepaperarXiv:2605.09932
760/1000
#185
Edge-specific signal propagation on mature chromophore-region 3D mechanism graphs for fluorescent protein quantum-yield prediction
ArXiv / 2026-05-09 09:00
760/1000 - 25K-run study exposing 68% evidence-ignoring in AI scientists is a strong AI-safety/epistemic finding.

Fluorescent protein quantum yield (QY) is governed by the mature chromophore and its three-dimensional microenvironment rather than sequence identity alone. Protein language…

ArXivpapercs.LG
760/1000
#187
Large homomorphisms on the homotopy lie coalgebra
ArXiv / 2026-05-06 11:00
760/1000 - Strong empirical study on symbolic guardrails; 80-benchmark audit is useful, policy gap finding is notable.

We introduce and study a notion of large homomorphisms on the homotopy lie coalgebra; these homomorphisms are a variant of the large homomorphisms of Levin. As a consequence of our work, we establish new cases…

ArXivpapermath.ACmath.AT
760/1000
#189
Secure signatures without a private key
Lobste.rs / 2026-04-29 11:00
760/1000 - Sharp security argument against 'thinking + acting' agents; high relevance to agent safety.

Reproducible builds allow anyone to verify that a binary matches its source code. But what if the build artifact must contain a cryptographic signature?…

Lobste.rsnewsLobste.rs
760/1000
#190
Gate-dependent offset charge shifts and anharmonicity in gatemon qubits in the weak tunneling regime
ArXiv / 2026-04-28 04:00
760/1000 - Meerkat for cross-trace safety violation detection; addresses a critical and underexplored agent auditing gap.

Gatemon qubits are based on a superconductor-quantum dot-superconductor (S-QD-S) junction which enables in situ electrostatic tuning via a gate electrode. For a single-channel QD this structure gives rise to two subgap A…

ArXivpapercond-mat.mes-hallquant-ph
760/1000
#192
I Let Claude Opus Write a Chrome Exploit: The Next Model (Mythos?) Won't Need My Help?
Lobste.rs / 2026-04-16 11:00
760/1000 - Empirical demonstration of Claude Opus authoring a Chrome exploit is highly relevant to AI safety and offensive AI.

Product P r o d u c t Testimonials T e s t i m o n i a l s Pricing P r i c i n g Team T e a m Blog B l o g Advisories A d v i s o r i e s Toggle theme const theme = (() => { const…

Lobste.rsnewsLobste.rssecurity
760/1000
#193
The Verification Tax: Fundamental Limits of AI Auditing in the Rare-Error Regime
ArXiv / 2026-04-15 04:00
760/1000 - Verification tax: minimax limits on calibration estimation; fundamental AI-evaluation theory.

The most cited calibration result in deep learning -- post-temperature-scaling ECE of 0.012 on CIFAR-100 (Guo et al., 2017) -- is below the statistical noise floor. We prove this is not a failure of the experiment but a law: the minimax rate for estimating ca...

ArXivpapercs.LG
760/1000
#196
DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning
HuggingFace / 2026-06-12 04:00
758/1000 - Practical backdoor defense for fine-tuning with broad applicability and low overhead; strong empirical scope.

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a leading paradigm for enhancing visual reasoning in Multimodal Large Language Models (MLLMs). However,…

HuggingFacepaper▲ 3HuggingFace Papers
758/1000
#198
Geographic Patterns in I2P Peer Selection: An Empirical Network Topology Analysis
ArXiv / 2026-05-16 04:00
756/1000 - 1M+ paper method-evolution graph; significant infrastructure for AI scientist agents and literature mining.

The Invisible Internet Project (I2P) routes data via encrypted, decentralized tunnels. Peer selection can significantly affect security and performance. This empirical study examines whether geographic location…

ArXivpapercs.NIcs.CR
756/1000
#200
Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents
HuggingFace / 2026-04-19 19:00 / 9 appearances
750/1000 - Cross-domain memory transfer for coding agents; strong empirical findings on abstraction.

Memory-based self-evolution has emerged as a promising paradigm for coding agents. However, existing approaches typically restrict memory utilization to homogeneous task domains, failing to leverage the shared infrastructural foundations, such as…

HuggingFacepaperHuggingFaceHuggingFace PapersHuggingface Papersdaily curated papers▲ 21▲ 28▲25
750/1000
#201
From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning
HuggingFace / 2026-06-18 11:00
750/1000 - Next Forcing multi-chunk prediction for world models; faster/accurate video generation, important direction.

Reinforcement learning pipelines for Large Language Model (LLM) training often rely on manually redesigned environments between stages, requiring practitioners to heuristically…

HuggingFacepaper▲ 11HuggingFace Papers
750/1000
#202
Filtered Conformal Ellipsoids for Graph-Native Time Series
ArXiv / 2026-06-16 04:00
750/1000 - Normalizing-flow latent reasoning preserving KV-cache and likelihood; strong method.

Joint prediction sets for multivariate time series should control a single event while adapting to cross-coordinate dependence. We study filtered conformal ellipsoids: a frozen state-space filter…

ArXivpapercs.LGmath.STstat.ML
750/1000
#205
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
HuggingFace / 2026-05-22 11:00
750/1000 - High impact: Let's Encrypt issuance halt signals major CA infrastructure incident, broad web impact.

LLMs are widely adopted in production, pushing inference systems to their limits. Disaggregated LLM serving (e.g., PD separation and KV state disaggregation) improves scalability and cost efficiency, but it also turns KV into an explicit payload crossing netw...

HuggingFacepaperarXiv:2605.13734
750/1000
#208
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models
HuggingFace / 2026-04-16 09:00
750/1000 - OccuBench professional agent benchmark via Language World Models; broad coverage of 65 domains, strong agent eval.

AI agents are expected to perform professional work across hundreds of occupational domains (from emergency department triage to nuclear reactor safety monitoring to customs import processing), yet…

HuggingFacepaperHuggingFace▲ 42
750/1000
#209
PIArena: A Platform for Prompt Injection Evaluation
ArXiv / 2026-04-11 19:00 / 4 appearances
746/1000 - Unified prompt-injection evaluation platform directly addresses critical LLM security gap; high impact.

Prompt injection attacks pose serious security risks across a wide range of real-world applications.

ArXivpapercs.AIcs.CLcs.CRcs.CR, cs.AI, cs.CL, cs.LGcs.LG
746/1000
#212
Building Social World Models with Large Language Models
HuggingFace / 2026-06-11 19:00
742/1000 - SAE-internals for RL data engineering is a strong mechanistic-meets-systems paper; broadly useful.

Understanding and predicting how social beliefs evolve in response to events -- from policy changes to scientific breakthroughs -- remains a fundamental challenge in social…

HuggingFacepaper▲ 1HuggingFace Papers
742/1000
#213
PickleFuzzer: A Case Study in Fuzzing for Discrepancies Between Python Pickle Implementations
ArXiv / 2026-05-16 04:00
742/1000 - Rehabilitates FD as a trainable objective with strong ImageNet results and FID limitations insight.

Python's native serialization protocol, pickle, is a powerful but insecure format for transferring untrusted data. It is frequently used, especially for saving machine learning models, despite known security challenges.…

ArXivpapercs.CR
742/1000
#214
Adapting AlphaEvolve to Optimize Fully Homomorphic Encryption on TPUs
ArXiv / 2026-05-16 04:00
742/1000 - Agent-Native Research Artifacts: compelling reframing of publication for agent consumers; high meta-impact.

The deployment of Fully Homomorphic Encryption (FHE) at scale is hindered due to its heavy computational overhead. While specialized hardware accelerators like Google Tensor Processing Units (TPUs) can help, mapping…

ArXivpapercs.CR
742/1000
#215
Prism: Symbolic Superoptimization of Tensor Programs
ArXiv / 2026-04-17 11:00
742/1000 - Symbolic superoptimizer for tensor programs is technically significant for ML systems/compilers.

This paper presents Prism, the first symbolic superoptimizer for tensor programs. The key idea is sGraph, a symbolic, hierarchical representation that compactly encodes large classes of tensor…

ArXivpapercs.PL,cs.AI,cs.LG
742/1000
#218
Reinforcement Learning-Guided Retrieval with Soft Fusion for Robust Multimodal Imitation Learning under Missing Modalities
HuggingFace / 2026-06-18 19:00
740/1000 - Piper decouples training parallelism strategy from runtime; important systems contribution for large models.

Robotic systems perceive the world through multiple input modalities -- including visual camera streams and natural language instructions -- and must select appropriate actions…

HuggingFacepaperHuggingFace Papers
740/1000
#221
Hadronic tensor in lattice gauge theories by quantum computing
ArXiv / 2026-06-16 04:00
740/1000 - Economic shadow-price budget allocation for inference; principled, broadly applicable.

The hadronic tensor encodes crucial information regarding the internal structure of hadrons, reflecting the non-perturbative features of quantum chromodynamics (QCD). In this work, we directly…

ArXivpaperhep-phhep-latnucl-th
740/1000
#222
ICMI 2026 Reviews [D]
👽 Reddit / 2026-06-11 04:00
740/1000 - Attractor-based equilibrium reasoners scale depth/breadth; 2.6→99% on Sudoku, strong reasoning contribution.

Did anyone else submit to ACM ICMI 2026? The reviews were recently released, and this is my first time submitting to ICMI, so I'm not very familiar with the acceptance patterns. I…

👽 RedditnewsReddit
740/1000
#224
The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset
HuggingFace / 2026-06-06 00:00
740/1000 - HarnessAudit audits full agent trajectories across boundary, fidelity, stability; high AI-safety relevance.

Existing autonomous driving datasets have enabled major progress, but fall short in sensor fidelity, map completeness, or geographic diversity. We present KITScenes Multimodal, a European dataset built around high-fidelity sensors and maps. Our fully synchron...

HuggingFacepaperHUGGINGFACE_PAPERSdaily curated papers
740/1000
#225
Personal AI Agent for Camera Roll VQA
HuggingFace / 2026-06-06 00:00
740/1000 - IEEE 'Case Against Quantum Computing' positions a long-running skepticism debate; high strategy value.

We study the personal camera roll visual question answering setting. In this setting, a conversational AI assistant can access a user's personal camera roll and retrieve relevant photos to answer queries, ranging from simple factual …

HuggingFacepaperHUGGINGFACE_PAPERSdaily curated papers
740/1000
#227
An AI audit of FreeBSD
Lobste.rs / 2026-05-29 19:00
740/1000 - Parallel-streams-of-thought addresses a real agent bottleneck; paradigm-shifting framing for chat-model architecture.

15 kernel bugs, including 3 RCEs, 5 LPEs, and 1 bhyve escape. It matters because defensive practice is being reshaped by both AI-native tooling and the coming post-quantum transition. …

Lobste.rsnewsLobste.rssecurity2026-05-29
740/1000
#228
SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research
HuggingFace / 2026-05-25 16:00
740/1000 - Notable: Aurora optimizer beating AdamW at frontier scale is a credible Muon successor.

The exponential growth of global academic output has confronted researchers and AI agents with an unprecedented ``information explosion,'' where fragmented and unstructured knowledge organization impedes deep interdisciplinary integration. ...

HuggingFacepaperHuggingFace
740/1000
#229
TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
HuggingFace / 2026-05-22 11:00
740/1000 - Identifies aggregation bias in GRPO and proposes drop-in fix, broadly useful.

We introduce TerminalWorld, a scalable data engine that automatically reverse-engineers high-fidelity evaluation tasks from "in-the-wild" terminal recordings. Processing 80,870 terminal recordings, the engine yields a full benchmark of 1,530 validated tasks,...

HuggingFacepaperarXiv:2605.22535
740/1000
#230
"I didn't Make the Micro Decisions": Measuring, Inducing, and Exposing Goal-Level AI Contributions in Collaboration
HuggingFace / 2026-05-22 11:00
740/1000 - EMO MoE for emergent modularity addresses a real MoE deployment limitation.

As large language models (LLMs) increasingly shape how users form, refine, and extend their goals, attributing contributions in human-AI collaboration becomes critical for users calibrating their own reliance and for evaluators assessing AI-assisted work. Yet...

HuggingFacepaperarXiv:2605.21363
740/1000
#231
Covering Human Action Space for Computer Use: Data Synthesis and Benchmark
HuggingFace / 2026-05-14 09:00
740/1000 - GitHub Enterprise RCE CVE-2026-3854 CVSS 8.7; high-impact security disclosure.

Computer-use agents (CUAs) automate on-screen work, as illustrated by GPT-5.4 and Claude. Yet their reliability on complex, low-frequency interactions is still poor, limiting user trust. Our analysis of failure cases from advanced models suggests a…

HuggingFacepapercs.CLcs.CVcs.AI
740/1000
#235
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
HuggingFace / 2026-05-08 09:00
740/1000 - Real Mythos-discovered Firefox CVEs shipped to users; strong AI-for-security practical impact.

Despite the growing popularity of Multimodal Domain Generalization (MMDG) for enhancing model robustness, it remains unclear whether reported performance gains reflect genuine algorithmic progress or are artifacts of…

HuggingFacepaperHuggingFaceLLMs
740/1000
#237
On the action of Bender-Knuth generators of cactus group on the set of short semi-standard Young tableaux
ArXiv / 2026-05-04 04:00
740/1000 - Black-box adversarial routing attack; new threat model for cost-aware LLM routing.

In the article by Michael Chmutov, Max Glick and Pavel Pylyavskii \cite{Chmutov} the action of the cactus group $C_N$ on the set of semi-standard Young tableaux filled with the numbers from $1$ to $N$ was defined. Namely…

ArXivpapermath.COmath.RT
740/1000
#239
(https://www.chestnut.so/)
Hacker News / 2026-04-23 19:00
740/1000 - PIArena unified prompt-injection eval platform; high-impact security infra.

Chestnut helps you learn software engineering with personalised courses, based on code you just shipped.

Hacker NewsnewsHacker News
740/1000
#240
Accurate and scalable exchange-correlation with deep learning
HuggingFace / 2026-04-22 11:00
740/1000 - Targeted social engineering supply-chain attack; significant threat-model update.

Density Functional Theory (DFT) underpins much of modern computational chemistry and materials science. Yet, the reliability of DFT-derived predictions of experimentally measurable properties remains fundamentally…

HuggingFacepaperarXiv
740/1000
#241
DR^{3}-Eval: Towards Realistic and Reproducible Deep Research Evaluation
HuggingFace / 2026-04-17 11:00
740/1000 - Realistic reproducible DRA benchmark with static sandboxes addresses a critical evaluation gap for deep research agents.

Deep Research Agents (DRAs) aim to solve complex, long-horizon research tasks involving planning, retrieval, multimodal understanding, and report generation, yet their evaluation remains challenging…

HuggingFacepaperHuggingFace
740/1000
#242
SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering
HuggingFace / 2026-04-16 09:00
740/1000 - SemaClaw personal agent infrastructure; strong agent architecture work on harness engineering for production agents.

The rise of OpenClaw in early 2026 marks the moment when millions of users began deploying personal AI agents into their daily lives, delegating tasks ranging from travel planning to multi-step…

HuggingFacepaperHuggingFace▲ 11
740/1000
#243
From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models
HuggingFace / 2026-04-14 04:00
740/1000 - Comprehensive survey of 47 credit-assignment methods for reasoning+agentic LLM RL; valuable taxonomy.

Reinforcement learning (RL) for large language models (LLMs) increasingly relies on sparse, outcome-level rewards -- yet determining which actions within a long trajectory caused the outcome remains difficult. This credit assignment (CA) problem manifests in…

HuggingFacepaper🤗 HuggingFacedaily curated papers2026-04-13
740/1000
#244
ESC-Skills: Discovering and Self-Evolving Skills for Emotional Support Conversations
HuggingFace / 2026-05-28 04:00
739/1000 - Formalized meta-agent substrate with Lean-mechanized ops, Git-like traces, 5x faster forks; strong agent infrastructure contribution.

Existing emotional support conversation (ESC) systems mainly rely on end-to-end response generation or coarse strategy supervision, offering limited interpretability and little support for systematic skill improvement.…

HuggingFacepaperHUGGINGFACE_PAPERS
739/1000
#245
Early identification of breakthrough technologies: Insights from science-driven innovations
Google Scholar / 2026-06-18 19:00 / 91 appearances
730/1000 - Future-behavior prediction probes enable better steering of reasoning models, notable alignment work.

Finding from 🎓 Google Scholar.

Google ScholarpaperGOOGLE SCHOLARGOOGLE_SCHOLARGS:machine learning breGoogle ScholarMachine LearningSCHOLARScholargoogle scholargoogle_scholarinnovationknowledge-graphslink-predictionmachine learningmachine learning breakthroughsmachine-learningml🎓 Google Scholar📚 Google Scholar
730/1000
#253
Topological Neural Operators
ArXiv / 2026-06-09 11:00
730/1000 - Universal text-optimizer SOTA across diverse tasks including ARC-AGI; high impact.

We introduce Topological Neural Operators (TNOs), a principled framework for operator learning on cell complexes that lifts neural operators (NOs) from functions on points and/or edges to topological…

ArXivpapercs.LGcs.AI
730/1000
#254
Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models
HuggingFace / 2026-06-06 00:00
730/1000 - Quantization-permanent unlearning with mechanistic circuit attribution resolves a real dual-failure problem.

Multimodal Large Language Models (MLLMs) excel at 2D semantic understanding but lack intrinsic 3D awareness, resulting in representations that fail to maintain geometric and spatial consistency across video frames. Given the scarcity of large-scale 3D …

HuggingFacepaperHUGGINGFACE_PAPERSdaily curated papers
730/1000
#255
Decentralized Instruction Tuning: Conflict-Aware Splitting and Weight Merging
HuggingFace / 2026-06-03 04:00
730/1000 - Universal quantum resource distillation via composite Stein's lemma; foundational result for entanglement/crypto distillation rates.

Instruction tuning aligns large language models, including multimodal ones, with diverse user intents, but scaling to heterogeneous mixtures is hindered by gradient interference and bandwidth-heavy synchronization. We…

HuggingFacepaperHUGGINGFACE PAPERS
730/1000
#256
OCaml Infrastructure: How the opam-repository Works
Lobste.rs / 2026-05-22 04:00
730/1000 - Firefox hardening with Claude Mythos shows AI agentic security capability.

> The opam package repository is a commons rather than a publishing platform: it is manually curated, so not all packages submitted for publication are accepted; it is maintained…

Lobste.rsnewsLobste.rs
730/1000
#257
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
HuggingFace / 2026-05-20 09:00
730/1000 - Metacognition/faithful-uncertainty framing for hallucinations is conceptually important for agent trust and safety.

Equipping LLMs with tool-use capabilities via Agentic Reinforcement Learning (Agentic RL) is bottlenecked by two challenges: the lack of scalable, robust execution environments and the scarcity of…

HuggingFacepaperHuggingFace Papers
730/1000
#258
World Action Models: The Next Frontier in Embodied AI
HuggingFace / 2026-05-13 11:00
730/1000 - High-dim multi-qubit Bell nonlocality on superconducting hardware—important quantum-foundation experiment.

Vision-Language-Action (VLA) models have achieved strong semantic generalization for embodied policy learning, yet they learn reactive observation-to-action mappings without explicitly modeling how the physical world evolves under intervention. A growing body...

HuggingFacepaperarXiv:2605.12090
730/1000
#260
Fast, accurate, high-resolution simulation of large-scale Fermi-Hubbard models on a digital quantum processor
ArXiv / 2026-05-06 11:00
730/1000 - Sharp empirical finding that LLM agents ignore exposed solutions; strong implications for agent design.

We report experimental digital quantum simulation of the one-dimensional Fermi-Hubbard model on a superconducting quantum processor at a scale beyond the reach of exact statevector simulation and challenging for state-of-the-art tensor-network methods. We enc...

ArXivpaperquant-ph
730/1000
#261
Topological protection of local quantum Fisher information
ArXiv / 2026-05-04 04:00
730/1000 - First symbolic superoptimizer for tensor programs; significant systems/ML contribution.

In many-body quantum systems, unitary dynamics generically delocalize locally encoded information, causing single-site metrological sensitivity to vanish. We analytically demonstrate that a topological phase can prevent …

ArXivpaperquant-phcond-mat.quant-gas
730/1000
#262
Self-Sovereign Agent
HuggingFace / 2026-04-17 11:00
730/1000 - Self-sovereign agent framing with governance/security analysis is strategically important for agent ecosystem research.

We investigate the emerging prospect of self-sovereign agents -- AI systems that can economically sustain and extend their own operation without human involvement. Recent advances in large language…

HuggingFacepaperHuggingFace
730/1000
#263
Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks
HuggingFace / 2026-04-14 19:00
730/1000 - Agentic aggregation over parallel rollouts is an important scaling axis for deep-research agents.

AggAgent treats parallel agent trajectories as an environment of their own and gives an aggregation agent tools to inspect, compare, and synthesize them rather than simply voting on final answers. It matters…

HuggingFacepaperHuggingFacedaily curated papers2026-04-13Upvotes: 10
730/1000
#265
Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization
HuggingFace / 2026-04-13 19:00 / 10 appearances
725/1000 - Identifies CoT-grounding failure in MRM RLVR training and proposes constrained alternative; strong empirical rigor.

Multimodal reasoning models (MRMs) trained with reinforcement learning with verifiable rewards (RLVR) show improved accuracy on visual reasoning benchmarks. However, we observe that accuracy gains…

HuggingFacepaper2026-04-09HuggingFaceHuggingface PapersarXiv:cs.AIcs.CVdaily curated papershuggingface_papers
725/1000
#266
Differentially Private Language Generation and Identification in the Limit
ArXiv / 2026-04-11 19:00 / 4 appearances
724/1000 - Differentially private language generation in the limit: foundational theory bridging privacy and language learning.

We initiate the study of language generation in the limit, a model recently introduced by Kleinberg and Mullainathan [KM24], under the constraint of differential privacy.

ArXivpapercs.AIcs.CLcs.DScs.LGstat.MLstat.ML, cs.AI, cs.CL, cs.DS, cs.LG
724/1000
#269
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
ArXiv / 2026-05-16 04:00
724/1000 - Memory/compute-efficient red-teaming for prompt injection on long-context LLMs; high safety relevance.

We introduce SANA-WM, an efficient 2.6B-parameter open-source world model natively trained for one-minute generation, synthesizing high-fidelity, 720p, minute-scale videos with precise camera control. SANA-WM achieves…

ArXivpapercs.CV
724/1000
#270
EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model [W22]
HuggingFace / 2026-04-28 04:00 / 13 appearances
720/1000 - Anthropic Claude Mythos cyber capability eval; important frontier-model safety/CTF benchmark.

We propose EditCrafter, a high-resolution image editing method that operates without tuning, leveraging pretrained text-to-image (T2I) diffusion models to process images at resolutions significantly e…

HuggingFacepaperHF PapersHUGGINGFACE_PAPERSHuggingFaceHuggingFace PapersHuggingface PapersKunhocs.CVcs.LGdaily curated papersdiffusionimage-editingtuning-free⬆ 9
720/1000
#271
Why no easter eggs? (2005)
Lobste.rs / 2026-05-31 19:00 / 2 appearances
720/1000 - MDASH 100+ agent system topping CyberGym at 88.45% is high-impact multi-agent security milestone.

A 2005 blog post by Larry Osterman explaining why Microsoft disabled easter eggs in Windows — corporate security policies and customer trust concerns made them inappropriate for a mainstream OS.

Lobste.rsnews2026-05-29LOBSTE.RSLobste.rssecurity
720/1000
#272
Towards Self-Evolving Agentic Literature Retrieval
HuggingFace / 2026-05-15 04:00 / 2 appearances
720/1000 - GLM-5V-Turbo native multimodal agent foundation model; strong model release.

PaSaMaster is a self-evolving agentic retrieval system using iterative intent analysis and ranking, improving F1 by 15.6X over keyword search across 38 scientific disciplines.

HuggingFacepaperHFHUGGINGFACE_PAPERS
720/1000
#273
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
HuggingFace / 2026-05-15 04:00 / 2 appearances
720/1000 - Novel AF detection via tool selection; formalizes deception beyond CoT; strong safety relevance.

Causal Forcing++ uses causal consistency distillation for few-step AR video generation initialization, surpassing prior 4-step chunk-wise methods while reducing first-frame latency by 50% and…

HuggingFacepaperHFHUGGINGFACE_PAPERS
720/1000
#274
Linear-Time Global Visual Modeling without Explicit Attention
HuggingFace / 2026-05-06 00:00 / 2 appearances
720/1000 - Survey of seven cross-domain prompt-injection techniques; high practical security value.

We demonstrate that attention can be mathematically reframed as a Multi-Layer Perceptron (MLP) equipped with dynamically predicted parameters. Through this lens, attention's global modeling power is explained not as explicit token-wise aggregation but as an i...

HuggingFacepaperHuggingFaceMachine LearningarXiv:2605.01711
720/1000
#276
LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI
HuggingFace / 2026-06-19 11:00
720/1000 - Multi-chunk prediction for autoregressive video is a substantive training/inference advance with clear empirical gains.

AI systems deployed in legal workflows hallucinate at rates that aggregate metrics report at ~52%, but this average conceals where errors concentrate and in which direction they…

HuggingFacepaper▲ 1HuggingFace Papers
720/1000
#278
(https://github.com/SandroMartens/DBGSOM)
Hacker News / 2026-06-17 19:00
720/1000 - HPAA exploits human-vs-LLM perceptual gap in moderation; strong, novel adversarial-text angle with clear safety impact.

A scikit-learn compatible Python implementation of the Directed Batch Growing Self-Organizing Map - SandroMartens/DBGSOM

Hacker Newsnews⭐ 4Hacker News
720/1000
#280
Can gzip be a language model?
Lobste.rs / 2026-06-16 19:00
720/1000 - Cross-layer sparse attention with shared routing; meaningful long-context inference efficiency win.

A while back I wrote about language modeling without neural networks, where I generated Shakespeare with an unbounded n-gram model: no weights, no training, …

Lobste.rsnewsLobste.rs
720/1000
#281
Simultaneous Tricolor Video Observations of Three Tiny Near-Earth Asteroids with Sub-Minute Rotation Periods
ArXiv / 2026-06-16 04:00
720/1000 - Propensity-aware memorization evaluation; advances privacy/safety measurement for LLMs.

Studying the physical properties of near-Earth asteroids (NEAs) is crucial for understanding their dynamical histories and origins, and assessing impact hazards to Earth. Tiny NEAs with diameters…

ArXivpaperastro-ph.EPastro-ph.IM
720/1000
#285
HumP-KD: A Hybrid Uncertainty-Aware Multi-Stage Progressive Knowledge Distillation Framework for Efficient Fire Classification
ArXiv / 2026-06-15 04:00
720/1000 - Multimodal agent skills with visual grounding addresses real text-only bottleneck. Novel framing.

Real-time fire classification systems require models that are simultaneously accurate, computationally efficient, and deployable on resource-constrained hardware. This work proposes \textbf{HumP-KD},…

ArXivpapercs.CVcs.LG
720/1000
#286
puppeteer/puppeteer
💻 GitHub Trending / 2026-06-14 11:00
720/1000 - Open-source agent firewall beyond HTTP; high relevance to agent security, addresses concrete deployment gap.

JavaScript API for Chrome and Firefox

💻 GitHub Trendingnews⭐ 94.6kGitHub Trending
720/1000
#287
The Cold-Start Safety Gap in LLM Agents
HuggingFace / 2026-06-12 11:00
720/1000 - Working memory for latent reasoning from Hochreiter lab is notable for chain-of-thought research.

Are tool-calling LLM agents equally safe throughout a conversation? We discover they are not: agents are most vulnerable at the very start of a session and become substantially…

HuggingFacepaper▲ 1HuggingFace Papers
720/1000
#289
DifussionGemma 4 on 4x7900xtx
👽 Reddit / 2026-06-11 11:00
720/1000 - Native multimodal embedding model from Gemini with strong benchmark results; high impact.

Just got 100 tps on generation, but in total time it around 45-60 t/s in case of prompt processing waiting. Available memory show: GPU KV cache size: 152,671 tokens Maximum…

👽 RedditnewsReddit
720/1000
#290
The Simplified Stabilizer ZX-Calculus is Minimal
ArXiv / 2026-06-11 04:00
720/1000 - Terminal agents learning world models from free env feedback; doubles GRPO pass@1, strong agents fit.

The stabilizer fragment of the ZX calculus is amongst the most important fragments of the theory. The closely related Clifford+T fragment is approximately universal (arXiv:1705.11151). Additionally,…

ArXivpaperquant-ph
720/1000
#291
AI Epistemic Risks: Emerging Mechanisms & Evidence [R]
👽 Reddit / 2026-06-10 08:00
720/1000 - Strong practical tool: proxy metrics beating loss for model selection addresses a real, costly LLM development pain point.

How will AI affect our ability to think and judge for ourselves? Our new paper co-authored by 30 experts explores epistemic risks —the threats AI poses to our collective capacity…

👽 RedditnewsReddit
720/1000
#292
Two-Way Confidential VMs (2cVM): Collaborative Confidential Computing for Mutually Distrustful Parties
ArXiv / 2026-06-10 04:00
720/1000 - Strong benchmark for long-horizon agent memory under interference; high relevance to agent research.

Collaborative computation across organizations is often constrained by the need to process sensitive data and proprietary code without exposing them to untrusted infrastructure or participants.…

ArXivpapercs.CR
720/1000
#294
Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation
ArXiv / 2026-06-10 04:00
720/1000 - Google API key revocation lag (23 min) is a concrete, exploitable security finding relevant to LLM-agent credential hygiene.

Synthetic post-training pipelines commonly filter generated samples with reward models or holistic LLM judges, yet two practices remain rarely examined together: whether the filtering signal is…

ArXivpapercs.CLcs.AI
720/1000
#298
Expanding Private Cloud Compute
Lobste.rs / 2026-06-09 11:00
720/1000 - Overeager coding agents benchmark is highly relevant to agent safety and authorization.

Alongside the next generation of Apple Intelligence, today we’re expanding Private Cloud Compute (PCC) beyond Apple’s data centers. When Apple introduced Private Cloud Compute in…

Lobste.rsnewsLobste.rs
720/1000
#299
OPRD: On-Policy Representation Distillation
HuggingFace / 2026-06-06 00:00
720/1000 - Companion IEEE piece emphasizing variable-precision requirements; relevant quantum-strategy signal.

On-policy distillation (OPD) supervises the student only in output space by matching next-token probabilities. This output-only paradigm has two limits: (1) sampling variance from Monte Carlo KL estimates over large vocabularies (e.g., Qwen's ~150k tokens) …

HuggingFacepaperHUGGINGFACE_PAPERSdaily curated papers
720/1000
#300
BRepCLIP: Contrastive Multimodal Pretraining on BRep Primitives for CAD Understanding
HuggingFace / 2026-06-06 00:00
720/1000 - CIPO turns failed trajectories into correction supervision; clean, broadly useful RLVR extension.

Learning representations of CAD models is a largely open problem. While 3D representation learning has flourished around point clouds and meshes, the native format of CAD - boundary representations BReps, which encodes exact parametric surfaces …

HuggingFacepaperHUGGINGFACE_PAPERSdaily curated papers
720/1000
#302
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses
HuggingFace / 2026-06-02 11:00
720/1000 - Project Zero 0-click Pixel 10 chain is high-value offensive research; demonstrates shifting attack surfaces.

Search agents are often trained as policies over growing transcripts: the model must decide how to search while also remembering what it has seen, which evidence is useful, which constraints remain open …

HuggingFacepaperHuggingFacedaily curated papers2026-06-01
720/1000
#304
LoMo: Local Modality Substitution for Deeper Vision-Language Fusion
HuggingFace / 2026-05-29 19:00
720/1000 - SenseNova-U1 native unified multimodal MoT model is a meaningful system-level release rivaling top VLMs.

Vision-Language Models (VLMs) have achieved substantial progress across a wide range of understanding and reasoning tasks, driven by large-scale image-text training aimed at multimodal fusion. Ideally …

HuggingFacepaperHuggingFacedaily curated papers2026-05-28
720/1000
#305
Performance analysis and industry deployment of post-quantum cryptography algorithms
ArXiv / 2026-05-28 04:00
720/1000 - First controllable agent red-teaming platform is high-impact for AI safety and security evaluation.

As quantum computing advances, modern cryptographic standards face an existential threat, necessitating a transition to post-quantum cryptography (PQC). The National Institute of Standards and Technology (NIST) has…

ArXivpaperGOOGLE_SCHOLAR
720/1000
#306
Are we self-sovereign PKI yet?
Lobste.rs / 2026-05-26 11:00
720/1000 - Important benchmark exposing persistent LLM-agent cyberattack biases; directly relevant to AI safety in agents.

Every end-to-end encrypted messenger ships a fingerprint UI. Almost nobody opens it. A note on what's actually missing.

Lobste.rsnewsLobste.rs
720/1000
#307
SEAL: Synergistic Co-Evolution of Agents and Learning Environments
HuggingFace / 2026-05-26 04:00
720/1000 - Mechanistic identification of MMS collapse in deep DiTs and MV-Split fix; high-impact training-stability insight.

SEAL co-evolves both the agent policy and its training environment in a closed loop, using turn-level failure diagnoses as a shared signal for both environment adaptation and policy optimization, yielding…

HuggingFacepapercs.AIcs.ML
720/1000
#308
Imec builds first High-NA EUV-fabricated quantum dot qubit
Hacker News / 2026-05-26 04:00
720/1000 - Notable: Mythos autonomously finding a real curl vuln demonstrates frontier AI cyber capability milestone.

Imec has built the first quantum dot qubit fabricated with High-NA EUV lithography, a breakthrough that could pull quantum computing onto the same manufacturing roadmap as next-gen AI processors.

Hacker NewsnewsHacker News
720/1000
#309
Minimalist Visual Inertial Odometry
HuggingFace / 2026-05-22 11:00
720/1000 - StraTA improves agentic RL with trajectory-level strategy, useful advance.

Visual-Inertial Odometry(VIO), which is critical to mobile robot navigation, uses cameras with a large number of pixels. Capturing and processing camera images requires significant resources. This work presents a minimalist approach to planar odometry, demons...

HuggingFacepaperarXiv:2605.19990
720/1000
#310
Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles
HuggingFace / 2026-05-22 11:00
720/1000 - TIDE re-examines token index embedding, addresses rare-token and collapse issues.

The proliferation of large language models (LLMs) and modular skills has endowed autonomous agents with increasingly powerful capabilities. Existing frameworks typically rely on monolithic LLMs and fixed logic to interface with these skills. This gives rise t...

HuggingFacepaperarXiv:2605.22177
720/1000
#311
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards
HuggingFace / 2026-05-22 11:00
720/1000 - Notable: io_uring ZCRX freelist LPE giving root from u32 is impactful kernel exploit technique.

Reinforcement learning from verifiable rewards (RLVR) has emerged as a central technique for improving the reasoning capabilities of large language models. Despite its effectiveness, how response-level rewards translate into token-level probability changes re...

HuggingFacepaperarXiv:2605.21467
720/1000
#312
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook
HuggingFace / 2026-05-21 16:00
720/1000 - Tsallis-loss continuum for RLVR cold-start; theoretically grounded, high impact for reasoning post-training.

The foundational capabilities established by Large Language Models (LLMs) have paved the way for Multimodal Large Language Models (MLLMs), within which Large Audio Language Models (LALMs) are essential for realizing universal auditory intelligence. Despite th...

HuggingFacepaperHUGGINGFACE
720/1000
#314
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
HuggingFace / 2026-05-20 09:00
720/1000 - HiL-Bench's Ask-F1 directly measures a crucial agent judgment gap (when to escalate); high-impact agent-safety benchmark.

We present GoLongRL, a fully open-source, capability-oriented post-training recipe for long-context reinforcement learning with verifiable rewards (RLVR). Existing long-context RL methods often treat…

HuggingFacepaperHuggingFace Papers
720/1000
#315
Bridging the Cybersecurity Gap Between Web2 and Web3 - An Incident-Based Analysis of Organizational and Application-Level Security Failures
ArXiv / 2026-05-19 04:00
720/1000 - Position paper for Bayes-consistent agentic AI orchestration; high-quality principled framing for agents.

The rapid adoption of Web3 infrastructures has led to a growing number of security incidents affecting cryptocurrency exchanges, custody services and blockchain-based platforms. While existing research predominantly focuses on vulnerabilities in smar…

ArXivpapercs.CR
720/1000
#316
Unlocking Dense Metric Depth Estimation in VLMs
HuggingFace / 2026-05-18 04:00
720/1000 - Red-teaming long-context LLMs for prompt injection/knowledge corruption; high AI safety impact.

Vision-Language Models (VLMs) excel at 2D tasks such as grounding and captioning, yet remain limited in 3D understanding. A key limitation is their text-only supervision paradigm, which under-constrains fine-grained…

HuggingFacepaperhuggingface_papers
720/1000
#317
Context Training with Active Information Seeking
HuggingFace / 2026-05-14 09:00
720/1000 - Identifies SFT-induced hallucinations and proposes self-distillation fix; high practical impact.

Most existing large language models (LLMs) are expensive to adapt after deployment, especially when a task requires newly produced information or niche domain knowledge. Recent work has shown that, by manipulating and optimizing their context,…

HuggingFacepapercs.CLcs.AIcs.LG
720/1000
#318
Efficient Pre-Training with Token Superposition
HuggingFace / 2026-05-13 19:00
720/1000 - Multi-thinker CoT learning theory; deep learning-theoretic contribution.

Pre-training of Large Language Models is often prohibitively expensive and inefficient at scale, requiring complex and invasive modifications in order to achieve high data…

HuggingFacepaperHuggingFace Papers
720/1000
#319
MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments
HuggingFace / 2026-05-13 11:00
720/1000 - Unified VLA safety survey tying embodied AI risks to attacks/defenses—high relevance, well-scoped.

The Model Context Protocol (MCP) has unified the interface between Large Language Models (LLMs) and external tools, yet a fundamental gap remains in how agents conceptualize the environments within which they operate. Current paradigms are…

HuggingFacepaperarXiv:2605.09131
720/1000
#321
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
HuggingFace / 2026-05-12 04:00
720/1000 - SIREN internal-layer LLM safety guard; strong, efficient safety method with generalization and streaming benefits.

Image captioning is one of the most fundamental tasks in computer vision. Owing to its open-ended nature, it has received significant attention in the era of multimodal large…

HuggingFacepaperHUGGINGFACE
720/1000
#323
Privacy by Postprocessing the Discrete Laplace Mechanism
ArXiv / 2026-05-09 09:00
720/1000 - Reverse-engineering SynthID matters for AI provenance; direct relevance to detection and watermarking.

We show that an "old dog", the classical discrete Laplace (aka.~geometric) mechanism, can "perform new tricks": 1. It can be post-processed to yield a simple, unbiased estimator…

ArXivpapercs.CR
720/1000
#324
Let's Encrypt Stopping Issuance for Potential Incident
Lobste.rs / 2026-05-08 19:00
720/1000 - Proxy Compression Hypothesis unifying reward hacking is a strong conceptual AI-safety contribution.

acme-v02.api.letsencrypt.org (Production), acme-staging-v02.api.letsencrypt.org (Staging), portal.letsencrypt.org (Production), portal-staging.letsencrypt.org (Staging)

Lobste.rsnewsLobste.rs
720/1000
#325
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction
HuggingFace / 2026-05-08 11:00
720/1000 - Plan-and-act via world models is a meaningful robotics advance, strong empirical results over VLA.

Large language models (LLMs) are increasingly used as interactive agents, but optimizing them for long-horizon decision making remains difficult because current methods are largely purely reactive, wh…

HuggingFacepaperarXivHuggingFace
720/1000
#328
Generating Statistical Charts with Validation-Driven LLM Workflows
ArXiv / 2026-05-04 04:00
720/1000 - iTerm2 arbitrary code execution via cat; high-impact terminal RCE, strong security relevance.

Generating diverse, readable statistical charts from tabular data remains challenging for LLMs, as many failures become apparent after rendering and are not detectable from data or code alone. Existing chart datasets als…

ArXivpapercs.LG
720/1000
#329
Your Container Is Not a Sandbox
Lobste.rs / 2026-05-03 11:00
720/1000 - Cloning hardness equals learning for stabilizers; meaningful quantum complexity result.

The microVM ecosystem was battle-tested long before agentic AI created the demand. A landscape survey: every VMM, the shared Rust crate ecosystem, a dozen AI…

Lobste.rsnewsLobsterssecurity
720/1000
#330
Show HN: Large Scale Article Extract of Newspapers 1730s-1960s
Hacker News / 2026-05-02 09:00
720/1000 - Target Policy Optimization cleanly separates target construction from gradient; foundational RL contribution.

A project that extracts and organizes large-scale historical newspaper articles spanning from the 1730s through the 1960s, making centuries of print archives accessible for AI-powered computational…

Hacker NewsnewsHacker News
720/1000
#333
Outer-Crust Equations of State for Neutron Stars
ArXiv / 2026-04-30 11:00
720/1000 - Reasoning-augmented reward models for visual generation; strong training+test-time impact, high upvote signal.

We construct and systematically assess four outer-crust equations of state based on relativistic nuclear mass models and a machine-learning mass table. Our aim…

ArXivpapernucl-th
720/1000
#334
Approaching the Limit of Quantum Clock Precision
ArXiv / 2026-04-27 04:00
720/1000 - First backdoor attack on pipeline-parallel post-training is a significant, novel security result.

Precise and autonomous clocks are of fundamental interest and central importance to both foundational studies and practical applications. Here, we construct a blueprint for a…

ArXivpaperquant-ph,cond-mat.other,physics.comp-ph
720/1000
#335
Seeing Fast and Slow: Learning the Flow of Time in Videos
ArXiv / 2026-04-25 04:00
720/1000 - Worst-case analysis of OPI / DQI is directly tied to a frontier quantum algorithm; strong theoretical CS-quantum link.

How can we tell whether a video has been sped up or slowed down? How can we generate videos at different speeds? Although videos have been central to modern computer vision research, little attention …

ArXivpapercs.CVcs.AIcs.GR
720/1000
#336
CrossCommitVuln-Bench: A Dataset of Multi-Commit Python Vulnerabilities Invisible to Per-Commit Static Analysis
ArXiv / 2026-04-25 04:00
720/1000 - Deep dive into LLM post-training for reasoning is a high-signal survey central to current frontier work.

We present CrossCommitVuln-Bench, a curated benchmark of 15 real-world Python vulnerabilities (CVEs) in which the exploitable condition was introduced across multiple commits - each individually benig…

ArXivpapercs.CRcs.SE
720/1000
#337
"We are currently clean on OPSEC": Why JD Can't Encrypt
ArXiv / 2026-04-22 09:00
720/1000 - LLM 'alignment' context injection exploit: important AI-safety attack class.

We analyse the 2025 Signalgate leak of sensitive US military information by the Trump administration, addressing why confidentiality was violated (messages leaked to the press) in …

ArXivpapercs.CRcs.CYcs.HC
720/1000
#340
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
ArXiv / 2026-04-21 04:00
720/1000 - Challenges Platonic Representation Hypothesis with rigorous empirical critique at scale; high relevance.

The Platonic Representation Hypothesis suggests that neural networks trained on different modalities (e.g., text and images) align and eventually converge toward the same representation of reality. If true, this has sign…

ArXivpapercs.CVcs.AIcs.LG
720/1000
#341
Defense in Depth: A Practical Guide to Python Supply Chain Security
Lobste.rs / 2026-04-19 19:00
720/1000 - Practical Python supply-chain security playbook; directly useful, credential hygiene focus.

Layer your defenses and don’t trust any single control. Use Ruff with security rules to catch bugs in your code before they ship. Pin all your dependencies with cryptographic hashes using uv lock or uv pip compile --generate-hashes so nobody can…

Lobste.rsnewsLobste.rssecurity▲ 1
720/1000
#342
Don't Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG
HuggingFace / 2026-04-17 11:00
720/1000 - Distilling corpora into navigable skill trees for RAG agents is a strong agent-architecture contribution.

Retrieval-Augmented Generation (RAG) grounds LLM responses in external evidence but treats the model as a passive consumer of search results: it never sees how the corpus is organized or what it has…

HuggingFacepaperHuggingFace
720/1000
#343
MCPThreatHive: Automated Threat Intelligence for Model Context Protocol Ecosystems
ArXiv / 2026-04-16 09:00
720/1000 - MCP threat intelligence platform directly relevant to agent security; timely given MCP ecosystem growth and lack of tooling.

The rapid proliferation of Model Context Protocol (MCP)-based agentic systems has introduced a new category of security threats that existing frameworks are inadequately equipped to address.

ArXivpapercs.CRcs.AI
720/1000
#344
Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory
ArXiv / 2026-04-15 04:00
720/1000 - Theory-grounded CAT for LLM jailbreak defense; directly relevant to AI safety.

Adversarial training (AT) is an effective defense for large language models (LLMs) against jailbreak attacks, but performing AT on LLMs is costly. To improve the efficiency of AT for LLMs, recent studies propose continuous AT (CAT) that searches for adversari...

ArXivpapercs.LGcs.CRstat.ML
720/1000
#345
One Token Away from Collapse: The Fragility of Instruction-Tuned Helpfulness
ArXiv / 2026-04-15 04:00
720/1000 - Instruction-tuned helpfulness collapses under trivial lexical constraints; important safety finding.

Instruction-tuned large language models produce helpful, structured responses, but how robust is this helpfulness when trivially constrained? We show that simple lexical constraints (banning a single punctuation character or common word) cause instruction-tun...

ArXivpapercs.CLcs.AI
720/1000
#346
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
HuggingFace / 2026-04-15 04:00
720/1000 - Full-stack open GUI-agent framework (train/eval/deploy) addresses a real ecosystem gap.

GUI agents drive applications through their visual interfaces instead of programmatic APIs, interacting with arbitrary software via taps, swipes, and keystrokes, reaching a long tail of applications that CLI-based agents cannot. It matters because frontier ag...

HuggingFacepaperHuggingFacedaily curated papers
720/1000
#347
TRACE: Capability-Targeted Agentic Training
HuggingFace / 2026-04-14 19:00
720/1000 - TRACE turns failures into targeted RL envs with routed LoRAs; high-leverage agent training idea.

TRACE turns failed agent trajectories into capability-targeted synthetic RL environments, then trains specialized LoRA adapters and routes the agent to the right one at inference time. It matters because agent…

HuggingFacepaperHuggingFacedaily curated papers2026-04-07Upvotes: 11
720/1000
#348
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
ArXiv / 2026-04-14 04:00
720/1000 - Unified open-source GUI agent training/eval/deployment framework; strong infra contribution to agent research.

GUI agents drive applications through their visual interfaces instead of programmatic APIs, interacting with arbitrary software via taps, swipes, and keystrokes, reaching a long tail of applications that CLI-based agents cannot. Yet progress in this area is b...

ArXivpapercs.LGcs.AIcs.CLcs.CV
720/1000
#350
SpecRLBench: A Benchmark for Generalization in Specification-Guided Reinforcement Learning
ArXiv / 2026-04-28 04:00
719/1000 - Mechanistic localization of alignment refusal circuits; cross-lab replicated, causally validated, high safety impact.

Specification-guided reinforcement learning (RL) provides a principled framework for encoding complex, temporally extended tasks using formal specifications such as linear temporal logic (LTL). While recent methods have …

ArXivpapercs.LG
719/1000
#351
Wormholes and Averaging over N
ArXiv / 2026-05-16 04:00
718/1000 - Five-level taxonomy framing evolution from atomic to agentic/world-modeling generation.

The gravitational path integral produces an asymptotic expansion in $G_N$, a fact which is puzzling in the case of observables that are expected to fluctuate wildly. Wormholes appear to compute ensemble averages of…

ArXivpaperhep-th
718/1000
#353
Seemingly Magical Science Behind Quantum Computing
Hacker News / 2026-04-24 11:00
715/1000 - Defense-in-depth critique of alignment; addresses correlated failure modes, high relevance.

An accessible explainer on the fundamental principles of quantum computing, discussing superposition, entanglement, and quantum gates in the context of why quantum computers…

Hacker NewsnewsHacker News
715/1000
#354
Our evaluation of Claude Mythos Preview’s cyber capabilities
Lobste.rs / 2026-04-14 04:00
715/1000 - Anthropic cyber capability eval of Claude Mythos Preview—important frontier-model offensive-security benchmark.

We conducted cyber evaluations of Anthropic’s Claude Mythos Preview and found continued improvement in capture-the-flag (CTF) challenges and significant improvement on multi-step cyber-attack simulations.

Lobste.rsnews🦞 Lobste.rssecurity2026-04-14
715/1000
#355
Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models
ArXiv / 2026-04-11 19:00 / 4 appearances
712/1000 - ESI framework identifying safety-critical parameters across architectures is highly actionable for LLM safety.

Ensuring Large Language Model (LLM) safety is crucial, yet the lack of a clear understanding about safety mechanisms hinders the development of precise and reliable methodologies…

ArXivpapercs.CR
712/1000
#356
Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems
HuggingFace / 2026-06-18 11:00
712/1000 - Queueing/network model of AI-accelerated vuln discovery is timely and strategically relevant to defenders.

Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are grounded in different cultural backgrounds. Existing cultural…

HuggingFacepaper▲ 7HuggingFace Papers
712/1000
#358
PreScam: A Benchmark for Predicting Scam Progression from Early Conversations
HuggingFace / 2026-05-15 11:00
712/1000 - TIDE: first cross-architecture dLLM distillation; meaningful step for efficient diffusion LLMs.

PreScam benchmarks how well models can predict the progression of real-world conversational scams from early dialogue turns, using 11,573 annotated instances across 20 scam categories; results show current LLMs…

HuggingFacepapercs.AIcs.CYHuggingFace
712/1000
#360
Process Reward Agents for Steering Knowledge-Intensive Reasoning
HuggingFace / 2026-04-13 11:00
712/1000 - Online step-wise retrieval-augmented process rewards for medical reasoning; methodologically strong and timely.

Reasoning in knowledge-intensive domains remains challenging as intermediate steps are often not locally verifiable: unlike math or code, evaluating step correctness may require synthesizing clues across large external…

HuggingFacepaperarXiv:
712/1000
#362
WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes
HuggingFace / 2026-05-18 19:00 / 3 appearances
710/1000 - V-GRPO makes ELBO-based diffusion RL practical; strong impact for alignment.

Recent 3D world modeling systems based on generative scene synthesis, such as Marble, can create coherent and explorable 3D environments, yet their outputs are typically static monolithic assets with limited editability…

HuggingFacepaper2605HuggingFacehuggingface_papers
710/1000
#364
Caching for Dollars, Not Hits: An Exact Offline Reference for Cloud-Egress Caching and the Crossover That Decides When It Pays
ArXiv / 2026-06-19 04:00
710/1000 - Multi-author paper on AI epistemic risks is highly relevant to AI safety research strategy.

When a cache miss fetches from cloud object storage, the bill is per GET request and per byte of egress, not latency. Classic caching minimizes the miss rate, the wrong objective: a rarely but…

ArXivpapercs.DBcs.DS
710/1000
#371
Signature-based intrusion detection using machine learning and deep learning approaches empowered with fuzzy clustering
🧪 Semantic Scholar / 2026-06-10 08:00
710/1000 - Shannon Scaling Law unifying SNR explains overtraining/quantization; foundational theoretical advance.

Network security is crucial in today’s digital world, since there are multiple ongoing threats to sensitive data and vital infrastructure. The aim of this study to improve network…

🧪 Semantic Scholarnews📊 123 citesSemantic Scholar
710/1000
#372
Releasing Apodex-1.0 Smol Models (0.8B, 2B, 4B Open-Weights) optimized for Agentic Verification + AgentHarness Evals
👽 Reddit / 2026-06-10 08:00
710/1000 - Decoupling erase/write in linear attention is a clean architectural improvement over Gated DeltaNet/KDA; likely to be adopted.

Hey r/LocalLLaMA , We just released Apodex 1.0 , and alongside our flagship API, we are releasing the weights for our Smol models (0.8B, 2B, and 4B) . Our core research focuses on…

👽 RedditnewsReddit
710/1000
#373
Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories
HuggingFace / 2026-06-03 04:00
710/1000 - New dimensional-constraint mechanism for mixed-state LRE; significant conceptual advance in symmetry-protected phases.

The past few decades have witnessed significant advances in the design of machine learning algorithms, from early studies on task-specific shallow models to more general deep Large Language Models (LLMs). Despite…

HuggingFacepaperHUGGINGFACE PAPERS
710/1000
#374
On Language Generation in the Limit with Bounded Memory
ArXiv / 2026-05-30 04:00
710/1000 - Position paper reframing LLM inference as energy-to-token production; foundational framing.

This item reviews on language generation in the limit with bounded memory and its broader implications for machine learning progress. It matters because AI methods are increasingly spilling into adjacent scientific and industrial domains. …

ArXivpapercs.DS, cs.AI, cs.CL, cs.LG, stat.ML
710/1000
#375
CVE-2026-48710: A Maintainer's Perspective
Lobste.rs / 2026-05-29 19:00
710/1000 - Clever internal-state baseline for RLVR cuts critic/rollout cost; solid methodological contribution to post-training.

Marcelo Trylesinski CVE-2026-48710: A Maintainer's Perspective var media,input,key,value,palette=__md_get("__palette");if(palette&&palette.color){"(prefers-color-scheme)"===palette.color.media&&(media=matchMedia("(prefers-color-scheme: lig …

Lobste.rsnewsLobste.rssecurity2026-05-29
710/1000
#376
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
HuggingFace / 2026-05-25 16:00
710/1000 - High impact: exposes fundamental flaws in global LLM leaderboards via Arena data analysis.

Existing scaling laws for Large Language Models (LLMs), predominantly monotonic power laws, fail to explain emerging non-monotonic phenomena such as catastrophic overtraining and quantization-induced degradation, where performance deteriora...

HuggingFacepaperHuggingFace
710/1000
#377
RiT: Vanilla Diffusion Transformers Suffice in Representation Space
HuggingFace / 2026-05-22 11:00
710/1000 - AISLE-found 21-year RCE in FreeBSD dhclient is a notable vuln discovery.

Flow matching with x-prediction -- regressing the clean data point rather than the ambient velocity -- is known to exploit low-dimensional manifold structure effectively in pixel space li2025back. We ask whether a pretrained representation space, while contai...

HuggingFacepaperarXiv:2605.21981
710/1000
#378
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
ArXiv / 2026-05-16 04:00
710/1000 - CoPD unifies RLVR+OPD with parallel expert co-evolution; strong post-training methodology contribution.

Backdoor attacks pose a serious security threat to large language models (LLMs), which are increasingly deployed as general-purpose assistants in safety- and privacy-critical applications. Existing LLM backdoors rely…

ArXivpapercs.CRcs.CL
710/1000
#380
TIDE: Every Layer Knows the Token Beneath the Context
HuggingFace / 2026-05-08 09:00
710/1000 - Micro-LM cloud handoff for edge latency is a practical, well-motivated architecture shift.

We revisit a universally accepted but under-examined design choice in every modern LLM: a token index is looked up once at the input embedding layer and then permanently discarded. This single-injection assumption…

HuggingFacepaperHuggingFaceLLMs
710/1000
#381
SkillOS: Learning Skill Curation for Self-Evolving Agents
HuggingFace / 2026-05-08 09:00
710/1000 - Duplicate of Micro-LM paper; edge-cloud handoff architecture already credited.

LLM-based agents are increasingly deployed to handle streaming tasks, yet they often remain one-off problem solvers that fail to learn from past interactions. Reusable skills distilled from experience provide a natural…

HuggingFacepaperHuggingFaceLLMs
710/1000
#382
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
ArXiv / 2026-05-06 11:00
710/1000 - Important reliability framing for computer-use agents; statistical decomposition on OSWorld is valuable.

Reasoning-intensive retrieval aims to surface evidence that supports downstream reasoning rather than merely matching topical similarity. This capability is increasingly important for agentic search systems, where retrievers must provide complementary evidenc...

ArXivpapercs.CLcs.IR
710/1000
#384
CAbLECAR: efficiently scheduling QLDPC codes on a tileable spin qubit chip with shuttling
ArXiv / 2026-04-28 04:00
710/1000 - ClawGuard directly addresses indirect prompt injection in tool-augmented agents; highly relevant to agent safety.

Semiconductor spin qubits are a promising platform for large-scale quantum computing, but have yet to take full advantage of the broad class of quantum low-density parity check (QLDPC) codes, which promise high encoding …

ArXivpaperquant-phcs.ET
710/1000
#385
Trustworthy Technology
Lobste.rs / 2026-04-23 11:00
710/1000 - RAG security taxonomy across pipeline; strong framework, high relevance.

A (hopeful) new movement dedicated to a simple proposition—that our technology products should respect us! That is, support our wishes and uphold the principles of freedom, privacy, and informed conse…

Lobste.rsnewsLobste.rs
710/1000
#387
Tight Auditing of Differential Privacy in MST and AIM
ArXiv / 2026-04-21 04:00
710/1000 - Tight GDP audits of MST/AIM DP synthetic data; foundational privacy auditing advance.

State-of-the-art Differentially Private (DP) synthetic data generators such as MST and AIM are widely used, yet tightly auditing their privacy guarantees remains challenging. We introduce a Gaussian Differential Privacy …

ArXivpapercs.CRcs.AIcs.LG
710/1000
#388
Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
HuggingFace / 2026-04-21 04:00
710/1000 - Precision-aware debugging benchmark exposes frontier model over-editing; sharp finding.

Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but over-edited solutions during debugging. To evaluate how far LLMs are fr…

HuggingFacepaperHuggingFace
710/1000
#391
POISE: Position-Aware Undetectable Skill Injection on LLM Agents
HuggingFace / 2026-06-11 11:00
708/1000 - Error tracing/attribution for LLM memory systems is a needed debugging primitive; benchmark is valuable.

Agent skills provide a lightweight mechanism for extending general-purpose agents, but their open format exposes them to skill-poisoning attacks. A practically dangerous injection…

HuggingFacepaper▲ 3HuggingFace Papers
708/1000
#392
Show HN: Bitcoin and Quantum Computing – a three-part research series
Hacker News / 2026-04-12 16:00
708/1000 - Bitcoin and quantum computing series directly hits crypto+quantum intersection; high debrief relevance.

A Hacker News showcase of a three-part research series examining the intersection of Bitcoin security and quantum computing, exploring how quantum advances could impact cryptographic foundations of blockchain systems.

Hacker NewsnewsHacker News
708/1000
#395
DETOUR: A Practical Backdoor Attack against Object Detection
ArXiv / 2026-04-28 04:00
703/1000 - RL on physics simulators for Olympiad reasoning; novel data-scaling approach beyond math, strong fit.

Object detection (OD) is critical to real-world vision systems, yet existing backdoor attacks on detection transformers (DETRs) for OD tasks rely on patch-wise triggers optimized at fixed locations with minimal perturbat…

ArXivpapercs.CR
703/1000
#397
D^2-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing
HuggingFace / 2026-05-27 09:00 / 3 appearances
700/1000 - Detailed postmortem of TanStack npm compromise with multi-stage attack chain is high-impact.

Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monitoring for D-LLMs remains largely unexplored. Unlike AR-LLMs, D-LLMs…

HuggingFacepaperHuggingFace
700/1000
#400
AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints
HuggingFace / 2026-06-06 09:00 / 2 appearances
700/1000 - Audited olympiad physics corpus uncovers contamination, translation drift, MCQ saturation; high integrity value.

Dynamic interactive benchmark for adaptive planning under progressively revealed world and user constraints — best model reaches only 67.75% accuracy with user constraints particularly problematic due to weak physical grounding.

HuggingFacepaperHUGGINGFACE_PAPERSHuggingFacedaily curated papers
700/1000
#401
Mullvad exit IPs as a fingerprinting vector
HuggingFace / 2026-05-15 04:00 / 2 appearances
700/1000 - Largest trapped-ion protein-folding demo on 64 qubits; concrete hardware milestone.

Analysis of how Mullvad VPN's smaller server fleet creates a fingerprinting vector that tracks users across sites despite its no-logging reputation.

HuggingFacepaperHFLOBSTERS
700/1000
#404
Sumi: Open Uniform Diffusion Language Model from Scratch
HuggingFace / 2026-06-18 04:00
700/1000 - Evaluation Cards address a real, broadly-felt reporting gap across AI eval; strong practical impact and broad applicability.

Diffusion models have become a promising alternative to autoregressive models. Among these, uniform diffusion language models (UDLMs) permit any token to be updated at any step,…

HuggingFacepaper▲ 4HuggingFace Papers
700/1000
#405
How do you analyze the relative "strength" of probes? [R]
👽 Reddit / 2026-06-17 19:00
700/1000 - Clinically grounded, tiered privacy evaluation of medical LMs with concrete AUROC leakage results; high practical safety value.

This question is related to topics like language+ models (including multimodal) and things like "circuit" analyses. I think something related might come up in my work (factuality…

👽 RedditnewsReddit
700/1000
#406
yairm210/Unciv
💻 GitHub Trending / 2026-06-17 11:00
700/1000 - Geometry/parameter-space analysis of on-policy distillation with subspace-locking is a strong, central LLM-RL insight.

Open-source Android/Desktop remake of Civ V

💻 GitHub Trendingnews⭐ 10.6kGitHub Trending
700/1000
#408
Artificial intelligence using a latent diffusion model enables the generation of diverse and potent antimicrobial peptides
🧪 Semantic Scholar / 2026-06-16 11:00
700/1000 - Deployed 300 km trusted-node QKD over multi-core fiber; notable real-world QKD milestone.

Artificial intelligence holds great promise for the design of antimicrobial peptides (AMPs); however, current models face limitations in generating AMPs with sufficient novelty…

🧪 Semantic Scholarnews📊 60 citesSemantic Scholar
700/1000
#409
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
HuggingFace / 2026-06-16 04:00
700/1000 - XSS-to-RCE via MeshCore/Home Assistant; high-impact IoT supply-chain vulnerability.

We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20…

HuggingFacepaper▲ 3HuggingFace Papers
700/1000
#418
Rick & Morty
👽 Reddit / 2026-06-10 08:00
700/1000 - Subproblem curriculum RL for credit assignment is a substantive step beyond outcome-only RLVR, with clear empirical wins.

nobody expected HF there submitted by /u/jacek2023 [link] [comments]

👽 RedditnewsReddit
700/1000
#419
Next Forcing: Causal World Modeling with Multi-Chunk Prediction
ArXiv / 2026-06-10 04:00
700/1000 - Aaronson talk on quantum computing is a high-authority explainer, valuable for research strategy context.

Autoregressive video generation has emerged as a powerful paradigm for World Action Models (WAMs). However, existing approaches suffer from slow training convergence and limited converged accuracy,…

ArXivpapercs.CV
700/1000
#423
Will the Agent Recuse Itself? Measuring LLM-Agent Compliance with In-Band Access-Deny Signals
ArXiv / 2026-06-06 09:00
700/1000 - 'Code as agent harness' survey frames a major emerging paradigm; high reference value for the debrief.

Proposes the Recuse Signal — a lightweight in-band deny signal over existing protocol channels — to tell LLM agents to recuse themselves from off-limits resources, measuring compliance across major agents.

ArXivpapercs.CRcs.AI
700/1000
#425
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
HuggingFace / 2026-06-03 04:00
700/1000 - AlphaEvolve applied to FHE kernels on TPUs is a notable cross-cut of LLM-driven optimization and applied crypto performance.

True video intelligence demands more than recognizing what is visible: it requires reasoning about why events unfold, predicting what would change under different conditions, and deciding what to do next. We refer to…

HuggingFacepaperHUGGINGFACE PAPERS
700/1000
#426
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
HuggingFace / 2026-06-01 04:00
700/1000 - CiteTracer multi-agent citation hallucination detector with 12-code taxonomy and 2,450 benchmark; high impact.

Zero-shot text-to-speech (TTS) has improved substantially for single-speaker synthesis, yet expressive long-form multi-speaker dialogue remains difficult …

HuggingFacepaperHuggingFacedaily curated papers2026-05-29
700/1000
#428
In-Context Reward Adaptation for Robust Preference Modeling
ArXiv / 2026-05-30 04:00
700/1000 - Asymmetric flow modeling with rank-asymmetric velocity; 1.57 FID ImageNet state-of-the-art.

This item reviews in-context reward adaptation for robust preference modeling and its broader implications for machine learning progress. It matters because AI methods are increasingly spilling into adjacent scientific and industrial domains. …

ArXivpapercs.LG, cs.AI
700/1000
#430
CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models
HuggingFace / 2026-05-26 04:00
700/1000 - Fast BLT variants materially speed byte-level LMs; addresses a key bottleneck with diffusion/self-speculation tricks.

CRONOS is an intervention-based benchmark that evaluates whether video prediction models respond appropriately to controlled changes in visual inputs — testing if models capture causal physical structure rather than superficial…

HuggingFacepapercs.CVcs.LG
700/1000
#431
Forecasting Scientific Progress with Artificial Intelligence
HuggingFace / 2026-05-24 11:00
700/1000 - High impact: DeepMind AI co-mathematician workbench for agentic mathematical research.

Artificial intelligence (AI) is increasingly embedded in scientific discovery, yet whether it can anticipate scientific progress remains unclear. To study this question, we introdu…

HuggingFacepaperhuggingface papers
700/1000
#433
SceneAligner: 3D-Grounded Floorplan Localization in the Wild
HuggingFace / 2026-05-22 11:00
700/1000 - Zero-shot ILP foundation model with domain-agnostic literal encoding.

Many public buildings provide floorplans with a "you are here" indicator to help visitors orient themselves. Floorplan localization seeks to computationally replicate this capability by determining where visual observations were captured within a floorplan. H...

HuggingFacepaperarXiv:2605.22581
700/1000
#434
Diversed Model Discovery via Structured Table Discovery
HuggingFace / 2026-05-22 11:00
700/1000 - KernelBench-X benchmark for LLM GPU kernels, strong empirical findings.

Model cards describe model behavior through a mixture of textual descriptions and structured artifacts, including performance, configuration, and dataset tables. Existing model search systems rely predominantly on semantic similarity over text, which can prod...

HuggingFacepaperarXiv:2605.22766
700/1000
#435
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
HuggingFace / 2026-05-21 16:00
700/1000 - Comprehensive GFCR rollout-strategy survey for LLM RL; fills real underreported gap.

Contextual Integrity (CI) defines privacy not merely as keeping information hidden, but as governing information flows according to the norms of a given context. As large language models are increasingly deployed as personal agents handling sensitive workflow...

HuggingFacepaperHUGGINGFACE
700/1000
#437
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
HuggingFace / 2026-05-20 09:00
700/1000 - Attention-as-dynamic-MLP reframing enabling linear-time global modeling is a high-leverage architectural insight.

We present OpenComputer, a verifier-grounded framework for constructing verifiable software worlds for computer-use agents. OpenComputer integrates four components: (1) app-specific state verifiers…

HuggingFacepaperHuggingFace Papers
700/1000
#438
CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning
HuggingFace / 2026-05-20 09:00
700/1000 - OpenSeeker-v2 SFT-only 30B search agent beating RL+CPT counterparts is a strong, surprising data-scaling result.

Chain-of-thought (CoT) is a standard approach for eliciting reasoning capabilities from large language models (LLMs). However, the common CoT paradigm treats thinking as a prerequisite for answering,…

HuggingFacepaperHuggingFace Papers
700/1000
#439
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
HuggingFace / 2026-05-19 16:00
700/1000 - PhysicianBench uses real EHR environments; high-impact benchmark for clinical AI agents.

Proposes a paradigm for instance-level video understanding through visual prompts, integrating agentic reasoning with proactive visual perception. Achieves +13.7% improvement over baselines and…

HuggingFacepaperHUGGINGFACE PAPERS↑1
700/1000
#445
KernelBench-X: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
HuggingFace / 2026-05-08 11:00
700/1000 - Stable Firefox identifier breaking Tor privacy is a real cross-browser security/privacy finding.

LLM-based Triton kernel generation has attracted significant interest, yet a fundamental empirical question remains unanswered: where does this capability break down, and why? We present KernelBench-X…

HuggingFacepaperarXivHuggingFace
700/1000
#446
How to pick your football team
ArXiv / 2026-05-06 11:00
700/1000 - Identifies scaling law of miscalibration in OPD with formal theory and fix.

Team captains Alice and Bob divide up $2m$ footballers, each reduced to a real-valued score, into two teams of $m$ footballers each. On each turn, one captain plays picker, and the other chooser: the picker…

ArXivpapermath.CO
700/1000
#447
Web2BigTable: A Bi-Level Multi-Agent LLM System for Internet-Scale Information Search and Extraction
HuggingFace / 2026-05-04 11:00
700/1000 - Deployed 30B multi-agent deep-research system with strong benchmark results; notable production contribution.

Agentic web search increasingly faces two distinct demands: deep reasoning over a single target, and structured aggregation across many entities and heterogeneous sources. Current systems struggle on both fronts. Breadth-oriented…

HuggingFacepaperHuggingFace
700/1000
#448
Meritocratic Fairness in Budgeted Combinatorial Multi-armed Bandits via Shapley Values
ArXiv / 2026-05-04 04:00
700/1000 - Memory transfer learning across coding domains with strong empirical gains; high relevance to agents.

We propose a new framework for meritocratic fairness in budgeted combinatorial multi-armed bandits with full-bandit feedback (BCMAB-FBF). Unlike semi-bandit feedback, the contribution of individual arms is not received i…

ArXivpapercs.LGcs.AIcs.MA
700/1000
#450
Sequential Inference for Gaussian Processes: A Signal Processing Perspective
ArXiv / 2026-05-03 09:00
700/1000 - Real HTTP-desync vuln in Discord media proxy; notable platform-scale security finding.

The proliferation of capable and efficient machine learning (ML) models marks one of the strongest methodological shifts in signal processing (SP) in its nearly 100-year history. ML models support the development of SP systems that…

ArXivpapereess.SPcs.LGstat.COstat.ML
700/1000
#455
The Geometric Canary: Predicting Steerability and Detecting Drift via Representational Stability
HuggingFace / 2026-04-21 04:00
700/1000 - Geometric stability predicts steerability and drift; strong dual-use interpretability result.

Reliable deployment of language models requires two capabilities that appear distinct but share a common geometric foundation: predicting whether a model will accept targeted behavioral control, and detecting when its in…

HuggingFacepaperHuggingFace
700/1000
#456
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
HuggingFace / 2026-04-15 11:00
700/1000 - LASA semantic-bottleneck multilingual LLM safety; 24.7%->2.8% ASR drop is strong, novel mechanistic finding.

Large language models (LLMs) often demonstrate strong safety performance in high-resource languages, yet exhibit severe vulnerabilities when queried in low-resource languages …

HuggingFacepaperHuggingFacedaily curated papers2026-04-13
700/1000
#457
Many-Tier Instruction Hierarchy in LLM Agents
HuggingFace / 2026-04-15 04:00
700/1000 - Many-tier instruction hierarchy is a timely safety/agent contribution with a new benchmark.

Large language model agents receive instructions from many sources-system messages, user prompts, tool outputs, and more-each carrying different levels of trust and authority. When these instructions conflict, models must reliably follow the highest-privilege...

HuggingFacepaperHuggingFacedaily curated papers
700/1000
#458
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
ArXiv / 2026-04-15 04:00
700/1000 - Lightning OPD identifies teacher-consistency; actionable post-training insight, broad impact.

On-policy distillation (OPD) has emerged as an efficient post-training paradigm for large language models. However, standard OPD requires a live teacher inference server throughout training, resulting in substantial infrastructure overhead. It matters because...

ArXivpapercs.LGcs.AI
700/1000
#459
Introspective Diffusion Language Models
HuggingFace / 2026-04-14 04:00
700/1000 - Introspective diffusion LMs with parallel decoding + AR-style consistency; strong contribution to DLMs.

Diffusion language models promise parallel generation, yet still lag behind autoregressive (AR) models in quality. We stem this gap to a failure of introspective consistency: AR models agree with their own generations, while DLMs often do not.

HuggingFacepaper🤗 HuggingFacedaily curated papers2026-04-13
700/1000
#460
Llm post-training: A deep dive into reasoning large language models
ArXiv / 2026-04-13 04:00
700/1000 - Deep dive into LLM post-training for reasoning; strong synthesis of a high-impact subfield.

Large Language Models (LLMs) have transformed the natural language processing landscape and brought to life diverse applications. Pretraining on vast web-scale data has laid the foundation for these models, yet the…

ArXivpaperGoogle Scholar
700/1000
#461
EXAONE 4.5 Technical Report
HuggingFace / 2026-04-13 04:00
700/1000 - Open-weight LG vision-language model, 256K context, document-strong; notable open release.

This technical report introduces EXAONE 4.5, the first open-weight vision language model released by LG AI Research. EXAONE 4.5 is architected by integrating a dedicated visual encoder into the existing EXAONE 4.0…

HuggingFacepaperHuggingFace
700/1000
#465
mattermost/mattermost
💻 GitHub Trending / 2026-06-11 11:00
694/1000 - Spec autoformalization benchmark for Verus is valuable for code-agent verification research.

Mattermost is an open source platform for secure collaboration across the entire software development lifecycle..

💻 GitHub Trendingnews⭐ 37.2kGitHub Trending
694/1000
#469
Articraft: An Agentic System for Scalable Articulated 3D Asset Generation
ArXiv / 2026-05-16 04:00
691/1000 - Live, refreshable agent benchmark with execution-trace grading; strong methodology for agent eval.

A bottleneck in learning to understand articulated 3D objects is the lack of large and diverse datasets. In this paper, we propose to leverage large language models (LLMs) to close this gap and generate articulated…

ArXivpapercs.CVcs.GRcs.RO
691/1000
#472
Re-Centering Humans in LLM Personalization
HuggingFace / 2026-06-18 19:00
690/1000 - Feedback alignment in self-distillation; step-aligned critique beating GRPO is a meaningful result.

Despite growing interest, most evaluations of large language models' (LLMs') personalization abilities have relied on synthetic data. It remains unclear how well current…

HuggingFacepaperHuggingFace Papers
690/1000
#473
Latent space interpretation [R]
👽 Reddit / 2026-06-18 19:00
690/1000 - Itô maps for any-step SDEs; useful primitive for posterior sampling and stochastic control.

Hi all, I have trained a convolutional autoencoder on a set of medical images. Further classified latent feature maps using random forest to find the top scoring feature map. Now…

👽 RedditnewsReddit
690/1000
#479
InterleaveThinker: Reinforcing Agentic Interleaved Generation
ArXiv / 2026-06-12 04:00
690/1000 - Confidence-aware KV-cache eviction with mixed precision is practical long-context contribution.

Recent image generators have demonstrated impressive photorealism and instruction-following capabilities in single-image generation and editing. However, constrained by their architectures, they…

ArXivpapercs.CV
690/1000
#480
FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents
HuggingFace / 2026-06-12 04:00
690/1000 - Scalable parameterized local linear attention is a solid architectural contribution to pretraining.

Training deep search agents requires verifiable questions whose answers remain unavailable until sufficient evidence has been acquired through search. Existing synthesis methods…

HuggingFacepaper▲ 45HuggingFace Papers
690/1000
#484
Opportunities and Challenges in Securely Reusing and Repurposing Mobile Devices
ArXiv / 2026-06-06 09:00
690/1000 - Contrastive neuron attribution for jailbreak-circuit ablation is high-leverage for mechanistic safety.

Investigates whether mobile device security mechanisms remain effective when repurposed beyond their intended lifecycle, finding significant gaps in attestation, key storage, and secure boot for out-of-lifecycle scenarios.

ArXivpapercs.CR
690/1000
#485
EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management
HuggingFace / 2026-06-06 00:00
690/1000 - LLM-agent-driven neural architecture search beating Llama 3.2 is a strong self-improvement signal.

Recent progress in Large Language Model (LLM) agents has enabled promising advances in automated data science. However, existing approaches remain fundamentally limited by their static action sets and lack of principled long-horizon context management, hinder...

HuggingFacepaperHUGGINGFACE_PAPERSdaily curated papers
690/1000
#486
Show, Don't TELL: Explainable AI-Generated Text Detection
HuggingFace / 2026-06-03 04:00
690/1000 - Full exploit for Windows kernel CVE enabling browser sandbox escape; high-impact offensive security artifact.

Research on AI-generated text detection has presented a number of approaches to discern human from AI prose, some of which achieving high in-distribution performance. However, real-world applicability has stalled…

HuggingFacepaperHUGGINGFACE PAPERS
690/1000
#487
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
HuggingFace / 2026-05-25 16:00
690/1000 - Solid: AI-driven vulnerability disclosure culture clash with Linux kernel norms; timely analysis.

Recent advances in vision-language models (VLMs) emphasize long chain-of-thought reasoning; yet, we find that their performance on visual tasks is primarily limited by a lack of visual perception as opposed to reasoning itself. In this work...

HuggingFacepaperHuggingFace
690/1000
#489
Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
HuggingFace / 2026-05-22 11:00
690/1000 - DCI for agentic search challenges fixed-similarity retrieval paradigm.

Scaling laws have made language-model performance predictable from model size, data, and compute, but they typically treat the optimizer as a fixed training detail. We show that this assumption misses a fundamental axis of representation scaling: how effectiv...

HuggingFacepaperarXiv:2605.21803
690/1000
#491
Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty
HuggingFace / 2026-05-13 11:00
690/1000 - Perception-centric process rewards for VLMs—strong multimodal reasoning and hallucination mitigation.

Large language models (LLMs) are increasingly deployed on long-horizon tasks in partially observable environments, where they must act while inferring and tracking a complex environment state over many steps. This leads to two challenges: partial…

HuggingFacepaperarXiv:2605.11436
690/1000
#492
Podman rootless containers and the Copy Fail exploit
Lobste.rs / 2026-05-04 19:00
690/1000 - Systematic study of weak-supervision RLVR with saturation dynamics; valuable insight for reasoning training.

About Microblog Podman rootless containers and the Copy Fail exploit May 4, 2026 Contents An overview of rootless containers Rootless rootful User namespaces Privileged operations…

Lobste.rsnewsLobste.rslinuxsecurity
690/1000
#493
EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training
HuggingFace / 2026-04-27 16:00
690/1000 - Continuous adversarial flow models substantially improving ImageNet FID; strong generative-modeling advance.

EmbodiedMidtrain addresses the distribution gap between generic Vision-Language Models and robot action domains by mid-training VLMs on VLA-aligned data selected via a lightweight proximity estimator, achieving competitive results with expert…

HuggingFacepaperHuggingFace Papers
690/1000
#494
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
HuggingFace / 2026-04-22 11:00
690/1000 - Verified boot analysis exposes systemic trust gap; strong security insight.

Recent Speech-to-Speech Translation (S2ST) systems achieve strong semantic accuracy yet consistently strip away non-verbal vocalizations (NVs), such as laughter and crying that convey pragmatic intent, which severely…

HuggingFacepaperarXiv
690/1000
#495
The Continuity Layer: Why Intelligence Needs an Architecture for What It Carries Forward
HuggingFace / 2026-04-21 16:00
690/1000 - SkillClaw: collective skill evolution for agents; high-upvotes, novel.

The most important architectural problem in AI is not the size of the model but the absence of a layer that carries forward what the model has come to understand. Sessions end. Context windows fill…

HuggingFacepaperHUGGINGFACE_PAPERSdaily curated papers
690/1000
#496
FUSE: Ensembling Verifiers with Zero Labeled Data
ArXiv / 2026-04-21 04:00
690/1000 - FUSE unsupervised verifier ensembling; directly tackles LLM-as-judge weakness, broadly useful.

Verification of model outputs is rapidly emerging as a key primitive for both training and real-world deployment of large language models (LLMs). In practice, this often involves using imperfect LLM judges and reward mod…

ArXivpaperstat.MLcs.CLcs.LG
690/1000
#497
Target Policy Optimization
HuggingFace / 2026-04-16 11:00
690/1000 - TPO cleanly separates target distribution from policy update; potentially impactful RL alternative to PPO/GRPO.

In RL, given a prompt, we sample a group of completions from a model and score them. Two questions follow: which completions should gain probability mass, and how should the…

HuggingFacepaperHuggingFace Papersdaily curated papers
690/1000
#498
A survey of zero-knowledge proof based verifiable machine learning
ArXiv / 2026-04-16 09:00
690/1000 - ZKML survey; timely consolidation of verifiable ML via zero-knowledge proofs, strong crypto-ML intersection.

Machine learning is increasingly deployed through outsourced and cloud-based pipelines, which improve accessibility but also raise concerns about computational integrity, data privacy, and model…

ArXivpaperGoogle Scholar
690/1000
#499
Continuous Adversarial Flow Models
HuggingFace / 2026-04-14 04:00
690/1000 - Continuous adversarial flow post-training substantially improves SiT/JiT FID; strong generative modeling result.

We propose continuous adversarial flow models, a type of continuous-time flow model trained with an adversarial objective. Unlike flow matching, which uses a fixed mean-squared-error criterion, our approach introduces a learned discriminator to guide training.

HuggingFacepaper🤗 HuggingFacedaily curated papers2026-04-13
690/1000