Carlos's Debrief

April 21, 2026 04:00
38ArXiv Papers
47Web Findings
85Total Sources
Topics searched: Artificial Intelligence Machine Learning Large Language Models Security & Cybersecurity Cryptography Zero Knowledge Quantum Computing Crypto & Blockchain AI Agents & Reasoning AI Safety & Alignment
April 21, 2026 04:00 — Curated by Hermes Research Scout

⚡ Quick Summary

📰 News Headlines

📄 ArXiv Papers

🧠 LLMs 13

Shaden Alshammari et al.
Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and task diversity. We introduce MathNet, a h…
cs.AIcs.DLcs.IRcs.LG
Liubomyr Horbatko
Modern sequence models are dominated by Transformers, where self-attention mixes information from the visible context in an input-dependent way. However, when retrieval is not sharp and attention remains diffuse over an …
cs.LGcs.AIcs.CL
Yunke Ao et al.
Proximal Policy Optimization (PPO) has become the predominant algorithm for on-policy reinforcement learning due to its scalability and empirical robustness across domains. However, there is a significant disconnect betw…
cs.LGcs.AI
Kevin Murphy
We present BLF (Bayesian Linguistic Forecaster), an agentic system for binary forecasting that achieves state-of-the-art performance on the ForecastBench benchmark. The system is built on three ideas. (1) A Bayesian ling…
cs.AI
Salman Rahman et al.
Large language models have achieved significant reasoning improvements through reinforcement learning with verifiable rewards (RLVR). Yet as model capabilities grow, constructing high-quality reward signals becomes incre…
cs.LGcs.AI
A. Sophia Koepke et al.
The Platonic Representation Hypothesis suggests that neural networks trained on different modalities (e.g., text and images) align and eventually converge toward the same representation of reality. If true, this has sign…
cs.CVcs.AIcs.LG
Andrew Zhang et al.
Modern medicine generates vast multimodal data across siloed systems, yet no existing model integrates the full breadth and temporal depth of the clinical record into a unified patient representation. We introduce Apollo…
cs.LGcs.AIcs.CL
Manan Gupta, Dhruv Kumar
Large language models frequently commit unrecoverable reasoning errors mid-generation: once a wrong step is taken, subsequent tokens compound the mistake rather than correct it. We introduce $\textbf{Latent Phase-Shift R…
cs.LGcs.AIcs.CL
Terry Leitch
We present a systematic evaluation of large language model families -- spanning both proprietary cloud APIs and locally-hosted open-source models -- on two purpose-built benchmarks for System Dynamics AI assistance: the …
cs.AIcs.HCcs.LG
Xirui Li et al.
Constructing environments for training and evaluating claw-like agents remains a manual, human-intensive process that does not scale. We argue that what is needed is not just a dataset, but an automated pipeline capable …
cs.AIcs.CL
Heming Zhu, Guoxing Sun, Marc Habermann
Building photorealistic, animatable full-body digital humans remains a longstanding challenge in computer graphics and vision. Recent advances in animatable avatar modeling have largely progressed along two directions: i…
cs.CV
Joshua T. Roth et al.
The T16 project has produced a uniformly detrended and systematics-corrected set of 83,717,159 TESS Cycle 1 full-frame image light curves for stars observed by TESS in its primary mission down to T=16 mag, enabling sensi…
astro-ph.EP
Aditya Arora et al.
Story Visualization aims to generate a sequence of images that faithfully depicts a textual narrative that preserve character identity, spatial configuration, and stylistic coherence as the narratives unfold. Maintaining…
cs.CV

📄 ArXiv Papers

🧠 Machine Learning 2

Maria-Eleni Sfyraki, Jun-Kun Wang
In this work, we revisit the problem of active sequential prediction-powered mean estimation, where at each round one must decide the query probability of the ground-truth label upon observing the covariates of a sample.…
stat.MLcs.LG
Minji Lee et al.
Models from the AlphaFold (AF) family reliably predict one dominant conformation for most well-ordered proteins but struggle to capture biologically relevant alternate states. Several efforts have focused on eliciting gr…
q-bio.BMcs.LG

📄 ArXiv Papers

🧠 Security & Crypto 16

Zhiyuan Chen et al.
Privacy policies are intended to inform users about how software systems collect and handle data, yet they often remain vague or incomplete. This paper presents an empirical study of patterns in log-related statements wi…
cs.CRcs.SE
Md Rysul Kabir, Zoran Tiganj
Open-weight language models can be rendered unsafe through several distinct interventions, but the resulting models may differ substantially in capabilities, behavioral profile, and internal failure mode. We study behavi…
cs.CRcs.AIcs.CL
Bowen Cai et al.
Smart contracts extended blockchain functionality beyond simple transactions, powering complex applications like decentralized finance (DeFi). However, this complexity introduces serious security challenges, including pr…
cs.CR
Georgi Ganev, Meenatchi Sundaram Muthu Selva Annamalai, Bogdan Kulynych
State-of-the-art Differentially Private (DP) synthetic data generators such as MST and AIM are widely used, yet tightly auditing their privacy guarantees remains challenging. We introduce a Gaussian Differential Privacy …
cs.CRcs.AIcs.LG
Jan Menz et al.
To ensure programs do not leak private data, we often want to be able to provide formal guarantees ensuring such data is handled correctly. Often, we cannot keep such data secret entirely; instead programmers specify how…
cs.PLcs.CR
Freddy Lendé Metouké et al.
This paper investigates subcodes of lambda-Gabidulin codes, viewed as rank-metric analogues of generalized Reed--Solomon codes, and their applications to compact-ciphertext cryptosystems. We first analyze subspace and ge…
cs.CRcs.IT
Thamilvendhan Munirathinam
Current open-source prompt-injection detectors converge on two architectural choices: regular-expression pattern matching and fine-tuned transformer classifiers. Both share failure modes that recent work has made concret…
cs.CRcs.CL
Sina Abdollahi et al.
Large Language Model (LLM) agents provide powerful automation capabilities, but they also create a substantially broader attack surface than traditional applications due to their tight integration with non-deterministic …
cs.CRcs.OS
Arthur Loureiro et al.
Smokescreen is an open-source Python library for data-vector concealment (blinding) in cosmological analyses. Data-vector blinding works by applying cosmology-dependent shifts to the observed data vector, moving it away …
astro-ph.IMastro-ph.CO
Alexandros Chatzinikolaou et al.
We introduce an operator-algebraic framework for Morita equivalence of quantum graphs based on $Δ$-equivalence of operator systems introduced by Eleftherakis, Kakariadis and Todorov. Adopting the perspective of Weaver, w…
math.OAmath-ph
Haojie Gu, Zhihao Zhu, Jun Zhang
The Schur square of linear codes over a finite field has emerged as a fundamental operation in both classical and quantum coding theory. In this paper, we investigate the Schur square problem of Hyperderivative Reed-Solo…
cs.IT
Shozo Saeki, Minoru Kawahara, Hirohisa Aman
A nearest-neighbor framework is a fundamental tool for various applications involving Large Language Models (LLMs) and Visual Language Models (VLMs). Vectors used for nearest-neighbor searches have richer information for…
cs.CR
Shannon Whitlock
AtomTwin.jl is an open-source Julia package for developing and simulating quantum protocols, hardware configurations and building digital twins for neutral-atom quantum processors and related atomic quantum devices. Atom…
quant-phphysics.atom-ph
Mihai Turinici
The 2015 fixed point result on rs-relational metric spaces due to Alam and Imdad [J. Fixed Point Th. Appl., 17 (2015), 693-702] is equivalent with the classical Banach Contraction Principle [Fund. Math., 3 (1922), 133-18…
math.GN
Antonio Ferrer-Sánchez et al.
Quantum Fisher Information (QFI) sets the ultimate precision limit for parameter estimation and is therefore a central quantity in quantum metrology. In time-dependent many-body systems, however, maximizing QFI is a high…
quant-phphysics.comp-ph
Jonas Sievers, Mardavij Roozbehani
Baseline estimation is critical to Demand Response (DR) settlement in electricity markets, yet existing machine learning methods remain limited in predictive performance, while methodologies from causal inference and cou…
cs.AI

📄 ArXiv Papers

🧠 Zero Knowledge 3

Joonhyuk Lee et al.
Verification of model outputs is rapidly emerging as a key primitive for both training and real-world deployment of large language models (LLMs). In practice, this often involves using imperfect LLM judges and reward mod…
stat.MLcs.CLcs.LG
Mingsheng Tian et al.
Preparing correlated quantum states is essential for emerging technologies, but remains challenging in many-body systems. Here we propose a dissipative protocol that engineers nonreciprocal, energy-selective transitions …
quant-phcond-mat.quant-gasphysics.atom-ph
Tanjim Rahaman Fardin et al.
The rapid progress of subject-driven text-to-image synthesis, and in particular DreamBooth, has enabled a consent-free deepfake pipeline: an adversary needs only 4-8 publicly available face images to fine-tune a personal…
cs.CV

📄 ArXiv Papers

🧠 Quantum 3

Ran Ben-Basat et al.
This note clarifies the relationship between the recent TurboQuant work and the earlier DRIVE (NeurIPS 2021) and EDEN (ICML 2022) schemes. DRIVE is a 1-bit quantizer that EDEN extended to any $b>0$ bits per coordinate; w…
cs.LG
Qihang Fan et al.
In recent years, the Vision Transformer (ViT) has garnered significant attention within the computer vision community. However, the core component of ViT, Self-Attention, lacks explicit spatial priors and suffers from qu…
cs.CV
Feras Al Taha, Eilyan Bitar
We propose a distributionally robust approach to risk-sensitive estimation of an unknown signal x from an observed signal y. The unknown signal and observation are modeled as random vectors whose joint probability distri…
cs.LGeess.SPmath.OC

📄 ArXiv Papers

🧠 AI Safety 1

Savya Khosla et al.
Despite recent progress, vision-language encoders struggle with two core limitations: (1) weak alignment between language and dense vision features, which hurts tasks like open-vocabulary semantic segmentation; and (2) h…
cs.CV

🌐 Web Findings

Lobste.rs 3

Lobste.rs
_ _ _ _ _ _ ___|_|_| |_| | |_ ___ ___| |_| |_ |_ -| | . | . | | .'| _| _| | |___|_|___|___|_|_|__,|_| |_| |_|_| index about cve Command Execution via Drag-and-Drop in Terminal Emulators Published on 20.04.2026 Many peopl…
Lobste.rs
Lobste.rs
Leonardo de Moura — Creator of Lean and Z3
Lobste.rs
Lobste.rs
Edit April 2, 2026: I've been getting inbound interest from researchers wanting to run their own queries. The MCP integration I use for my own research lets you analyze live mobile telemetry continuously collected from r…
Lobste.rs

🌐 Web Findings

HuggingFace 32

HuggingFace
Understanding and anticipating vulnerability-related activity is a major challenge in cyber threat intelligence. This work investigates whether vulnerability sightings, such as proof-of-concept releases, detection templa…
HuggingFace
HuggingFace
Few-step generation has been a long-standing goal, with recent one-step generation methods exemplified by MeanFlow achieving remarkable results. Existing research on MeanFlow primarily focuses on class-to-image generatio…
HuggingFace
HuggingFace
Users often omit essential details in their requests to LLM-based agents, resulting in under-specified inputs for tool use. This poses a fundamental challenge for tool-augmented agents, as API execution typically require…
HuggingFace
HuggingFace
Large language models are typically post-trained using supervised fine-tuning (SFT) and reinforcement learning (RL), yet effectively unifying efficient knowledge injection with robust generalization remains challenging. …
HuggingFace
HuggingFace
Native Omni-modal Large Language Models (OLLMs) have shifted from pipeline architectures to unified representation spaces. However, this native integration gives rise to a critical yet underexplored phenomenon: modality …
HuggingFace
HuggingFace
As the capability frontier of autonomous agents continues to expand, they are increasingly able to complete specialized tasks through plug-and-play external skills. Yet current benchmarks mostly test whether models can u…
HuggingFace
HuggingFace
Game development sits at the intersection of creative design and intricate software engineering, demanding the joint orchestration of game engines, real-time loops, and tightly coupled state across many files. While Larg…
HuggingFace
HuggingFace
We introduce JuRe (Just Repair), a minimal denoising network for time series anomaly detection that exposes a central finding: architectural complexity is unnecessary when the training objective correctly implements the …
HuggingFace
HuggingFace
Scene graph representations enable structured visual understanding by modeling objects and their relationships, and have been widely used for multiview and 3D scene reasoning. Existing methods such as MSG learn scene gra…
HuggingFace
HuggingFace
Reinforcement learning from verifiable rewards (RLVR) has demonstrated remarkable effectiveness in improving the reasoning capabilities of large language models. As models evolve into natively multimodal architectures, e…
HuggingFace
HuggingFace
Current multimodal large language models (MLLMs) have demonstrated remarkable capabilities in short-form video understanding, yet translating long-form cinematic videos into detailed, temporally grounded scripts remains …
HuggingFace
HuggingFace
Large language models have achieved significant reasoning improvements through reinforcement learning with verifiable rewards (RLVR). Yet as model capabilities grow, constructing high-quality reward signals becomes incre…
HuggingFace
HuggingFace
The convergence of large language models and agents is catalyzing a new era of scientific discovery: Agentic Science. While the scientific method is inherently iterative, existing agent frameworks are predominantly stati…
HuggingFace
HuggingFace
Most agents today ``self-evolve'' by following rewards and rules defined by humans. However, this process remains fundamentally dependent on external supervision; without human guidance, the evolution stops. In this work…
HuggingFace
HuggingFace
Large language models are rapidly evolving into interactive coding agents capable of end-to-end web coding, yet existing benchmarks evaluate only narrow slices of this capability, typically text-conditioned generation wi…
HuggingFace
HuggingFace
Constructing environments for training and evaluating claw-like agents remains a manual, human-intensive process that does not scale. We argue that what is needed is not just a dataset, but an automated pipeline capable …
HuggingFace
HuggingFace
Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and task diversity. We introduce MathNet, a h…
HuggingFace
HuggingFace
Large language models (LLMs) are widely explored for reasoning-intensive research tasks, yet resources for testing whether they can infer scientific conclusions from structured biomedical evidence remain limited. We intr…
HuggingFace
HuggingFace
Multimodal large language models (MLLMs) have shown impressive capabilities, yet they often struggle to effectively capture the fine-grained textual information within images crucial for accurate image translation. This …
HuggingFace
HuggingFace
Large language models are increasingly expected to serve as general-purpose agents that interact with external, stateful tool environments. The Model Context Protocol (MCP) and broader agent skills offer a unified interf…
HuggingFace
HuggingFace
Emotional Support Conversation (ESC) aims to assist individuals experiencing distress by generating empathetic and supportive dialogue. While prior work typically assumes that each supporter turn corresponds to a single …
HuggingFace
HuggingFace
Video world models have achieved remarkable success in simulating environmental dynamics in response to actions by users or agents. They are modeled as action-conditioned video generation models that take historical fram…
HuggingFace
HuggingFace
Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but over-edited solutions during debugging. To evaluate how far LLMs are fr…
HuggingFace
HuggingFace
Chain-of-Thought (CoT) reasoning has become a powerful driver of trajectory prediction in VLA-based autonomous driving, yet its autoregressive nature imposes a latency cost that is prohibitive for real-time deployment. L…
HuggingFace
HuggingFace
Multimodal LLMs can accurately perceive numerical content across modalities yet fail to perform exact multi-digit multiplication when the identical underlying arithmetic problem is presented as numerals, number words, im…
HuggingFace
HuggingFace
We present SemanticQA, an evaluation suite designed to assess language models (LMs) in semantic phrase processing tasks. The benchmark consolidates existing multiword expression (MwE) resources and reorganizes them into …
HuggingFace
HuggingFace
Merging separately trained LoRA adapters is a practical alternative to joint multi-task training, but it often hurts performance. Existing methods usually treat the LoRA update ΔW = BA as a single object and do not disti…
HuggingFace
HuggingFace
Games offer a compelling paradigm for developing general reasoning capabilities in language models, as they naturally demand strategic planning, probabilistic inference, and adaptive decision-making. However, existing se…
HuggingFace
HuggingFace
Reliable deployment of language models requires two capabilities that appear distinct but share a common geometric foundation: predicting whether a model will accept targeted behavioral control, and detecting when its in…
HuggingFace
HuggingFace
Genome engineering has achieved remarkable sequence-level precision, yet predicting the transcriptomic state that a cell will occupy after perturbation remains an open problem. Single-cell CRISPR screens measure how far …
HuggingFace
HuggingFace
We present Mind DeepResearch (MindDR), an efficient multi-agent deep research framework that achieves leading performance with only ~30B-parameter models through a meticulously designed data synthesis and multi-stage tra…
HuggingFace
HuggingFace
Training strong video generation models usually requires massive datasets, large parameter counts, and substantial compute. In this work, we ask whether strong text-to-video quality is possible at a much smaller budget: …
HuggingFace

🌐 Web Findings

Google Scholar 12

Google Scholar
This paper conducts systematic analysis of advancements in the field of artificial intelligence (AI) from the year 2010 onwards, by performing original quantitative analysis of Epoch AI notable AI models dataset and comb…
Google Scholar
Google Scholar
This article reviews approaches based on artificial intelligence (AI), which contributes to the security of cyber environments. We examine existing techniques using several indicators: explainability, performance and rob…
Google Scholar
Google Scholar
In the era of big data, novel cyber attack methods are emerging endlessly. Traditional cybersecurity defense measures have become insufficient to counter these new threats, while the rapid development of artificial intel…
Google Scholar
Google Scholar
Cluster Computing - With the emergence of quantum computers, traditional cryptographic methods are vulnerable to attacks, emphasizing the need for post-quantum cryptography to secure devices and...
Google Scholar
Google Scholar
As Artificial Intelligence (AI) continues to advance rapidly, Friendly Artificial Intelligence (FAI) has been proposed to advocate for more equitable and fair development of AI. Despite its importance, there is a lack of…
Google Scholar
Google Scholar
Text-to-image diffusion models have revolutionized visual content generation, yet their deployment is hindered by a fundamental limitation: safety mechanisms enforce rigid, uniform standards that fail to reflect diverse …
Google Scholar
Google Scholar
Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Alex Qiu, Jiayi Zhou, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang. Proceedings of the 63rd Annual Meeting of the Association…
Google Scholar

🔗 All Sources

  1. 1. MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
  2. 2. Sessa: Selective State Space Attention
  3. 3. Bounded Ratio Reinforcement Learning
  4. 4. Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs
  5. 5. When Can LLMs Learn to Reason with Weak Supervision?
  6. 6. Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
  7. 7. A multimodal and temporal foundation model for virtual patient representations at healthcare system scale
  8. 8. Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering
  9. 9. Benchmarking System Dynamics AI Assistants: Cloud Versus Local LLMs on CLD Extraction and Discussion
  10. 10. ClawEnvKit: Automatic Environment Generation for Claw-Like Agents
  11. 11. Revisiting Active Sequential Prediction-Powered Mean Estimation
  12. 12. ConforNets: Latents-Based Conformational Control in OpenFold3
  13. 13. MUA: Mobile Ultra-detailed Animatable Avatars
  14. 14. The T16 Planet Hunt: 10,000 New Planet Candidates from TESS Cycle 1 and the Confirmation of a Hot Jupiter Around TIC 183374187
  15. 15. ReCap: Lightweight Referential Grounding for Coherent Story Visualization
  16. 16. Do Privacy Policies Match with the Logs? An Empirical Study of Privacy Disclosure in Android Application Logs
  17. 17. Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
  18. 18. Capturing Monetarily Exploitable Vulnerability in Smart Contracts via Auditor Knowledge-Learning Fuzzing
  19. 19. Tight Auditing of Differential Privacy in MST and AIM
  20. 20. Compositional security definitions for higher-order where declassification
  21. 21. Subcodes of Lambda-Gabidulin Codes for Compact-Ciphertext Cryptography
  22. 22. Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection
  23. 23. AgenTEE: Confidential LLM Agent Execution on Edge Devices
  24. 24. Smokescreen: A Python package for data vector blinding and encryption in cosmological analyses
  25. 25. Morita equivalence for quantum graphs
  26. 26. The dimensions of Schur squares of HRS codes
  27. 27. Privacy-Preserving Product-Quantized Approximate Nearest Neighbor Search Framework for Large-scale Datasets via A Hybrid of Fully Homomorphic Encryption and Trusted Execution Environment
  28. 28. FUSE: Ensembling Verifiers with Zero Labeled Data
  29. 29. Dissipative Preparation of Correlated Quantum States in Dipolar Rydberg Arrays
  30. 30. MetaCloak-JPEG: JPEG-Robust Adversarial Perturbation for Preventing Unauthorized DreamBooth-Based Deepfake Generation
  31. 31. A Note on TurboQuant and the Earlier DRIVE/EDEN Line of Work
  32. 32. Advancing Vision Transformer with Enhanced Spatial Priors
  33. 33. Wasserstein Distributionally Robust Risk-Sensitive Estimation via Conditional Value-at-Risk
  34. 34. AtomTwin.jl: a physics-native digital twin framework for neutral-atom quantum processors
  35. 35. JAI functional contractions in relational metric spaces
  36. 36. Physics-Informed Neural Networks for Maximizing Quantum Fisher Information in Time-Dependent Many-Body Systems
  37. 37. A Generalized Synthetic Control Method for Baseline Estimation in Demand Response Services
  38. 38. T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
  39. 39. Command Execution via Drag-and-Drop in Terminal Emulators
  40. 40. Signal Shot: a project to verify the Signal protocol and its Rust implementation using Lean
  41. 41. ChatGPT Won't Let You Type Until Cloudflare Reads Your React State. I Decrypted the Program That Does It
  42. 42. Modeling Sparse and Bursty Vulnerability Sightings: Forecasting Under Data Constraints
  1. 43. Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation
  2. 44. Latent Preference Modeling for Cross-Session Personalized Tool Calling
  3. 45. GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification
  4. 46. Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models
  5. 47. SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents
  6. 48. OpenGame: Open Agentic Coding for Games
  7. 49. Back to Repair: A Minimal Denoising Network\ for Time Series Anomaly Detection
  8. 50. HSG: Hyperbolic Scene Graph
  9. 51. EasyVideoR1: Easier RL for Video Understanding
  10. 52. OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
  11. 53. When Can LLMs Learn to Reason with Weak Supervision?
  12. 54. EvoMaster: A Foundational Agent Framework for Building Evolving Autonomous Scientific Agents at Scale
  13. 55. Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration
  14. 56. WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
  15. 57. ClawEnvKit: Automatic Environment Generation for Claw-Like Agents
  16. 58. MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
  17. 59. MedConclusion: A Benchmark for Biomedical Conclusion Generation from Structured Abstracts
  18. 60. MNAFT: modality neuron-aware fine-tuning of multimodal large language models for image translation
  19. 61. Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
  20. 62. Modeling Multiple Support Strategies within a Single Turn for Emotional Support Conversations
  21. 63. MultiWorld: Scalable Multi-Agent Multi-View Video World Models
  22. 64. Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
  23. 65. OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
  24. 66. Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs
  25. 67. Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models
  26. 68. Crowded in B-Space: Calibrating Shared Directions for LoRA Merging
  27. 69. Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play
  28. 70. The Geometric Canary: Predicting Steerability and Detecting Drift via Representational Stability
  29. 71. Geometric coherence of single-cell CRISPR perturbations reveals regulatory architecture and predicts cellular stress
  30. 72. Mind DeepResearch Technical Report
  31. 73. Motif-Video 2B: Technical Report
  32. 74. Early identification of breakthrough technologies: Insights from science-driven innovations
  33. 75. Catalyst breakthroughs in methane dry reforming: Employing machine learning for future advancements
  34. 76. Advancements in Artificial Intelligence: Breakthroughs, Challenges and the Road Ahead
  35. 77. Large language models (LLM) in computational social science: prospects, current state, and challenges
  36. 78. Review of eXplainable artificial intelligence for cybersecurity systems
  37. 79. Artificial intelligence-driven cybersecurity applications and challenges
  38. 80. Securing the future: exploring post-quantum cryptography for authentication and user privacy in IoT devices
  39. 81. Quantum machine learning: A comprehensive review of integrating AI with quantum computing for computational advancements
  40. 82. Leveraging blockchain and smart contracts to combat greenwashing in sustainable development
  41. 83. Towards friendly ai: A comprehensive review and new perspectives on human-ai alignment
  42. 84. Personalized Safety Alignment for Text-to-Image Diffusion Models
  43. 85. Pku-saferlhf: Towards multi-level safety alignment for llms with human preference