HuggingFace Papers · 2026-07-27 07:00
In multilingual retrieval augmented generation, a retriever can retrieve relevant documents written in multiple languages, which are subsequently reranked befor
Reasoning / Mirror / Other · 2026-07-25 07:00
While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logic
Unda / Unpaired / Domain · 2026-07-25 07:00
Multimodal based approaches often outperform single modality approaches in downstream tasks as the different modalities provide complementary information, yet a
Domain / Adaptation / Online · 2026-07-23 07:00
This paper studies the problem of stochastic variance reduction (SVR) for the maximum mean discrepancy (MMD) and correlation alignment (CORAL) loss functions. A
Domain / Adaptation / Online · 2026-07-23 07:00
Correlation alignment and the maximum mean discrepancy are two widely used distribution-matching frameworks for unsupervised domain adaptation (UDA). However, h
Lkvalues / Aligning / Sri · 2026-07-23 07:00
Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in
HuggingFace Papers · 2026-07-23 07:00
Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action
HuggingFace Papers · 2026-07-23 07:00
Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) off
Google Scholar · 2026-07-23 07:00
AI alignment is a human problem
Methodology / Auditable / Trustworthiness · 2026-07-22 19:00
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has be
Quantum / Robustness / Variational · 2026-07-22 19:00
The Variational Quantum Eigensolver (VQE) is a leading algorithm for estimating molecular ground-state energies on near-term quantum hardware, with applications
Methodology / Auditable / Trustworthiness · 2026-07-22 07:00
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has be
Quantum / Robustness / Variational · 2026-07-22 07:00
The Variational Quantum Eigensolver (VQE) is a leading algorithm for estimating molecular ground-state energies on near-term quantum hardware, with applications
Methodology / Auditable / Trustworthiness · 2026-07-21 19:00
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has be
HuggingFace Papers · 2026-07-21 19:00
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforc
HuggingFace Papers · 2026-07-21 19:00
The prevailing inference framework for diffusion models formulates generation fundamentally as a problem of numerical integration. This perspective casts the mo
Methodology / Auditable / Trustworthiness · 2026-07-21 07:00
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has be
HuggingFace Papers · 2026-07-21 07:00
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforc
HuggingFace Papers · 2026-07-21 07:00
The prevailing inference framework for diffusion models formulates generation fundamentally as a problem of numerical integration. This perspective casts the mo
Methodology / Auditable / Trustworthiness · 2026-07-20 19:00
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has be
Lobste.rs · 2026-07-20 19:00
Platform overview AI Discovery & Posture Red Teaming & Attack Surface Exposure Runtime Guardrails Governance & Compliance Latest Announcement Pillar Security Na
Google Scholar · 2026-07-20 19:00
Against the Manhattan project framing of AI alignment
Google Scholar · 2026-07-20 19:00
Developing AI systems with a human-like understanding of everyday concepts is a key step towards developing safe, reliable systems whose behavior makes sense to
Hacker News · 2026-07-20 19:00
Multipath derivation of a formal algebraic framework — the Algebra of Four-Fold Distinction — which is then applied across physics, cognition, and AI alignment.
Methodology / Auditable / Trustworthiness · 2026-07-20 07:00
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has be
Google Scholar · 2026-07-20 07:00
Against the Manhattan project framing of AI alignment
Google Scholar · 2026-07-20 07:00
Developing AI systems with a human-like understanding of everyday concepts is a key step towards developing safe, reliable systems whose behavior makes sense to
Hacker News · 2026-07-20 07:00
Multipath derivation of a formal algebraic framework — the Algebra of Four-Fold Distinction — which is then applied across physics, cognition, and AI alignment.
Hacker News · 2026-07-19 19:00
Multipath derivation of a formal algebraic framework — the Algebra of Four-Fold Distinction — which is then applied across physics, cognition, and AI alignment.
Reddit · 2026-07-18 19:00
How to submit to AI alignment track? I can only see these at openReview: AAAI 2027 AAAI 2027 Artificial Intelligence for Social Impact Track AAAI 2027 Conferenc
Security & Cybersecurity · 2026-07-18 07:00
Fine-tuning large language models (LLMs) on domain-specific datasets has become a standard paradigm for adapting LLMs to specialized applications. However, rece
HuggingFace Papers · 2026-07-17 19:00
CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, enabling applications in rob
Google Scholar · 2026-07-16 19:00
Xinyue Lou, You Li, Jinan Xu, Xiangyu Shi, Chi Chen, Kaiyu Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Google Scholar · 2026-07-16 19:00
The field of AI safety considers whether and how AI development can be safe and beneficial for humans and other animals, and the field of AI welfare considers w
Google Scholar · 2026-07-16 19:00
As frontier AI systems advance toward transformative capabilities, we need a parallel transformation in how we measure and evaluate these systems to ensure safe
HuggingFace Papers · 2026-07-16 07:00
Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployme
Silent / Alarm / J-space · 2026-07-15 07:00
Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach. Such evaluations, however, are sensitive to
Agent / When / Local · 2026-07-14 07:00
Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands, and workspace state, making safety failures directly action
HuggingFace Papers · 2026-07-09 19:00
Touch supplies the physical grounding needed to perceive intrinsic material properties, such as friction and compliance, that vision alone often cannot resolve.
Artificial Intelligence · 2026-07-09 07:00
We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fi
Security & Cybersecurity · 2026-07-09 07:00
Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that this binary
Crypto & Blockchain · 2026-07-09 07:00
This study adopts a behavioural bottom-up approach to AI value alignment to investigate whether an implicitly conveyed user identity shifts the moral evaluation
Google Scholar · 2026-07-09 07:00
Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, Peter Henderson
Google Scholar · 2026-07-09 07:00
As large language models (LLMs) become more powerful and pervasive across society, ensuring these systems are beneficial, safe, and aligned with human values is
Reddit · 2026-07-08 19:00
Most safety alignment work treats "detect the attack" as a text classification problem — does the prompt contain language the model's safety guardrails should c
Artificial Intelligence · 2026-07-08 07:00
Vision-language models (VLMs) struggle to generalize in interactive physical reasoning, particularly under unseen tasks and environments. Two key failure modes
HuggingFace Papers · 2026-07-08 07:00
Significant disparities exist in the diagnosis and clinical presentation of depression across different linguistic populations. Speech-based depression detectio
Reddit · 2026-07-07 19:00
late night thoughts as I was working on my paper that is about specific behavior that arises from RHLF, it got me thinking what if train a model in an environme
HuggingFace Papers · 2026-07-06 07:00
The fast growth of open-source AI infrastructure, from model serving engines and agent platforms to the Model Context Protocol (MCP) ecosystem and the language
HuggingFace Papers · 2026-07-06 07:00
Cloud removal (CR) is essential for optical remote sensing, serving as a prerequisite for reliable downstream interpretation, such as semantic segmentation and
Google Scholar · 2026-07-06 07:00
Star-1: Safer alignment of reasoning llms with 1k data
Google Scholar · 2026-07-06 07:00
Taming artificial intelligence: A theory of control-accountability alignment among AI developers and users
Google Scholar · 2026-07-06 07:00
Introduction to AI safety, ethics, and society
Artificial Intelligence · 2026-07-03 19:00
Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no
Large Language Models · 2026-07-03 19:00
Generative diffusion models excel at synthesizing high-quality images, videos, and 3D content under multimodal control. However, arbitrary user-defined modality
AI Safety & Guardrails · 2026-07-03 19:00
Representation alignment has become an effective way to accelerate diffusion transformer training and improve generation quality. Recent self-alignment methods,
HuggingFace Papers · 2026-07-03 19:00
Representation alignment has become an effective way to accelerate diffusion transformer training and improve generation quality. Recent self-alignment methods,
Google Scholar · 2026-07-03 19:00
The proliferation of artificial intelligence (AI) systems as super-capable assistants in everyday life has revolutionized productivity and decision-making acros
Google Scholar · 2026-07-03 19:00
The rapid adoption of artificial intelligence (AI) systems in public sector organizations raises a question that traditional governance frameworks were not desi
Google Scholar · 2026-07-03 19:00
Borong Zhang, Yuhao Zhang, Jiaming Ji, Yingshan Lei, Juntao Dai, Yuanpei Chen, Yaodong Yang
Artificial Intelligence · 2026-07-03 07:00
Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no
Large Language Models · 2026-07-03 07:00
Generative diffusion models excel at synthesizing high-quality images, videos, and 3D content under multimodal control. However, arbitrary user-defined modality
AI Safety & Guardrails · 2026-07-03 07:00
Representation alignment has become an effective way to accelerate diffusion transformer training and improve generation quality. Recent self-alignment methods,
HuggingFace Papers · 2026-07-03 07:00
Representation alignment has become an effective way to accelerate diffusion transformer training and improve generation quality. Recent self-alignment methods,
Google Scholar · 2026-07-03 07:00
The proliferation of artificial intelligence (AI) systems as super-capable assistants in everyday life has revolutionized productivity and decision-making acros
Google Scholar · 2026-07-03 07:00
The rapid adoption of artificial intelligence (AI) systems in public sector organizations raises a question that traditional governance frameworks were not desi
Google Scholar · 2026-07-03 07:00
Borong Zhang, Yuhao Zhang, Jiaming Ji, Yingshan Lei, Juntao Dai, Yuanpei Chen, Yaodong Yang
HuggingFace Papers · 2026-07-02 19:00
Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrast
HuggingFace Papers · 2026-06-30 19:00
Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations
HuggingFace Papers · 2026-06-30 19:00
Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment updates, unlearned capabili
HuggingFace Papers · 2026-06-30 07:00
In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to application-specific safety policies, rather t
HuggingFace Papers · 2026-06-29 19:00
Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale.
Lobste.rs · 2026-06-29 19:00
[NIST article](https://www.nist.gov/news-events/news/2026/06/nist-mathematical-proof-supports-transition-continuous-monitor-and-update) covers this paper well
Machine Learning · 2026-06-29 07:00
Preference-based alignment often struggles to capture the reasoning that underlies human judgments. Many evaluations rely on multiple interacting criteria, yet
Security & Cybersecurity · 2026-06-29 07:00
Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood. We provide evidence that attacks do not comprehensively eliminate
HuggingFace Papers · 2026-06-29 07:00
Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades perceptual quality in ways that are not cap