Generation / Minevalicoder / Reliable · 2026-07-27 07:00
Rust's ownership model and type system offer strong memory safety guarantees, but unsafe code and runtime panics still present significant risks. Formal verific
HuggingFace Papers · 2026-07-27 07:00
Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances.
Reddit · 2026-07-24 19:00
I've been chasing the question of what algorithms a transformer can actually express -- separate from what it can learn. So I built a compiler: define a computa
Toolsciver / Multimodal / Scientific · 2026-07-22 19:00
Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tab
Toolsciver / Multimodal / Scientific · 2026-07-22 07:00
Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tab
Toolsciver / Multimodal / Scientific · 2026-07-21 19:00
Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tab
HuggingFace Papers · 2026-07-21 19:00
We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning. F
HuggingFace Papers · 2026-07-21 19:00
Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned only if k verifier calls all accept it. Under condit
Toolsciver / Multimodal / Scientific · 2026-07-21 07:00
Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tab
HuggingFace Papers · 2026-07-21 07:00
We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning. F
HuggingFace Papers · 2026-07-21 07:00
Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned only if k verifier calls all accept it. Under condit
Toolsciver / Multimodal / Scientific · 2026-07-20 19:00
Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tab
HuggingFace Papers · 2026-07-20 19:00
We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning. F
Lobste.rs · 2026-07-20 19:00
LLMs have grown alarmingly capable at finding bugs in production software. This accentuates an already severe risk: much of our critical infrastructure is media
Lobste.rs · 2026-07-20 19:00
Triton language and compiler for SAIL. Contribute to t-head/triton-for-sail development by creating an account on GitHub.
Google Scholar · 2026-07-20 19:00
Quantum software engineering: algorithm design, error mitigation, and compiler optimization for fault-tolerant quantum computing
Hacker News · 2026-07-20 19:00
The Continuous Verification and Optimization layer for self-hosted LLMs - colomalabs/coloma
Hacker News · 2026-07-20 19:00
Show HN: Open-source verification and tuning layer for self-hosted LLMs
Toolsciver / Multimodal / Scientific · 2026-07-20 07:00
Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tab
Lobste.rs · 2026-07-20 07:00
Triton language and compiler for SAIL. Contribute to t-head/triton-for-sail development by creating an account on GitHub.
Hacker News · 2026-07-20 07:00
The Continuous Verification and Optimization layer for self-hosted LLMs - colomalabs/coloma
Hacker News · 2026-07-20 07:00
Show HN: Open-source verification and tuning layer for self-hosted LLMs
Lobste.rs · 2026-07-19 19:00
Triton language and compiler for SAIL. Contribute to t-head/triton-for-sail development by creating an account on GitHub.
Hacker News · 2026-07-19 19:00
The Continuous Verification and Optimization layer for self-hosted LLMs - colomalabs/coloma
Hacker News · 2026-07-19 19:00
Show HN: Open-source verification and tuning layer for self-hosted LLMs
Lobste.rs · 2026-07-19 07:00
Triton language and compiler for SAIL. Contribute to t-head/triton-for-sail development by creating an account on GitHub.
Hacker News · 2026-07-19 07:00
The Continuous Verification and Optimization layer for self-hosted LLMs - colomalabs/coloma
Hacker News · 2026-07-19 07:00
Show HN: Open-source verification and tuning layer for self-hosted LLMs
HuggingFace Papers · 2026-07-16 19:00
Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their strictness makes generation more difficult. Off
Google Scholar · 2026-07-16 19:00
The Solana blockchain was created by Anatoly Yakovenko of Solana Labs and was introduced in 2017, employing a novel transaction verification method. However, at
Visualrepair / Dynamic / Tool · 2026-07-16 07:00
Automated Program Repair (APR) has witnessed significant progress with the advent of Large Language Models (LLMs). However, as modern software systems increasin
HuggingFace Papers · 2026-07-16 07:00
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics
HuggingFace Papers · 2026-07-15 19:00
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics
HuggingFace Papers · 2026-07-15 07:00
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics
HuggingFace Papers · 2026-07-14 19:00
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics
HuggingFace Papers · 2026-07-14 07:00
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics
Reddit · 2026-07-13 19:00
TLDR: I trained Qwen3.5-4B and gemma-4-12b on self-distilled, compressed reasoning traces; compression was section-aware (compute and verification spans remain,
Security & Cybersecurity · 2026-07-10 07:00
QR codes are a ubiquitous part of daily life, widely trusted by millions. However, their lack of inherent security features has given rise to critical attack ve
Security & Cybersecurity · 2026-07-10 07:00
The NIS-2 Directive increases the need for continuous, auditable compliance evidence and motivates a shift from document-based compliance toward machine-readabl
Reddit · 2026-07-09 19:00
I want to share something I've been building as an independent researcher from Indonesia. TL;DR: Face verification model that replaces cosine similarity with sl
HuggingFace Papers · 2026-07-08 07:00
Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification. While recent parallel drafters ef
Artificial Intelligence · 2026-07-07 07:00
Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of LLMs. In this work, we identify v
Security & Cybersecurity · 2026-07-07 07:00
Neural network verification and data privacy are inherently in tension: verification demands full access to model parameters and input data, yet both are increa
HuggingFace Papers · 2026-07-07 07:00
Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of LLMs. In this work, we identify v
HuggingFace Papers · 2026-07-07 07:00
LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing target
HuggingFace Papers · 2026-07-07 07:00
Repository-level vulnerability reproduction is a demanding software engineering (SE) task: an agent must inspect a codebase, infer the input grammar that reache
Google Scholar · 2026-07-06 07:00
Quantum Software Engineering: Developing Algorithms for IBM Quantum Systems
HuggingFace Papers · 2026-07-01 19:00
We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks. Existing frontier SWE
Artificial Intelligence · 2026-07-01 07:00
We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We present AxDafny, a veri
Security & Cybersecurity · 2026-07-01 07:00
As ML-KEM is adopted as a post-quantum cryptographic standard, resilience against physical side-channel attacks has become essential. Among the constituent step
Bug Bounty & Vulns · 2026-07-01 07:00
The tendency of large generative models to memorize training data makes sample verification critical for privacy auditing and copyright enforcement. Current mem
Crypto & Blockchain · 2026-07-01 07:00
Standard quantum verification and certification protocols often assume that experimental sources emit independent and identically distributed (i.i.d.) states. I
HuggingFace Papers · 2026-06-30 19:00
Multi-agent large language model (LLM) systems often rely on verifier and critic agents to suppress hallucinations, but verification is delayed. During this del
Machine Learning · 2026-06-30 07:00
We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks. Existing frontier SWE
HuggingFace Papers · 2026-06-29 19:00
LLM-based agents for program repair are increasingly built on a "generate-run-revise" paradigm, iteratively executing tests to evaluate and refine patches. This
Lobste.rs · 2026-06-29 19:00
Auto What's wrong with EU age verification? Mon, 29 June 2026 20:01 • Clear 27°C • Kifissia, GR • #en #privacy #eu Nothing really, to be honest.