Yan Cui, Jacob S. Leiby, Wenhui Lei, Dokyoon Kim, Yanxiang Deng
Here we present Haiku, a tri-modal contrastive learning model trained on multiplexed immunofluorescence (mIF). It comprises 26.7 million spatial proteomics patches from 3,218 tissue sections across 1,606 patients spanning 11 organ types, with matched hematoxylin and eosin (H&E) histology and…
HuggingFace
Mohamed Elfeki, Tu Trinh, Kelvin Luu, Guangze Luo, Nathan Hunt
Frontier coding agents solve complex tasks when given complete context but collapse when specifications are incomplete. The bottleneck is not raw capability but judgment: knowing when to act autonomously and when to ask for help. Current benchmarks are blind to…
HuggingFace
Siqi Zhu
This position paper argues that agentic AI systems should be designed and evaluated as marginal token allocation economies rather than as text generators priced by the unit. Following a single request -- a developer asking a coding agent to fix…
HuggingFace
M. Riera-Marin, O. K. Sikha, J. Rodriguez-Comas, M. S. May, T. Kirscher
Surgical resection is the only potentially curative treatment for pancreatic ductal adenocarcinoma (PDAC), and eligibility depends on accurate assessment of vascular invasion (VI). We introduce CURVAS-PDACVI, an open benchmark for uncertainty-aware AI in PDAC staging based on a densely annotated…
HuggingFace
Stefanos Pasios
Video game engines generate large volumes of visual synthetic datasets for training computer vision algorithms, but a notable sim2real appearance gap limits real-world applicability. This letter investigates using FLUX.2-4B Klein (a diffusion model) and REGEN (an image-to-image translation model) to…
HuggingFace
Ruize He, Dongchen Han, Gao Huang
We demonstrate that attention can be mathematically reframed as a Multi-Layer Perceptron (MLP) equipped with dynamically predicted parameters. Through this lens, attention's global modeling power is explained not as explicit token-wise aggregation but as an implicit process where dynamically generated…
HuggingFace
Tianxiang Dai, Jonathan Fan
We introduce Stable Counting Capacity, an assay in which models count repeated symbols until failure. The assay removes knowledge dependencies, semantics and ambiguity from evaluation, avoids lexical and tokenization confounds, and provides a direct measure of procedural reliability. Across more…
HuggingFace
Gal Yona, Mor Geva, Yossi Matias
Most factuality gains in LLM research have come from expanding the model's knowledge boundary rather than improving awareness of that boundary. We conjecture that distinguishing known from unknown is inherently difficult: models may lack discriminative power to perfectly separate truths…
HuggingFace
Massimo Rondelli, Francesco Pivi, Maurizio Gabbrielli
Automatic generation of executable Blender code from natural language remains challenging, with state-of-the-art LLMs producing frequent syntactic errors and geometrically inconsistent objects. We present BlenderRAG, a retrieval-augmented generation system operating on a curated multimodal dataset of 500 expert-validated examples across…
HuggingFace
Laure Berti-Equille
Tabular Foundation Models (TFMs) achieve state-of-the-art zero-shot accuracy on small tabular datasets by meta-learning over synthetic data-generating processes. However, their in-context learning assumes approximately clean inputs: missing values, outliers, and duplicates create a prior mismatch that degrades both accuracy and…
HuggingFace
Zhen Ye, Xu Tan, Aoxiong Yin, Hongzhan Lin, Guangyan Zhang
Joint audio-video generation models yield stronger cross-modal coherence than cascaded approaches, but existing models couple modalities throughout denoising via pervasive attention, treating high-level semantics and low-level details in a fully entangled manner. We propose Talker-T2AV, an autoregressive diffusion framework where…
HuggingFace