DiagramBank provides 89,422 schematic diagrams curated from top-tier scientific publications, paired with captions, abstracts, and figure-reference pairs for multimodal retrieval and exemplar-driven scientific figure generation — bridging the 'AI scientist'…
HuggingFace Papers
EmbodiedMidtrain addresses the distribution gap between generic Vision-Language Models and robot action domains by mid-training VLMs on VLA-aligned data selected via a lightweight proximity estimator, achieving competitive results with expert…
HuggingFace Papers
Memanto introduces a universal typed memory schema (13 categories), automated conflict resolution, and temporal versioning backed by a no-indexing semantic database achieving sub-90ms deterministic retrieval — reaching 89.8% on LongMemEval…
HuggingFace Papers
This work introduces a Semantic Progress Function — a 1D representation of how meaning evolves across video frames — enabling analysis, correction, and steering of semantic pacing in both generated…
HuggingFace Papers
This work introduces CHAI, a critique-based human-AI oversight framework for precise video captioning — trained experts revise model-generated pre-captions using structured visual primitives from professional filmmakers, and the resulting preference…
HuggingFace Papers
EditCrafter enables high-resolution (arbitrary aspect ratio) image editing without fine-tuning by combining tiled inversion that preserves identity with a noise-damped manifold-constrained classifier-free guidance (NDCFG++), extending pretrained T2I diffusion models beyond…
HuggingFace Papers
Abstain-R1 uses a clarification-aware RLVR reward to train a 3B model that abstains from unanswerable queries AND explains what information is missing, achieving unanswerable-query behavior competitive with DeepSeek-R1 — showing…
HuggingFace Papers