Lobste.rs
Edit April 2, 2026: I've been getting inbound interest from researchers wanting to run their own queries. The MCP integration I use for my own research lets…
Lobste.rs
HuggingFace
Modern Text-to-SQL systems generate multiple candidate SQL queries and rank them to judge a final prediction. However, existing methods face two limitations.…
HuggingFace
HuggingFace
Scaling Diffusion Transformers (DiTs) to hundreds of layers introduces a structural vulnerability: networks can enter a silent, mean-dominated collapse state…
HuggingFace
HuggingFace
Fast Weight Programmers (FWPs) encode temporal dependencies through dynamically updated parameters rather than recurrent hidden states. Quantum FWPs (QFWPs)…
HuggingFace
HuggingFace
AI agents are increasingly deployed across diverse domains to automate complex workflows through long-horizon and high-stakes action executions. Due to their…
HuggingFace
HuggingFace
The theory of state tracking in recurrent architectures has predominantly focused on expressive capacity: whether a fixed architecture can theoretically…
HuggingFace
HuggingFace
Long-context inference in decoder-only language models is costly because long prompts are processed during Prefill, cached at every layer, and repeatedly…
HuggingFace
HuggingFace
Reinforcement learning (RL) has substantially improved the ability of large language model (LLM) agents to interact with environments and solve multi-turn…
HuggingFace
HuggingFace
Speculative decoding accelerates LLM inference by drafting a tree of candidate continuations and verifying it in one target forward. Existing drafters fall…
HuggingFace
HuggingFace
Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next step…
HuggingFace
HuggingFace
Existing multimodal search agents process target entities sequentially, issuing one tool call per entity and accumulating redundant interaction rounds whenever…
HuggingFace
HuggingFace
Training multimodal large language models has long been limited by the scarcity of high-quality paired multimodal data. Recent studies show that the shared…
HuggingFace
HuggingFace
Existing Flow Matching (FM) text-to-image models suffer from two critical bottlenecks under multi-task alignment: the reward sparsity induced by scalar-valued…
HuggingFace
HuggingFace
Tokenizers are a crucial component of latent diffusion models, as they define the latent space in which diffusion models operate. However, existing tokenizers…
HuggingFace
HuggingFace
Modern sensors generate rich, high-fidelity data, yet applications operating on wearable or remote sensing devices remain constrained by bandwidth and power…
HuggingFace
HuggingFace
Self-distillation (SD) offers a promising path for adapting large language models (LLMs) without relying on stronger external teachers. However, SD in…
HuggingFace
HuggingFace
Linear Attention (LA) offers a promising paradigm for scaling large language models (LLMs) to long sequences by avoiding the quadratic complexity of…
HuggingFace
HuggingFace
As large language models (LLMs) continue to advance rapidly, they are becoming increasingly capable while simultaneously demanding ever-longer context lengths.…
HuggingFace
HuggingFace
Large Language Model (LLM)-based agents have fundamentally reshaped artificial intelligence by integrating external tools and planning capabilities. While…
HuggingFace
HuggingFace
Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input or…
HuggingFace
HuggingFace
Deep generative models have advanced rapidly across text and vision, motivating unified multimodal systems that can understand, reason over, and generate…
HuggingFace
HuggingFace
Diffusion-based models decompose sampling into many small Gaussian denoising steps -- an assumption that breaks down when generation is compressed to a few…
HuggingFace
HuggingFace
Recent byte-level language models (LMs) match the performance of token-level models without relying on subword vocabularies, yet their utility is limited by…
HuggingFace
HuggingFace
Test-time scaling (TTS) has become an effective approach for improving large language model performance by allocating additional computation during inference.…
HuggingFace
HuggingFace
A natural intuition about the economics of AI agents is that, because agents can be replicated at very low marginal cost, agent labor may be supplied highly…
HuggingFace
HuggingFace
Synthesizing consistent and coherent long video remains a fundamental challenge. Existing methods suffer from semantic drift and narrative collapse over long…
HuggingFace
HuggingFace
Dynamic spatial reasoning from monocular video is essential for bridging visual intelligence and the physical world, yet remains challenging for…
HuggingFace
HuggingFace
Large language models (LLMs) have become a central foundation of modern artificial intelligence, yet their lifecycle remains constrained by a rigid separation…
HuggingFace
HuggingFace
Reinforcement learning with verifiable rewards (RLVR) has become a standard approach for large language models (LLMs) post-training to incentivize reasoning…
HuggingFace
HuggingFace
Accurately understanding the intent behind speech, conversation, and writing is crucial to the development of helpful Large Language Model (LLM) assistants.…
HuggingFace
HuggingFace
DeepSeek Sparse Attention (DSA) sets the state of the art for fine-grained inference-time sparse attention by introducing a learned token-wise indexer that…
HuggingFace
Google Scholar
Large language models (LLM) in computational social science: prospects, current state, and challenges
Google Scholar