Krzysztof Ociepa, Łukasz Flis, Remigiusz Kinas, Krzysztof Wróbel, Adrian Gwoździej
This report shows how a Polish-specific tokenizer and adapted pretraining pipeline improve efficiency and effective context use in the Bielik v3 series. It matters because multilingual models still under-serve…
HuggingFacedaily curated papers
Duy Le Dinh Anh, Patrick Amadeus Irawan, Tuan Van Vo
The authors show that VLMs fail simple counting tasks not just because perception is weak, but because visual evidence gets diluted as reasoning shifts into language layers. It matters because grounding failures on…
HuggingFacedaily curated papers
Gregory N. Frank
This mechanistic study localizes sparse policy-routing heads that detect disallowed content and amplify refusal behavior deeper in the network. It matters because it suggests alignment behavior can be controlled through…
HuggingFacedaily curated papers
Nick Stracke, Kolja Bauer, Stefan Andreas Baumann, Miguel Angel Bautista, Josh Susskind
This work compresses long-horizon motion into a learned latent space and uses conditional flow matching to generate realistic trajectories from text or spatial prompts. It matters because it points to a cheaper…
HuggingFacedaily curated papers
Zhixin Lin, Jungang Li, Dongliang Xu, Shidong Pan, Yibo Shi
TIPO trains mobile GUI agents to follow privacy-first or utility-first preferences by using preference intensity and trajectory-aware optimization. It matters because agent systems are moving beyond generic task…
HuggingFacedaily curated papers
Ivan Sedykh, Nikita Sorokin, Valentin Malykh
The authors show masked diffusion language models do not need the same model capacity at every denoising step, allowing smaller models to replace the full model in less sensitive regions. It matters because a simple…
HuggingFacedaily curated papers
Muhammad Kamran Janjua, Abdul Wahab, Bahador Rashidi
The paper proposes Distortion Graphs, a region-level representation for comparing paired images and reasoning about distortion type, severity, and relative quality. It matters because current multimodal models still…
HuggingFacedaily curated papers
Binbin Zheng, Xing Ma, Yiheng Liang, Jingqing Ruan, Xiaoliang Fu
SCOPE routes correct and incorrect rollouts into different supervision paths and adaptively weights distillation or MLE signals based on teacher/student confidence. It matters because reasoning post-training still…
HuggingFacedaily curated papers
João Gonçalves, Sonia de Jager, Petr Knoth, David Pride, Nick Jelicic
SHARE introduces domain-specific base models for the social sciences and humanities plus a MIRROR interface designed to aid review without auto-writing prose for the user. It matters because it offers a narrower but…
HuggingFacedaily curated papers
Yang Liu, Enxi Wang, Yufei Gao, Weixin Zhang, Bo Wang
MEDS stores historical rollout representations, clusters recurring failure modes, and penalizes policies that keep revisiting the same bad behaviors. It matters because RL-trained language models often lose diversity or…
HuggingFacedaily curated papers