hugging papers
30 articles about hugging papers in AI news
Recursive Multi-Agent Systems Top Hugging Papers; Eywa Bridges LLMs and Scientific Models
Recursive Multi-Agent Systems leads Hugging Papers with 242 upvotes. Eywa and OneManCompany signal a move from chat-based to structural agent collaboration.
Hugging Face weekly papers: Monotonic inference policy overtakes training optimization
Hugging Face's top papers July 6-12 include a paper arguing monotonic inference policies are the true LLM RL objective, and Vidu S1 for real-time interactive video generation.
Hugging Face Papers: 35B Agent Matches Trillion-Parameter Performance
Hugging Face Daily Papers featured eight AI papers, including Orca (world model), Dockerless (62% SWE-bench), and a 35B agent matching trillion-parameter performance.
Hugging Face OCRs 27,000 arXiv Papers to Markdown with Open 5B Model
Hugging Face CEO Clement Delangue announced the OCR conversion of 27,000 arXiv papers to Markdown using an open 5B-parameter model and 16 parallel jobs on L40S GPUs. This demonstrates a scalable, open-source pipeline for large-scale academic document processing.
ACE-Data-0: 150h Embodied Dataset Released on Hugging Face
ACE-Data-0, a 150-hour multimodal embodied dataset, released on Hugging Face for robotics research, featuring egocentric video, motion, and tactile data.
Meta's AskChem Turns 147K Papers Into 2.4M Cited Claims
Meta's AskChem converts 147,000 chemistry papers into 2.4M DOI-grounded claims, shifting search from documents to atomic assertions.
100+ Papers Surveyed: LLMs' Metacognition Gap
A systematic survey of 100+ papers reveals gaps in LLM metacognition, including 10-30% miscalibration in top models like GPT-4 and Claude 3.
LOCUS-v1: 2.2M US Laws Hit HuggingFace via AI Pipeline
LOCUS-v1, a dataset of 2.2M US laws built via AI pipeline, released on HuggingFace. First comprehensive legal database of its kind, but quality and validation metrics remain undisclosed.
Hugging Face Launches 'Kernels' Hub for GPU Code, Like GitHub for AI Hardware
Hugging Face has launched 'Kernels,' a new section on its Hub for sharing and discovering optimized GPU kernels. This treats performance-critical code as a first-class artifact, similar to AI models.
754B-Parameter AI Model Hits Hugging Face, Weighs 1.51TB
An unidentified 754-billion-parameter AI model has been uploaded to the Hugging Face platform, consuming 1.51TB of space. This represents one of the largest publicly accessible model repositories by size.
Cohere Transcribe: 2B-Parameter Open-Source ASR Model Achieves 5.42% WER, Topping Hugging Face Leaderboard
Cohere released Transcribe, a 2B-parameter open-source speech recognition model. It claims a 5.42% average word error rate, beating OpenAI Whisper v3 and topping the Hugging Face Open ASR Leaderboard.
Black Forest Labs Unleashes FLUX.2 klein: Sub-Second AI Image Generation Hits Hugging Face
Black Forest Labs has released FLUX.2 klein on Hugging Face, delivering state-of-the-art image generation and editing in under a second. The model runs on consumer GPUs with just 13GB VRAM, making high-speed AI art creation dramatically more accessible.
ClBench-V: New Benchmark Tests Multimodal Contextual Learning in 3 Dimensions
ClBench-V benchmark from @HuggingPapers tests multimodal contextual learning across three dimensions: grounding, application, and knowledge learning. No results disclosed yet.
Alibaba Open-Sources Qwen-AgentWorld for Generalist Agent Training
Alibaba open-sourced Qwen-AgentWorld and Wan-Streamer v0.1 on Hugging Face, targeting generalist agent training and real-time streaming. The releases include 8 additional papers on agent benchmarks and architectures.
EvoEmbedding Beats Static Embedders 3× Larger via Latent Memory Queue
EvoEmbedding uses a latent memory queue to beat static embedders 3× its size on long-context retrieval, per @HuggingPapers.
ByteDance SwanTale: Unified Speech-Audio Model Hits HF
ByteDance released SwanTale, a unified speech-audio model on Hugging Face, covering voice cloning, style control, and scene synthesis. No benchmarks or technical details disclosed.
Alibaba Releases RynnBrain 1.1 Embodied AI Models at 2B-122B Scales
Alibaba released RynnBrain 1.1 on Hugging Face with 2B, 9B, and 122B-A10B MoE models for robot manipulation, but disclosed no benchmarks.
Microsoft Fara1.5-27B Open-Source Agent Scores 72.3% on Web Tasks
Microsoft released Fara1.5-27B, a vision-only web browsing agent scoring 72.3% on Online-Mind2Web, open-source on Hugging Face.
239-Paper Survey Maps How AI Agents Self-Improve via Scaffold Updates
A survey of 239 papers shows 68% of AI agent self-improvement methods focus on scaffold updates rather than model retraining, raising evaluation quality concerns.
NVIDIA Releases FP4 Quantized Kimi-K2.7-Code with 1T Parameters
NVIDIA released FP4 quantized Kimi-K2.7-Code on Hugging Face, a 1T-parameter model for Blackwell GPUs with claimed accuracy retention.
Google Releases Magenta RealTime 2 for Open-Weight Music Generation
Google released Magenta RealTime 2 on Hugging Face, the only open-weights model for real-time continuous music generation on device with ~200ms latency.
30B-A3B Reasoning Model Hits Gold Medal on Physics, Math Olympiads
30B-A3B reasoning model from @stingning achieves gold-medal level on physics and math Olympiads, released on Hugging Face.
Massive Open-Source Dataset of Computer Screen Recordings Released to Train AI Agents
Researchers have released the world's largest open-source dataset of computer-use recordings on Hugging Face. The collection contains 48,478 screen recording videos totaling approximately 12,300 hours of professional software usage, licensed under CC-BY-4.0 for AI training and evaluation.
NVIDIA's Kimi-K2.5 Eagle Head: Supercharging Moonshot's Reasoning with Speculative Decoding
NVIDIA has released the Kimi-K2.5 Eagle head on Hugging Face, implementing Eagle-3 speculative decoding to dramatically accelerate inference for Moonshot's reasoning models. This breakthrough promises blazing-fast performance while maintaining accuracy.
Microsoft's VibeVoice-ASR Shatters Transcription Limits with 60-Minute Single-Pass Processing
Microsoft has released VibeVoice-ASR on Hugging Face, a revolutionary speech recognition model that transcribes 60-minute audio in one pass with speaker diarization, timestamps, and multilingual support across 50+ languages without configuration.
AI Research Breakthroughs: From Video Reasoning to Self-Stopping Models
This week's top AI papers reveal major advances in video understanding, reasoning efficiency, and agent training. Researchers introduced a massive video reasoning dataset, models that know when to stop thinking, and techniques for improving AI agents without full retraining.
Tencent's WorldClaw Generates Explorable 3D Worlds from Text
Tencent's WorldClaw generates explorable 3D worlds from text via agentic planning. No benchmarks or release date disclosed, but the agentic decomposition approach marks a structural shift.
OSReward: Open VLM Judges Match Commercial at 30-60x Lower Cost
OSReward is a human-gold benchmark for VLM judges scoring computer-use agents. Open OS-Shepherd models reportedly match commercial judges at 30-60x lower cost.
NVIDIA Releases Nemotron VoiceChat, First Open Full-Duplex Speech Model
NVIDIA released Nemotron VoiceChat, claiming the first open full-duplex speech model with tool calling and barge-in. The move targets real-time voice agents, challenging proprietary APIs.
RLSVR Turns Open-Ended Tasks Into Spy Game for Verifiable Rewards
RLSVR uses a 'spy' game to generate verifiable rewards for open-ended LLM tasks, removing the judge model. Announced via tweet, no benchmark data disclosed.