[State of AI Papers 2025] Fixing Research with Social Signals, OCR & Implementation — Team AlphaXiv @LatentSpacePod
[State of AI Papers 2025] Fixing Research with Social Signals, OCR & Implementation — Team AlphaXiv  @LatentSpacePod
Uploaded December 2025 | Updated September 2026, 3 weeks ago
From late-night dorm-room hacking sessions at Stanford to building the most-used platform for navigating AI research, the founders of *AlphaXiv* have spent the last few years watching the archive firehose explode from thousands to 30,000+ papers per month—and building the intelligent layer on top that researchers, applied engineers, and even VCs now rely on to make sense of it all. We caught up with *Raj, Rayhahn, and the AlphaXiv team* live at *NeurIPS 2025* to dig into the origin story (spoiler: it started as a class project with a "view comment" button next to archive paragraphs), why they beat Hugging Face Papers by obsessing over UI and getting Laura's authors to comment directly on their own work, how they've evolved from comments to feeds to benchmarks to AI assistants that answer paper questions using Gemini's 1M context window, what's broken in the academic review process (20% AI-generated reviews at ICLR, quality collapsing under exponential submission growth), why *papers are becoming less important* than implementations, Docker containers, and real-world usability, their vision for making it trivial to spin up any trending paper's code directly in your browser, and the power law they see every day: archive has 2.4M papers, but applied researchers only care about the top 0.1% that are actually implementable—and AlphaXiv is building the ranking, tooling, and social signals to surface exactly those.
We discuss:

* The *origin story:* from a late-night web dev class project ("view comment" buttons on archive PDFs) to going viral on LinkedIn and becoming a full-time startup backed by Stanford advisors
* Why AlphaXiv beat *Hugging Face Papers:* better UI, direct commenting on the paper itself, and early traction with high-signal authors like the Laura/DPO/Llama teams
* The evolution: *comments → feed → benchmarks → AI assistant,* using social signal (views, comments, Twitter) to filter the archive firehose and surface what actually matters
* How they handle *PDF parsing and OCR:* DeepSeek Coder for cost/accuracy, Mistral OCR APIs, and passing diagrams directly into multimodal models like Claude
* Why *semantic search alone fails:* 3M papers, buzzword overload, and the need to weight by social signal (views, engagement) to return actually relevant results
* The *state of academic publishing:* exponential submission growth (30k/month in CS), 20% AI-generated reviews at ICLR, review quality collapsing, and the rise of AI-written papers flooding the system
* Why *papers are becoming obsolete:* the future is implementations, Docker containers, and interactive sandboxes—not PDFs describing what code does
* The *power law:* 2.4M papers on archive, but applied researchers (Spotify, Expedia, Nintendo) only care about the top 0.1% that are actually implementable and useful
* AlphaXiv's roadmap: *Docker containers for papers,* agent-automated setup, ranking by "implementation ease," and making it trivial to spin up trending work directly in your browser
* Papers of the year: *Tiny Recursive Models (TRMs),* evolutionary strategies at hyperscale, Agent R1 (RL-trained agents for complex reasoning), and the rise of AI for science (Agent Laboratory, Bionemo)
* The *Qwen vs. DeepSeek* dynamic: Qwen's breakout year, DeepSeek's continued influence (OCR models triggering a flurry of competitive releases), and the open-source fine-tuning ecosystem around both
* Why *AI reviewers* could help (Stanford's Agent Reviewer, using AI as a linter for clarity and quality) but assessing novelty and lit review remains hard without search + social signal
* The vision: AlphaXiv as the *tool layer for research*—not a social network, but the best place to discover, understand, and implement the ideas that actually matter in AI

—
AlphaXiv Team

* AlphaXiv: alphaxiv.org
* X: https://x.com/alphaxiv

Where to find Latent Space

* X: https://x.com/latentspacepod
* Substack: https://www.latent.space/

00:00:00 Introduction: The Origin Story of AlphaXiv
00:03:12 Beating Hugging Face Papers: What Made AlphaXiv Different
00:04:17 From Comments to Discovery: Building the Research Feed
00:04:55 PDF Parsing and OCR: The Technical Infrastructure
00:06:01 Becoming a Company: Beyond Papers to Research Artifacts
00:06:37 Benchmarks and State of the Art: Replacing Papers with Code
00:07:54 Personalization and Search: Building the Best Research Assistant
00:08:54 Docker Containers and Reproducibility: Making Papers Runnable
00:10:52 Favorite Papers: Tiny Recursive Models and Evolutionary Strategies
00:13:15 AI for Science: Agent Laboratory and the Future of Research Automation
00:18:41 Agent R1 and RL for Long-Horizon Tasks
00:20:54 Qwen's Year and the Open Source Landscape
00:22:43 The Crisis in Academic Publishing: AI Slop and Review Quality
00:25:04 Search and Social Signals: Why Semantic Search Isn't Enough
00:28:15 The Future of Research: Beyond the PDF
00:31:17 Implementation Ease as the New Ranking Signal
[State of AI Papers 2025] Fixing Research with Social Signals, OCR & Implementation — Team AlphaXiv‘You guys are so inefficient’ #substack #shortsWhen AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon LabsWill AI Kill Language LearningOptimizing Hyperparameters TogetherDreamer: the Agent OS for Everyone — David SingletonThe Truth Behind Cursors Biggeset Model LaunchLeveraging Strengths for Model SuccessFPV Drones -The Next War Is Already Here — Yaroslav Azhnyuk, The Fourth Law & Noah Smith, NoahpinionDylan Patel Explains the AI War While Cooking | In-Context CookingThe Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO🔬Max Welling: Materials Underlie Everything
Latent Space |

[State of AI Papers 2025] Fixing Research with Social Signals, OCR & Implementation — Team AlphaXiv

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER