ClaudeMost teams shipping AI products can't build evals that predict how a model will actually perform in production. Michele Catasta, President & Head of AI at Replit, shares how his team closed that gap with ViBench — a public vibe-coding benchmark that scores whether the generated app works — and the offline/online evaluation loop behind Replit Agent that turns weeks of engineering into compounding overnight gains. Anthropic's Hannah Moran joins to share what separates evals that look rigorous from ones that actually help teams adopt new models with confidence.
Evaluating and improving Replit Agent at scaleClaude2026-05-08 | Most teams shipping AI products can't build evals that predict how a model will actually perform in production. Michele Catasta, President & Head of AI at Replit, shares how his team closed that gap with ViBench — a public vibe-coding benchmark that scores whether the generated app works — and the offline/online evaluation loop behind Replit Agent that turns weeks of engineering into compounding overnight gains. Anthropic's Hannah Moran joins to share what separates evals that look rigorous from ones that actually help teams adopt new models with confidence.Projects are now a conversation with ClaudeClaude2026-09-17 | Projects in Claude used to be a folder of chats. Now a project is one conversation: you say what you need as it comes to you, and Claude picks up each request and reports back as it goes. A few quick messages about a pricing test, a slow cold start and a checkout bug become three threads running at once. Projects show what's ready for review and what's waiting on you, and a follow-up in the conversation goes straight to the thread it's about. The work runs in the cloud, so it keeps going after you close your laptop.
Available in beta for select Pro and Max users, and rolling out in stages across Claude Code, chat and Cowork. Existing projects stay as they are for now, and on Pro and Max we'll upgrade them later. Pro and Max subscribers can join the waitlist: claude.com/form/projects
Read the announcement → claude.com/blog/projects-redesignedMeet Claude Slides, Claude Design and Claude DocsClaude2026-09-16 | A deck, a set of social images and a field one-pager from one launch brief, without leaving the conversation. Claude Slides, Claude Design and Claude Docs are now in beta. Comment on anything and send it to Claude to change it.
Try it: claude.aiClaude Cowork and chat are now one ClaudeClaude2026-09-16 | You don't have to choose between chat and Claude Cowork anymore. Ask Claude a quick question or hand it a bigger task, and it brings in whatever the task needs. Meaghan Choi, who leads design for Claude apps, on why keeping bigger work in a separate place stopped making sense, and what stays put: your chats, tasks, skills and memories are where you left them. Hand Claude last week's numbers and ask for the deck, and it builds it in the same chat. The work comes back as something you can edit. Claude does the work. You keep the final say.
Read more: claude.com/blog/cowork-is-now-claude Try it: claude.aiFrontier Day | Claude for startupsClaude2026-09-15 | Hear from founders attending Frontier Day on what it’s like building a startup with Claude. Caitlin Colgrove of Hex, Jacob of Spawn, Sherwood of Sazabi and Tianwei of Phylo talk about taking on projects that used to be too big to attempt, compressing biology research that took years into hours, and getting to a working game loop in an afternoon.
Learn more about the Claude Startups Program: claude.com/programs/startupsSalesforce in ClaudeClaude2026-09-15 | Salesforce in Claude is a plugin built with Salesforce that brings a seller's accounts, opportunities, and pipeline into Claude, with 37 sales skills for the work account executives do daily. Prep a call, review a deal, build a pipeline dashboard, or send your forecast, without leaving your conversation. Claude runs under the Salesforce permissions your organization already put in place, and by default asks the seller to approve each proposed change before it's written to Salesforce. Available in beta today on all paid Claude plans.
Learn more: claude.com/blog/salesforce-in-claudeHow data retention works when using ClaudeClaude2026-09-14 | Starting with Claude Fable 5, the last thirty days of prompts and Claude’s outputs are stored for safety monitoring. Learn why this window exists, who can read it, and what Enterprise Frontier Safeguards (EFS) changes for organizations.
How Claude Code uses your data: code.claude.com/docs/en/data-usageWhy Claude works better inside SlackClaude2026-09-09 | Anthropic engineers on why Claude Tag makes better decisions with more context.How founders build on Claude Managed AgentsClaude2026-09-08 | We sat down with three founders to chat about what they learned building and scaling agents with Claude Managed Agents. Wispr shipped the first version of their meeting assistant in a day and scaled it in a few weeks. Actively built a cross-account sales agent in two weeks. Pendo’s analytics platform now reads customers' codebases and suggests fixes.
They get into verifying agent work with outcomes, running memory in production, sandboxing, evals, cost, and deciding what to build yourself.
Learn more about Claude Managed Agents: platform.claude.com/docs/en/managed-agents/overviewAnthropic engineers on what Claude changed for themClaude2026-09-08 | Three Anthropic engineers on how coding with Claude changes the job. What gets faster, what they miss, and where their attention goes now.How the Claude Code team uses Claude CodeClaude2026-09-02 | A year ago, using Claude Code meant prompting, giving feedback, and accepting permission prompts. What does it look like now? Thariq Shihipar, Sid Bidasaria, and Robert Boyce on the Claude Code team share how they use Claude Code in their day-to-day work, including why they do most of their coding through Claude Tag, why they give Claude goals rather than tasks, and why they delete harness features as the models outgrow them. They trace how the product expanded from the terminal into a multi-surface tool that incorporates primitives like auto mode, workflows, and routines. They also share what they miss most about their work as developers pre-Claude Code.
0:00 Intro 0:35 - Working through Claude Tag: from tool calls to goals 2:17 - Building on technology that evolves every two months 4:48 - How the Claude Code team uses Claude Code: AskUserQuestion, artifacts, and Claude Tag 6:41 - Running loops and routines remotely 8:52 - How code review inspired dynamic workflows 14:04 - Building Claude Tag with Claude Tag: verification and feedback loops in Slack 18:37 - What they miss about the old way of software engineering
Follow ClaudeDevs on X for product updates and best practices from the Claude Code team: https://x.com/ClaudeDevsFable 5.1 is hereClaude2026-09-02 | Better judgment on every task. Fable 5.1 writes plainly, checks every number, and shows its sources. Included with Max.Debugging across the whole stack with Claude Fable 5.1Claude2026-09-01 | Claude Fable 5.1 is particularly strong on long-running engineering work, and it's smart enough to fix the root cause of an issue rather than the symptoms.
In this example, Fable 5.1 works through months of vehicle data, support tickets, and several codebases at once to track down a bug the team couldn't reproduce internally. It proposes a cause, and the developer confirms it, decides what's safe to ship, and approves the fix.
More on Claude Fable 5.1: anthropic.com/claude/fableClaude Fable 5.1 builds the ops review in SlackClaude2026-09-01 | Claude Fable 5.1 can take a job from a folder of raw data to a finished deck, and it checks the numbers as it goes.
In this example, a team puts together an ops review for leadership in Slack, using Claude Tag. Fable 5.1 builds the deck from the team's inputs. Partway through, it catches a number that doesn't reconcile, flags it, and reworks the calculation once someone weighs in. The team reviews the deck together before it goes to leadership.
More on Claude Fable 5.1: anthropic.com/claude/fableClaude Fable 5.1 runs the forecast overnightClaude2026-09-01 | Claude Fable 5.1 can build a forecast from raw data on its own, and it checks its numbers as it goes.
In this example, one person at a B2B company maintains the revenue forecast for a business with a mix of contract and consumption revenue. Claude updates the forecast overnight, running unattended on our API. In the morning, the analyst asks how accurate past forecasts have been, and Claude backtests itself on the spot. The analyst reviews and approves before anything goes out.
More on Claude Fable 5.1: anthropic.com/claude/fableClaude turns flight data into look behind youClaude2026-09-01 | Claude built a Mac app that tells you when to look up, and where. Type your address, sketch the trees and rooftops around your yard, and it watches the live air traffic overhead, does the physics (including how long the sound takes to reach you), and counts you down out loud: "Look east... in view in five, four, three..." A plane only counts if it will actually clear your treeline; one hidden behind your house gets "listen," not "look." Every aircraft on screen is real, live traffic.
The data: - Live aircraft positions: crowd-sourced ADS-B via the community aggregators adsb.lol and adsb.fi, the public radio broadcasts every transponder-equipped aircraft transmits - Flight routes: adsb.lol's community route database - Ground elevation: Open-Meteo - Geocoding, maps, and the voice: Apple's on-device frameworks
Thanks to the volunteer ADS-B community, thousands of hobbyist antennas, for keeping this data open to everyone.Claude codes turn-by-turn directions for the MoonClaude2026-09-01 | Claude built a working navigation app for the Moon from scratch. It plans drives over real NASA terrain, steering around craters too steep to enter, cutting across ones shallow enough to drive, and explaining every turn, all on a hand-built 3D globe in the browser with no libraries.
Every number on screen comes from public, freely available data: - Terrain and crater depths: NASA's LOLA laser altimeter aboard the Lunar Reconnaissance Orbiter (NASA PDS) - Imagery: the LROC camera's global mosaic, via NASA's Scientific Visualization Studio (NASA/GSFC/Arizona State University) - Place names: the Gazetteer of Planetary Nomenclature (IAU/USGS) - Landing sites: published NASA mission coordinates
All US-government public domain. Thanks to the LRO, LOLA, and LROC teams and the USGS Astrogeology Science Center for making this data open to everyone.Claude codes a watercolor engineClaude2026-09-01 | Claude built a working watercolor simulator from scratch. It computes how water moves, how pigment settles, and how light mixes color, all live in the browser with no libraries.Claude designs proteins that bind in the labClaude2026-09-01 | Claude designed new proteins from scratch, and the lab confirmed they work. Given a target and 24 hours on its own, it generated, folded, and scored them using open-source biomolecular models, checked their diversity and novelty with bioinformatics tools, and chose the designs most likely to grip the target. Those designs were then made and measured by two independent labs, Adaptyv Bio and GenScript. Across the 12 targets shown, nearly half of Claude's designs bound, and most targets got a high-affinity binder.
What you're watching: A co-folding diffusion trajectory rendered by Claude for a validated binder against each of 12 targets. Every protein on screen is one confirmed to bind experimentally.Claude turns calculus into combustionClaude2026-09-01 | Claude built a working car engine from scratch, out of equations. It computes how the fuel burns, how pressure pushes the pistons, and how the exhaust pulses become sound, all live in the browser with no libraries. The calculus is on screen: the area inside the engine's pressure loop is the work of every cycle, integrated live.Claude codes a sentence traveling through a brainClaude2026-09-01 | Claude built an interactive model of the human brain from scratch. It sculpts every fold of the cortex in code, then follows one sentence, "Could you pass the salt?", as it moves through: the brainstem at 5 milliseconds, the auditory cortex at 20, meaning at 400, and finally the frontal lobe deciding what to do about it. All live in the browser, with no scans, models, or downloaded assets. The timings come from published EEG, MEG and intracranial recordings, slowed about a hundredfold so you can watch it happen.Building Enterprise Frontier Safeguards with our customersClaude2026-09-01 | The most capable AI models raise the stakes for the businesses that deploy them. We worked with leaders at Salesforce, Visa, Uber, and KPMG to understand what it takes to run frontier models on their highest-stakes work.
Their feedback shaped Enterprise Frontier Safeguards, a solution that combines the privacy of zero data retention (ZDR) with state-of-the-art safeguards for detecting misuse.
Read more: anthropic.com/news/enterprise-frontier-safeguardsClaude for Word: Turn a draft into a finished documentClaude2026-08-26 | See exactly how Claude works inside Microsoft Word. Reading a document, resolving reviewer comments, fact-checking claims, cutting length, and copy editing, all as tracked changes you approve along the way.
Chapters: 0:00 Claude in the Word panel 0:18 Summarize the open comments 1:12 Turn on tracked changes 1:35 Check claims against the source doc in Box 2:26 Fix formatting and rewrite for the customer 3:18 Get under the word count 4:57 Run a final copy-edit pass with a skillHow to choose the right Claude model for any taskClaude2026-08-19 | Choose a Claude model based on the complexity of the task. Then adjust the effort level to move between faster or more thorough results.What does AI actually know about you?Claude2026-08-14 | The information you share with AI only travels as far as you let it. Zoe from the Anthropic education team explains what happens to your information when you chat with an AI, how long it stays there, and how to take control of where it goes.What does AI actually know about you?Claude2026-08-13 | The information you share with AI only travels as far as you let it. Zoe from the Anthropic education team explains what happens to your information when you chat with an AI, how long it stays there, and how to take control of where it goes.
Chapters 0:00 What does AI know about you? 1:05 The four places your data can go 1:25 Use 1: The conversation itself 1:48 Use 2: Product memory 2:19 Use 3: The provider's systems 2:48 Use 4: Training future models 3:33 Habits for staying in controlClaude Cowork is now your Chrome side panelClaude2026-08-12 | Claude sees the page you're already signed in to and works on it: it reads, clicks, types, and fills forms. Your skills, plugins, and connectors work in the browser for the first time. Every conversation saves to your history, and sessions live with your account, not the machine. Start a task in a browser tab, then pick it up on your desktop or phone.
Try it: claude.com/chromeClaude FM 🎵 music for thinking and buildingClaude2026-08-11 | Press play and keep thinking. Made and curated by musicians.Can you trust what AI tells you?Claude2026-08-11 | How much you can trust an AI depends on what you’re asking. Kyra from the Anthropic education team breaks down the two most common reasons for an AI to be confidently wrong: hallucination and sycophancy.Can you trust what AI tells you?Claude2026-08-11 | How much you can trust an AI depends on what you’re asking. Kyra from the Anthropic education team breaks down the two most common reasons for an AI to be confidently wrong: hallucination and sycophancy.
Chapters 0:00 Can you trust an AI's answer? 1:39 Hallucination and sycophancy, explained 3:04 Trust is a dial, not a switch 3:36 Four habits for checking AIWhat happens when you talk to AI?Claude2026-08-08 | An AI model writes one word at a time, but it doesn't think one word at a time. Jane from Anthropic’s user experience team breaks down the prediction process behind every AI output, and how to better interpret the responses you get back.
Have a question? Let us know in the comments.
Learn more at Claude Academy: http://academy.claude.comHow Ramp engineers work with AI agents at every stepClaude2026-08-06 | Ramp runs AI agents across its entire engineering lifecycle: writing code, reviewing it, watching production, and root-causing incidents. Boris sat down with Austin Ray and Rahul Sengottuvelu of Ramp to talk about how they got there. Building for the models that are coming rather than the ones that exist, giving every engineer uncapped access to intelligence, and the guardrails that make it work. They compare notes on Claude Code setups, loops versus dynamic workflows, and what Claude Fable 5 unlocked.
Chapters 0:00 "Fix all our import cycles" 0:32 Stress-testing Fable on Ramp's Python modules 1:33 Fable and dynamic workflows cut CI time 66% 3:36 Loops vs. dynamic workflows for long-horizon tasks 5:15 Claude Code setups: vanilla vs. background-heavy 6:49 AI agents across the engineering lifecycle 7:23 Building for future models, not today's 9:11 AI agent guardrails and least privilege 12:00 Cost controls and AI code review 13:08 Ramp's culture of experimentation 13:52 Glass and Inspect: Ramp's AI coworkers 16:05 On-call assistant: an AI SRE on Claude Code 17:13 More agent sessions from automations than humans 18:44 No token budgets for engineers 20:48 Advice for CTOs adopting AI agentsWhat happens when you talk to AI?Claude2026-08-05 | An AI model writes one word at a time, but it doesn't think one word at a time. Jane from Anthropic’s user experience team breaks down the prediction process behind every AI output, and how to better interpret the responses you get back.
Chapters 0:00 What happens when you talk to AI? 1:07 How AI training works 2:32 How the model thinks 3:37 Four habits for better resultsHow auto mode works with Claude CodeClaude2026-08-04 | Auto mode lets Claude Code complete long-running work with fewer interruptions, with a separate classifier screening each action instead of you. This video covers how the classifier works, why Claude never approves its own actions, and how to configure auto mode for your environment and team.
0:00 Intro 0:44 How auto mode works 3:33 Configuring auto mode 4:49 OutroRecord A Skill With ClaudeClaude2026-08-03 | Some tasks you only need to do once. There's now a Record a skill option in the + menu in Claude Cowork.What do AI models actually know?Claude2026-07-24 | AI models don't know everything. Their training gives them remarkable depth in some areas, and creates blind spots in others. Here's how to tell the difference.Why does AI hallucinate?Claude2026-07-23 | Ask an AI for a specific statistic and it might just invent one. These errors are called hallucinations and are the result of an AI trying hard to be helpful when it doesn’t know the answer.How does AI get its character?Claude2026-07-22 | AI models are grown, not built. They learn their behaviors from human text, which are then shaped further through curated examples during the fine-tuning process.What is sycophancy?Claude2026-07-21 | A model that always agrees with you isn’t a helpful model. See why sycophancy matters, how you can spot it when chatting with AI, and what we’re doing to reduce its presence in our models.Making New York City miniature with ClaudeClaude2026-07-16 | Danny Cortes makes miniatures of the New York City details most people walk right past: bodegas, mailboxes, dumpsters, storefronts. To him, every rusted corner and faded sign is worth preserving. With the help of Claude, he turns a single photo into a 1:12-scale blueprint, getting every dimension just right.Regenerative beekeeping with ClaudeClaude2026-07-14 | When a swarm of bees shows up in someone's yard, Onyx Baird gets the call. Ten years of regenerative beekeeping taught her to rely on both instinct and a wealth of information. With the help of Claude, she holds it all at once, turning years of collected questions into a one-page guide she can hand to any client.Plan smarter with Claude for TeachersClaude2026-07-14 | Elementary teacher Karina shares how she uses Claude for Teachers to get daily feedback and build that coaching into her lesson plans automatically, all grounded in real standards by connecting to TeachFX and the Learning Commons knowledge graph.The Briefing: AI for ScienceClaude2026-07-13 | Some of the world's leading researchers and pharmaceutical companies are now using Claude to accelerate their science. Last month, the people behind that work joined us to share what they've built -- and to launch Claude Science, an AI workbench for scientists.
If you missed it, watch the session here:: anthropic.com/events/the-briefing-ai-for-science-virtual-eventBuilding the future of agentic infrastructureClaude2026-07-10 | Agents are moving from tools you prompt to infrastructure that runs your business. But what does it take to run them in production? Jess Yann (Product Manager, Claude Managed Agents), Katelyn Lesse (Head of Engineering, Claude Platform), and Angela Jiang (Head of Product, Claude Platform) discuss how teams are building agentic infrastructure, including identity, permissions, memory, and agent-to-agent communication. They also share how organizations should think about agentic ROI and designing human-agent teams that adapt to evolving model intelligence.
0:00 Intro 1:00 - Building Claude Managed Agents in production 2:15 - How agents talk to each other 3:00 - The future of agentic infrastructure: thinner harnesses and adversarial agent pairs 8:20 - Barriers to agentic adoption: security, compliance, and evals 9:15 - How to measure agent ROI 12:45 - Failure modes: hyper independence and sprawl 13:30 - The future: agents as an invisible substrate 15:15 - What's next for the Claude PlatformThere’s hope in hard questionsClaude2026-07-09 | We don’t get the benefits of AI without addressing the hard questions. Share your own: claude.com/hard-questions
All voices featured in this film are from real people we’ve spoken with. You can learn more about the initiative here: anthropic.com/news/hard-questionsThere’s hope in hard questionsClaude2026-07-09 | We don’t get the benefits of AI without addressing the hard questions. Share your own: claude.com/hard-questions
All voices featured in this film are from real people we’ve spoken with. You can learn more about the initiative here: anthropic.com/news/hard-questionsMaking New York City miniature with ClaudeClaude2026-07-08 | Danny Cortes makes miniatures of the New York City details most people walk right past: bodegas, mailboxes, dumpsters, storefronts. To him, every rusted corner and faded sign is worth preserving. With the help of Claude, he turns a single photo into a 1:12-scale blueprint, getting every dimension just right.Working at the Frontier: Thomson ReutersClaude2026-07-08 | Thomson Reuters has spent more than 150 years serving professions where being right is non-negotiable, like law, tax, and compliance. CTO Joel Hron explains how legal research went from a tedious, manual search problem to agentic deep research that retrieves relevant case law while also verifying every citation. That means faster answers lawyers can actually stand behind.Photographing the stars with ClaudeClaude2026-07-07 | The Milky Way is invisible to the naked eye but that didn't stop Shane Auckland from chasing it. With the help of Claude, he mastered everything from exposure times to stitching the panorama together, driving deep into Death Valley to shoot a panorama of the Milky Way arching over the desert.Regenerative beekeeping with ClaudeClaude2026-07-07 | When a swarm of bees shows up in someone's yard, Onyx Baird gets the call. Ten years of regenerative beekeeping taught her to rely on both instinct and a wealth of information. With the help of Claude, she holds it all at once, turning years of collected questions into a one-page guide she can hand to any client.Claude Cowork: coming to mobile and webClaude2026-07-07 | Your work goes everywhere with you, and keeps going without you. Claude Cowork is now rolling out to mobile and web: start a task at your desk, check on it from your phone, and pick up the finished output anywhere.
Hand Claude a real job, like prepping for Monday's client meeting. Claude works through the email threads, transcripts, and client updates, builds the briefing doc, and leaves the follow-up drafted but unsent. Scheduled tasks now run with no device online, and when a decision needs your judgment, the question comes to your phone.
Beta is rolling out over the next several weeks starting with the Max plan, with more plans to follow.