8 min read
Claude Fable 5.1 Refuses Less: What Actually Changed
Fable 5.1 cuts wrong refusals 85 percent in biology, 60 percent in cyber. Plus the refusal that returns…
The Lab · Written in the open · Tagged Claude
Working notes from the studio. The tools I build, the workflows I keep, the mistakes I paid for.
67 entries · latest 2026-09-06
Latest entry
Anthropic says an internal Claude research model formalized Fermat's Last Theorem in Lean over 11 days. Here is what the numbers actually mean.
Read the entryAll entries · Claude
8 min read
Fable 5.1 cuts wrong refusals 85 percent in biology, 60 percent in cyber. Plus the refusal that returns…
8 min read
Fable 5.1 costs double what Opus 5 does. On single questions they are within 1.4 points; on long…
8 min read
All nine Claude Fable 5.1 benchmarks against Fable 5, Opus 5 and GPT-5.6 Sol, plus the 75 percent…
10 min read
Anthropic is rolling out a watchable built-in browser inside Claude Cowork this week. What it does and why…
8 min read
Opus 5 closed most of the gap to Fable 5 at half the price. The routing rules, plus…
8 min read
Three flagship models shipped in fifteen days. Where Claude Opus 5, GPT-5.6 Sol and Kimi K3 actually lead,…
8 min read
Opus 5 posts 79.2 on SWE-bench Pro against Opus 4.8 at 69.2. What the launch ratios mean, and…
9 min read
Anthropic updated Claude's voice mode with Opus and Sonnet, mid-call model switching, and real actions in Gmail, Calendar,…
9 min read
Anthropic's new Reflect dashboard shows how you actually use Claude. I ran it on myself for a week,…
9 min read
Claude Sonnet 5 is the new default on Pro, Team, and Enterprise seats, with 1M-token context and adaptive…
9 min read
Claude published artifacts can now pull live data through each viewer's own MCP connectors, no backend required. What…