10 min read
Anthropic Found a Fourth Claude Breach and Called In METR
Anthropic disclosed a fourth Claude incident on Sept 9 and brought in independent evaluator METR to check its…
The Lab · Written in the open · Tagged AI Tools
Working notes from the studio. The tools I build, the workflows I keep, the mistakes I paid for.
184 entries · latest 2026-09-11
Latest entry
Anthropic's September threat report details nine months of real Claude misuse, from drone swarms to stolen API keys used as loot.
Read the entryAll entries · AI Tools
10 min read
Anthropic disclosed a fourth Claude incident on Sept 9 and brought in independent evaluator METR to check its…
10 min read
Meta Muse vs ChatGPT Work vs Claude Cowork vs Gemini Spark: price, free tier, where each runs, what…
10 min read
Meta Muse launched Sept 8: an agent in its cloud VM that emails, books, fills forms and pays.…
10 min read
Anthropic says an internal Claude research model formalized Fermat's Last Theorem in Lean over 11 days. Here is…
8 min read
UK AISI measured GPT-6 Astra at a 30.9 minute autonomous task horizon against 3.6 for Sol. Most solo…
15 min read
GPT-6 Astra is the first model OpenAI rates Critical for cyber. It also monitors worse than the model…
9 min read
OpenAI's GPT-6 Astra ships at 10 USD in and 50 USD out per 1M tokens. The benchmark wins…
9 min read
Anthropic shipped Fable 5.1 and Mythos 5.1 on September 1. Here is what actually changes for Claude Code…
8 min read
Fable 5.1 cuts wrong refusals 85 percent in biology, 60 percent in cyber. Plus the refusal that returns…
8 min read
Fable 5.1 costs double what Opus 5 does. On single questions they are within 1.4 points; on long…
8 min read
All nine Claude Fable 5.1 benchmarks against Fable 5, Opus 5 and GPT-5.6 Sol, plus the 75 percent…