🗞️ AI Daily Briefing — 2026-06-18

🔥 Top Story

Anthropic leaders in emergency Washington talks as Fable 5 / Mythos 5 remain suspended under export controls. Six days after the Commerce Department ordered Anthropic to cut off all foreign nationals from Claude Fable 5 and Mythos 5, the company’s leadership flew to D.C. for emergency negotiations. Anthropic chose to shut both models down entirely rather than selectively block its own foreign-born staff. The standoff now threatens Anthropic’s planned ~$1T IPO — Fortune reports the government’s ability to kill-switch frontier models is a material risk factor. Meanwhile, Dario Amodei’s June 10 policy blueprint calling for mandatory third-party testing of frontier models reads differently in the wake of his own model getting pulled. (Fortune) (Quartz)

🚀 Model & Research News

  • NVIDIA ENPIRE enables AI robots to run their own research on real hardware: Jim Fan’s GEAR Lab (with CMU and UC Berkeley) released ENPIRE, a closed-loop framework where AI coding agents autonomously reset physical scenes, run trials, verify outcomes, and rewrite code — achieving 99% pass@8 success on contact-rich tasks including GPU installation into motherboards. Fan called it “AutoResearch in the physical world for the first time.” (TechTimes)
  • OpenAI releases LifeSciBench: A 750-task benchmark for evaluating AI on real-world life science research, developed with 173 scientists. The strongest model scored only 36.1% — a sobering reminder of how far frontier models are from expert-level science. (MarkTechPost)
  • MiniMax M3 — open-weight model punches above its weight: First open-weight model combining frontier coding, 1M context, and native multimodality. Scores 59.0% on SWE-bench Pro (beating GPT-5.5’s 58.6%), priced at $0.60/$2.40 per million tokens. Architecture cuts compute to 1/20th of previous levels. (MiniMax)
  • Microsoft reportedly evaluating DeepSeek V4 to replace OpenAI/Anthropic in Copilot Cowork: A notable signal of Microsoft hedging its OpenAI dependency. V4-Pro (1.6T params, 49B active) uses only 27% of V3.2’s inference FLOPs. (GuruFocus)
  • AI disproves 80-year-old math conjecture: An Automated Conjecture Resolution framework pairing a reasoning agent with a Lean 4 formalization agent overturned a long-standing conjecture with machine-checked verification. (BuildThisNow)

🛠️ Tools & Developer Updates

  • CoreWeave completes $1.7B acquisition of Weights & Biases: W&B’s 1M+ developers and enterprise customers (OpenAI, Meta, NVIDIA, Snowflake, Toyota) are now part of CoreWeave’s end-to-end AI cloud. A major consolidation move in the MLOps space. (CoreWeave)
  • DSPy 3.3.0b1 beta ships ReActV2: New module supports parallel and multi-turn native tool calls, a typed provider-neutral LM system, and reduced dependencies. MIPROv2 is now the default optimizer, delivering 10-40% quality lifts over hand-written prompts. (GitHub)
  • dbt Wizard launches as the recommended AI agent for data development: Handles investigation, building, validation, and shipping grounded in project lineage. dbt Core v2.0 ships a Rust-based Fusion engine. Announced at Snowflake Summit. (dbt docs)
  • OpenCode enters at #1 in AI dev tool rankings: 160K+ GitHub stars and 7.5M MADs, disrupting the tools category for the first time since Cursor 3’s rebuild. (LogRocket)

💰 Funding & Business

  • GitHub infrastructure crisis deepens — availability at 88.4%: AI coding agents overwhelmed GitHub’s infrastructure, forcing Microsoft to route traffic through AWS. Agent-opened PRs surged from 4M (Sept 2025) to 17M+ (March 2026). A securities class-action lawsuit was filed June 12. (TechTimes)
  • OpenAI audited financials paint a stark picture: $34B in spending, $13B revenue, $38.5B net loss in 2025. Q1 2026: $3.7B loss on $5.7B revenue. This is the company filing for an IPO. (unrot.co)
  • Nebius acquires Eigen AI: The AI cloud company completed its acquisition of Eigen AI, a leading inference and model optimization company, to strengthen its full-stack AI infrastructure play. (BusinessWire)
  • IPO season is real: SpaceX ($1.75T), Anthropic (~$965B), and OpenAI (~$852B) all have S-1 filings in. Cerebras already trading after its $5.5B raise in May. The AI-adjacent public market is about to get very crowded.

🐦 Notable from the Timeline

  • Harrison Chase (LangChain) on the Sequoia podcast: “Context engineering is the new AI moat” — the critical path to reliable long-horizon agents is mastering execution traces and feedback loops, not just improving base models. (Sequoia)
  • Jerry Liu (LlamaIndex) at Databricks Data + AI Summit this week, arguing “the framework era is over” and context quality is the only remaining competitive advantage. Both he and Chase are converging on the same thesis from different angles.
  • Ilya Sutskever quietly self-appointed as CEO of SSI after co-founder Daniel Gross departed for Meta. The lab has $6B raised, $32B valuation, 27 employees, and zero products.
  • Marc Andreessen on Joe Rogan: “AGI has already arrived… we crossed that about 3 months ago.” Separately defended targeted AI regulation while criticizing bureaucratic approaches. (TechRadar)
  • François Chollet’s Ndea ($43M via YC W2026) continues pure research into program synthesis as the missing ingredient for AGI. ARC-AGI-3 frontier systems still score below 1%.

📊 Benchmark Watch

LMArena crossed 6.9M human preference votes across 367 models. Three models have broken the 1500 Elo barrier. Claude Opus 4.8 holds #1 at 1510+ Elo, leading the Artificial Analysis Intelligence Index at 61.4% (first to break 60). Agent Arena leaderboard launched June 4. MLPerf Training v6.0 results released June 16 with two new benchmarks. Meanwhile, OpenAI’s LifeSciBench shows frontier models topping out at 36.1% on expert-level science tasks — a useful counter-narrative to the “AGI is here” discourse. (Arena.ai)

🎙️ Podcast Highlights

  • Hard Fork #201 (June 17): Live show with Dylan Field (Figma CEO) discussing the “Design Is Dead” campaign, plus the sudden resignation of Anthropic’s Mike Krieger from Figma’s board. Robot dogs on stage. (Apple Podcasts)
  • Pivot (June 16): Kara and Scott covered the Anthropic/Trump admin clash, SpaceX public debut, state AGs opening formal process against OpenAI, and Fox buying Roku. (Apple Podcasts)
  • All-In (June 13): “Anthropic’s Fable Backlash” — the besties discussed undisclosed performance downgrades on dev tasks and a 30-day prompt retention policy. Also: AI regulatory capture and inflation hitting 3+ year highs. (allin.com)
  • TBPN (June 16-18): SpaceX valuation discussion. Earlier: full conversation with Alex Karp (Palantir CEO) on “taste” as the most important competitive advantage in AI. (tbpn.com)

🔗 Worth Reading