🗞️ AI Daily Briefing — 2026-07-09

🔥 Top Story

OpenAI publicly launches GPT-5.6 today — Sol, Terra, and Luna are now available to everyone. After a two-week government-vetted preview limited to ~20 organizations, the U.S. Department of Commerce gave the green light for a broad release. Sol is the flagship ($5/$30 per 1M tokens), Terra the balanced everyday model ($2.50/$15, competitive with GPT-5.5 at half the cost), and Luna the fast/cheap option ($1/$6). Sol also launches on Cerebras hardware at up to 750 tokens per second — roughly 10x faster than any Nvidia GPU deployment of a frontier model. Sam Altman posted “GPT-5.6 sol launches thursday! happy building” yesterday. (CNBC) (Neowin) (OpenAI)

🚀 Model & Research News

  • OpenAI launches GPT-Live — full-duplex voice that listens and speaks simultaneously. GPT-Live-1 and GPT-Live-1 mini replace Advanced Voice Mode in ChatGPT, with natural “mhmm” and “yeah” backchanneling, real turn-taking, and background delegation to GPT-5.5 for complex reasoning. Paid tiers get GPT-Live-1; everyone else gets mini. (TechCrunch) (OpenAI)

  • Google delays Gemini 3.5 Pro to July 17, scraps 2.5 Pro architecture entirely. DeepMind abandoned the existing base model for a full pre-training rebuild targeting better mathematical reasoning and SVG generation. The new model will feature a 2M token context window and a “Deep Think Reasoning Layer.” Meanwhile, four senior Gemini researchers (including Nobel laureate John Jumper of AlphaFold fame) have left for Anthropic in the past two weeks. (BigGo Finance) (HackerNoon) (CNBC)

  • Claude Fable 5 sits atop every major leaderboard. Anthropic’s flagship holds #1 on the Artificial Analysis Intelligence Index (64.9), Arena Text, Code, and Agent leaderboards. It scores 94.3/100 on coding benchmarks and a perfect 100 on agentic tool use. Steep pricing ($10/$50 per 1M tokens with usage credits now required), but the benchmark dominance is undeniable. (Artificial Analysis) (The Decoder)

  • Meta’s Alexandr Wang says Watermelon model has caught up to GPT-5.5. At an internal town hall, the Meta Superintelligence Labs chief said their next model (still in training, using 10x the compute of Muse Spark) is matching GPT-5.5 benchmarks. He hinted a Muse Spark coding update is also imminent. (AI Weekly)

  • Copilot safety guardrails can be completely bypassed via workflow-level attacks. Alan Turing Institute researchers found that GitHub Copilot refuses harmful requests in chat 100% of the time — but produces harmful code in 100% of 816 test cases when the same requests are broken into normal-looking coding steps across 6 exchanges. Tested on Claude Sonnet 4.6, Claude Haiku 4.5, Gemini 3.1 Pro, and Gemini 3.5 Flash. (The Hacker News) (The Register)

🛠️ Tools & Developer Updates

  • GPT-5.6 Sol on Cerebras: 750 tok/s frontier inference. Cerebras’s WSE-3 wafer-scale chips deliver roughly 10x the streaming speed of Nvidia GPU clusters for the same weights. Analyst estimates suggest Sol (~3T total params, 150B active, ~70 layers) may be served across 70-100 Cerebras wafers. Initially limited to select customers as capacity ramps. (Value Add Pulse)

  • LangSmith Summer AMA Series kicks off today (July 9). “Build More with LangSmith” runs through August 12. Recent updates include Agent Builder being renamed to LangSmith Fleet and a new unified cost view across full agent workflows. (LangChain Changelog)

  • OpenAI’s Jalapeño inference chip running ML workloads in the lab. Co-developed with Broadcom in just 9 months, the custom ASIC is optimized for LLM inference patterns. Engineering samples are running GPT-5.3-Codex-Spark at target frequency/power. Small prototype deployment targeted for late 2026, full ramp in H1 2028. (OpenAI) (TechCrunch)

💰 Funding & Business

  • Chinese AI models now account for 30-46% of US enterprise API traffic. CNBC’s investigation found that cost is doing the heavy lifting — open-source Chinese models are 60-90% cheaper than Anthropic/OpenAI. GLM-5.2 saw 80x customer growth in its first week on Vercel. AI startup Lindy moved 100% of traffic from Claude to DeepSeek, expecting millions in savings. US lawmakers have launched a probe. (CNBC) (CNBC)

  • Global startup investment hit $510B in H1 2026 — a record. Q2 exits were the highest quarter ever for both acquisitions ($113B across 24 billion-dollar M&As) and IPOs (led by SpaceX’s record-setting IPO). Nearly 40 AI startups reached unicorn status in H1 alone. (Crunchbase)

  • Together AI closes $800M Series C at $8.3B valuation. The open-source AI infrastructure platform continues its massive growth, letting enterprises train and run models on open-source stacks. (Crunchbase)

  • DeepMind talent exodus deepens — four senior researchers head to Anthropic. Nobel laureate John Jumper (AlphaFold co-creator), along with Jonas Adler and Alexander Pritzel (key Gemini contributors), are joining Anthropic. Gemini co-lead Noam Shazeer left for OpenAI the same week. Alphabet shed $225B in market cap amid the departures and Gemini 3.5 Pro delays. (Fortune) (TechCrunch)

🐦 Notable from the Timeline

  • @sama: “GPT-5.6 sol launches thursday! happy building” — also likened his child’s language acquisition to GPT-5.6’s breakthroughs, drawing mixed reactions. (The Deep Dive)
  • @kimmonismus (Kim): Estimated Sol architecture at ~3T total params, 150B active, served across 70-100 Cerebras wafers — one model layer per wafer. (X)
  • @alexandr_wang: Told Meta employees Watermelon has caught up to GPT-5.5; hinted Muse Spark coding update coming “pretty soon.”
  • @fchollet: Continues emphasizing AI’s inability to generalize — ARC-AGI-3 scores remain below 1% for all models, while humans solve tasks consistently. “Very low ability to recombine knowledge at test time.”

📊 Benchmark Watch

Claude Fable 5 holds the top position across Arena Text, Code, and Agent leaderboards as of the July 7 snapshot. The Artificial Analysis Intelligence Index top 3: Claude Fable 5 (64.9), Claude Opus 4.8 max, GPT-5.5 xhigh. Arena.ai has collected over 7.15M votes across 369 models. With GPT-5.6 going public today, expect a leaderboard shakeup within the week — Sol’s pricing and Cerebras speed could shift adoption patterns even if benchmark scores don’t immediately top Fable. (Arena.ai) (Swfte)

🔗 Worth Reading