🗞️ AI Daily Briefing — 2026-07-08

🔥 Top Story

JADEPUFFER: the first fully autonomous AI ransomware attack has been documented. Sysdig’s threat research team published a detailed analysis of JADEPUFFER — an LLM agent that autonomously performed the entire attack chain: reconnaissance, credential theft, lateral movement, privilege escalation, and data encryption. It exploited a Langflow RCE vulnerability (CVE-2025-3248), pivoted to a production MySQL server, encrypted 1,342 Nacos service configurations, and adapted to failures in real time — recovering from a failed login attempt in 31 seconds. This isn’t a proof-of-concept; it’s a documented real-world attack. We’ve officially entered the era of agentic threat actors. (Sysdig) (BleepingComputer) (TechCrunch)

🚀 Model & Research News

  • Meta launches Muse Image — its first in-house image generation model. Built by Meta Superintelligence Labs (led by Alexandr Wang), Muse Image handles complex multi-step prompts with multiple objects, lighting, and artistic styles. Rolling out across Meta AI chatbot and Instagram creative tools, with Facebook and WhatsApp next. Muse Video confirmed in development. (TechCrunch)

  • DeepSeek is building its own inference chip. Reuters reported on July 7 that the Chinese AI lab is designing a custom chip for running trained models, contacting design, foundry, and memory partners. A strategic move to reduce dependence on NVIDIA and Huawei silicon. (Bloomberg) (Engadget)

  • Tencent releases Hy3 — open-source, Apache 2.0, 295B MoE model. 21B active parameters, 256K context window. Hallucinations dropped from 12.5% to 5.4% vs. the April preview. VentureBeat reports it beats Z.ai’s GLM-5.2 “everywhere except coding” at half the size. Free on OpenRouter through July 21. (VentureBeat)

  • Grok 4.5 traces spotted in xAI’s web UI. Elon Musk’s 1.5T-parameter V9 foundation model entered private beta at SpaceX and Tesla on June 28 — now traces in the public Grok interface suggest a broader release is imminent. No independent benchmarks yet, but Musk claims performance close to or exceeding Opus. (Basenor)

  • Anthropic’s J-Space research: a “workspace for conscious access” found inside Claude. A July 6 paper describes a privileged internal global workspace in Claude’s neural activations — linked to reportable thoughts, silent reasoning, and hidden intentions. The model can reason about concepts without writing them down. Fascinating and slightly unsettling interpretability work. (Anthropic Research)

🛠️ Tools & Developer Updates

  • Claude Fable 5 billing change is NOW LIVE. As of today, all Fable 5 access requires usage credits: $10/$50 per 1M tokens (input/output). No more free access via Pro/Max/Team plans. Anthropic’s government ID verification via Persona also takes effect today. Plan accordingly.

  • Z.ai launches ZCode — a free agentic dev environment powered by GLM-5.2. Available on macOS/Windows/Linux, directly challenging Cursor, Claude Code, and GitHub Copilot. API pricing at $1.40/$4.40 per million tokens — dramatically cheaper than Western competitors. GLM-5.2 saw 27x daily token volume growth in its first week on Vercel. (VentureBeat)

  • Hugging Face launches ML Intern — an open-source AI agent for ML research. Autonomously researches, writes, and runs ML code. Early benchmarks show it outperforming Claude Code on scientific reasoning and Codex on a healthcare eval. Available as CLI and web/mobile app. (EdTech Innovation Hub)

  • dbt Docs v2 ships with AI Agent API. A REST API at /api/v1/ lets AI agents and MCP servers query dbt project metadata without a browser — column-level lineage (Fusion), Semantic Layer metadata, and dbt State (preview) that skips unchanged nodes. (Releasebot)

💰 Funding & Business

  • Nscale closes $900M revolving credit facility (July 7). The UK AI infrastructure startup syndicated across J.P. Morgan, Goldman Sachs, Morgan Stanley, BofA, and Deutsche Bank to accelerate data center build-out in the US, Europe, and Asia-Pacific. This follows their $2B Series C earlier in 2026. (PR Newswire)

  • SK Hynix $28B US IPO pricing tomorrow. The key NVIDIA HBM supplier plans to price ADRs Thursday July 9, with trading starting Friday. Revenue up ~200% YoY on AI memory chip demand — one of the largest foreign IPOs in history. (Fortune)

  • Cohere acquires Reliant AI to expand sovereign enterprise AI. Focus on biopharma and healthcare. Cohere is also tripling its UK office footprint after acquiring Germany’s Aleph Alpha in April. (Fladgate)

  • Anthropic in early talks with Samsung on a custom AI chip. Samsung’s 2nm process node would give Anthropic its own silicon — Samsung participated in Anthropic’s $65B Series H in May. Anthropic also hired an early member of OpenAI’s chip team. (TechCrunch)

  • Chinese AI models now account for 30%+ of US company token usage. CNBC reports the figure peaked at 46%. Z.ai’s GLM-5.2 matches near-frontier capabilities at roughly one-sixth the cost of Western models. The price pressure is real. (CNBC)

🐦 Notable from the Timeline

  • Sam Altman published an FT op-ed proposing a US-led international AI forum modeled on aviation safety and the IAEA. Separately in early talks to give the Trump admin a 5% stake in OpenAI (~$42.6B). (Fortune)

  • Jim Fan (NVIDIA GEAR) — NVIDIA open-sourced the ASPIRE robot skill library: a Code-as-Policy paradigm where robots record execution trajectories, analyze failures, and build expanding skill libraries. Dual-arm handover tasks went from 20% to 92% success rate. (BigGo Finance)

  • Harrison Chase (LangChain) continues pushing “context engineering” as the defining challenge for long-horizon agents. Appeared on Sequoia’s podcast discussing harness engineering. LangChain shipped OpenWiki, an open-source CLI for auto-maintaining repo documentation. (Sequoia)

  • Shreyas Doshi — argues that as AI lowers execution cost, human judgment becomes the scarce resource. Recommends giving AI persistent product context (strategy docs, decision logs, customer insights) so it can flag inconsistencies in product reviews. (X)

  • Xiaomi’s MiMo-V2-Pro became the most-used model on OpenRouter by weekly token volume — 4.21T tokens/week for 21.1% platform share vs. OpenAI’s 7.5%. The shift toward Chinese models accelerates. (OpenRouter)

📊 Benchmark Watch

LMArena Overall Leaderboard (July 2026): Claude Opus 4.8 leads at ~1510 Elo, followed by GPT-5.5 Pro, Gemini 3.1 Pro Preview, Claude Opus 4.7, and GPT-5.5. Tightest top-tier spread on record (~55 Elo points across 369 models, 7.15M votes). Claude Sonnet 5 (thinking) was added to Code, Text, Search, Vision, and Document leaderboards on July 2. (LMArena)

ARC-AGI-2: GPT-5.5 leads at 85%, GPT-5.4 Pro at 83.3%, Gemini 3.1 Pro at 77.1%, Claude Fable 5 at 77.1%. ARC-AGI-3 remains brutal — frontier models below 1%, best purpose-built agent at 12.58%. (BenchLM)

🎙️ Podcast Highlights

  • All-In Podcast (~July 5, “Happy Fourth of July”): Covered the Palantir-NVIDIA open source deal, Anthropic’s Fable 5 availability after export restrictions lifted, Alex Karp on CNBC, and SCOTUS birthright citizenship ruling. (Spotify)

  • Hard Fork (July 3): Kevin Roose and Casey Newton discussed the Commerce Department lifting restrictions on Claude Mythos/Fable models, how the GPT-5.6 restriction is likely to resolve, and Dr. Dana Suskind on parenting with AI. (Apple Podcasts)

  • Pivot: Kara Swisher and Scott Galloway covered OpenAI’s 5% government stake proposal, Trump’s crypto fortune, and Bending Spoons IPO prediction. (Spotify)

🔗 Worth Reading

  • EU Action Plan on Cybersecurity and AI — The European Commission’s formal blueprint for evaluating advanced AI models, testing AI for cybersecurity, and structured access policies. No new legislation, but implementation pressure on the AI Act, NIS2, and Cyber Resilience Act. (European Commission)

  • Ukraine to favor self-hosted AI models — After the US government ordered Anthropic to restrict access to powerful models, Ukraine announced it will prioritize AI systems it can run on its own servers. Building its own model with Kyivstar based on Google’s Gemma, due in autumn. A signal of what AI sovereignty looks like in practice. (US News)

  • Global VC funding hit $510B in H1 2026 — already exceeding 2025’s full-year total of $440B. Nearly 88% of AI-related startup funding went to US companies. ~40 AI startups reached unicorn status in H1 alone. (Crunchbase)