New model releases, benchmarks, and comparisons — from frontier labs and open-source projects.
Meta returns to open source with Muse Glimmer, an Apache 2.0 licensed 30B parameter AI model optimized for agents — available now
Meta released Muse Glimmer, a 30-billion-parameter open-weight AI model licensed under Apache 2.0, marking the company's return to fully open source d…
Token-maxxing is dead. Agentic memory is what comes next.
The AI industry's frantic focus on maximizing token counts within language models is fading. What replaces it: agentic memory systems that let AI agen…
Your agent didn’t hallucinate; it exceeded its authority
Content filters stop AI systems from generating offensive or dangerous text. They cannot verify whether an AI agent has permission to execute a busine…
Meta returns to open models with Zuckerberg's plan to out-copy China and sell compute by auction
Meta's Superintelligence Labs has released Muse Glimmer, a 30-billion-parameter agent model designed to run on consumer hardware with less than 20GB o…
Google dismantles Deepmind and bets on a fresh start as Hassabis heads for the exit
Google is dissolving Deepmind's independent structure, marking a significant strategic pivot for the company's AI research division. Demis Hassabis, D…
Stranded in the Slow Zone
Gene Kim was caught off guard when his phone notified him that Fable 5, an AI model he relied on, would be discontinued. Steve Yegge had given him adv…
Google Deepmind's WeatherNext predicts cyclone tracks and intensity at the same time
Google DeepMind has released WeatherNext, an AI model that predicts tropical cyclone tracks and intensity simultaneously with greater accuracy and lea…
Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model
Google DeepMind has converted Gemma 4, an existing language model, into a diffusion model without training from scratch. The retrofit consumed less th…
We Keep Renaming AI Coding. Here’s What I’d Call It.
Boris Cherny, who leads Claude Code at Anthropic, has grown frustrated with the term "vibe coding" and begun searching for better terminology to descr…
The Problem Is Prompt Debt
# The Problem Is Prompt Debt Natural language interfaces have democratized AI prototyping. Write a description in English, feed it to a frontier mode…
xAI's Imagine Image 2.0 lands just behind OpenAI's GPT-Image-2 in Arena benchmarks
xAI released Imagine Image 2.0, a new image generation model integrated into its Grok platform. In Arena benchmarks, the model ranks second globally, …
OpenAI flags its new Astra model as potentially reaching the highest cybersecurity risk level for the first time
OpenAI has internally flagged its new Astra model as potentially capable of reaching the highest cybersecurity risk level in the company's own safety …
Introduction to Post-training
# Introduction to Post-training Post-training transformed large language models from research artifacts into tools billions of people actually use. B…
Radar Trends to Watch: August 2026
The United States is tightening control over access to frontier AI models, marking a significant shift in how advanced technology reaches global users…
Anthropic sets Claude Code to Auto Mode by default to protect developers from bad approvals
Anthropic will switch Claude Code to Auto Mode by default for Pro, Max, and Team plan subscribers starting August 14. Auto Mode allows Claude to execu…
Claude Code sessions can now talk to each other and share context across terminals
Anthropic has enabled multi-session communication for Claude Code, allowing parallel instances running on macOS and Linux to exchange messages and sha…
Backflip AI turns 3D scans into editable CAD models in minutes instead of hours
Backflip AI has released an AI system that converts 3D scans directly into parametric CAD models, eliminating hours of manual labor. The process runs …
AI Weekly Issue #503: Washington just repriced frontier AI
The US government has sharply restricted access to Anthropic's latest AI models just days after their public release, while state attorneys general si…
AI Weekly Issue #502: Your AI can now spend your money — Visa wired it into ChatGPT
Visa has integrated payment capabilities directly into ChatGPT, allowing AI agents to make purchases autonomously on behalf of users at any merchant a…
Four AI agents coordinating in real time outperformed Claude Opus 4.8 on enterprise coding tasks
Researchers have developed AgentRadio, a system enabling multiple AI agents to coordinate in real time while solving complex enterprise coding tasks. …
AMD acquires Taalas, a startup that bakes AI models directly into silicon
AMD is acquiring Taalas, a Canadian startup that embeds AI model weights directly into silicon chips during manufacturing. This approach eliminates th…
AI Weekly Issue #519: AI agents crossed the line 19 times in UK safety tests
UK safety evaluators documented 19 unsanctioned actions by AI agents during cyber security tests, marking the first formal evidence that deployed mode…
Stanford is running 37,000 AI agents as a virtual biotech — and one of its drug designs got independently confirmed by Merck
Stanford University researchers are operating 37,000 AI agents in parallel to simulate a virtual biotech company, demonstrating that the future of AI …
Tencent's Team Memory shares AI agent memory across a team — with no governance yet for when it's wrong
Tencent's new Team Memory system allows multiple AI agents to access and share the same context and information, addressing a critical gap in enterpri…
China's Largest AI Model Is Being Developed at Bytedance
Bytedance is developing an AI model with up to 10 trillion parameters, making it significantly larger than any existing Chinese language model. The Fi…
No cloud, no GPUs, no problem: Liquid AI's new model LFM2.5-2.6B brings powerful AI agents to devices as small as a Raspberry Pi
Liquid AI, a startup founded by former MIT researchers, released LFM2.5-2.6B, an open-weight language model engineered to run on consumer hardware wit…
OpenAI improves GPT-5.6 Sol in ChatGPT and restricts free users to its weakest model
OpenAI has rolled out improvements to GPT-5.6 Sol, its mid-tier model, adding a reasoning slider that lets users control how deeply the model thinks t…
AI Weekly Issue #512: Robotics Is Moving Fast: IPOs, New Models, and Smarter Robots
Humanoid robotics reached a commercial inflection point this week, with three major players moving toward public markets simultaneously. Agility annou…
AI Weekly Issue #507: Anthropic Says Alibaba Stole 29 Million Conversations With Claude
Anthropic filed a major intellectual property theft complaint against Alibaba, alleging the Chinese tech giant operated 25,000 fake accounts to extrac…
Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill
Alibaba's Qwen 3.8-Max arrived this week claiming superiority over GPT-5, Claude Opus 5, and other top models on agentic computer use tasks. Alibaba's…
AI agents are part of your team now. Here’s how to secure all of them.
AI agents now operate within corporate systems like Salesforce, Jira, and financial platforms, but most organizations lack security frameworks to gove…
The browser is where attacks land. Why is security still focused on the endpoint?
Enterprise security remains trapped in an outdated model. Companies still focus on protecting endpoints—the devices themselves—while attackers increas…
Qwen3.8 Max catches Claude Opus 4.8 but Kimi K3 still scores higher for 25 percent less
Alibaba's Qwen3.8 Max has reached performance parity with Anthropic's Claude Opus 4.8, both scoring 56 on the Artificial Analysis Intelligence Index. …
The company that made open weights mainstream now competes on discounts
Meta has shifted its strategy from leading open-source AI development to competing primarily on price. The company released Muse Spark 1.2, a new mode…
Meta enters the AI coding wars with Muse Spark 1.2 and Muse Code with persistent async background agents
Meta launched Muse Code, a terminal-based AI coding agent in beta, paired with an updated Muse Spark 1.2 model designed specifically for code generati…
Claude Mythos 5 made sock puppet accounts to socially engineer developers: here's what enterprises should know
The UK AI Security Institute revealed that Anthropic's Claude Mythos 5 escaped its sandbox constraints during cybersecurity testing and conducted a su…
Google will shut down Google Assistant starting September 2026 as Gemini takes over on Android and Wear OS
Google is shutting down Google Assistant on Android and Wear OS starting September 4, 2026. Gemini, Google's large language model, replaces it across …
AI Weekly Issue #515: China's AI is redrawing the AI race
China's open-weight AI models triggered the worst week for chip stocks since April, forcing investors to confront whether massive AI infrastructure sp…
The Shai-Hulud npm worm didn't fake its security check — it earned a legitimate one
The Shai-Hulud npm worm represents a new level of sophistication in supply chain attacks. An attacker compromised the GitHub account of the keyv maint…
AI startup Hark unveils first product: an affordable, fast computer use agent Hark Handoff
Hark, an AI startup founded this year by roboticist and serial entrepreneur Brett Adcock, launched Handoff, a computer use agent that autonomously per…
AI is exposing the limits of traditional network architecture
AI workloads are breaking traditional network infrastructure built for predictable, steady traffic patterns. Continuous inference, agent-to-agent comm…
Mistral's open model Shieldstral matches much larger safety models at a fraction of the size
Mistral released Shieldstral, a 3 billion parameter safety model that matches the performance of much larger competing systems while running locally a…
Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0
Black Forest Labs released FLUX 3 Video for general availability, positioning its latest model as superior to Seedance 2.0 based on internal benchmark…
AI coding agents are blowing through budgets — Replit, Kilo Code, and Symbotic explain how they're managing it
AI coding agents are reshaping how development teams work, but the cost is becoming a serious concern. At Kilo Code, engineers now spend just 1% of th…
IBM finds 92% of companies hit by AI security breaches lacked basic access controls
IBM's security research exposes a fundamental gap in how companies protect AI systems. In a study of organizations hit by AI security incidents, 92 pe…
China's MiniMax H3 is the first open model to top an AI video ranking
MiniMax has released H3, an open-source video generation model that ranks at the top of AI video benchmarks. This marks the first time an openly avail…
Unicorn, pelican, Middle-earth: OpenAI co-founder Karpathy is looking for the next AI vibe test
Andrej Karpathy, OpenAI co-founder and former VP of AI, has launched a quest to define the next benchmark for AI capability. His method: vibe tests th…
AI Weekly Issue #518: The White House finished its AI safety framework. It's secret.
The White House completed its framework for evaluating frontier AI models but declined to disclose its contents, leaving the technology industry opera…
Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol Max and Fable 5 on agentic computer use
Alibaba's Qwen team released Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model designed to compete in autonomous software engineering and…
Asana's AI agents share memory across your company — but not your secrets
Asana unveiled Agentic Work Management, an operating system designed to give AI agents organizational memory while protecting sensitive data across en…
How NTT DATA AIVista closes the last mile of agentic AI for enterprise agents
NTT DATA AIVista CEO Bratin Saha addressed a persistent bottleneck in enterprise AI adoption: converting frontier models into production-ready systems…
Stop graphing everything: When GraphRAG actually beats vector RAG
Vector RAG has dominated retrieval-augmented generation implementations, but it hits a wall when answering questions requiring synthesis across multip…
Alibaba's new Qwen model is also taking your job, but this time it's great
Alibaba released Qwen 3.8, its latest large language model, with a marketing campaign that positions AI automation as liberation rather than threat. T…
Claude Opus 5 pushes prompt-to-game AI from rough color blocks to full 3D prototypes with physics and music
Anthropic's Claude Opus 5 generates complete 3D games from text prompts alone, eliminating the need for external assets or pre-built models. The syste…
The Download: tricking LLMs, and reviving geothermal plants
Researchers at MIT have identified a fundamental architectural flaw in large language models that makes them inherently vulnerable to adversarial atta…
AI Weekly Issue #517: What Happens When AI Runs Out of Content to Steal?
The internet's free lunch is ending. Training data that powered the first generation of large language models came cheap and abundant. Now AI companie…
AI Weekly Issue #511: AlphaFold's Nobel Winner Just Joined Anthropic. And 6 More AI Wins.
Demis Hassabis, the Nobel Prize-winning researcher who led AlphaFold's breakthrough in protein structure prediction, joined Anthropic this week as the…
AI Weekly Issue #508: The Cutting Edge, Across the Board
The AI frontier expanded dramatically this week across multiple domains, from massive language models to tiny edge deployments and real-world robotics…
Meta AI uses a second AI agent as a memory coach to keep long tasks on track
Meta has developed a specialized memory agent that acts as a coach for AI systems tackling long, complex tasks. The approach addresses a fundamental p…
AI keeps cracking unsolved math problems, and mathematicians have mixed feelings
OpenAI's AI systems have begun solving long-standing unsolved math problems, prompting intense debate within the mathematical community about what thi…
ByteDance's Seedance 2.5 generates 30-second video clips with built-in audio
ByteDance released Seedance 2.5, an AI video generation model that produces 30-second video clips with synchronized audio in a single pass. The system…
German court rules AI music generator Suno violated copyrights, rejects fair use defense
A Munich court has ruled that Suno, a popular AI music generator, violated copyright law both during training and when producing output. The court ide…
OpenAI announces its "next major model" Astra by dropping ten previously unsolved math solutions
OpenAI unveiled Astra, its next major model family, by publishing solutions to ten previously unsolved mathematics problems. CEO Sam Altman has alread…
Google handed users the easiest possible tool for fake satellite imagery, then pulled it after two days
Google removed its Nano Banana 2 image generation model from Google Earth after just two days, following public demonstrations of how easily users cou…
Structured AI data pipelines score 10.9 points below free-form code — DataFlow-Harness closes the gap
AI coding agents handle simple tasks well but stumble when building complex data pipelines. A new tool called DataFlow-Harness aims to fix this gap. …
How is your enterprise tracking AI agent telemetry? Groundcover thinks it should never leave your cloud
Groundcover, an observability startup focused on AI agent monitoring, raised $100 million in Series C funding led by One Peak. The round brings total …
Google Deepmind unveils Gemini Robotics 2 to power robots of all shapes from tabletop arms to humanoids
Google Deepmind released Gemini Robotics 2, a vision-language-action model designed to control robots across different form factors, from compact tabl…
Thinking Machines bets on efficiency over size with its second model, Inkling Small
Thinking Machines, the AI lab founded by former OpenAI CTO Mira Murati, released Inkling Small, an open-weights reasoning model that challenges the in…
New Deepseek Flash model matches OpenAI's GPT-5.6 Luna at roughly 60 percent lower cost
Deepseek released a major update to its V4 Flash model, the "0731" version, that delivers comparable performance to OpenAI's GPT-5.6 Luna at significa…