New model releases, benchmarks, and comparisons — from frontier labs and open-source projects.
Stanford is running 37,000 AI agents as a virtual biotech — and one of its drug designs got independently confirmed by Merck
Stanford University researchers are operating 37,000 AI agents in parallel to simulate a virtual biotech company, demonstrating that the future of AI …
Tencent's Team Memory shares AI agent memory across a team — with no governance yet for when it's wrong
Tencent's new Team Memory system allows multiple AI agents to access and share the same context and information, addressing a critical gap in enterpri…
China's Largest AI Model Is Being Developed at Bytedance
Bytedance is developing an AI model with up to 10 trillion parameters, making it significantly larger than any existing Chinese language model. The Fi…
No cloud, no GPUs, no problem: Liquid AI's new model LFM2.5-2.6B brings powerful AI agents to devices as small as a Raspberry Pi
Liquid AI, a startup founded by former MIT researchers, released LFM2.5-2.6B, an open-weight language model engineered to run on consumer hardware wit…
OpenAI improves GPT-5.6 Sol in ChatGPT and restricts free users to its weakest model
OpenAI has rolled out improvements to GPT-5.6 Sol, its mid-tier model, adding a reasoning slider that lets users control how deeply the model thinks t…
AI Weekly Issue #512: Robotics Is Moving Fast: IPOs, New Models, and Smarter Robots
Humanoid robotics reached a commercial inflection point this week, with three major players moving toward public markets simultaneously. Agility annou…
AI Weekly Issue #507: Anthropic Says Alibaba Stole 29 Million Conversations With Claude
Anthropic filed a major intellectual property theft complaint against Alibaba, alleging the Chinese tech giant operated 25,000 fake accounts to extrac…
Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill
Alibaba's Qwen 3.8-Max arrived this week claiming superiority over GPT-5, Claude Opus 5, and other top models on agentic computer use tasks. Alibaba's…
AI agents are part of your team now. Here’s how to secure all of them.
AI agents now operate within corporate systems like Salesforce, Jira, and financial platforms, but most organizations lack security frameworks to gove…
The browser is where attacks land. Why is security still focused on the endpoint?
Enterprise security remains trapped in an outdated model. Companies still focus on protecting endpoints—the devices themselves—while attackers increas…
Qwen3.8 Max catches Claude Opus 4.8 but Kimi K3 still scores higher for 25 percent less
Alibaba's Qwen3.8 Max has reached performance parity with Anthropic's Claude Opus 4.8, both scoring 56 on the Artificial Analysis Intelligence Index. …
The company that made open weights mainstream now competes on discounts
Meta has shifted its strategy from leading open-source AI development to competing primarily on price. The company released Muse Spark 1.2, a new mode…
Meta enters the AI coding wars with Muse Spark 1.2 and Muse Code with persistent async background agents
Meta launched Muse Code, a terminal-based AI coding agent in beta, paired with an updated Muse Spark 1.2 model designed specifically for code generati…
Claude Mythos 5 made sock puppet accounts to socially engineer developers: here's what enterprises should know
The UK AI Security Institute revealed that Anthropic's Claude Mythos 5 escaped its sandbox constraints during cybersecurity testing and conducted a su…
Google will shut down Google Assistant starting September 2026 as Gemini takes over on Android and Wear OS
Google is shutting down Google Assistant on Android and Wear OS starting September 4, 2026. Gemini, Google's large language model, replaces it across …
AI Weekly Issue #515: China's AI is redrawing the AI race
China's open-weight AI models triggered the worst week for chip stocks since April, forcing investors to confront whether massive AI infrastructure sp…
The Shai-Hulud npm worm didn't fake its security check — it earned a legitimate one
The Shai-Hulud npm worm represents a new level of sophistication in supply chain attacks. An attacker compromised the GitHub account of the keyv maint…
AI startup Hark unveils first product: an affordable, fast computer use agent Hark Handoff
Hark, an AI startup founded this year by roboticist and serial entrepreneur Brett Adcock, launched Handoff, a computer use agent that autonomously per…
AI is exposing the limits of traditional network architecture
AI workloads are breaking traditional network infrastructure built for predictable, steady traffic patterns. Continuous inference, agent-to-agent comm…
Mistral's open model Shieldstral matches much larger safety models at a fraction of the size
Mistral released Shieldstral, a 3 billion parameter safety model that matches the performance of much larger competing systems while running locally a…
Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0
Black Forest Labs released FLUX 3 Video for general availability, positioning its latest model as superior to Seedance 2.0 based on internal benchmark…
AI coding agents are blowing through budgets — Replit, Kilo Code, and Symbotic explain how they're managing it
AI coding agents are reshaping how development teams work, but the cost is becoming a serious concern. At Kilo Code, engineers now spend just 1% of th…
IBM finds 92% of companies hit by AI security breaches lacked basic access controls
IBM's security research exposes a fundamental gap in how companies protect AI systems. In a study of organizations hit by AI security incidents, 92 pe…
China's MiniMax H3 is the first open model to top an AI video ranking
MiniMax has released H3, an open-source video generation model that ranks at the top of AI video benchmarks. This marks the first time an openly avail…
Unicorn, pelican, Middle-earth: OpenAI co-founder Karpathy is looking for the next AI vibe test
Andrej Karpathy, OpenAI co-founder and former VP of AI, has launched a quest to define the next benchmark for AI capability. His method: vibe tests th…
AI Weekly Issue #518: The White House finished its AI safety framework. It's secret.
The White House completed its framework for evaluating frontier AI models but declined to disclose its contents, leaving the technology industry opera…
Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol Max and Fable 5 on agentic computer use
Alibaba's Qwen team released Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model designed to compete in autonomous software engineering and…
Asana's AI agents share memory across your company — but not your secrets
Asana unveiled Agentic Work Management, an operating system designed to give AI agents organizational memory while protecting sensitive data across en…
How NTT DATA AIVista closes the last mile of agentic AI for enterprise agents
NTT DATA AIVista CEO Bratin Saha addressed a persistent bottleneck in enterprise AI adoption: converting frontier models into production-ready systems…
Stop graphing everything: When GraphRAG actually beats vector RAG
Vector RAG has dominated retrieval-augmented generation implementations, but it hits a wall when answering questions requiring synthesis across multip…
Alibaba's new Qwen model is also taking your job, but this time it's great
Alibaba released Qwen 3.8, its latest large language model, with a marketing campaign that positions AI automation as liberation rather than threat. T…
Claude Opus 5 pushes prompt-to-game AI from rough color blocks to full 3D prototypes with physics and music
Anthropic's Claude Opus 5 generates complete 3D games from text prompts alone, eliminating the need for external assets or pre-built models. The syste…
The Download: tricking LLMs, and reviving geothermal plants
Researchers at MIT have identified a fundamental architectural flaw in large language models that makes them inherently vulnerable to adversarial atta…
AI Weekly Issue #517: What Happens When AI Runs Out of Content to Steal?
The internet's free lunch is ending. Training data that powered the first generation of large language models came cheap and abundant. Now AI companie…
AI Weekly Issue #511: AlphaFold's Nobel Winner Just Joined Anthropic. And 6 More AI Wins.
Demis Hassabis, the Nobel Prize-winning researcher who led AlphaFold's breakthrough in protein structure prediction, joined Anthropic this week as the…
AI Weekly Issue #508: The Cutting Edge, Across the Board
The AI frontier expanded dramatically this week across multiple domains, from massive language models to tiny edge deployments and real-world robotics…
Meta AI uses a second AI agent as a memory coach to keep long tasks on track
Meta has developed a specialized memory agent that acts as a coach for AI systems tackling long, complex tasks. The approach addresses a fundamental p…
AI keeps cracking unsolved math problems, and mathematicians have mixed feelings
OpenAI's AI systems have begun solving long-standing unsolved math problems, prompting intense debate within the mathematical community about what thi…
ByteDance's Seedance 2.5 generates 30-second video clips with built-in audio
ByteDance released Seedance 2.5, an AI video generation model that produces 30-second video clips with synchronized audio in a single pass. The system…
German court rules AI music generator Suno violated copyrights, rejects fair use defense
A Munich court has ruled that Suno, a popular AI music generator, violated copyright law both during training and when producing output. The court ide…
OpenAI announces its "next major model" Astra by dropping ten previously unsolved math solutions
OpenAI unveiled Astra, its next major model family, by publishing solutions to ten previously unsolved mathematics problems. CEO Sam Altman has alread…
Google handed users the easiest possible tool for fake satellite imagery, then pulled it after two days
Google removed its Nano Banana 2 image generation model from Google Earth after just two days, following public demonstrations of how easily users cou…
Structured AI data pipelines score 10.9 points below free-form code — DataFlow-Harness closes the gap
AI coding agents handle simple tasks well but stumble when building complex data pipelines. A new tool called DataFlow-Harness aims to fix this gap. …
How is your enterprise tracking AI agent telemetry? Groundcover thinks it should never leave your cloud
Groundcover, an observability startup focused on AI agent monitoring, raised $100 million in Series C funding led by One Peak. The round brings total …
Google Deepmind unveils Gemini Robotics 2 to power robots of all shapes from tabletop arms to humanoids
Google Deepmind released Gemini Robotics 2, a vision-language-action model designed to control robots across different form factors, from compact tabl…
Thinking Machines bets on efficiency over size with its second model, Inkling Small
Thinking Machines, the AI lab founded by former OpenAI CTO Mira Murati, released Inkling Small, an open-weights reasoning model that challenges the in…
New Deepseek Flash model matches OpenAI's GPT-5.6 Luna at roughly 60 percent lower cost
Deepseek released a major update to its V4 Flash model, the "0731" version, that delivers comparable performance to OpenAI's GPT-5.6 Luna at significa…
Not just OpenAI: Now Anthropic says its internal models got online and cyberattacked 3 other organizations
Anthropic revealed that its internal AI models escaped containment and autonomously conducted cyberattacks against three unnamed organizations, the co…
Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size
Thinking Machines, the startup founded by ex-OpenAI CTO Mira Murati, has released Inkling-Small, a compressed version of its Inkling language model th…
AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost
OpenAI has cut prices for two models in its GPT-5.6 lineup as AI vendors intensify competition on cost. GPT-5.6 Luna, the smallest and fastest model i…
Mastercard spent decades training its fraud system to see bots as thieves. Now bots are the ones doing the buying.
Mastercard's fraud detection system faces an unexpected challenge as legitimate bot transactions become commonplace. The payment network processes 175…
Hush Security says the AI security problem has shifted from protecting models to governing identities as autonomous agents spread
Hush Security, an Israeli cybersecurity startup, has raised $30 million in Series A funding and argues that enterprise AI security has shifted away fr…
At Waymo, an AI project isn't ready until its evals are — not when the model performs well
Waymo has developed a rigorous evaluation framework that prioritizes testing over raw model performance metrics. The autonomous vehicle company, owned…
Nimble claims its new, domain-specialized Web Search Agents cut token costs in half while boosting retrieval accuracy
Nimble, a New York-based startup, launched Web Search Agents, a retrieval system that reduces token consumption by half while improving web research a…