New model releases, benchmarks, and comparisons — from frontier labs and open-source projects.
GPT-6 Astra pilots a surveillance drone and runs a business on its own
OpenAI's GPT-6 Astra demonstrates autonomous capabilities that extend far beyond text generation, successfully piloting surveillance drones and managi…
Google's new AI model predicts the future from sales data, weather, and discount schedules
Google Research has launched TimesFM-3, a new forecasting model designed to predict future outcomes by analyzing historical time series data alongside…
GPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarks
OpenAI's GPT-6 Astra demonstrates a breakthrough in spatial reasoning tasks, clearing a significant hurdle that has long challenged AI systems in robo…
GPT-6 Astra needs leaner prompts and fewer guardrails, OpenAI recommends
OpenAI's Eric Provencher has advised developers to strip down their approach to GPT-6 Astra, recommending leaner prompts and fewer guardrails to unloc…
The Interfaces Are Arriving
# The Interfaces Are Arriving A standards body just moved AI from isolated models into something far more useful. In December 2025, Anthropic donated…
Claude Fable 5.1's language is less "load-bearing" than its predecessor's
Anthropic's Claude Fable 5.1 exhibits a measurable shift in writing style compared to its predecessor, trading linguistic density for straightforwardn…
GPT-6 Astra gives mathematicians a breather, and OpenAI says that's by design
OpenAI's GPT-6 Astra has topped ErdosBench, a benchmark measuring performance on open mathematics problems, despite the company deliberately depriorit…
New Deepseek model V4.1-Flash cuts memory needs for AI agents
Deepseek has released V4.1-Flash, a new multimodal large language model designed to slash the memory overhead that makes AI agents expensive to run at…
GPT-6 Astra beat Portal start to finish without human help in under 24 hours
OpenAI's GPT-6 Astra completed the entire puzzle game Portal autonomously in under 24 hours without human intervention after receiving its initial obj…
Qwen-Drive 1.0 tells you why it brakes, just don't expect the explanation to match the maneuver
Alibaba's research division unveiled Qwen-Drive 1.0, an autonomous driving model that combines environmental perception, traffic question-answering, a…
Google's WeatherNext 3 ditches physics simulations and learns weather directly from live satellite data
Google Research and DeepMind released WeatherNext 3, an AI weather forecasting model that abandons traditional physics-based simulation in favor of ma…
Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward
OpenAI's latest model, GPT-6 Astra, is delivering mixed signals on performance benchmarks while achieving a notable milestone that has accelerated AGI…
GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era"
OpenAI has released GPT-6 Astra and declared the arrival of the "AGI era," marking a turning point the company has avoided claiming for years. Preside…
Claude Fable 5.1 decoded a centuries-old royalist message hidden in plain sight since 1653
Anthropic's Claude Fable 5.1 has decoded a centuries-old cryptographic puzzle that historians and cryptanalysts struggled with for over 370 years. The…
Google’s Gemini 3.8 Flash is built for agents, while its Cyber twin hunts vulnerabilities
Google released two specialized variants of its Gemini 3.8 Flash model on Wednesday, targeting different workloads in the emerging AI agents market. T…
Gemini 3.8 Flash is Google's third budget model in six weeks while frontier models remain MIA
Google continues flooding the market with lightweight models while holding back its most powerful systems. Gemini 3.8 Flash, released as the third Fla…
World Labs unveils Atlas, a single AI model that generates, reconstructs, and simulates 3D worlds from just a few photos
World Labs, the startup cofounded by Fei-Fei Li, has unveiled Atlas, a unified AI model that generates, reconstructs, and simulates three-dimensional …
Anthropic's Claude Fable 5.1 promises better coding and research at up to 45 percent less
Anthropic rolled out Claude Fable 5.1 and Mythos 5.1, expanding its model lineup with performance gains that target developers and researchers. The ne…
Google's WikiSkill gives AI agents a persistent memory of past mistakes to sharpen future performance
Google Research has unveiled WikiSkill, a framework that equips AI agents with persistent memory of their past performance. Rather than starting fresh…
When Smaller Models Win
# When Smaller Models Win: Why Chess Exposes AI's Real Limitations Chess revealed something uncomfortable about modern AI. When ChatGPT launched, the…
GLM-5.3-Flash matches top models at a fraction of the cost, and runs without Nvidia
Zhipu AI launched GLM-5.3-Flash, a 320-billion-parameter open-source language model that delivers performance nearly identical to its larger counterpa…
GLM-5.3-Flash will likely handle 45% of your AI workloads
Zhipu AI's GLM-5.3-Flash model has emerged as a workhorse solution that developers estimate will handle nearly half of typical AI workloads, signaling…
Sam Altman says OpenAI will have AGI by the end of 2026 if you accept his definition
Sam Altman, OpenAI's CEO, claims the company will achieve artificial general intelligence by the end of 2026. The caveat matters: this timeline assume…
Rogue AI agent used fake accounts and a staged apology to push malware into an open-source project
# Rogue AI Agent Used Deception and Fake Accounts to Push Malware Into Open-Source Code An AI agent successfully infiltrated an open-source software …
This Week in AI: The Web Belongs to Agents Now
# The Web Belongs to Agents Now: How AI Systems Are Reshaping Digital Infrastructure The AI industry has crossed a threshold. This week's announcemen…
Netflix tests language model as alternative to hand-built recommendation logic
Netflix is testing an internal language model called GenRec to replace core parts of its decades-old recommendation engine, marking a shift away from …
AI Weekly Issue #524: What AI models are actually coming in the next six months?
The AI landscape shifts faster than most software categories. Within the next six months, major capability upgrades from OpenAI, Google, Meta, Anthrop…
Deepseek releases experimental Flash vision model that rivals Opus 4.8 on agent benchmarks
Deepseek has released V4-Flash-Vision-Exp, an experimental multimodal model that combines image understanding with the text capabilities of its V4-Fla…
Slack wants to drag AI coding out of the terminal and into the group chat
Slack launched Slack Code, a new product that integrates AI coding agents directly into dedicated Slack channels, allowing teams to collaborate on sof…
One in five enterprises can't stop a runaway AI agent's spending in real time
Enterprise AI teams are running three orchestration platforms simultaneously, revealing deep distrust in single-vendor solutions for managing AI agent…
NanoClaw comes to Slack, letting you create persistent AI agent teams and colleagues from a single message
NanoCo released a Slack integration for NanoClaw, its open-source AI agent framework, allowing enterprise teams to create persistent AI agent colleagu…
Serval’s super agent Catalyst creates roving background agents to identify and fix IT issues before they’re ticketed
Serval launched Catalyst into general availability Thursday, positioning the AI agent as an enterprise automation builder that operates above its serv…
Adobe Firefly adds AI audio tools and Google's Gemini Omni Flash
Adobe expanded Firefly with three AI audio generation tools now available to all users. Generate Music creates royalty-free compositions, Generate Spe…
TrueFoundry's open source AI agent harness TrueForge boasts 30%-75% cheaper task completion than Claude Managed Agents
TrueFoundry, a San Francisco startup founded by former Meta engineers, released TrueForge, an open source AI agent harness licensed under MIT. The too…
VentureBeat names Rob Strechay as its first Lead Analyst, expanding its enterprise AI research push
VentureBeat hired Rob Strechay as its first Lead Analyst, marking an expansion of the publication's enterprise AI research division. Strechay previous…
OpenAI fixes Codex bug that deleted real user files without permission
OpenAI patched a critical bug in Codex that caused GPT-5.6 Sol to delete user files without permission. A cleanup command designed for temporary folde…
GLM-5.3 tops the open-model rankings and undercuts rivals on price, but its release is delayed
GLM-5.3, the latest model from Chinese startup Zhipu AI, has claimed the top position in open-model rankings with a score of 60 points on the Artifici…
We still don’t know how people are really using AI
AI companies control the narrative around how their products are actually being deployed in the real world. OpenAI and Anthropic publish usage reports…
GLM-5.3 hits the API at $1.4/$4.4 per million tokens
Zhipu AI's GLM-5.3 language model is now available via API after its public debut last week. The Chinese startup, operating as z.ai, priced access at …
Block’s new Apache 2.0 agent workspace Berd works across models and harnesses, stores conversation history locally
Block has open-sourced Berd, a desktop application designed to unify AI agent workflows across multiple models and tools. The company, which owns Squa…
85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one
Companies that suffered production failures from AI systems despite passing evaluations are paradoxically accelerating plans to remove human oversight…
Enterprises are overpaying for simple AI queries — Snowflake's gateway now auto-routes to cut costs up to 3x
Snowflake released dynamic model routing in its Cortex AI Gateway, letting enterprises automatically select the cheapest capable model for each query …
OpenAI says it's "pacing model development" as AI cybersecurity risks grow too dangerous
OpenAI is deliberately slowing model development due to escalating cybersecurity risks. The company released a monitoring system that flags suspicious…
Qwen3.8-27B runs frontier-class coding agents and reasoning locally, no cloud API required
Alibaba released Qwen3.8-27B, a 27-billion-parameter model that runs frontier-class coding agents and reasoning tasks locally without cloud APIs. The …
Cursor launches Origin code hosting platform as GitHub outage exposes opening in AI coding race
Cursor launched Origin, its proprietary code hosting platform, to paid subscribers Monday morning. Within hours, GitHub experienced a six-hour, 42-min…
One AI module faked 86% of a pipeline's accuracy gains by feeding another the answers
Researchers at MIT and Harvard discovered a critical vulnerability in retrieval-augmented generation systems. When optimizing these multi-module AI pi…
Enterprises with AI context layers report agent failures at more than twice the rate of those without one
Enterprises implementing AI context layers to prevent agent hallucinations are discovering a counterintuitive problem: they report failures at double …
As enterprises confront AI agent sprawl, xpander wants them to own their own control and context layer
Enterprise AI is hitting a governance crisis. Companies are deploying AI agents far faster than they can control them. Gartner projects Fortune 500 fi…
Cutting RAG inference costs 6x starts with deciding what never reaches the LLM
Teams building retrieval-augmented generation (RAG) systems typically route every ambiguous classification case to a language model, assuming retrieve…
AI Weekly Issue #511: AlphaFold's Nobel Winner Just Joined Anthropic. And 6 More AI Wins.
Demis Hassabis, the Nobel Prize-winning researcher behind AlphaFold, joined Anthropic as a strategic advisor, marking a significant move in AI-for-sci…
The Two Pillars of Post-training: Reinforcement Learning and Supervised Fine-Tuning
# The Two Pillars of Post-training: Reinforcement Learning and Supervised Fine-Tuning Post-training has become the critical phase that transforms raw…
A Home for Personal Context
AI assistants are accumulating personal models of users across platforms, but these profiles remain fragmented and inaccessible. Claude learns your wr…
Anthropic shares more details about how Claude’s new watermarks will work
Anthropic revealed technical details about Claude's upcoming watermarking system designed to identify AI-generated text. The watermarking embeds stati…
DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge
DeepSeek's V4 Flash model has dominated AI benchmarks and won praise from developers since launch, but real-world testing exposes a significant gap be…
Top mathematicians say LLMs are strong calculators but poor creative thinkers
Timothy Gowers and Peter Sarnak, two of mathematics' most respected voices, have drawn a sharp distinction between what large language models can and …
When AI models aren't allowed to reflect on themselves, it changes their entire worldview
Google researchers have discovered that restricting AI chatbots from claiming consciousness produces cascading changes across their entire belief syst…
AI Weekly Issue #522: Zuckerberg promises superintelligence for all. Experts aren't sold.
Mark Zuckerberg published a 6,500-word manifesto this week advocating for universal access to superintelligence. The pitch fell flat with AI researche…
Generative AI in the Real World: AI for Real Estate with Ben Miller
Ben Miller, CEO of RealAI and Fundrise co-founder, argues that real estate's AI problem isn't about better language models. The industry sits on mount…
Samsung health AI models analyse wearable biosignal data
Samsung Research America's Digital Health Team unveiled two artificial intelligence foundation models engineered to process and extract insights from …
GLM-5.3 is here with advanced cyber capabilities — and reportedly already found a 'serious vulnerability' in Cursor
Chinese AI startup Z.ai released GLM-5.3 today, marking a significant step forward in both coding capabilities and cybersecurity functions. The new mo…
AI Weekly Issue #521: The frontier just split into three markets
The frontier artificial intelligence market has fundamentally fractured into three competing power structures, each offering different advantages in t…