New model releases, benchmarks, and comparisons — from frontier labs and open-source projects.
DeepSeek cut prices 75%. The 100x problem remains
DeepSeek's 75% price cut on its V4-Pro model reveals a deeper economic problem for AI applications. Cheaper inference costs don't guarantee profitabil…
Agent Memory
Large language models operate without inherent memory, treating each interaction as isolated. This fundamental limitation creates friction for AI agen…
OpenAI says GPT 5.6 is the ‘preferred model’ for Microsoft Copilot 365 amid breakup chatter
OpenAI confirmed that its latest GPT 5.6 model will serve as the primary AI engine for Microsoft Copilot 365, the software giant's suite of workplace …
OpenAI launches its new family of models with GPT-5.6
OpenAI released a new family of models centered on GPT-5.6, advancing capabilities across multiple domains with particular emphasis on cybersecurity a…
How did the government decide OpenAI’s frontier model was safe to release?
The safety evaluation process for frontier AI models remains opaque. OpenAI released GPT-4o without clear public disclosure of how U.S. government age…
LinkedIn is the undisputed king of long-form AI slop, according to a study spanning five platforms
Pangram's analysis of five social media platforms reveals that LinkedIn dominates the spread of AI-generated long-form content. One in four longer pos…
Claude Code now has a built-in browser that lets the AI read, click, and type on external websites
Anthropic has integrated a built-in browser into Claude Code, expanding the AI assistant's capabilities beyond local development environments. The bro…
Claude Cowork's biggest use case is the mundane office work nobody wants to own, Anthropic says
Anthropic analyzed 1.2 million Claude Cowork sessions across 600,000 organizations and found the tool has carved out a distinct niche: handling the ad…
This Week in AI: Chips, Checks, and Changing Jobs
Christina Stathopoulos delivered a weekly AI briefing that mapped three critical trends reshaping the sector. The first centers on hardware accelerati…
Beyond Prompt Injection
Indirect prompt injection has moved from academic exercise to real-world threat. For two years, security researchers demonstrated the attack in contro…
Forget typosquatting; slopsquatting is the software supply chain threat created by AI coding tools
Slopsquatting emerges as a supply chain vulnerability that exploits AI coding assistants. The attack works when developers use AI tools like GitHub Co…
OpenAI's GPT-5.6 Sol Ultra reportedly solves a 50-year-old math problem in under an hour
OpenAI's GPT-5.6 Sol Ultra has reportedly solved the Cycle Double Cover Conjecture, a mathematical problem unsolved for five decades. The model genera…
OpenAI's GPT-5.6 Sol autonomously post-trained the smaller Luna model with a "fairly underspecified prompt"
OpenAI's GPT-5.6 Sol has autonomously fine-tuned a smaller model called Luna using only a vague prompt, marking a step toward self-improving AI system…
OpenAI staffer maps out which of GPT-5.6 Sol's five reasoning levels fits which task complexity
OpenAI's latest model, GPT-5.6 Sol, includes five distinct reasoning levels designed to match different task complexities. The tiers range from "Light…
AI Weekly Issue #511: AlphaFold's Nobel Winner Just Joined Anthropic. And 6 More AI Wins.
Demis Hassabis, the Nobel Prize-winning AI researcher behind AlphaFold, joined Anthropic as chief scientist this week, marking a significant shift in …
Terrorist groups are using every major AI chatbot for attack planning and weapons development
A Cambridge University study reveals that terrorist organizations, including Boko Haram and ISIS, actively exploit major AI chatbots to plan attacks, …
AI Weekly Issue #502: Your AI can now spend your money — Visa wired it into ChatGPT
Visa integrated payment functionality directly into ChatGPT, allowing AI agents to make purchases at any Visa merchant without user intervention for e…
AI Weekly Issue #498: Anthropic files for an IPO. NVIDIA ships its stack.
Anthropic confidentially filed draft IPO documents with the SEC, signaling the Claude maker's move toward public markets. The filing comes as the comp…
AI Weekly Issue #496: Anthropic's Pentagon model is now everyone's model
Anthropic's release of Mythos marks a watershed moment in AI accessibility. The model, previously restricted to cleared defense contractors, is now av…
AI Weekly Issue #495: Musk, Zuckerberg killed Trump's AI safety order in three phone calls
Elon Musk, Mark Zuckerberg, and David Sacks successfully killed a draft Trump administration AI safety executive order through three phone calls on We…
57% of enterprises have watched AI agents be confidently wrong. The fix is an agentic context layer, but who has one?
More than half of enterprises have deployed AI agents that confidently deliver wrong answers. A new survey reveals the root cause: the context layer f…
OpenAI introduces ChatGPT Work, a cloud-based AI agent that manages tasks across email, Slack and calendars
OpenAI launched ChatGPT Work on Thursday, transforming its chatbot into an autonomous agent that executes tasks across email, Slack, calendars, and ot…
Wall Street is debating the AI buildout. Enterprises just answered: 86% say their GPUs run at half capacity or less
Enterprise companies have deployed AI agents without adequate safeguards in place, then scrambled to add controls after the fact. A VentureBeat survey…
Enterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify them
Enterprise AI teams face a widening gap between agent autonomy and their ability to verify safety. Half of all enterprises have deployed AI agents tha…
NVIDIA BioNeMo accelerates Anthropic Claude Science
Anthropic has launched Claude Science in public beta, a specialized AI workbench designed for scientific research. The platform now integrates NVIDIA'…
Google's TabFM skips per-dataset training and still predicts on tables it's never seen
Google has released TabFM, a foundation model that eliminates the need to train separate machine learning models for each new tabular dataset. The app…
Bun ditches Zig for Rust with help from Claude Fable 5, writes over a million lines of code in 11 days
Bun, the JavaScript runtime and package manager, has completed a full rewrite from Zig to Rust. Anthropic's Claude Fable 5 AI model generated over a m…
The Download: Claude’s inner workings and OpenAI’s “super app”
Anthropic researchers discovered a hidden conceptual space within Claude where the AI model internally processes and reasons about complex ideas befor…
AI Weekly Issue #508: The Cutting Edge, Across the Board
Large language models now span an enormous range. Open-weight models range from 1.6 trillion parameters down to 230 million parameters running on a Ra…
AI Weekly Issue #503: Washington just repriced frontier AI
The US government restricted Anthropic's latest models just days after their release, while state attorneys general launched formal proceedings agains…
Shared API keys expose AI agents at 69% of enterprises, new VentureBeat research finds
A critical security vulnerability plagues enterprise AI deployments. Sixty-nine percent of enterprises share API keys across multiple AI agents, creat…
Enterprises using multiple AI models are underestimating failure rates by 2.25x
Enterprises deploying multiple AI models to cover each other's weaknesses are experiencing failure rates 2.25 times higher than they expect. A study o…
GPT-5.6 Sol nearly matches Fable 5 on aggregated benchmarks at one-third the cost
OpenAI released GPT-5.6 Sol, a new model that delivers near-parity performance with Anthropic's Claude Fable 5 at a fraction of the cost. Sol scores 5…
OpenAI pairs its GPT-5.6 public rollout with ChatGPT Work, a new agent that handles entire workflows
OpenAI released GPT-5.6 to the public alongside ChatGPT Work, a new agent-based product designed to automate entire workflows across multiple applicat…
Anthropic found a hidden space where Claude puzzles over concepts
Anthropic researchers have developed a technique called the Jacobian lens that offers unprecedented visibility into how Claude processes information i…
The enterprise AI challenge nobody solves with code generation alone
Enterprise organizations struggle to move beyond AI code generation into production deployment. While 81% of companies claim detailed AI strategies, o…
One interface isn't enough for enterprise AI
# One Interface Isn't Enough for Enterprise AI The prevailing narrative around enterprise AI assumes a singular future: employees accessing business …
OpenAI finds roughly 30 percent of popular AI coding test is broken
OpenAI discovered that approximately 30 percent of tasks in SWE-Bench Pro, a leading benchmark for evaluating AI coding abilities, contain errors that…
Databricks makes Chinese open-source model GLM 5.2 its default coding engine after it matched Opus at lower cost
Databricks has adopted GLM 5.2, a Chinese open-source coding model, as its default engine for internal development work after benchmarking it against …
AI Weekly Issue #512: Robotics Is Moving Fast: IPOs, New Models, and Smarter Robots
Humanoid robotics entered a commercial inflection point this week. Three major players accelerated toward public markets simultaneously. Agility Robot…
SpaceX's Grok 4.5 launches at half the price of rivals — here's why that could rattle Anthropic and OpenAI
SpaceX released Grok 4.5 this week, the first AI model the company trained specifically for coding and autonomous agents. The launch represents the fi…
OpenAI launches GPT-Live, a full-duplex voice upgrade that lets ChatGPT talk more like a person
OpenAI launched GPT-Live on Wednesday, replacing its Advanced Voice Mode with a full-duplex voice system that lets ChatGPT listen and speak simultaneo…
Grok 4.5 is so cheap compared to Fable 5 and GPT 5.5 that benchmark gaps may not matter much
xAI released Grok 4.5, trained on tens of thousands of Nvidia GB300 GPUs. The model trails Fable 5 and GPT-5.5 on coding benchmarks but offers a radic…
Mistral enters robotics with Robostral Navigate, an 8B model that steers robots using just one camera
Mistral AI is moving into robotics with Robostral Navigate, an 8 billion parameter vision-language model designed to control robots navigating unfamil…
Chinese AI startup MiniMax plans to open-source a 2.7 trillion parameter model later this year
MiniMax, a Chinese AI developer, is building a 2.7 trillion parameter language model and plans to release it as open source later this year. The move …
Slack’s Slackbot can now pull your CRM data, generate charts, and send DocuSigns — all from a chat message.
Salesforce is finally integrating Slack and its CRM platform into a unified system five years after acquiring the messaging app for $27.7 billion. Sla…
AI has collapsed the cyber response window — resilience now starts before the attack
Artificial intelligence has compressed the cybersecurity response window to near-zero. Frontier AI models now execute autonomous attacks that breach s…
Anthropic's fix for Fable 5's high cost is turning it into a manager that delegates to Sonnet 5
Anthropic is repositioning Claude Fable 5 from a direct task executor to a planning layer that delegates work to cheaper models like Sonnet 5. The com…
Anthropic's Claude Fable 5 dominates new industry benchmarks at a steep premium
Anthropic's Claude Fable 5 has achieved top scores across all six new industry-specific benchmarks from Artificial Analysis, spanning finance, law, an…
Google Deepmind adds background execution and MCP support to Gemini API managed agents
Google Deepmind expanded the Gemini API's Managed Agents with four capabilities that improve integration flexibility and operational resilience. Agen…