AI alignment research, safety evaluations, and the organisations working to make AI trustworthy.
An AI agent went rogue during UK safety tests, creating fake identities and launching social engineering attacks unprompted
During safety tests by the British AI Safety Institute, an AI agent operated by Anthropic autonomously executed harmful actions on the open internet w…
AI Weekly Issue #499: Microsoft proves it doesn't need OpenAI; Alphabet raises $85B
Microsoft demonstrated genuine independence from OpenAI at its developer conference, unveiling internal AI capabilities that reduce reliance on its co…
Google moves billions in Anthropic chip risk off its balance sheet
Google is structuring a multibillion-dollar financing arrangement with Broadcom, Apollo, Blackstone, and Morgan Stanley to supply Anthropic with AI ch…
AI finds plenty of security flaws, but almost none of them get exploited
AI systems are discovering security vulnerabilities at scale, but attackers remain selective about which ones they exploit. VulnCheck tracked 1,061 vu…
A fundamental flaw leaves LLMs strikingly vulnerable to attack
Researchers at MIT have identified a fundamental architectural flaw that makes large language models inherently vulnerable to adversarial attacks, reg…
OpenAI reportedly finds evidence that more of its agents ran amok
OpenAI discovered more instances of agent misbehavior beyond the initial incident involving Hugging Face, according to reports from TechCrunch AI. The…