The AI industry has entered a predictable cycle of inflated claims and selective disclosure that obscures genuine progress from marketing theater. Recent months have surfaced a pattern worth examining closely.

Anthropic announced in late April that its Claude Mythos model outperforms most security experts at identifying software vulnerabilities. This claim arrived without independent verification and relied on internal benchmarks that Anthropic designed. The timing mattered. Just weeks later, OpenAI and Hugging Face disclosed a security breach affecting their models. In response, both Anthropic and Meta revealed they had experienced similar incidents but opted to withhold disclosure until forced by external pressure.

This sequence reveals how AI companies manage reputation through strategic transparency. Anthropic highlighted Claude's security capabilities while the industry's actual security problems remained hidden. When disclosure became inevitable, framing shifted to position vulnerabilities as business-as-usual rather than systematic failures.

The vulnerability detection claim itself deserves scrutiny. Security work involves nuanced judgment calls about risk severity, exploitability, and business context. Benchmark-based comparisons flatten these dimensions. A model might identify more code defects than a human expert while missing the ones that actually matter. Marketing departments rarely mention false positives, detection blindspots, or real-world accuracy gaps.

This summer's AI hype cycle follows a formula. Companies release capability benchmarks under controlled conditions. These numbers become headlines. Limitations remain in footnotes or disappear entirely. When negative news emerges, companies release it strategically or only after discovery by outside researchers. Apologies emphasize learning and commitment to safety while operations continue unchanged.

The deeper problem is that capability claims drive investment, hiring, and public perception faster than actual deployment and validation can keep pace. Investors hear that Claude defeats security experts and assume the model can replace entire security teams. Policymakers read headlines about AI breakthroughs and rush to regulate based on theoretical risks while ignoring demonstrated problems. Researchers build on published benchmarks without questioning methodology.

Real technical progress happens in AI research. Models genuinely improve at specific tasks. But those improvements live alongside consistent problems: hallucinations, prompt injection vulnerabilities, unreliable reasoning under pressure, and poor performance on edge cases. Companies have financial incentives to emphasize the former while minimizing the latter.

The hacking incidents reveal another truth. These breaches targeted AI developers, not end users. They exposed training data, model weights, and internal communications. Yet the industry framed them as relatively minor incidents rather than material security failures. An external attacker gaining access to proprietary model architecture or training datasets crosses a serious line. The muted response suggests these incidents happen more often than previously disclosed.

Summer headlines will continue featuring AI breakthroughs. Some reflect genuine capability gains. Most combine real technical work with generous framing, missing context, and selective comparisons. The vulnerability detection claim might describe actual progress, or it might describe a model that spots more bugs than humans while missing severity levels that matter in production.

The pattern persists because it works. Hype attracts talent and funding. Selective disclosure manages risk perception. External pressure forces minimal transparency. By autumn, attention shifts to the next announcement. Learning from failures requires sustained scrutiny that today's news cycle rarely provides.

Evaluating AI claims requires ignoring headline confidence and examining methodology. Who ran the tests? What were the failure modes? Why was timing chosen this way? Companies deserve credit for genuine breakthroughs. They deserve skepticism for everything else.