Apple's bug bounty program has become so overwhelmed with AI-generated submissions that legitimate security researchers struggle to report real vulnerabilities. Italian startup Bynario discovered a serious macOS flaw worth up to $200,000 on the black market but faced barriers to reporting it through official channels.
Apple has capped the number of submissions each researcher can make per day because fabricated bug reports from AI tools are flooding the review pipeline. This throttling creates a perverse incentive structure. Researchers with genuine discoveries face submission limits while the company wastes resources sorting through synthetic noise.
The Bynario case exposes a fundamental problem with bug bounty programs in the AI era. When submission systems lack proper filtering, automation lowers the cost of spam enough to degrade service for everyone. Researchers who would otherwise follow responsible disclosure find the friction unbearable and may sell vulnerabilities to other buyers instead.
This isn't a minor inconvenience. A $200,000 vulnerability represents serious risk to macOS users. The vulnerability should have reached Apple's security team quickly through official channels. Instead, it hit friction from an overflow of garbage reports. The startup eventually reported the flaw through other means, but the program's dysfunction created unnecessary delay.
Apple faces a scaling problem that many platforms encounter when AI tools become cheap enough to weaponize against their own systems. Distinguishing legitimate submissions from AI-generated nonsense requires human review, which defeats the purpose of automation. Adding CAPTCHA-style challenges or verification requirements works but adds friction for real researchers.
The broader lesson applies across the security industry. Bounty programs work when they maintain low friction for legitimate researchers. Once that friction rises due to spam, researchers vote with their feet and take vulnerabilities elsewhere. Apple will need better filtering mechanisms, whether through ML-based detection, reputation systems, or researcher verification, to restore the program's credibility. Until then,