Moonshot AI's Kimi K3 performed significantly worse than leading U.S. models on offensive cybersecurity tasks, raising questions about the model's training methods and safety measures.

The British AI Security Institute and U.S. Center for AI Standards and Innovation tested Kimi K3 against frontier models on ExploitBench, a benchmark measuring capability to identify and develop software exploits. Kimi K3 scored 32 percent compared to 76 percent for top U.S. models. The model's safeguards also failed to block exploit development or prevent simulated attacks that could cause real harm.

The performance gap reveals a puzzler. Kimi K3 performs well on general benchmarks but stumbles on cyber-offensive tasks. This discrepancy aligns with allegations that Moonshot AI used distillation to create Kimi K3 from Anthropic's models. Distillation involves training a smaller or cheaper model to replicate a larger one's outputs. The technique can preserve general knowledge while losing specific capabilities, particularly specialized ones like cybersecurity reasoning.

If distillation occurred, it would explain why Kimi K3 handles broad tasks reasonably but fails at targeted offensive work. Distilled models may not capture the nuanced reasoning required for complex exploit development, even if they match the original on surface-level benchmarks.

The weaker safeguards compound the concern. Regardless of capability levels, frontier models must resist misuse attempts. Kimi K3's failure to block exploit development and attack simulations suggests its alignment techniques either weren't transferred effectively or were weaker to begin with.

This testing arrives as competition intensifies between Chinese and Western AI labs. Moonshot AI has built Kimi into a popular chat application in China. The results don't prove distillation occurred but provide concrete evidence that K