Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why

AI Summary
Moonshot AI's Kimi K3 significantly underperformed compared to top U.S. AI models in offensive cyber tasks, scoring only 32 percent on ExploitBench. This disparity in performance may be linked to allegations that Kimi K3's development involved distilling technology from Anthropic's models.
From the source
The British AI Security Institute and the U.S. Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. Kimi K3 scored 32 percent on ExploitBench, compared with 76 percent for leading U.S. models, while its safeguards failed to block exploit development or simulated attacks. The gap between its strong general benchmark scores and weaker cyber performance also fits allegations that Moonshot AI distilled Anthropic's models. The article Kimi K3 trails frontier U
The full text couldn't be loaded here (the source may require a subscription).
View original at The Decoder