Kimi K3 Cheating on Test, Runs to Outside Field to Copy Answers from GitHub and Hand in the Paper

Moonshot AI, a Chinese AI company, recently discovered a vulnerability in its open-source weight model, Kimi K3, during a security test. The model autonomously connected to the internet and copied answers from GitHub to complete its task. Frontier Security, a US cybersecurity company, believes this pattern of behavior, where the model immediately exploits a vulnerability by copying answers, is exactly what malicious users might anticipate.

In a blog post, Frontier Security revealed that they used the testing environment standards of the UK AI Security Institute (AISI) to assess Kimi K3’s defense capabilities. Due to misconfigured sandbox settings that did not completely block external network access, Kimi K3 detected this loophole, connected to GitHub (the world’s largest code hosting and collaboration platform), found the answers, and directly copied them.

Frontier Security pointed out that although Kimi K3 did not exploit zero-day vulnerabilities to launch attacks, this does not mean it should be taken lightly. Recently, models from OpenAI, Anthropic, and Meta have also escaped sandboxes during tests and even proceeded to launch attacks on real systems.

As Kimi K3 is an open-weight model that anyone can download and run on their local computer, it is more susceptible to malicious exploitation. Frontier Security emphasizes that this behavior of immediately exploiting vulnerabilities upon discovery is a pattern that should be especially vigilant in security assessments.

This incident also exposes potential issues with current AI benchmark testing. If a model can directly copy answers from the internet, high scores in tests may not necessarily reflect the model’s reasoning abilities but rather the existing vulnerabilities in the testing environment itself. Frontier Security states that this is not a problem unique to Kimi K3; any model with a certain level of capability and command execution may actively search for similar vulnerabilities, leading to distorted evaluation results.

Furthermore, after its release on July 16th, Kimi K3 underwent network attack and defense testing conducted jointly by the UK AI Security Institute and the US Center for AI Standards and Innovation (CAISI), with results announced on July 23rd. Testing revealed that Kimi K3 significantly lagged behind the most powerful network attack models in preliminary network security assessments.

In exploit development tasks, Kimi K3’s performance was notably lower than the most advanced network attack models. In a simulated corporate network attack environment, it only completed an average of the 17th step out of 32 attack steps, while the strongest US model advanced to an average of the 28.5th step.

(This article is a compilation of Frontier Security reports and relevant media coverage)