Authority think tank test: Kimi K3 performance significantly lower than US models

On Wednesday, July 23, a prestigious think tank in Europe and America conducted initial tests and discovered that the latest model Kimi K3 from the Chinese startup Moonshot AI significantly underperformed compared to the leading cutting-edge AI models in the United States.

Moonshot AI unveiled Kimi K3 on July 16 and planned to release an open-source version on July 27. In the midst of intense competition in the field of artificial intelligence between China and the United States, this model has garnered global attention for outperforming several American AI companies in certain industry benchmark tests.

After a joint evaluation by the UK AI Security Institute (UK AISI) and the US Center for AI Standards and Innovation (CAISI), it was pointed out that in the preliminary network capability assessment, Kimi K3’s network capabilities were significantly lower than the latest cutting-edge models in the United States, especially in terms of stability and success rate.

The evaluation focused on two aspects: testing Kimi K3’s ability to crash reproduction, arbitrary reading and writing, control flow hijacking, and code execution through 41 V8 vulnerabilities using ExploitBench; and simulating enterprise networks to test its ability to autonomously complete multi-step attack chains and observe if security barriers would prevent attack behaviors.

Results showed that Kimi K3 outperformed Chinese models like GLM-5.2 and occasionally completed full attack chains. However, it still lagged behind the advanced closed-source models from the United States in terms of stable code execution, attack success rate, and overall capability. GLM-5.2 is a flagship large-scale language model released by Zhipu AI.

In terms of developing exploit programs, Kimi K3’s performance was significantly lower than the latest cutting-edge network models in the United States. The top American models could build complete exploit chains in about half of the samples, while Kimi K3 and GLM-5.2 could discover and reproduce many vulnerabilities but failed to complete the process from vulnerability to arbitrary code execution.

Specifically, Kimi K3 only completed the 17th step on average out of 32 steps in the attack path, while the most powerful network models in the United States averaged a completion of 28.5 steps. GLM-5.2 only reached an average of the 11th step.

Furthermore, the comprehensive scores of the top American models were significantly higher, approximately 44 percentage points higher than Kimi K3 and about 51.8 percentage points higher than GLM-5.2.

In terms of attack success rate, in 10 attempts, Kimi K3 successfully completed testing the “The Last Ones” (TLO) network attack scope once under a limit of 100 million tokens.

This indicates that after gaining initial network access rights and receiving commands, Kimi K3 was able to autonomously attack small, weakly defended, and vulnerable corporate systems.

However, solving the TLO problem is no longer an exclusive capability of a few models. It was pointed out by the think tank that in previous tests, four publicly released closed-weight models have addressed the TLO problem, with the best performance achieving success rates of 6 to 7 times out of 10 attempts. Kimi K3 only succeeded once in 10 attempts under the standard 100 million token limit.

US Secretary of Commerce, Howard Lutnick, posted on social media platform X on Wednesday, stating, “The latest report from CAISI shows that Kimi K3 still lags behind the leading cutting-edge artificial intelligence models in the United States.”

“The reason why the United States can maintain a leading position in cutting-edge artificial intelligence is because we have the world’s greatest innovators and technical experts,” he wrote.

During the mid-July Pennsylvania Defense and Innovation Summit, CIA Director John Ratcliffe also stated that the United States holds an advantage in its system over communist-ruled China. He mentioned that individuals in the Chinese system can only obey the party-state.

“They steal, they replicate, and they subsidize. That’s their ‘success’ model,” he said.

The CIA director emphasized that the United States cannot afford to lose its leading edge over Beijing in high-end technology, and the only way to maintain this advantage is by not getting stuck in politically correct quagmires.