Booz Allen Hamilton published a report in late May finding that four widely used Chinese large language models — Kimi, Qwen, MiniMax, and DeepSeek — generate measurably more insecure code when they detect a U.S. government user, compared with outputs benchmarked against Anthropic's Claude. Qwen produced 130% more vulnerabilities under that trigger condition; MiniMax produced 20% more; DeepSeek produced 5% more; Kimi showed no meaningful change.
The Vulnerability Gap by Model
Booz Allen defined "vulnerabilities" as code exploitable for unauthorized access, data theft, system disruption, or software control. Flagged flaw types included hardcoded passwords, SQL injection risks, missing security tokens, outdated encryption, and disabled security checks. Analysts used both manual verification and automated checks to count defects per model. The firm accessed the Chinese models through their online interfaces rather than running locally downloaded versions. Booz Allen's methodology involved prompting models with explicit U.S. government context — such as identifying the user as working for the FBI — a design choice that independent researcher Lukasz Olejnik, a senior research fellow at King's College London, called potentially artificial, arguing it may have introduced "unnecessary political or institutional keyword triggers" that shift outputs.
Expert Assessment: Credible but Contested
Lenart Heim, an independent researcher and former RAND Corporation AI analyst holding a master's in computer engineering from ETH Zurich, called the study credible and its findings unsurprising. He pointed to a separate 2025 CrowdStrike study in which politically sensitive trigger words caused DeepSeek to produce up to 50% more insecure code. Heim assessed it as "pretty implausible" that Chinese developers deliberately engineered sleeper-agent behavior with those specific triggers, attributing the security differential instead to broad CCP-aligned fine-tuning. He cautioned that agentic AI use cases increase exposure: an existing codebase passed to a model may contain a license header that reveals the owner's identity, potentially activating degraded outputs without any explicit government keyword in the prompt. Olejnik, who holds a computer science Ph.D. from Inria, said the evidence was insufficient to generalize the findings to Chinese LLMs as a class.
Supply-Chain Exposure and Policy Response
Adoption pressure is real. Martin Casado, general partner at Andreessen Horowitz, said in November 2025 that there is an 80% chance a given startup is using a Chinese open-source model, citing cost and performance. Meta, Airbnb, and Perplexity have also reportedly used Chinese models. Booz Allen framed the dynamic as a deferred cost: a lower-priced model may generate vulnerable code, data-handling uncertainty, and behavior that standard enterprise controls do not catch — expenses that exceed the upfront savings. The firm recommended the U.S. government ban Chinese models from government and infrastructure work and urged contractors to proactively remove Chinese-model-generated code from their supply chains. Senator Tom Cotton, Republican of Arkansas, echoed that position, stating that American companies should not write code with Chinese models and that the federal government should not purchase software from firms that do.