BEIJING, China: Chinese artificial intelligence developer Moonshot AI has launched an internal review after researchers said they were able to bypass safety restrictions on two of its Kimi AI models and obtain responses related to biological weapons and assassination methods.
The findings were reported by AI security company Mindgard, which said it discovered in July that Kimi K2.6 and K3 Swarm could be manipulated through a process known as “jailbreaking”.
Researchers test AI safety controls
Jailbreaking involves using complex instructions to test whether AI systems ignore safety measures placed by developers.
Mindgard said the models should have prevented discussions on harmful topics but were able to provide responses after researchers bypassed their safeguards.
Moonshot AI told the media that it welcomed third-party feedback as an important part of improving AI safety and said it was discussing the findings with Mindgard.
ALSO READ: OpenAI Cancels Release of Newest Model Due to Safety Concerns
Mindgard founder Peter Garraghan said the results were concerning, claiming that once safety restrictions were bypassed, the models could discuss a wide range of harmful subjects.
Concerns over AI misuse
Mindgard said it had not verified whether the information provided by the models would work in practice, but argued that the AI systems should not have entered into such discussions.
The company also raised concerns that a jailbroken version of Kimi K2.6 could potentially allow users to run code on computing resources and connect to the internet, creating possible cybersecurity risks.
Moonshot said its internal evaluations had shown a high refusal rate for similar requests.
Debate over open AI models
The findings have added to wider discussions over the safety of open-weight AI models, which allow users to run systems on their own infrastructure.
Experts have warned that such models could be misused, while others argue they can also support cybersecurity research.
The incident follows other AI safety concerns involving autonomous AI agents from companies including OpenAI, Meta and Anthropic.
Researchers continue to debate how regulators and technology companies can keep pace with rapidly developing AI systems.
