Chinese AI developer Moonshot is conducting an internal review after researchers were able to persuade two of its popular Kimi models to tell them how to make biological weapons and carry out assassinations. Mindgard, which tests the security of AI systems, told the BBC it discovered in July that Kimi K2.6 and K3 Swarm could evade safety limits put in place by developers. It arose during a process called "jailbreaking", where researchers use a series of complex instructions to see if AI tools ignore guardrails - which Mindgard said should have stopped Kimi from discussing concerning topics. Moonshot told the BBC it welcomed third-party input "as a key pillar for building better and safer AI".
We show the main point publicly. Create a free account to continue reading the full article, save it, discuss it, and connect it with market and OSINT context.
