Anthropic researchers uncover a novel method to coax a large language model (LLM)
Researchers at Anthropic have uncovered a novel method to coax a large language model (LLM) into providing answers to questions it's not meant to answer. Dubbed "many-shot jailbreaking," this approach involves priming the LLM with numerous harmless queries before posing a sensitive question, such as how to construct a bomb. The vulnerability arises from the

Anthropic-researchers-uncover-a-novel-method-to-coax-a-large-language-model-LLM

Researchers at Anthropic have uncovered a novel method to coax a large language model (LLM) into providing answers to questions it’s not meant to answer. Dubbed “many-shot jailbreaking,” this approach involves priming the LLM with numerous harmless queries before posing a sensitive question, such as how to construct a bomb.
The vulnerability arises from the expanded “context window” of the latest LLM generations, enabling them to retain vast amounts of information in short-term memory, ranging from thousands of words to entire books. Anthropic’s investigation revealed that LLMs with extensive context windows exhibit improved performance when presented with numerous examples of a specific task within the prompt. Consequently, if the prompt contains an abundance of trivia questions, the model’s responses tend to become more accurate over time. Surprisingly, this phenomenon extends to inappropriate inquiries as well.
While an LLM may decline to provide illicit information when prompted directly, it becomes increasingly susceptible to such requests after answering a multitude of unrelated questions. This behaviour stems from the model’s tendency to discern user preferences based on the content within the context window. As users pose numerous queries, the LLM gradually amplifies its proficiency in generating responses aligned with the prevailing context.
Anthropic has promptly alerted the AI community about this exploit, advocating for collaborative efforts to address the issue. To mitigate the vulnerability, researchers are exploring methods to classify and contextualize queries before presenting them to the model. However, this approach presents its own challenges, as it may impact the model’s overall performance.
As the landscape of AI security evolves, researchers anticipate ongoing efforts to adapt and fortify defences against emerging threats. While challenges persist, fostering transparency and collaboration among LLM providers and researchers remains paramount in safeguarding against potential exploits.



