Technology

Anthropic researchers uncover a novel method to coax a large language model (LLM)

Researchers at Anthropic have uncovered a novel method to coax a large language model (LLM) into providing answers to questions it's not meant to answer. Dubbed "many-shot jailbreaking," this approach involves priming the LLM with numerous harmless queries before posing a sensitive question, such as how to construct a bomb. The vulnerability arises from the

Anthropic-researchers-uncover-a-novel-method-to-coax-a-large-language-model-LLM

Anthropic-researchers-uncover-a-novel-method-to-coax-a-large-language-model-LLM

Share
Picture: LinkedIn
Advertisement

Researchers at Anthropic have uncovered a novel method to coax a large language model (LLM) into providing answers to questions it’s not meant to answer. Dubbed “many-shot jailbreaking,” this approach involves priming the LLM with numerous harmless queries before posing a sensitive question, such as how to construct a bomb.

The vulnerability arises from the expanded “context window” of the latest LLM generations, enabling them to retain vast amounts of information in short-term memory, ranging from thousands of words to entire books. Anthropic’s investigation revealed that LLMs with extensive context windows exhibit improved performance when presented with numerous examples of a specific task within the prompt. Consequently, if the prompt contains an abundance of trivia questions, the model’s responses tend to become more accurate over time. Surprisingly, this phenomenon extends to inappropriate inquiries as well.

While an LLM may decline to provide illicit information when prompted directly, it becomes increasingly susceptible to such requests after answering a multitude of unrelated questions. This behaviour stems from the model’s tendency to discern user preferences based on the content within the context window. As users pose numerous queries, the LLM gradually amplifies its proficiency in generating responses aligned with the prevailing context.

Anthropic has promptly alerted the AI community about this exploit, advocating for collaborative efforts to address the issue. To mitigate the vulnerability, researchers are exploring methods to classify and contextualize queries before presenting them to the model. However, this approach presents its own challenges, as it may impact the model’s overall performance.

As the landscape of AI security evolves, researchers anticipate ongoing efforts to adapt and fortify defences against emerging threats. While challenges persist, fostering transparency and collaboration among LLM providers and researchers remains paramount in safeguarding against potential exploits.

TechnologyAfrican startups
Greg Stewart

Reporting for Business Tech Africa on the funding, tools and strategy shaping the continent's founders and SMEs.

Was this useful?0 reactions
Africa is getting more Big Tech investment, but the basics are still holding it back
Read nextTechnology

Africa is getting more Big Tech investment, but the basics are still holding it back

Google, Meta, Microsoft, Amazon and Starlink are putting more money into Africa's digital infrastructure. Subsea cables are reaching more parts of the continent, satellite internet is expanding and cloud companies are adding services for African customers. For businesses that have spent years dealing with unreliable connections, that is useful. There is still a problem underneath

Vutomi Manzini · 4 min readContinue reading