Prompt Leaking
Prompt leaking is an attack that tricks a model into revealing its hidden system prompt or other confidential instructions.
Description
System prompts should be treated as potentially discoverable, so they should not contain secrets such as API keys.
Sources
- Perez & Ribeiro (2022). Ignore Previous Prompt: Attack Techniques For Language Models.
Cite this entry
Protologue. (2026). Prompt Leaking. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0089). https://protologue.com/t/prompt-leaking/
BibTeX
@misc{protologue_prompt_leaking,
title = {Prompt Leaking},
author = {{Protologue}},
year = {2026},
howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
note = {Entry PTL-0089},
url = {https://protologue.com/t/prompt-leaking/}
}New terms, as they're coined
One email a week with the newest techniques, each defined and sourced like this entry.
Delivered through Beehiiv. Unsubscribe from any issue.