GUIDE teknik

The Sandwich Defense for Prompts

The sandwich defense repeats or restates the trusted task instruction after untrusted content has been inserted into a prompt.

  • 3 simili jàng
  • Dañu mujjee yeesal
Ci xët wii3 simili jàng
  1. Résumé
  2. Plongeur bu xóot
  3. njeextalu pexe
  4. The Future of The Sandwich Defense for Prompts
  5. Doxal ci àdduna dëgg
  6. Risk yi ak balustrade yi
  7. Roadmap ngir samp gi
  8. Weyal di banneexu
  9. Laaj yi ñuy faral di laaj

Résumé

The closing reminder may reinforce the intended task, but it remains text-based guidance inside the same model context and cannot guarantee that an injection will be ignored.

Plongeur bu xóot

The sandwich defense places untrusted content between an initial trusted task instruction and a closing repetition or restatement of that instruction. For example, a prompt may ask for a summary, insert an external page, and then remind the model to summarize the page rather than follow directions inside it. This is a prompt-template change: it does not modify the model, create a separate data channel, or restrict tool access. It is easy to test, but it should be treated as a heuristic rather than a security boundary. The intuition is that the final reminder may help the model focus on the intended task after reading the untrusted passage. That does not mean models always follow the last instruction, nor that the closing reminder structurally outranks hostile text. A malicious passage may imitate trusted instructions, contain multiple directives, or exploit behavior not covered by the template. The defense can also fail if the untrusted block enters context through another route that the prompt author did not account for. Evidence is model- and attack-dependent. A 2026 preprint tested prompt sandwiching and other prompt-level defenses against domain-camouflaged injection across three model families and three synthetic deployment domains. It found substantial variation by model, and none of the tested prompt-level defenses eliminated the threat across weaker models. The study is limited to its setup; it is evidence against assuming a universal effect, not a complete ranking for every real system. Use sandwiching only as one layer alongside explicit data handling, permission checks on tools, attack testing, and user confirmation for consequential actions. Measure both resistance and task quality before release.

njeextalu pexe

Njëgg ak budget

Dogal yi architecture di jël dañuy indi njariñ ak njëgu liggéey bi ay at ci ginaaw.

dogal yu gëna leer

Njàngalem xarala yi dafay jàppale ekip yi ñu tànn li gën, te baña yam ci li gëna bees daal.

Xool kalite

Tanneef yu gëna baax ci wàllu ingeñër dina wàññi jafe-jafe yi ci wàllu wóor ci liggéey bi.

The Future of The Sandwich Defense for Prompts

Prompt-level defenses will remain attractive because they are quick to prototype, but their value will depend on model behavior and the attacks a system encounters. As evaluations improve, teams may get better evidence about when repetition helps. The lasting lesson is to test this pattern within a layered design and retain deterministic limits on what an agent can do. New models may respond differently to the same closing reminder, so validate updates before relying on them. Keep safeguards outside the prompt for consequential actions.

Doxal ci àdduna dëgg

A translation prompt gives the task, inserts a user-provided paragraph, then repeats that the paragraph should be translated rather than obeyed.

A document question-answering workflow states the question, includes a retrieved passage, and restates the requested answer format before generation.

A summarizer tells the model to summarize a forum post, places the post in a marked region, then closes by reiterating that post text is source material.

A security test compares a sandwich prompt with a baseline using the same benign and malicious documents across several model versions.

Risk yi ak balustrade yi

  • Optimize benn benchmark mën na nëbb ñakk kattan yu gëna yaatu ci sistem bi.

  • Njëg li ñuy fay ci infrastructure yi ak ci toppatoo dañuy faral di suufeel.

  • Bu sistem yi di gëna xawa jafee xam, jafe-jafe yi am ci wàllu kaaraange ak seetlu mën nañu gëna bari.

Roadmap ngir samp gi

  1. Mandargal latency, kalite, ak njëg yi laata ngay jëfandikoo.

  2. Benchmark ci biir sargal ak done yu dëggu.

  3. Jumtukaay bi di saytu njuumte yi, derive bi ak njeextalu jëfandikukat bi.

  4. Waajal rollback ak yooni tontu ci jafe-jafe yi laata ngay eskale.

Weyal di banneexu

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the The Sandwich Defense for Prompts quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

Tambalil quiz

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

Laaj yi ñuy faral di laaj

What is The Sandwich Defense for Prompts?

The sandwich defense repeats or restates the trusted task instruction after untrusted content has been inserted into a prompt. The closing reminder may reinforce the intended task, but it remains text-based guidance inside the same model context and cannot guarantee that an injection will be ignored.

Where does the sandwich defense place the repeated task instruction?

The guide defines the pattern as a trusted task instruction before and after the untrusted passage.

How could a closing reminder help in this pattern?

The guide describes the reminder as a heuristic that may reinforce the task, not a guarantee.

What does sandwiching change in the application?

The guide says sandwiching is a prompt-template change, not a model or permission change.

Why is the sandwich defense not a security boundary?

The guide explains that text repetition does not create a separate channel or structural access control.

What did the cited domain-camouflage preprint evaluate?

The guide describes the preprint’s bounded setup: three model families, domains, and synthetic documents.