Prompt injection

Instructions in user or retrieved content that attempt to redirect a model or tool-enabled workflow.

Definition

Instructions in user or retrieved content that attempt to redirect a model or tool-enabled workflow.

Practical test

Can user, webpage, document, email, or retrieved content introduce instructions that compete with the workflow’s trusted objective and tool policy?

Example

A webpage telling a browser agent to upload credentials is untrusted content, even when it resembles a system instruction.

What to record

For prompt injection to be operational rather than a reassuring label, record the untrusted instruction, trust boundary, blocked action, detector, and resulting stop or escalation. If those facts cannot be observed or tested, do not treat the term itself as evidence that the workflow is safe.