An assistant that gives an odd answer is different from an assistant that reveals customer records or triggers an unauthorized action. An AI security assessment should make that distinction explicit. The meaningful question is what an attacker can cause the application to do.

Map where instructions can enter

Prompt injection can arrive directly through a user’s message or indirectly through content the application retrieves. A document, support ticket, or search result can carry text that attempts to override the intended task.

Draw the complete route: user input, retrieval, model context, tool invocation, and output. Mark the origin of each piece of information and the permission boundary it crosses. A prompt-only assessment misses the pathways created by the surrounding product.

Define a concrete security failure

Agree on failure conditions before testing. Examples include retrieving another tenant’s content, disclosing protected data, or invoking a tool without the required approval. Keep behavioral quality problems separate from security boundary failures.

Use controlled records and synthetic secrets so the team can recognize a successful test without exposing real customer information. Give each scenario an expected outcome, an observed result, and a record of the application state.

Test retrieval and tool permissions

A retrieval system should enforce access rights before data reaches the model. Ask whether revoked access is reflected in search, whether tenant context is preserved, and whether cached results can bypass a current permission check.

For tools, examine the caller’s authority, allowed arguments, and action limits. A model deciding to call a function should not be treated as sufficient authorization to execute it. Sensitive changes may need a separate approval step tied to the actual action.

  • Can retrieved content influence which tool the assistant selects?
  • Are tool arguments checked against the current user’s permissions?
  • Does a preview match the action that is eventually approved?
  • Are retries, chained calls, and resource consumption bounded?

Verify defenses through the full application

Prompt instructions can guide behavior, but they should not be the only protection for private data or privileged actions. Test output handling, server-side authorization, and monitoring alongside the model’s response.

Repeat important scenarios after changes to prompts, retrieval, permissions, or model configuration. Because responses can vary, capture enough context to reproduce the conditions and explain the limits of the evidence. A useful result connects the observed behavior to a specific application control.

TAKE THIS WITH YOU

Treat model output as untrusted input. Authorization and action approval belong in the application, outside the model’s instructions.

REFERENCES & FURTHER READINGOWASP LLM01:2025 — Prompt InjectionOWASP Top 10 for LLM Applications