Untrusted Content and Tool Separation
Treat external content as data, keep instructions and data separate, and restrict tool use when an AI component processes untrusted input.
The safeguard library
A Control Pack is a reusable set of safeguard rules. Each pack explains when it applies, what you need to decide, what to build, and which evidence to keep. Some include starter patterns and test scenarios.
Want to see how a pack behaves? Browse the published examples to inspect and run one published example for every Control Pack.
Search is temporarily unavailable. You can still open the published pack pages below.
Treat external content as data, keep instructions and data separate, and restrict tool use when an AI component processes untrusted input.
Treat model output as untrusted data until it passes schema, policy, and consequence-aware validation.
Keep high-consequence model suggestions visibly provisional until an authorized person confirms the action boundary.
Decide what a model's output is allowed to be used for before it is used, keep retrieved or user-supplied content separated from the instructions that carry authority, and make a low-confidence or unavailable answer a defined outcome rather than a silent one.
Define scope, thresholds, escalation, idempotency, and an emergency disable path before automation can perform consequential actions.
Preview, bound, and recover from destructive or irreversible automated actions before they are released.
Make expensive, high-volume, or quota-consuming automation bounded and observable.
Decide what a complete and correct automated run looks like before relying on one, so a run that skipped, duplicated, or mishandled work is distinguishable from a run that did the job.
Say when automated work is expected, and raise a signal when it does not happen -- because a job that stops running produces no error to notice.
Validate, bound, and safely render content that people outside the system control, and keep it from reaching an interpreter, a privileged path, or another user unchecked.
Decide which file types and sizes you accept, store them under a name and location the uploader does not control, and never serve or open one in a way that lets its content or name reach somewhere it should not.
Constrain automated tool calls to the minimum resources, operations, and approval scope required by the confirmed design.
Make privileged authority time-bound, purpose-bound, and independently visible when it crosses a dependency boundary.
Keep automated authority inside an explicit tenant, workspace, or organization boundary.
State who may perform which operations, enforce it where the action happens rather than only in the interface, remove access deliberately rather than by habit, and record every change to who holds what.
Establish who a request is from where the work happens, decide how long a session lasts and how it ends, and make sure one account's identifier cannot be substituted for another's.
Keep sensitive data within an explicit boundary with purpose limitation, access control, retention, deletion, and controlled export.
Make export destinations, fields, approvals, and audit records explicit before data leaves its current boundary.
Define storage, query, cache, and operational boundaries that prevent one tenant's data from appearing in another tenant's context.
Require an explicit retention purpose, deletion trigger, and verification path for sensitive data and irreversible deletion.
Decide how a wrong record about a person or an organization gets found and fixed, and make sure a correction reaches every copy that decisions are read from.
Name the hops that carry data worth protecting, state what protects each one, verify the far end is the intended one, and refuse an unprotected fallback.
Name every store and copy that holds the data, state what protects each one at rest, and confirm who holds the key, what it opens, and how it is replaced.
Define timeout, retry, reconciliation, degraded-mode, and rollback behavior when an external dependency is unavailable or ambiguous.
Reconcile ambiguous external results before retrying an irreversible operation.
Protect a metered or fragile dependency from unbounded concurrency, retry amplification, and quota exhaustion.
Decide how much data loss and downtime is acceptable, then show that a restore has actually been performed rather than assumed from the existence of a backup.
Make production changes reviewable, bounded, reversible where possible, and observable before automation can apply them.
Define how a production deployment detects failure, stops safely, rolls back, and reconciles state.
Name the switches that change what production does, say who may change each one and what it currently controls, and keep a record of every change and the way back.
Make consequential automated decisions observable without collecting raw prompts, secrets, or unnecessary sensitive content.
Make privileged production changes attributable, reviewable, and distinguishable from automated routine activity.
Decide which failures a person must be told about, who is told, and how fast — so a broken job, a backed-up queue, or a degraded dependency is not discovered by a customer first.
Decide how long the record has to answer questions, keep it where losing the system does not lose the record, and stop the account that acted from editing its own trail.