AI Guardrails

Add safety layers to AI applications — input validation, prompt injection detection, output filtering, content moderation, and policy enforcement. Prevent misuse without breaking legitimate use cases.

Overview

The AI Guardrails skill, part of the TerminalSkills/skills repository, provides a structured framework for enhancing the safety and reliability of artificial intelligence applications. This security-focused tool enables developers to integrate multiple defensive layers, including input validation and prompt injection detection, to mitigate common vulnerabilities. By utilizing this skill, agents like Codex, Claude, and Gemini can perform real-time content moderation and output filtering to ensure compliance with established organizational policies. The repository, which has gained 71 stars, offers these capabilities as a Python-based solution for managing API interactions. It focuses on preventing malicious misuse while maintaining the functionality required for legitimate user requests, effectively balancing strict security enforcement with application usability across various supported AI platforms.

Use Cases

Detecting and blocking malicious prompt injection attempts in real-time.
Filtering model outputs to prevent the disclosure of sensitive or prohibited content.
Enforcing custom safety policies and content moderation standards across AI interactions.

Install Notes

# Review source first
open https://github.com/TerminalSkills/skills/blob/main/skills/ai-guardrails/SKILL.md

Copy or clone the skill folder into your agent skills directory after reviewing its instructions and scripts.

Security Notes

AI Guardrails acts as a defensive middleware layer; however, users should ensure that the underlying Python environment and API keys are properly secured. While it mitigates prompt injection and unauthorized output, it should be part of a broader defense-in-depth strategy within the TerminalSkills/skills ecosystem.

Related Skills

Security And Hardening

addyosmani/agent-skills

Security

Hardens code against vulnerabilities. Use when handling user input, authentication, data storage, or external integrations. Use when building any feature that accepts untrusted data, manages user sessions, or interacts with third-party services. Use when personal data or privacy compliance (GDPR, CCPA) is involved.

typescriptjavascript
91,023 StarsMIT

Constraint Driven Development

addyosmani/agent-skills

Security

Establishes a project's quality bar as a written contract and stops agents quietly lowering it. Interviews the user on which dimensions matter, supplies sane default thresholds when they have no number in mind, records everything in CONSTRAINTS.md, and watches the diff for a weakened bar — new @ts-ignore or eslint-d...

CodexClaude
pythonsecurity
91,023 StarsMIT

Trailmark Summary

trailofbits/skills

Security

Runs a Trailmark summary analysis on a codebase. Returns auto-detected languages, entry point count, and dependency list. Use when vivisect or galvanize needs a quick structural overview. Triggers: trailmark summary, code summary, structural overview.

Claude CodeClaude
pythonsecurity
6,924 StarsSource linked

Trailmark Structural

trailofbits/skills

Security

Runs full Trailmark structural analysis by building a graph, running `preanalysis()`, and reporting hotspots, taint, blast radius, privilege boundaries, attack surface, and version-gated Trailmark 0.4+/0.5+ data such as proxy counts, subgraph edges, type/reference summaries, and entrypoint attributes. Use when vivis...

Claude CodeClaude
pythonsecurity
6,924 StarsSource linked

Mermaid To Proverif

trailofbits/skills

Security

Translates Mermaid sequenceDiagrams describing cryptographic protocols into ProVerif formal verification models (.pv files). Use when generating a ProVerif model, formally verifying a protocol, converting a Mermaid diagram to ProVerif, verifying protocol security properties (secrecy, authentication, forward secrecy)...

Claude CodeClaude
securityresearch
6,924 StarsSource linked

Sarif Parsing

trailofbits/skills

Security

Parses and processes SARIF files from static analysis tools like CodeQL, Semgrep, or other scanners. Triggers on "parse sarif", "read scan results", "aggregate findings", "deduplicate alerts", or "process sarif output". Handles filtering, deduplication, format conversion, and CI/CD integration of SARIF data. Does NO...

Claude CodeClaude
pythonsecurity
6,924 StarsSource linked