AI Guardrails

Add safety layers to AI applications — input validation, prompt injection detection, output filtering, content moderation, and policy enforcement. Prevent misuse without breaking legitimate use cases.

概要

The AI Guardrails skill, part of the TerminalSkills/skills repository, provides a structured framework for enhancing the safety and reliability of artificial intelligence applications. This security-focused tool enables developers to integrate multiple defensive layers, including input validation and prompt injection detection, to mitigate common vulnerabilities. By utilizing this skill, agents like Codex, Claude, and Gemini can perform real-time content moderation and output filtering to ensure compliance with established organizational policies. The repository, which has gained 71 stars, offers these capabilities as a Python-based solution for managing API interactions. It focuses on preventing malicious misuse while maintaining the functionality required for legitimate user requests, effectively balancing strict security enforcement with application usability across various supported AI platforms.

ユースケース

Detecting and blocking malicious prompt injection attempts in real-time.
Filtering model outputs to prevent the disclosure of sensitive or prohibited content.
Enforcing custom safety policies and content moderation standards across AI interactions.

導入方法

# Review source first
open https://github.com/TerminalSkills/skills/blob/main/skills/ai-guardrails/SKILL.md

Copy or clone the skill folder into your agent skills directory after reviewing its instructions and scripts.

セキュリティ

AI Guardrails acts as a defensive middleware layer; however, users should ensure that the underlying Python environment and API keys are properly secured. While it mitigates prompt injection and unauthorized output, it should be part of a broader defense-in-depth strategy within the TerminalSkills/skills ecosystem.

関連Skills

Cargo Fuzz

trailofbits/skills

セキュリティ

cargo-fuzzは、Cargoを使用するRustプロジェクトにおける事実上の標準的なファジングツールです。libFuzzerバックエンドを使用したRustコードのファジングに使用します。

Claude CodeClaude
securityresearch
6,407 Starsソースあり

Yara Rule Authoring

trailofbits/skills

セキュリティ

マルウェア識別のための高品質な YARA-X 検知ルールの作成をガイドします。YARA ルールの記述、レビュー、または最適化時に使用します。命名規則、文字列の選択、パフォーマンスの最適化、レガシーな YARA からの移行、および誤検知の削減をカバーしています。トリガー:YARA、YARA-X、malware detection、threat hunting、IOC、signature、crx module、dex module。

Claude CodeClaude
designsecurity
6,407 Starsソースあり

Agentforce D360 Analyze

forcedotcom/sf-skills

セキュリティ

単一の Agentforce セッションの Data Cloud 360° ビュー。ユーザーがセッション ID(Agent Session UUID `019d…` または MessagingSession ID `0Mw…`)によって特定の Agentforce セッションの追跡、検査、要約、または説明を求めたときに TRIGGER。また、ユーザーがまだセッション ID を持っていない場合に、時間、エージェント、チャネル、結果、または会話テキストによるセッションの検出(検索、一覧表示、探索)でもトリガーされます。設計時のアーキテクチャに関する質問(代わりに agentforce-architecture-analyze を使用)や、ランタイムのパフォーマンス/l については NOT TRIGGER。

pythondata
783 StarsApache-2.0

Security Audit

TerminalSkills/skills

セキュリティ

OWASP Top 10の脆弱性スキャン、既知のCVEに関する依存関係のチェック、流出したシークレットやAPIキーの検出を行い、優先順位付けされた修正案を生成することで、コードベースの包括的なセキュリティ監査を実行します。このスキルは、静的解析パターンと依存関係監査ツールを組み合わせています。

CodexClaude Code
securityaudit
72 StarsApache-2.0