Awesome AI Security Tools
Section: Scanners, Evals & Guardrails · 🟢 — Programmable guardrails (input/output/dialog/retrieval rails) for LLM apps. (NVIDIA) · updated 2026-08-17)
Entry
Appears in 6 awesome lists
NVIDIA's programmable guardrails toolkit: define input, dialog, retrieval, execution, and output rails that intercept the agent loop at five distinct layers using the Colang DSL. The execution rail layer specifically governs what tools the LLM can invoke and what their inputs/outputs may contain —…
Section: Scanners, Evals & Guardrails · 🟢 — Programmable guardrails (input/output/dialog/retrieval rails) for LLM apps. (NVIDIA) · updated 2026-08-17)
Section: Security, Sandbox & Permissions · NVIDIA's programmable guardrails toolkit: define input, dialog, retrieval, execution, and output rails that intercept the agent loop at five distinct layers using the Colang DSL. The execution rail layer specifically governs what tools the LLM can invoke and what their inputs/outputs may contain —…
Section: Security and Sandboxing · an open-source toolkit from NVIDIA for easily adding programmable guardrails to LLM-based conversational systems
Section: 10. AI Safety, Alignment & Interpretability · Programmable guardrails toolkit for LLM-based conversational systems. Uses Colang DSL to define safety rules, dialog flows, and content boundaries. Integrates with LangChain, LangGraph, and LlamaIndex for production deployments. Apache 2.0 licensed.
Section: Privacy and Safety · NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.
Section: LLM and Inference · NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.
Test your prompts, models, RAGs. Evaluate and compare LLM outputs, catch regressions, and improve prompt quality. LLM evals for OpenAI/Azure GPT, Anthropic Claude, VertexAI Gemini, Ollama, Local & private models like Mistral/Mixtral/Llama with CI/CD
Firecracker microVM sandboxes purpose-built for agent tool loops: ~150ms cold start, Python/JS SDKs, open source. The clearest reference implementation of "code execution as a harness primitive" rather than a CI system bolted on.
Input/output validation framework for building reliable AI applications. Detects and mitigates risks through composable validators for PII, toxicity, prompt injection, and structured output validation. Features Guardrails Hub with 50+ pre-built validators. Apache 2.0 licensed.
Python framework for adversarial attacks, data augmentation, and model training in NLP. Augment datasets to increase model robustness and generate adversarial examples. MIT licensed.
🟢 — Multi-language toolkit for policy-enforced agent tool calls and audit records, with optional identity, MCP-gateway, sandboxing, reliability, and compliance components. (Microsoft) — note: official public preview; APIs and deployment patterns may change before general availability. · updated…;…
The LLM vulnerability scanner. Probes models for hallucinations, data leakage, prompt injection, misinformation, toxicity, and jailbreaks. Extensive plugin-based architecture with 100+ vulnerability probes. Apache 2.0 licensed.
Tencent Cloud's production-validated microVM sandbox for AI agents: sub-60ms cold start via snapshot cloning, <5MB per-instance overhead, and true kernel-level isolation with eBPF-enforced network policies. E2B-compatible drop-in replacement that demonstrates how hyperscale cloud infrastructure…
Open-source policy-driven sandbox runtime for autonomous AI agents, announced at GTC 2026. Enforces security constraints at the kernel level via Landlock LSM (filesystem), seccomp BPF (syscalls), and an OPA/Rego-evaluated HTTP CONNECT proxy (network) — constraints are enforced on the environment…