Skip to content

Entry

Rebuff

Appears in 4 awesome lists

🟒 β€” Archived prompt-injection detector (heuristics + LLM + vector DB + canary tokens). (Protect AI) β€” note: archived by the maintainer; retained as a historical prompt-injection defense reference. Β· updated 2024-01-25); Related: LLM Guard

Open github.comprotectai/rebuff

Found in these lists

Awesome Ai Agents 2026

Section: AI Safety and Guardrails Β· Prompt injection detection.

ActiveScore 74

Awesome AI Security Tools

Section: Scanners, Evals & Guardrails Β· 🟒 β€” Archived prompt-injection detector (heuristics + LLM + vector DB + canary tokens). (Protect AI) β€” note: archived by the maintainer; retained as a historical prompt-injection defense reference. Β· updated 2024-01-25); Related: LLM Guard

FreshScore 85

Awesome LLM Security

Section: Tools Β· a self-hardening prompt injection detector

SlowScore 53

SOCIAL MEDIA

Section: LLM Security & AI Security Β· Prompt injection detector for LLM applications.

FreshScore 86

promptfoo

Test your prompts, models, RAGs. Evaluate and compare LLM outputs, catch regressions, and improve prompt quality. LLM evals for OpenAI/Azure GPT, Anthropic Claude, VertexAI Gemini, Ollama, Local & private models like Mistral/Mixtral/Llama with CI/CD

In 11 listsDetails

Agentic Radar

Open-source CLI security scanner for agentic workflows. Scans your workflow’s source code, detects vulnerabilities, and generates an interactive visualization along with a detailed security report. Supports LangGraph, CrewAI, n8n, OpenAI Agents, and more.

In 6 listsDetails

Guardrails AI

Input/output validation framework for building reliable AI applications. Detects and mitigates risks through composable validators for PII, toxicity, prompt injection, and structured output validation. Features Guardrails Hub with 50+ pre-built validators. Apache 2.0 licensed.

In 6 listsDetails

NeMo Guardrails

NVIDIA's programmable guardrails toolkit: define input, dialog, retrieval, execution, and output rails that intercept the agent loop at five distinct layers using the Colang DSL. The execution rail layer specifically governs what tools the LLM can invoke and what their inputs/outputs may contain —…

In 6 listsDetails

TextAttack

Python framework for adversarial attacks, data augmentation, and model training in NLP. Augment datasets to increase model robustness and generate adversarial examples. MIT licensed.

In 6 listsDetails

garak

The LLM vulnerability scanner. Probes models for hallucinations, data leakage, prompt injection, misinformation, toxicity, and jailbreaks. Extensive plugin-based architecture with 100+ vulnerability probes. Apache 2.0 licensed.

In 4 listsDetails

LangKit

🟒 β€” LLM monitoring toolkit extracting safety/security signals such as jailbreak similarity, prompt-injection similarity, hallucination checks, PII patterns, toxicity, and refusal metrics. (WhyLabs) Β· updated 2024-11-22)

In 4 listsDetails

NeMo Guardrails

NeMo Guardrails is an open-source toolkit facilitating the integration of programmable guardrails, essential for steering and safeguarding AI agents' conversational outputs, into large language model-based applications github | research paper

In 4 listsDetails