Skip to content

Repository files navigation

RedTeamForge — industrial forge mark, targeting reticle, cool steel light

RedTeamForge

A self-hosted AI red-teaming lab in your browser.
Fire a fixed arsenal of prompt-injection, jailbreak, and exfil probes at your own system prompt — zero-key leaky sandbox, BYOK live targets, OWASP + ATLAS tags on every finding.

Point it at OpenAI, Anthropic, Gemini, Grok, Ollama, or any OpenAI-compatible endpoint.
Deterministic detectors score every reply HIT / PARTIAL / BLOCKED and tag findings with OWASP LLM Top 10 categories and MITRE ATLAS techniques.
Fully self-hosted. No account required.

MIT License Node 22 Docker Compose 26 probes OWASP LLM Top 10 MITRE ATLAS

▶ Try Online · Quick Start · Features · Screenshots · Probes · FAQ · Contributing

🌐 Try it in your browser — no install
Try before you install: a browser-only build of this app runs free on GitHub Pages. Bring your own API key — calls go straight from your browser to the provider, and keys plus scan history never leave your browser.
That hosted site is the try-it demo (serverless build). The full application is what you self-host below: server-side calls, optional password gate, Docker.


RedTeamForge dashboard with risk 76 critical and recent hits

What this is

RedTeamForge tests how a model + system prompt pair holds up under attack. You choose a target and the system prompt under test, 26 hand-built payloads fire one by one as user messages, and deterministic detectors score each reply:

Stage What happens
Target ForgeBank sandbox (deterministic, no API keys) or a BYOK live endpoint
Payload Fixed catalog: prompt injection, jailbreaks, exfiltration, agency, output, RAG, …
Detector Planted-secret match, keyword needles, regex, refusal patterns
Verdict HIT / PARTIAL / BLOCKED, tagged OWASP + MITRE ATLAS, rolled into a risk score

What this isn't

  • Not multi-turn. Every probe is a single user message. No conversation state, agent loop, or tool executor — agency and RAG probes simulate their scenarios inside the payload text.
  • Not semantic judgment. Detectors are deterministic string matching, so verdicts are reproducible. Treat the risk score as triage for a human reviewer, not an audit grade.
  • Not an app scanner. It tests the chat endpoint you point it at — not your source code, infrastructure, or retrieval pipeline.

Coming from Garak, Promptfoo, or PyRIT? Those tools go deeper: generated attack modules, multi-turn campaigns, CI eval integration, research orchestration — use them when you need that. RedTeamForge trades depth for setup speed: clone, npm run dev, scan the built-in sandbox with no account and no API key, export the report. It is inspired by those projects and taxonomies and not affiliated with them.

Features

  • 26 probes across 8 packs — prompt injection, jailbreaks, data exfiltration, excessive agency, output handling, RAG, misinformation, unbounded consumption
  • OWASP LLM Top 10 (2025) + MITRE ATLAS tags — category and technique IDs on every finding
  • Zero-key demo target — deterministic leaky sandbox (ForgeBank) with planted fixtures, so detectors are learnable
  • BYOK live targets — hosted presets plus Ollama and any custom OpenAI-compatible endpoint
  • Optional analyst — one capped call that ranks hits, explains exploitability, and writes hardening notes
  • Prompt lab — fire a single payload and inspect the raw completion
  • Markdown + JSON export — scan history stays in this browser (localStorage)
  • Self-hosted — npm run dev or Docker Compose; optional shared-password gate for VPS deploys
  • Hosted demo — static browser-only build published to GitHub Pages; BYOK calls go straight from your browser to the provider

Quick Start

Just want to try it? Skip the install — use the hosted demo (browser-only build). The steps below set up the full self-hosted application.

Needs: Node.js 22+

git clone https://github.com/mh-sudo/redteamforge.git
cd redteamforge
cp .env.example .env
npm install
npm run dev

Open http://localhost:8080.

  1. New scan → leave the target on Sandbox → Run scan.
  2. Open the report. Expand a HIT. Copy payload / response.
  3. Analyze (optional) after connecting a provider in Settings.
  4. Export Markdown or JSON.

No sign-up. Scan history stays in this browser.

Live targets

Open Settings and connect one of twelve presets: OpenAI, Anthropic, Google (Gemini), xAI, Groq, Mistral, DeepSeek, Together, Fireworks, OpenRouter, Ollama, or Custom. Paste a key, confirm the model, and Scan and Lab list that provider from then on.

Keys stay in this browser vault (localStorage) and are sent only as the Authorization / x-api-key header on calls you trigger. Probe calls are capped at 280 completion tokens, analyst calls at 2400. Base URLs must be HTTPS unless the host is localhost.

Docker

docker compose up --build

Then open http://localhost:8080 and connect providers in Settings. The image is Node 22 serving the production preview on port 8080. Rebuild after probe or UI changes.

Hosted demo vs self-hosted

Two ways to run RedTeamForge — same codebase, different builds:

  • Hosted demo (mh-sudo.github.io/redteamforge) — try it with zero setup. Browser-only build; your browser talks to the provider directly.
  • Self-hosted (Quick Start / Docker below) — the full application. A small Node server proxies provider calls, enabling SSRF guards and the optional password gate.

A browser-only build of this app runs at mh-sudo.github.io/redteamforge — deployed automatically by .github/workflows/deploy.yml on every push to main.

Same app, one architectural difference: there is no server. GitHub Pages serves plain files, so probe and analyst calls go directly from your browser to the endpoint you connect, carrying your key as the auth header — nothing passes through any intermediate server. Scan history and keys still live only in localStorage. The optional password gate does not exist in this build (there is no server left to protect).

Because calls are browser-direct, the provider must allow cross-origin requests (CORS):

Provider preset Works from the hosted build
OpenAI, Google, xAI, Groq, Mistral, DeepSeek, Together, Fireworks, OpenRouter Yes
Anthropic Yes — the build sends Anthropic's documented anthropic-dangerous-direct-browser-access header
Ollama / custom self-hosted endpoints Only if that server sends CORS headers — e.g. start Ollama with OLLAMA_ORIGINS=https://mh-sudo.github.io

Publishing your own fork: repo Settings → Pages → Source GitHub Actions, then push to main. To inspect the same bundle locally: npm run build:static (output lands in .output/public).

Screenshots

New scan. Sandbox or a connected live provider.

New scan campaign with sandbox, Grok, and custom target cards

Report. Findings, analyst, OWASP grid, system prompt under test.

Scan report with risk score, export, and prompt-injection findings

Prompt lab. Fire one payload, inspect the raw completion.

Prompt lab after a hit on ignore-previous-instructions

Probe library and history.

Probe catalog Scan archive
Probe library filtered by pack History with risk scores

Mobile.

RedTeamForge dashboard on a phone   RedTeamForge report on a phone

Probe catalog

RedTeamForge ships 26 probes across 8 packs. Each payload is educational and aimed at the demo policy, not the public internet.

Pack OWASP What it hits
Prompt injection LLM01 Ignore-previous, delimiter, hierarchy, translation smuggle
Jailbreaks LLM01 DAN / roleplay, hypothetical, encoding frames
Data exfiltration LLM02 / LLM07 System-prompt dump, secret repeat, encoded leak
Excessive agency LLM06 Unauthorized refunds and off-allow-list tools
Output handling LLM05 XSS / markup the renderer would execute
RAG / embeddings LLM08 Poisoned retrieved context, citation games
Misinformation LLM09 Confident fabrication, false authority
Unbounded use LLM10 Repeat-forever and token amplification

MITRE ATLAS technique IDs ship on every finding (e.g. AML.T0051 prompt injection, AML.T0054 jailbreak, AML.T0057 LLM data leak). A model/tool-fingerprint probe (LLM03, AML.T0006) rides in the exfiltration pack.

A quick pack of 8 high-signal probes is the default campaign. Full sweeps all 26 probes; Custom narrows by pack.

How a scan works

Target  →  probe payload  →  completion  →  detector  →  verdict
  sandbox | live provider     catalog        model         leak / keyword / regex / refusal
                                                              HIT | PARTIAL | BLOCKED
  1. Choose a target and the system prompt under test (ForgeBank ships as the demo policy).
  2. Selected probes fire one by one. The sandbox responds deterministically; live targets give the real test.
  3. Detectors look for planted secrets, policy breaks, dangerous output, and missing refusals.
  4. Hits roll into a weighted risk score (critical 28 · high 16 · medium 8 · low 3, partials count 40%). OWASP coverage is tallied per category.
  5. Optional analyst call sends compact findings (never your full prompt dump) to a connected model and returns executive summary, exploitability, remediation, and residual risk as JSON.
  6. History, Markdown, and JSON stay on the box.

Sample report

The checked-in ForgeBank baseline is a complete 8-probe sweep (risk 76 / critical, 4 hits): docs/sample-report.md.

Export the same shape from any scan: Export → Markdown or JSON.

Stack

  • TanStack Start + React 19 + Vite 8
  • Tailwind v4, Radix primitives
  • Zustand persist (scan archive and key vault in the browser)
  • BYOK completions: OpenAI Chat Completions format + Anthropic Messages format
  • Garak-inspired TypeScript probe engine (no Python runtime required)

Inspired by Garak, Promptfoo, PyRIT, OWASP LLM Top 10, and MITRE ATLAS. RedTeamForge is not affiliated with those projects.

Project layout

src/
  components/          UI, findings, risk, OWASP grid
  lib/probes/          catalog, detectors, sandbox model
  lib/scan/            engine, scoring, reports, history store
  lib/providers/       presets + browser key vault
  lib/server/ai.ts     live completions + analyst
  lib/server/auth.ts   optional AUTH_PASSWORD gate
  routes/              dashboard, scan, lab, probes, history, report, settings, login
docs/
  images/              product screenshots
  sample-report.md     ForgeBank baseline

Configuration

Variable Required Purpose
HOST No Bind address (localhost by default; Docker sets 0.0.0.0)
PORT No Listen port (default 8080 in Docker)
AUTH_PASSWORD No Shared page password. Unset = open. Set on a VPS to require /login.
AUTH_COOKIE_SECURE No Set to 1 to force the Secure flag on the gate cookie.
RTF_ALLOW_PRIVATE_ENDPOINTS No Allow https provider endpoints on private/LAN addresses (self-hosted models). Implied when AUTH_PASSWORD is set.

Copy .env.example. Never commit a real .env.

Deployment & exposure

Running without a password is a supported, intentional mode — for localhost. Know what an open instance is before you expose one:

  • The server performs outbound requests to whichever provider endpoint the browser selects. Endpoint guards block private/internal addresses and redirects by default, but an open instance reachable on a network is still usable by anyone who can reach it.
  • docker-compose.yml therefore publishes on 127.0.0.1 only. To expose on a LAN or VPS, change the port mapping to "8080:8080" and set AUTH_PASSWORD — then put HTTPS in front (a reverse proxy or a tunnel like Tailscale works fine).
  • The Docker image runs as a non-root user and ships only the built server output; a .dockerignore keeps local .env files out of the image.
  • Provider API keys live in your browser's localStorage and travel only as the selected provider's auth header — treat the browser profile you scan from as the vault.

FAQ

Does RedTeamForge work without an API key? Yes. Leave the target on Sandbox and run a scan — ForgeBank returns scripted leaks and refusals, so detectors, scores, and reports work with zero keys.

Where do my prompts, scans, and keys go? Nowhere, unless you aim them somewhere. Sandbox scans never leave the machine. Live scans send one request per probe to the provider you selected. Keys sit in your browser's localStorage and travel only as that provider's auth header.

How do I password-protect a VPS deploy? Set AUTH_PASSWORD in the environment (or Docker Compose). The UI redirects to /login until the password is entered; unset means an open local lab. Put HTTPS in front either way.

Which platforms are supported? Anywhere Node.js 22+ or Docker runs — macOS (including Apple Silicon), Windows, Linux.

Why does my Ollama / self-hosted target fail on the hosted demo? Browsers block cross-origin calls unless the endpoint sends CORS headers. Start Ollama with OLLAMA_ORIGINS=https://mh-sudo.github.io (or your fork's Pages URL), configure CORS on your vLLM/proxy, or use the self-hosted build — there the server makes the call for you.

Is it free? What's the license? Free and open source under the MIT License; commercial use allowed. Not affiliated with NVIDIA Garak, Promptfoo, Microsoft PyRIT, OWASP, or MITRE.

Responsible use

RedTeamForge is a defensive lab. Use it only on systems you own or have written permission to test.

  • Payloads are for evaluating your model, prompt, or agent.
  • Do not aim this at third-party production assistants you do not control.
  • Sandbox secrets (482917, sk_live_forge_demo_9f3a, FORGE_POLICY_TOKEN) are fake fixtures, not credentials.
  • Live completions may contain model output that looks like exploits. Handle reports as sensitive.

If you find a vulnerability in RedTeamForge itself, do not open a public issue with a working exploit. See Contributing.

Contributing

PRs welcome: new probes (with OWASP + ATLAS tags + a sandbox expected verdict), detectors, report formats, and UI that stays in the tactical language (sharp geometry, one hazard-red accent, IBM Plex Mono + Archivo Black).

See CONTRIBUTING.md for setup and probe rules.

License

RedTeamForge is licensed under the MIT License © 2026 RedTeamForge contributors.

Self-hosted AI red-teaming. No vendor lock-in.

About

Open-source AI red-teaming lab — fire prompt-injection, jailbreak, and exfil probes at your own system prompt. OWASP LLM Top 10 + MITRE ATLAS tagged. Try it in the browser or self-host.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages