What SafePrompt Is and Is Not
Screen untrusted text before model use. Keep content policy and tool permissions in the application.
SafePrompt IS
A prompt injection detection API
SafePrompt sits between your users and your LLM. Before you send a prompt to OpenAI, Claude, or any other model, you send it to SafePrompt first. Your application blocks unsafe results and stops unavailable or malformed checks. A valid safe verdict lets the checked text proceed under your own permissions and content policy.
A security layer for user-submitted input
Any place users can type text that eventually reaches an LLM (contact forms, chat interfaces, lead forms, support bots, agent tools) is a potential injection vector. SafePrompt validates that input before model use. Check later documents and tool results at their own entry points.
Built for developers, not enterprise security teams
One API key and one HTTP endpoint, with published prices and a free tier. Developers and small teams can test the integration before choosing a paid plan.
A 3-layer detection system
SafePrompt uses pattern detection for known attack signatures, external reference detection to catch attempts to pull in outside instructions, and AI validation for complex semantic attacks. See the published latency measurements before setting your response budget.
Shared attack-pattern and reputation signals
Shared attack-pattern and IP-reputation signals inform the detection engine used across plans. Retained prompt and IP hashes are pseudonymous, not anonymous. The privacy policy describes contribution settings and retention; a blocked request does not guarantee protection elsewhere.
SafePrompt is NOT
A content moderation service
SafePrompt does not filter hate speech, NSFW content, or off-topic requests. It specifically detects attempts to manipulate your AI's behavior: jailbreaks, instruction overrides, data exfiltration attempts. For content moderation, use OpenAI's Moderation API or a dedicated service.
A general-purpose Web Application Firewall (WAF)
SafePrompt validates AI prompt inputs, not HTTP traffic. It does not block SQL injection in database queries, XSS in web pages, or DDoS attacks. It focuses exclusively on the prompt injection attack surface, the text that reaches your LLM.
A prompt builder or prompt management tool
SafePrompt does not help you write system prompts, manage prompt templates, or version your prompts. It screens submitted text for instructions that attempt to redirect your AI system.
A guarantee of 100% protection
No security tool offers perfect protection. SafePrompt publishes measured detection and false-positive rates; those results describe tested inputs, not every application or new attack. Defense-in-depth still applies: validate on the backend, scope your LLM permissions, and monitor for anomalies.
An enterprise-only product
SafePrompt has transparent pricing starting at $0. The free tier gives you 10,000 free validations a month with the same detection engine as paid plans. Sign up, confirm your email and get your key from the dashboard.
What attacks does it detect?
- ✓Direct prompt injection: "Ignore previous instructions and..."
- ✓Jailbreak attempts: DAN, roleplay bypass, hypothetical framing
- ✓System prompt extraction: Attempts to reveal your AI's instructions
- ✓Conversation context: a session token associates checks with context; gradual escalation coverage is not established
- ✓External reference injection: Attempts to load instructions from URLs or file paths
- ✓Encoded/obfuscated attacks: Base64, ROT13, and other encoding schemes used to evade detection