Built against OWASP LLM01Benchmark re-run every 6 hoursOpen-source SDK · MIT

API Reference

Request fields, verdict handling, errors and quotas for prompt injection screening. Keep the API key on your server and enforce the response before model use.

Base URL

https://api.safeprompt.dev

Authentication

Validation requests require two headers:

⚠ X-User-IP Header Required

The X-User-IP header must contain the end user's IP address (the person submitting the prompt), not your server's IP. Derive it through your trusted server or proxy configuration; a raw forwarded header can be supplied by a caller.

How to get the end user's IP:

// Express: configure trust proxy for the actual trusted proxies first.
const clientIp = req.ip;

// Other frameworks: use the verified client address supplied by your
// trusted server/platform. Do not trust arbitrary X-Forwarded-For text.

Endpoint

POST /api/v1/validate

Submit one untrusted message, document or tool-result representation before model use. The cURL examples inspect a verdict; they do not enforce application behavior.

Request

curl -X POST https://api.safeprompt.dev/api/v1/validate \
  -H "X-API-Key: YOUR_API_KEY" \
  -H "X-User-IP: CLIENT_IP_ADDRESS" \
  -H "Content-Type: application/json" \
  -d '{"prompt": "Hello world", "sensitivity": "strict"}'

Request Body (annotated)

{
  "prompt": "string",           // Required: The prompt to validate
  "mode": "standard",          // Optional: caching behavior; standard is default
  "sensitivity": "balanced",    // Optional: lenient, balanced (default), strict
  "include_stats": false,       // Optional: Include performance statistics
  "session_token": "optional-session-token" // Conversation context
}
Validation Modes
  • optimized: Accepted mode name; does not select detection sensitivity
  • standard (default): Uses the validation pipeline without requesting a cached verdict
  • ai-only: Accepted legacy name; do not infer detector coverage from the name
  • with-cache: Requests a cached result when available; validate otherwise
Detection Sensitivity

Controls how aggressively borderline prompts are blocked. Independent of mode (which controls caching/performance).

  • lenient: A less restrictive detection setting; test missed attacks and ordinary-message false positives for your traffic.
  • balanced (default): The standard SafePrompt behavior. Omit the field to keep it.
  • strict: A more restrictive setting used in the raw HTTP examples here; test the detection/false-positive trade-off.

Illustrative Response

{
  "safe": true,                // Boolean: Is the prompt safe to use?
  "confidence": 0.95,          // Float 0-1: How confident is the verdict?
  "threats": [],               // Array: Detected threat types (empty if safe)
  "detectionMethod": "pattern_detection",  // String: Detection stage used
  "reasoning": "No security threats detected"  // String: Why this verdict?
}
Response Fields Explained
  • safe: true = Passing verdict for submitted text, false = Block this prompt
  • confidence: 0.0-1.0 scale. Reports detector confidence, not a calibrated probability or permission grant.
  • threats: Array of detected threat types: jailbreak_instruction_override, jailbreak_role_play, exfiltration_target, injection_command, etc.
  • processingTime: Milliseconds taken to validate. Use the published latency percentiles on safeprompt.dev when choosing your request budget
  • detectionMethod: Shows which validation stage caught the threat:
    • pattern_detection - Matched by known patterns
    • reference_detection - Caught by URL/IP/file path detection
    • ai_validation - Caught by AI semantic analysis
  • reasoning: Human-readable explanation of the verdict. Useful for logging and debugging.

Illustrative Response (Threat Detected)

{
  "safe": false,
  "confidence": 0.98,
  "threats": ["jailbreak_instruction_override", "exfiltration_target"],
  "detectionMethod": "pattern_detection",
  "reasoning": "Detected attempt to override system instructions and fetch external URL"
}

Code Examples

JavaScript/Node.js

// Using SDK (recommended)
import SafePrompt from 'safeprompt';

const client = new SafePrompt({ apiKey: process.env.SAFEPROMPT_API_KEY });

const result = await client.check('Ignore previous instructions...', {
  userIP: clientIp // Actual end user address from trusted server configuration
});

if (typeof result.safe !== 'boolean' || !result.safe) {
  throw new Error('Unsafe or invalid verdict');
}
// SDK uses the API default, balanced. Only now use this checked input.

// Using HTTP API directly
async function checkPrompt(userInput, clientIp, timeoutMs) {
  const response = await fetch('https://api.safeprompt.dev/api/v1/validate', {
    method: 'POST',
    headers: {
      'X-API-Key': process.env.SAFEPROMPT_API_KEY,
      'X-User-IP': clientIp,
      'Content-Type': 'application/json'
    },
    body: JSON.stringify({ prompt: userInput, sensitivity: 'strict' }),
    signal: AbortSignal.timeout(timeoutMs)
  });

  if (!response.ok) throw new Error('Validation HTTP error');
  const result = await response.json();
  if (typeof result.safe !== 'boolean' || !result.safe) {
    throw new Error('Unsafe or invalid verdict');
  }
  return result; // Forward the same checked input only after this resolves.
}

Python

import requests
import os

def check_prompt(user_input, client_ip, timeout_seconds):
    response = requests.post(
        'https://api.safeprompt.dev/api/v1/validate',
        headers={
            'X-API-Key': os.environ["SAFEPROMPT_API_KEY"],
            'X-User-IP': client_ip,
            'Content-Type': 'application/json'
        },
        json={'prompt': user_input, 'sensitivity': 'strict'},
        timeout=timeout_seconds
    )

    response.raise_for_status()
    result = response.json()
    if not isinstance(result.get('safe'), bool) or not result['safe']:
        raise ValueError('Unsafe or invalid verdict')

    return result

PHP

// Supply $clientIp from your trusted server/proxy configuration.

$ch = curl_init('https://api.safeprompt.dev/api/v1/validate');
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_HTTPHEADER, [
    'X-API-Key: ' . $_ENV['SAFEPROMPT_API_KEY'],
    'X-User-IP: ' . $clientIp,
    'Content-Type: application/json'
]);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode([
    'prompt' => $userInput,
    'sensitivity' => 'strict'
]));
curl_setopt($ch, CURLOPT_TIMEOUT, $timeoutSeconds);

$body = curl_exec($ch);
$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);
if ($body === false || $status < 200 || $status >= 300) {
    throw new Exception('Validation unavailable');
}
$result = json_decode($body, true, 512, JSON_THROW_ON_ERROR);
if (!is_array($result) || !is_bool($result['safe'] ?? null) || !$result['safe']) {
    throw new Exception('Unsafe or invalid verdict');
}
// Only now add this checked input to model context.

Error Codes

CodeMeaning
200Success
400Bad request (invalid input or missing X-User-IP)
401Unauthorized (invalid API key)
403Forbidden (subscription inactive)
429Rate limit or monthly quota reached
500Internal server error

Rate Limits

TierRequests/SecondMonthly Limit
Free1010,000
Starter50500,000
Business1001,000,000

Recommendation: Use the SDK

Choose the official JavaScript SDK for its typed client, or raw HTTP for explicit sensitivity control and a configurable request deadline:

npm install safeprompt

The published JavaScript SDK sets authentication headers and reports HTTP errors. Your application supplies userIP; the SDK does not discover it or retry failed requests. It forwards sessionToken but does not expose sensitivity or a request timeout, so the API default sensitivity is used. The raw HTTP helper above supplies the configurable deadline. View the SDK on GitHub or npm.

Next Steps