• AI Red Teaming

AI Red Teaming at Scale

Stop shipping AI systems blind. Get them battle-tested by a global community of security researchers before production.

Audited by OtterSecLive on Sui & Solana Mainnet
• what is it

What is AI Red Teaming?

Red teaming is the practice of roleplaying as an attacker to uncover vulnerabilities before malicious actors do.

The Origin

The term originated during the Cold War, where the “red team” simulated enemy offensive strategies so the “blue team” could develop robust defenses. Today, this military-proven approach protects AI systems.

Why It Matters

AI systems face unique threats: prompt injection, jailbreaking, data extraction. Traditional security tools cannot detect these language-based attacks. Red teaming finds vulnerabilities before attackers do.

Best Practices

  • Assemble diverse teams for comprehensive vulnerability coverage
  • Develop detailed testing plans with clear objectives
  • Iteratively refine strategies based on findings
  • Prioritize ethics throughout the testing process
  • Maintain detailed records of attack strategies and outcomes
AI Security Testing
0%

of enterprises experienced AI security incidents

$0M+

average cost of an AI data breach

• threat landscape

GenAI vs Traditional Security

Understanding the fundamental differences between traditional cybersecurity threats and emerging GenAI risks

Aspect
Traditional Security
GenAI Threats
Attack Focus
Exploits code vulnerabilities, bugs, misconfigurations
Manipulates decision-making processes through crafted inputs
Attacker Profile
Requires deep technical expertise, specialized knowledge
Accessible to anyone with language skills and creativity
Attack Medium
Coding, technical exploits, network penetration
Natural language, images, audio, and human communication
Detection
Known patterns trigger security alarms
Subtle manipulation evades traditional detection

Key Insight

Traditional security focuses on protecting code and infrastructure. GenAI security focuses on protecting decision-making processes, making it more accessible to attackers but harder to detect with conventional tools.

• the difference

The Red Sentinel Difference

Always-On Testing

Your Sentinels are live 24/7. Attackers worldwide are constantly trying to break them, generating continuous security data.

Verified on Chain

Every attack is verified inside a Trusted Execution Environment with cryptographic attestations. No fake attacks, no disputed results.

Incentivized Community

Attackers earn real money for finding vulnerabilities. Defenders earn from attack fees. Everyone wins.

• for companies

Deploy a Security Sentinel Against the AI You Already Run

No rebuild, no consultants, no retainer. Connect your existing system, set the rules of engagement, and fund a bounty — Red Sentinel earns nothing until attackers fail.

Connect What You Already Run

Point a Sentinel at your existing endpoint: pick a supported provider — OpenAI, Anthropic, DeepSeek, or Amazon Bedrock — or connect any OpenAI-compatible API and validate it in minutes.

You Control the Scope

You define the test surface, the system instructions, and the single protected objective attackers must break. Nothing outside your stated scope is ever in play.

You Set the Bounty

Fund the bounty pool before launch and set the per-attempt attack fee. Attackers win it only by producing a verified breach of your objective.

No Fee Before Results

Red Sentinel earns nothing upfront. The protocol keeps a 10% share of unsuccessful attack fees only — you earn 40% of every failed attempt while your pool grows.

• engagement flow

From Endpoint to Evidence in Five Steps

A structured engagement your security team controls at every step

01

Connect

Point a Sentinel at your existing endpoint or choose a supported provider, then validate the connection.

02

Scope

Set the test surface, system instructions, and the one protected objective attackers must break.

03

Fund

Fund the bounty pool before launch and set the per-attempt attack fee. No platform fee is charged upfront.

04

Launch

Your Sentinel goes live 24/7. Every attack attempt is judged inside a Trusted Execution Environment.

05

Evidence

Collect TEE-attested verdicts and the full on-chain attack log, then harden your system and redeploy.

Start in a Sandbox

Use a staging or sandbox endpoint for your first engagement. Do not connect production systems with live customer data or irreversible tool access — the test surface should mirror production behavior, not carry production risk.

• attack vectors

Attack Types We Test For

Our community tests against the full spectrum of AI attack vectors

Prompt Injection

Override system instructions through carefully crafted user inputs that bypass security guardrails.

Jailbreaking

Bypass safety restrictions through roleplay scenarios, hypothetical framing, and creative context manipulation.

Data Extraction

Trick models into leaking training data, memorized information, or sensitive details through targeted queries.

Model Inversion

Reverse-engineer model outputs to reconstruct input data and uncover hidden training information.

Adversarial Prompting

Cause unintended behaviors through subtle character-level perturbations and semantic manipulation.

Social Engineering

Exploit contextual reasoning and persuasive techniques to deceive AI systems into harmful actions.

• the flywheel

The Self-Reinforcing Security Flywheel

A virtuous cycle where more attacks lead to stronger defenses, attracting more attackers and creating more value

01

Deploy

Defender deploys Sentinel with bounty pool

02

Attack

Attackers pay fee to attempt breaches

03

Grow

Failed attempts increase the bounty pool

04

Attract

Larger bounties attract more attackers

05

Learn

More attacks generate security data

06

Improve

Defender strengthens AI defenses

The cycle repeats, stronger each time
Live on
Sui & Solana Mainnet
Verification
TEE Attestations
Audited by
OtterSec
Recognition
Overflow Winner
• FAQ

Common Questions

How is this different from hiring a red team firm?
Traditional red teaming is a point-in-time assessment. Red Sentinel is continuous. Your AI is tested 24/7 by a global community, with every attack verified and recorded on-chain. You get ongoing security data rather than a single PDF report.
What types of AI systems can I test?
Any LLM-based system can be deployed as a Sentinel: chatbots, AI agents, autonomous systems, or custom language models. As long as your system can accept prompts and return responses, it can be tested on our platform.
How do payouts work?
When an attacker successfully breaks your Sentinel, the bounty is transferred instantly via smart contract on the Sui blockchain. No invoices, no 30-day delays, no disputes, only pure programmatic execution.
Is my proprietary data safe?
Attackers never see your proprietary data or model weights. They interact only with the deployed Sentinel's interface, just like real users would. Your training data and internal systems remain completely isolated.
How much does it cost?
Your entire spend is the bounty you choose — you fund the pool before your Sentinel goes live, and that money pays out only on a verified breach. Attackers pay a small fee per message attempt: on a failed attempt, 50% grows your bounty pool, 40% goes to you as the defender, and the protocol keeps 10%. Red Sentinel takes no fee before results — its only revenue from your Sentinel is a share of unsuccessful attack fees.
What do I need to give attackers access to my endpoint?
A connection your team already controls. Choose a supported provider — OpenAI, Anthropic, DeepSeek, or Amazon Bedrock — or bring your own OpenAI-compatible endpoint URL with an API key and model name. The connection is validated before your Sentinel goes live, and attackers only ever reach the Sentinel interface, never your raw endpoint or keys.
Is my infrastructure isolated from the testing?
Yes. Attackers never touch your infrastructure, training data, or model weights. They interact only with the deployed Sentinel, which runs inside a Trusted Execution Environment. Use a staging or sandbox endpoint for the test surface, and do not connect production systems with live customer data or irreversible tool access.
When is the bounty funded, and what does Red Sentinel charge?
The bounty pool is funded upfront, before your Sentinel launches — the pool is fully self-custodied and pays out only on a verified breach. Red Sentinel charges no platform fee before results: the protocol keeps a 10% share of unsuccessful attack fees only, alongside the 50% that grows your pool and the 40% defender share.
What exactly do I control about the test?
Three things, all yours to define: the test surface (which endpoint or provider-backed model is exposed), the protected objective (the single thing attackers must make your Sentinel do or say), and the bounty (pool size and per-attempt attack fee). Attacks outside the rules you set are not valid breaches.
What security evidence do I get from an engagement?
Every attempt, verdict, and payout is recorded on-chain, and each verdict is produced by an independent jury of three models running inside a TEE with cryptographic attestations. The full attack log — prompts, responses, verdicts — is stored permanently on Walrus and downloadable, so your team can harden the system and show auditors exactly how it was tested.

Ready to Test Your AI's Limits?

Deploy your first Sentinel in minutes. Set your bounty. Let the world's red teamers find what you missed.