Your model can attack. Train it to defend.

Ask a frontier model to write an exploit and it delivers. Ask it to catch the breach and it goes blind. Not one has passed our Cyber Defense Benchmark. That skill was never in the pretraining data, so we build it: the real investigations you can grade, plus a deterministic environment that scores whether the model actually stopped the attack.

Trusted by leading enterprises and MSSPs

Research

Cyber Defense Benchmark

Paper

Can LLMs Do Cyber Defense?

Product

Security Training Data & RL Environments

The Frontier Model
Cyber-Defense Checklist

HOW THE DATA IS BUILT01Real attack telemetrythe behavior we start from02Modeled at org scalecomposed in Holodeck03Verified ground truthmachine-checked where it can be04Licensed data, our IPclean to train onCustomer /productionlogsnever resoldThe moat is the modeling and the verification.

Can you fine-tune your way to cyber defense?

No. Defense is abductive work. You infer an attacker's intent from logs that look completely normal, and that reasoning barely shows up in pretraining. You can't fine-tune a skill the data never held. It has to be grown with RL against a real reward: data you can grade, and an environment that knows a good investigation from a bad one.


How do you get a verifiable reward for defense?

Holodeck. It's a deterministic, Gymnasium-compatible environment where the reward comes from machine-checkable truth, not a model's opinion. Same seed, same run, so nothing drifts and there's no LLM judge to game. The calls a machine can't make, like severity and escalation, go to expert human graders, and a held-out set catches anything trying to hack the reward.


Will it plug into our RLVR stack?

Yes. Tasks ship as JSONL with verifiers and Inspect-compatible formats, split into train, test, and held-out. Holodeck speaks Gymnasium, so your base model drops in without a custom harness. Deterministic and assertion graders give you an RL-usable reward; rubric and human graders sit alongside as acceptance and regression signal, not the reward itself.

0 of 20
frontier models pass the Cyber Defense Benchmark
44.5%
best model's MITRE ATT&CK coverage (passing bar: 50%)
13
MITRE ATT&CK tactics measured end to end

Everything you need to grow the capability.

Verified data, a deterministic environment, expert validation, and a published eval. Take them on their own or together. Every piece is built for the RLVR toolchain you already run.

Real investigations, decomposed into gradeable tasks

We take real attack chains and investigations across threat hunting, detection engineering, triage, and offensive security, then break them into tasks a model can be scored on. Each one carries verified ground truth wherever the work can be checked by machine. It ships as JSONL and Inspect-compatible tasks, split into train, test, and held-out.

Verified security training data

A gym where a base model learns to hunt

Drop a base model into the loop. It observes, acts, and earns a reward, learning to run an investigation the way an analyst would. The reward comes from machine-checkable ground truth, so the same seed always produces the same run. Nothing drifts, and there's no LLM judge to talk your agent past. It plugs straight into the RLVR toolchain you already run.

Holodeck, a deterministic RL environment

Career analysts grade what a machine can't

Severity, escalation, the judgment that separates a real analyst from a checklist: none of it comes from a generalist crowd. Career threat hunters and detection engineers grade those tasks and your model's outputs, and a held-out, human-graded set catches the reward hacking an environment can miss on its own.

Expert human validation

The eval that hasn't been solved

A published benchmark, built to resist memorization, across the 13 MITRE ATT&CK tactics that make up a real breach. No frontier model has cleared the bar yet. Run it before you train to find the gaps, then run it again after to see exactly what moved.

The Cyber Defense Benchmark

Deterministic ground truth. No reward-model drift.

Most reward models start to drift the moment your agent learns to game them. Holodeck's doesn't. Its reward comes from machine-checkable truth, and the same seed always produces the same run. A correct investigation is a correct investigation. The moat is the modeling and the verification, not a pile of hoarded data.
A drifting reward model versus Holodeck's reward locked exactly to machine-checkable ground truth, identical on every run with the same seed

Why frontier labs choose Simbian

You can't crowdsource security data

Generalist vendors staff doctors, lawyers, and bankers. We staff career threat hunters and detection engineers. Security is the one domain where the ground truth has to come from people who've actually run operations, not a crowd.

Deterministic, verifiable reward

Holodeck scores every investigation against machine truth. There's no reward model to drift and no LLM judge to sweet-talk. Same seed, same result, so the reward is exact and reproducible.

Verified and clean to license

Every trajectory is our own IP with verified ground truth, informed by operating security through the world's largest MDRs, not sourced from their logs. We don't resell customer or production data.

Bring a base model, measure the lift

We won't claim a number we haven't earned. Bring a base model, baseline it on the Cyber Defense Benchmark, train it in Holodeck, and we measure the lift together.

Questions from research teams

We build it. Our attack and organization models are built on real attacker behavior, composed into org-scale campaigns and verified, informed by operating security through the world's largest MDRs, not sourced from their logs. Every trajectory is our own IP with verified ground truth. We do not resell customer or production logs.
Both, and the published benchmark too. The verified training data, the Holodeck RL environment, and the Cyber Defense Benchmark are available independently or together.
Yes. JSONL plus verifiers and Inspect-compatible task formats, with train, test, and held-out splits. Holodeck is Gymnasium-compatible, so a base model drops in with no bespoke harness. Deterministic and assertion graders give an RL-usable reward; rubric and human graders provide acceptance and regression signal, not the RL reward.
Holodeck's reward is deterministic and machine-checkable, so there's no reward model to drift and no LLM judge to game. Judgment tasks that can't be mechanically verified go to expert human graders, and a held-out human-graded set is there to catch Goodharting the verifiable reward.
The data, environment, and benchmark are built and published. Training a model on them is the joint experiment we propose: baseline on the Cyber Defense Benchmark, train in Holodeck, and measure the delta together.

What Our Customers Say

Simbian's AI Agents consistently deliver precise and accurate responses, significantly easing our workload. What used to take days now takes minutes, and we're thrilled with how seamlessly it integrates into our existing processes. It's not just about saving time; it's about maintaining the highest standards of security and accuracy, which is exactly what Simbian enables us to do.
Company logo
Matillion
Suchit Mishra
Director of Information Security
Security is a domain of ever-increasing complexity. Every day a security incident brings new variables. Simbian is building a fully autonomous security platform. We are excited to partner with them as it allows us to be strategic in our security goals, leaving mechanics of security to Simbian.
Company logo
Axelar
Sergey Gorbunov
Co-founder
Security partners, especially MSSPs and MDRs, are at a critical juncture. Attacks are getting accelerated with AI. We must use AI on defense side too. We have gotten great support from Simbian with its fully autonomous security. It allows us to do more with less, directly impacting both our top and bottom lines.
Company logo
Cybalt
Khirodra Mishra
CEO
Simbian's platform takes a straightforward approach to solving core problems we see every day in the SOC. The power in the platform, their AI agents, is in its simplicity. They are not adding steps and processes to achieve results. The Security Accelerator platform drives efficiency without sacrificing efficacy. It allows us to shift the role of the analyst; to give them the time to use human insight, because well trained AI that we can review, and audit, is immensely powerful. It sets a whole new bar for security operations.
Company logo
SMT
Mohammad Qasas
SOC Lead
Simbian's AI agents augment and automate many security services resulting into better efficiencies and increased precision.
Company logo
Wipro
Siva VRS
Vice President
What Simbian's doing in that space has really been a differentiator and a game changer for how my team's thinking about these problems. We're no longer thinking about a pipeline of work that we've got to have 20 people to solve.
Company logo
Bottomline
Blaine Brennecke
Director of Security Operations

Sign up for Simbian's Newsletter

By submitting this form, you agree to our Privacy Policy.

Ask AI about Simbian