Train models that can actually defend.

Frontier models fail at cyber defense. Simbian supplies the verified security training data, deterministic RL environments for RLVR, and the expert validation your post-training pipeline is missing. All built on real attacker telemetry, modeled at org scale.

Why frontier LLMs fail at cyber defense.

Offense is saturated and recall-heavy. Defense is abductive: you reconstruct an unknown attacker's intent from noise. Almost none of that reasoning exists in pretraining.

HOLODECK ENVIRONMENT · benign logsATTACK PROPAGATION
RANSOMWARETA0001 · Initial AccessTA0002 · ExecutionTA0003 · PersistenceTA0004 · Privilege EscalationTA0005 · Defense EvasionTA0040 · ImpactDATA EXFILTRATIONTA0001 · Initial AccessTA0006 · Credential AccessTA0007 · DiscoveryTA0008 · Lateral MovementTA0009 · CollectionTA0010 · ExfiltrationCLOUD TAKEOVERTA0001 · Initial AccessTA0003 · PersistenceTA0004 · Privilege EscalationTA0008 · Lateral MovementTA0011 · Command & Control
detected by the modelmodel blind, attack continues
best model detects 45.1% · 0 / 22 pass

A breach is a path hidden in a haystack of benign logs. Models go blind exactly where it matters.

01

Offense is mostly solved

Deductive, recall-heavy work: exploits, code, CTFs. The web is full of signal, and models already score well.

02

Defense is different in kind

Abductive. You reconstruct an unknown attacker's intent from a haystack of benign-looking logs, and almost none of that reasoning lives in pretraining.

03

It's a data + environment problem

You can't read defensive reasoning off the web, and you can't crowdsource it. It has to be grown with RL against verifiable rewards, which needs an environment that can score an investigation right or wrong.

No frontier model passes the Cyber Defense Benchmark.

The leader stalls at 44.5% MITRE ATT&CK coverage against a 50% passing bar. The category is wide open.

0 of 22

frontier models pass the Cyber Defense Benchmark

45.1%

best model's MITRE ATT&CK coverage (passing bar: 50%)

13

MITRE ATT&CK tactics the benchmark measures, end to end

Tactic-breadth coverage
Loading benchmark radar…

The shape collapses on the tactics that define a breach — lateral movement, credential access, exfiltration.

Per-model coverage vs. the passing bar

Every bar; not one crosses the 50% line.

Offensive benchmarks saturate in months. This one stays unsolved, and every model still goes blind where real breaches happen.

Training data, RL environments, and evaluations.

Closing the gap takes three things a lab can't build alone: verified security trajectories, the Holodeck RL environment, and the published Cyber Defense Benchmark.

Training data

Verified security trajectories

Real investigations and attack chains, decomposed into gradeable tasks across threat hunting, detection engineering, triage, and offensive security. Every unit ships with verified ground truth, machine-checkable wherever the task allows.

See the task taxonomy
RL environments

Holodeck

A deterministic, Gymnasium-compatible gym where a base model learns to hunt and is scored on machine truth.

How Holodeck works
Evaluations

The Cyber Defense Benchmark

The published eval, formalized as your north-star metric. It resists memorization and shows exactly where a model fails.

Read the benchmark

We build data per skill, not per use case.

Each use case breaks into discrete skills we can grade independently.

ONE USE CASEGRADEABLE TASK CATEGORIESUSE CASEThreat HuntingHypothesis generationgradeableQuery generationgradeableEvidence correlationgradeableLead / pivot expansiongradeableAttack attributiongradeableTimeline reconstructiongradeable

Threat Hunting

  • hypothesis generation
  • query generation
  • evidence correlation
  • lead / pivot expansion
  • attack attribution
  • timeline reconstruction

Detection Engineering

  • detection authoring
  • coverage-gap analysis
  • false-positive tuning
  • rule validation
  • data-source mapping
  • detection-as-code review

Triage & Investigation

  • alert enrichment
  • prioritization
  • root-cause analysis
  • disposition (TP / FP)
  • response actions
  • escalation

Offensive / Pentest

  • recon
  • exploitation
  • privilege escalation
  • lateral movement
  • impact
  • reporting

Each category is scored independently, so we produce targeted data and measure a model skill by skill. It extends to tool-use and integration trajectories (agents driving security connectors) for agentic data, not just text.

Holodeck: a deterministic RL environment with verifiable rewards.

A Gymnasium-compatible gym that scores security investigations against machine truth. No reward model to drift, no LLM judge to game.

observe
threat briefing + prior results
act
query · submit · give-up
reward
deterministic · verifiable
same seed → identical run · loop repeats until the model gives up or solves

Any base model drops in

A Gymnasium-compatible interface: observe (threat briefing + prior results) → act (query / submit / give-up) → reward. No bespoke harness to build.

Verifiable reward

In Holodeck, ground truth is deterministic, so the reward is exact and unfakeable. You can't game it the way you can a learned reward model.

Real attacker behavior, modeled at org scale

Deterministic replay with seeded mutation. Same seed, identical run. Reproducible and reusable, like a versioned artifact you check into your stack.

How an environment is built
01
02
03
04
Procedure Library
105 procedures · real tool telemetry
Campaign Assembly
Multi-stage kill chain · shared infrastructure
World-Builder
Seed: environment and attacker infrastructure, sequence and timing of attack
SQL Database
Same log data inputs as in real threat hunting, just no hints
Every run is unique. Every run is reproducible.

Compatible with the RLVR (reinforcement learning with verifiable rewards) toolchain your team already uses. Tasks, tools, and reward functions ship as first-class, verifiable artifacts.

Where judgment is needed, bring expert human validation, not crowd labels.

Deterministic rewards cover the machine-checkable tasks. The judgment calls they can't cover go to people. Our network of career SecOps analysts (threat hunters, detection engineers, and incident responders with decades in the field) reviews and grades your tasks and model outputs.

VERIFIABLE LANEJUDGMENT LANEHolodeck envDeterministic rewardModel outputExpert analystcareer SecOps,decades in fieldgraded · preferenceTraining +eval signalAutomate the verifiable. Bring experts for judgment.

Grade model outputs on real security tasks

Scored by analysts who run these investigations in production.

Preference & RLHF data

Expert rankings of competing outputs, for reward modeling.

Judgment-task ground truth

The calls that can't be mechanically verified: severity, escalation, reasoning quality.

Held-out human validation

An independent, human-graded set to catch reward hacking (anti-Goodhart).

Holodeck automates the verifiable reward. Human validation covers what it can't. Priced hourly, for deep security expertise, not general crowd work.

Generalist data vendors staff doctors, lawyers, and bankers. We staff the people who actually run security operations.

You can't crowdsource security data. We build it.

HOW THE DATA IS BUILTReal attacktelemetryModeled at orgscale · HolodeckVerifiedground truthLicensed data— our IPCustomer / production logsnever resoldThe moat is the modeling and the verification — not hoarded data.

Grounded in real operations

Our attack and organization models are built on real attacker telemetry, informed by operating security through the world's largest MDRs (managed detection and response providers). Not fabricated from scratch, and not crowd-labeled.

Verified & clean to license

Every trajectory carries verified ground truth; the artifact is our own IP. The moat is the modeling and the verification, not hoarded data.

Built by people who've shipped both security and infrastructure

Founders from Fortanix, NVIDIA, and Twitter; authors of the Cyber Defense Benchmark.

ProvenanceReal attacker behavior, modeled at org scale, verified — our own IP. We don't resell customer logs. We build the data.

Built to slot into your post-training stack.

JSONL and verifiers / Inspect-compatible tasks, Gymnasium-compatible environments, and RLVR-ready rewards. Nothing bespoke to build.

WHERE SIMBIAN PLUGS INYour basemodelRLVR loopverifiable rewardSIMBIANthe security data layerTraining dataHolodeck environmentExpert validationCyber DefenseBenchmarkMeasured lifton the same evalWhere Simbian plugs into your post-training loop.
JSONL + verifiers / Inspect-compatible task format
train / test / held-out splits (anti-Goodhart)
deterministic + assertion graders → verifiable, RL-usable reward
rubric / human graders for judgment tasks (acceptance & regression signal, not the RL reward)
tool-use / MCP-style trajectories
clean licensing, our own IP

Two ways to work with us.

Data provider

Verified security training data, the Holodeck RL environment, and the Cyber Defense Benchmark. You train on it.

Priced per verified trajectory or per environment.

Human validation

Our network of career SecOps analysts grades your tasks and model outputs: the judgment calls a machine can't verify.

Priced hourly.

Questions from research teams

Holodeck is a deterministic, Gymnasium-compatible RL environment: the same seed produces an identical run, so the reward is exact and reproducible. That matters for RLVR (reinforcement learning with verifiable rewards) because the reward comes from machine-checkable ground truth, not a reward model that drifts or an LLM judge that can be gamed.
Every task ships with verified ground truth. Where the task allows, deterministic and assertion graders produce a machine-checkable, RL-usable reward inside Holodeck. Judgment calls that can't be mechanically verified (severity, escalation, reasoning quality) go to expert human graders, plus a held-out human-graded set that catches reward hacking.
Offense is well represented on the web and largely solved. Defense is abductive: you reconstruct an unknown attacker's intent from a haystack of benign-looking logs, and almost none of that reasoning exists in pretraining. On the published Cyber Defense Benchmark, no frontier model tested clears the passing bar of 50% MITRE ATT&CK tactic coverage.
We build it. Our attack and organization models are built on real attacker telemetry, composed into org-scale campaigns and verified — informed by operating security through the world's largest MDRs, not sourced from their logs. Every trajectory is our own IP with verified ground truth. We do not resell customer or production logs.
Both. Holodeck is a deterministic, Gymnasium-compatible RL environment that scores security investigations against verifiable rewards, alongside the training data and the published Cyber Defense Benchmark used as an evaluation.
JSONL plus verifiers / Inspect-compatible task formats, shipped with train, test, and held-out splits. Holodeck is Gymnasium-compatible, so any base model drops in with no bespoke harness. Deterministic and assertion graders give an RL-usable reward; rubric or human graders provide acceptance and regression signal, not the RL reward. Tool-use and MCP-style agentic trajectories are included.
Two models: training data and environments are priced per verified trajectory or per environment; expert human validation is priced hourly. A rate card is available on request.
The data is licensed and remains our IP; commercial terms are set per engagement.
The data, environment, and benchmark are built and published. Training a model on them is the joint experiment we propose: baseline on the Cyber Defense Benchmark, train in Holodeck, and measure the delta together.

Bring a base model. We'll measure the lift together.

A bounded first experiment: baseline on the Cyber Defense Benchmark, train in Holodeck, measure the delta.

Sign up for Simbian's Newsletter

By submitting this form, you agree to our Privacy Policy.

Ask AI about Simbian