Loading...
Loading...
A threat hunting framework is a repeatable process for proactively searching for threats your detections missed, from a written hypothesis to a recorded verdict. PEAK (Splunk, 2023) and TaHiTI (2018) define the steps, MITRE ATT&CK supplies the techniques, and the Hunting Maturity Model grades the program. Every hunt ends in a verdict, even one that finds nothing.
Only 37% of organizations now run a formally defined threat hunting methodology, down from 51% in 2024, according to the SANS 2026 Threat Hunting Survey. Formal measurement of hunt outcomes fell faster, to 40% from 64%. Fewer teams write down what their hunts produced, and a program that can't show results is hard to defend when budgets get cut.
We build the AI Threat Hunt Agent at Simbian, and the six steps below are the loop it runs, so treat this as a vendor's view. You can run every step by hand.
A threat hunting framework is the written process your team follows for proactive threat hunting, from hypothesis to conclusion. It lets two hunters given the same question run the same hunt and report the same kind of result. Without one, hunting depends on whoever's on shift, and you can't compare this quarter's results with last quarter's.
Five names come up when people say framework, and two of them, PEAK and TaHiTI, walk you through a cyber threat hunt from start to finish. The Sqrrl Hunting Loop describes the same cycle at a higher level.
| Name | Origin | What it is |
|---|---|---|
| PEAK | David J. Bianco and Ryan Fetterman, Splunk SURGe, 2023 | Hunt process: Prepare, Execute, and Act with Knowledge, across hypothesis-driven, baseline, and model-assisted (M-ATH) hunts |
| TaHiTI (Targeted Hunting integrating Threat Intelligence) | Dutch financial institutions, 2018 | Intel-driven hunt process: Initiate, Hunt, Finalize |
| Sqrrl Hunting Loop | Sqrrl, 2015 | Four-stage cycle that turns hunts into automated analytics |
| MITRE ATT&CK | MITRE | Technique library for choosing hunts and mapping coverage |
| Hunting Maturity Model (HMM) | David J. Bianco, 2015 | Five-level scale for grading the program, HMM0 to HMM4 |
In practice you pick the behavior from ATT&CK, run the hunt with PEAK or TaHiTI, and grade the program against the HMM. Our threat hunting reference in the Learning Center covers each one in depth.
Picture the quarterly review. Your hunt lead reports a busy quarter with nothing alarming found, and your CISO asks what the hunt team's time bought.
If nobody wrote down a verdict or the gaps, there's no good answer.
The SANS survey points to part of the problem. Its 2026 edition found that, for the first time in the survey's history, data quality or quantity (50%) outranked a shortage of skilled staff (45%) as the top barrier to hunting. Ad hoc approaches climbed back to 39%.
Our read goes past what SANS measured. Hunters start from a hypothesis before anyone checks whether the logs can test it, so they chase data that isn't there. With no exit rule, hunts drift for weeks. Results go uncounted because nobody defined a verdict.
Then the framework gets the blame.
The intrusions hunting exists to catch are the ones that stay hidden longest. In Mandiant's M-Trends 2026, the median dwell time for cyber espionage and North Korean IT worker intrusions was 122 days, against 14 days overall.
TaHiTI already gets part of this right, because every hypothesis has to end proven, disproven, or inconclusive.
This threat hunting methodology has six steps: map the data, baseline normal, find anomalies, deep-dive candidates, try to prove it benign, and conclude. Each one is finished only when it produces something you could hand to another hunter.
| Step | What it produces |
|---|---|
| 1. Map the data you can query | The logs the hypothesis needs, marked available or missing |
| 2. Baseline normal | What normal looks like for this behavior |
| 3. Find the anomalies | A short candidate list |
| 4. Deep-dive each candidate | An evidence chain per candidate |
| 5. Try to prove it benign | Every alternative tested, and why it held or failed |
| 6. Conclude | A verdict, a hypothesis status, an ATT&CK map, and the gaps |
To see all six in action, take Kerberoasting (ATT&CK technique T1558.003), one of the classic threat hunting examples. The hypothesis: in the last 90 days, a compromised domain account requested Kerberos service tickets for service accounts to crack their passwords offline.
Kerberoasting shows up in Windows Security event 4769 (a Kerberos service ticket was requested), which records the ticket encryption type. The event only exists on domain controllers that audit Kerberos service ticket operations, and it doesn't say which process asked, so you'll want EDR process telemetry from workstations as well, whether it lives in your SIEM or your EDR console. Check that nobody filters 4769 at the forwarder to save volume, and that retention reaches back 90 days. If two branch domain controllers don't forward Security logs at all, write that down now. It's a disclosed gap, and step 6 has to account for it.
Count which accounts request service tickets, for which services, how often, and with which encryption. Drop service names that are computer accounts (ending in $) and krbtgt first, or they'll swamp everything else.
Look for a user account that asks for tickets to dozens of services within minutes. RC4 tickets (type 0x17) in a domain that otherwise negotiates AES are another flag, and so are requests from a workstation that has never touched those services. In Splunk, a first pass looks like this (field names vary by add-on):
index=wineventlog EventCode=4769 Ticket_Encryption_Type=0x17 Failure_Code=0x0 Service_Name!="*$" Service_Name!="krbtgt"
| bin _time span=15m
| stats dc(Service_Name) AS services BY _time, Account_Name, Client_Address
| where services >= 10
| sort - services
Tune the 10 and the 15-minute window against your baseline. Attackers can request AES tickets too, so treat the RC4 filter as a starting point and lean on the volume and new-workstation signals.
Start with who was logged into the host and which process made the requests (a known offensive tool or an unexpected PowerShell session is a strong signal), then follow what happened next. If one of those service accounts later signs in from somewhere new, treat the password as likely cracked and check for lateral movement.
When the clock runs short, this step gets cut first. A skeptical reviewer will ask about it anyway. Go looking for the boring explanation. A vulnerability scanner or Active Directory audit tool may be requesting tickets to test for weak service-account passwords. It could be a legacy application pinned to RC4, or a backup client that always pulls tickets in bulk. Write down each one you tested and whether it held.
The report carries a verdict and a hypothesis status. It maps ATT&CK techniques to the accounts and hosts involved, names any detection worth adding, and lists the gaps from step 1, each marked with whether it could change the verdict.
Already running PEAK or TaHiTI? Keep it. These steps sit inside PEAK's Prepare, Execute, and Act phases and TaHiTI's Initiate, Hunt, and Finalize. What changes is the emphasis, with a named output for every step and a step of its own for testing benign explanations.
A threat hunt is finished when every benign alternative has been tested, when the remaining questions need data you don't have, or when its time box runs out. Each one ends in a verdict and a note on how much of the hypothesis you tested. Put all three exit conditions in the hunt plan before the first query runs:
Then report the result in a fixed vocabulary. This one comes from our own product. Borrow it.
| Verdict | Meaning | What happens next |
|---|---|---|
| Confirmed Threat | Malicious, and every benign alternative failed | Hand to incident response |
| Suspicious Activity | Likely malicious, but the evidence is incomplete, often because logs are missing | A human digs deeper and the logging gets fixed |
| Detection Opportunity | Legitimate, but shaped like an attack your rules wouldn't catch | Detection engineering writes or tunes the rule |
| Benign | Explained, and the explanation held up | Record it so nobody re-hunts it next quarter |
Pair each verdict with a hypothesis status: Validated (fully tested), Partially validated (gaps blocked part of the test), or Invalidated (tested and refuted). TaHiTI's inconclusive outcome sits closest to Partially validated.
Back to Kerberoasting. Say the burst of RC4 tickets traces to a vulnerability scanner that security signed off on. The activity was approved, and your rules stayed silent while one user account pulled RC4 tickets for dozens of services.
Call it a Detection Opportunity.
The same requests from an unapproved account would sail through your rules today, and detection engineering needs to hear that. The status is Partially validated, because two branch domain controllers never forwarded logs and the hypothesis is refuted only where you could see. That gap gets its own fix in the report, next to a misconfiguration the hunt turned up along the way. Those service accounts still accept RC4, so enforcing AES goes on the list too.
A threat hunt that comes back clean proves as much as its coverage allows. It confirms one attack path was checked across specific systems over a specific window, and it shows which detections are missing and which data you can't see. Write those three down even when the verdict is Benign.
"You can have a successful hunt even if you didn't find what you were looking for, as long as you are satisfied that you would have found it, had it been present," Bianco wrote in Splunk's guide to turning hunts into detections with PEAK.
So report what each hunt produced. PEAK proposes five metrics:
Bianco frames the difference as "We performed nine hunts this quarter" against "We put twelve new automated detections into production this quarter." Bring the second sentence to your next quarterly review.
For executives, describe a clean hunt as a compromise assessment covering the path you checked, the systems, the window, and what you still can't see.
Those endings only work if the hypothesis was sharp to begin with. A good threat hunting hypothesis names a behavior, a scope, and a time window, and the data you already have can prove it false. A reusable template: in the last [time window], [account or actor] used [ATT&CK technique] against [scope], visible in [data source].
A scoped hypothesis knows what it's looking for: two credentials tied to the CFO appeared in a leak; did either authenticate anywhere after the leak date? It can skip the baseline and go straight to deep dives. A vague question like is any employee involved in, or a victim of, suspicious activity? can't be tested as written. It needs all six steps, and the baseline does the work of narrowing it into hypotheses the data can disprove.
| Source | Example hypothesis | ATT&CK | Data needed |
|---|---|---|---|
| Threat intel report | Non-admin users created scheduled tasks on servers in the last 30 days, matching an actor's documented persistence | T1053.005 | Process creation and task scheduler logs |
| Cluster of SOC alerts | Two weeks of low-severity impossible-travel alerts on finance users are one credential-stuffing campaign | T1110.004 | Identity provider and VPN sign-in logs |
| Pentest finding | Someone used the privilege-escalation path the pentest proved, before it was fixed | T1068 | EDR telemetry for affected hosts |
| New exposure | A user consented to a malicious OAuth app that now reads mail | T1528 | Cloud app consent and mailbox audit logs |
Sweeping for hashes and IP addresses from a report still has its place, but those indicators sit near the bottom of Bianco's Pyramid of Pain, where attackers change things cheaply. MITRE's 2019 TTP-Based Hunting report calls indicator-based detection "brittle." All four examples above hunt for a behavior.
The Hunting Maturity Model (HMM), published by David J. Bianco in 2015, grades a hunting program from HMM0 to HMM4 on three factors: the quality of the data it collects, the tools it gives analysts, and the skills of those analysts. If you're stuck at one level, look at whichever of the three is weakest.
| Level | Bianco's definition, condensed |
|---|---|
| HMM0 Initial | Relies on automated alerting; no hunting |
| HMM1 Minimal | Searches historical data for indicators from threat intel |
| HMM2 Procedural | Follows hunting procedures others developed |
| HMM3 Innovative | Creates and publishes its own procedures |
| HMM4 Leading | Turns successful hunting procedures into automated detection |
Bianco observed that "HMM2 is the most common level of capability among organizations that have active hunting programs." He also wrote that "you can't buy your way to HMM4," and that includes tools like ours. A tool can make each hypothesis cheaper to test, and moving up a level depends on whether your team turns the results into procedures and detections.
AI fits in the expensive middle of a threat hunting framework, where querying, baselining, and testing alternatives across months of logs eat most of the hours. The hunter owns both ends, choosing the hypothesis and signing off on the verdict and any detection it produces.
| Step | What the AI does | What the hunter decides |
|---|---|---|
| Hypothesis | Proposes hunts from threat intel and alert patterns | Which hypotheses to run |
| 1–2 Map and baseline | Inventories log sources, flags missing ones, computes normal | Whether the gaps are acceptable |
| 3–4 Anomalies and deep dives | Writes and runs queries across sources in parallel | Where to push deeper |
| 5 Prove it benign | Tests and records benign alternatives | Which rejected alternatives deserve a second look |
| 6 Verdict | Drafts the verdict, ATT&CK map, gaps, and detection suggestions | Approves the verdict and which detections ship |
A common objection is that an automated hunt is just a detection, and it's half right. A detection runs one fixed query forever, and a hunt runs a different investigation for each hypothesis and ends with a judgment. Once a hunt finds something worth catching every time, it should become a detection. Bianco's HMM4 level is that move.
In the SANS 2026 survey, the share of teams ranking AI or machine learning among their top planned hunting improvements fell to 39% from 48% a year earlier. Whatever's behind the dip, any AI verdict faces the same test: if a hunter can't trace it back to the queries behind it, approving it is guesswork.
We built the AI Threat Hunt Agent to pass that test. It runs all six steps against the logs you already have, in Microsoft Sentinel, Splunk, CrowdStrike, or Microsoft Defender, without moving data. Each report lists the benign explanations it tested and rejected, plus every gap it hit with a severity and a flag for whether it affected the verdict. You can download every query it ran and check the work yourself.
When disproving a hypothesis costs less analyst time, the broad questions you've been deferring get easier to justify.
Q: Is MITRE ATT&CK a threat hunting framework? MITRE ATT&CK is the technique library most threat hunting frameworks build on. It tells you what to hunt for and how to map coverage. For how a hunt runs and ends, a common pattern is to pair it with a process framework such as PEAK or TaHiTI and record each hunt's techniques against ATT&CK IDs.
Q: Does NIST have a threat hunting framework? NIST covers hunting as a control. NIST SP 800-53 Revision 5 includes control RA-10, Threat Hunting, which asks organizations to establish and maintain a cyber threat hunting capability. That capability should search for indicators of compromise and detect, track, and disrupt threats that evade existing controls, at a frequency the organization defines. PEAK or TaHiTI can supply the step-by-step method behind it.
Q: What goes in a threat hunting process document? Two parts. The hunt plan states the hypothesis, the data sources, the time box, and the exit conditions before the first query runs. The hunt report records the verdict, the hypothesis status, the ATT&CK techniques with the accounts and hosts involved, every benign explanation tested, any detection worth writing, and each data gap. Someone else should be able to rerun the hunt from it.
Q: Who should sign off on a threat hunt verdict? A lead or a hunter who did not run the hunt should approve every verdict before it closes, Confirmed Threat and Benign above all. The reviewer reads the rejected alternatives and the gaps before the conclusion. When an AI Agent runs the hunt, the execution log is what the reviewer audits.
Q: How long should a threat hunt take? No standard duration exists. PEAK suggests considering a maximum duration when you scope the hunt, sized to the hypothesis, so a scoped question needs a fraction of the time a vague one does. When the time box runs out, the hunt still ends in a verdict, and the untested parts go in the report as gaps.
Q: What is the difference between threat hunting and detection engineering? Threat hunting tests a hypothesis that something got past your rules and ends in a verdict. Detection engineering writes and maintains those rules. The two feed each other, since a hunt's Detection Opportunity verdicts become new rules and every new rule narrows what the next hunt has to cover.
Q: What should you ask of an agentic threat hunting framework? Ask it to show its work. An agentic threat hunting framework hands the query-heavy steps to an AI Agent, while a person chooses the hypothesis and approves the verdict. Before you trust one, check that it lists the queries it ran, the alternatives it rejected, and the data gaps it hit.
The quickest test of your threat hunting framework is the last hunt that found nothing. If you can't say what it proved, start by fixing the ending. To watch the six steps run against your own logs, book a demo and bring the hypothesis you keep putting off.