Skip to content

Industry-specific operational data

SOC alert triage decisions as AI training and evaluation data

Quick answer

A useful SOC alert triage dataset pairs each SIEM or EDR alert with the enrichment the analyst saw, the investigation steps and notes, and a confirmed disposition (true positive, benign true positive, false positive, duplicate), ideally linked to incident outcomes and ATT&CK techniques. Public sets are mostly derived from lab traffic, so real analyst-labeled data usually has to be licensed from security operators, with client identifiers tokenized, secrets removed and malware excluded.

By SourceX Editorial · Updated

Why public alert triage benchmarks fall short for agent training

Public benchmarks are useful for prototyping, but none fully reproduces real analyst decisions with investigation context. A 2026 survey of AI-driven alert screening found no public dataset that fully supports alert-prioritization evaluation with real analyst-disposition labels, investigation metadata and longitudinal context [1]. SALAD, a public benchmark, contains 2,778,424 alerts derived from the CIC-IDS2017 and UNSW-NB15 network datasets, with ATT&CK mapping, triage decision, priority and natural-language descriptions [2].

That synthetic origin matters. Labels in lab-derived sets come from known attack injection, not from an analyst weighing a noisy Sigma rule against an asset inventory at 3 a.m. Agent research such as CORTEX [3] and repositories such as SecAlertBench [4] show the direction of the field, but a triage agent sold into MSSPs and enterprise SOCs is judged on production alert mixes: vendor-specific rule names, tuned and untuned detections, and the long tail of benign administrative activity.

The record unit: one alert, its context and its final disposition

The unit buyers should specify is the alert lifecycle, not the raw log line. Each record should carry what triggered, what the analyst knew, what they did, and how it ended. Ask suppliers which source products feed the queue (for example Microsoft Sentinel, Splunk Enterprise Security, CrowdStrike Falcon, SentinelOne, Elastic) because rule naming, severity scales and entity fields differ by product.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "alert_id": "tok_a91f",
  "source_product": "EDR",
  "rule_name": "Suspicious encoded PowerShell command line",
  "vendor_severity": "high",
  "created_at": "2026-03-14T02:41:07Z",
  "entities": {"host": "host_tok_7c2", "user": "user_tok_19", "asset_role": "domain_controller", "user_role": "it_admin"},
  "enrichment": {"ti_matches": [], "prior_alerts_same_host_30d": 4, "change_ticket_linked": true},
  "attack_techniques": ["T1059.001"],
  "investigation_steps": [
    {"step": 1, "action": "query", "summary": "Process tree: parent is SCCM client"},
    {"step": 2, "action": "lookup", "summary": "Change ticket covers scheduled maintenance script"}
  ],
  "analyst_note": "Encoded command matches approved maintenance script; activity expected under change window.",
  "disposition": "benign_true_positive",
  "disposition_basis": "change_ticket_evidence",
  "escalated": false,
  "incident_id": null,
  "closed_by": "analyst_tier2",
  "bulk_closed": false
}

The fields that carry the most training signal are investigation_steps, analyst_note, disposition_basis and incident_id. They let you supervise the reasoning path, not just the label, which is what supervised fine-tuning on investigation notes needs.

Disposition taxonomies differ, so get the supplier's definitions

Disposition labels are only comparable once you have the supplier's written definitions. A benign true positive (the rule fired correctly on legitimate activity) is a different training target from a false positive (the rule logic was wrong), and both differ from "closed as duplicate" or "closed, no action." Merging them teaches an agent that a correct detection on an admin script is a broken rule.

Map every supplier code to a shared schema before you train, and keep the original code. Record whether the supplier escalates by severity, by asset criticality or by analyst judgment, because the escalation flag is only a label if the rule behind it is known.

Supplier codeMeaning to confirmTraining useRisk if merged
True positive, escalatedMalicious or policy-violating, handed to IRPositive class, escalation targetLow
Benign true positiveDetection correct, activity authorizedSuppression and tuning targetAgent learns to call good rules noisy
False positiveRule logic or data wrongDetection-quality labelConfused with authorized activity
Duplicate / correlatedCovered by another alert or caseDeduplication taskInflates benign counts
Closed, no action / auto-closedNo investigation evidenceWeak label or excludeLabel noise from alert fatigue

Alert fatigue makes many closures weak labels

Bulk-closed alerts are the largest source of label noise in operational data. When analysts face more alerts than they can investigate, a share of closures reflect queue pressure rather than evidence, which the alert-fatigue literature documents as a core SOC problem [1]. An agent trained on those closures learns to dismiss what tired humans dismissed.

Ask suppliers for fields that separate evidence-backed decisions from mass closure: closure method (manual, bulk, playbook, auto-close by SOAR), time-to-close, number of investigation queries run, and whether a later incident was linked to the same entity. Weight or filter training data on those fields, and reserve incident-confirmed alerts for your evaluation set. For a deeper treatment of checking outcome fields, see verifying outcome labels in operational records.

ATT&CK mapping lets you measure coverage by technique

Mapping alerts to MITRE ATT&CK techniques turns a flat accuracy number into a coverage map. The ATT&CK Enterprise matrix organizes tactics and techniques across Windows, macOS, Linux, cloud identity, SaaS, IaaS, network and container platforms, so a buyer can see whether a dataset is dominated by Execution and Credential Access alerts while Lateral Movement is barely present.

Ask whether mappings come from the detection rule's metadata, from analyst tagging, or from post-hoc tooling, since rule-level tags describe what the rule intends to catch, not what actually happened. Check technique and sub-technique granularity (T1059 versus T1059.001), and compute per-technique counts of confirmed true positives before you commit to a volume.

Security limits: what must be tokenized, stripped or excluded

Alert data is sensitive in two directions: it exposes the supplier's clients and it can carry live attack material. Before delivery, client names, hostnames, internal and external IP addresses, usernames, email addresses and victim-identifying file hashes should be consistently tokenized so that joins across alerts still work. Command lines, script blocks and analyst notes can contain credentials, API keys and connection strings; secrets in enterprise text need dedicated detection and remediation, which generic PII redaction does not provide [5].

Live malware samples, exploit code and full packet captures should be excluded or replaced with hashes and behavioral summaries. Ask for a data card or equivalent documentation describing sources, collection window, labeling process and known gaps [6], and keep the tokenization method on file for your own security review.

Illustrative example: invented to show structure; it does not describe an available dataset.

Buyer checklist for a SOC alert triage request

  • Source products and rule sets in scope (EDR, SIEM correlation, email, identity, cloud)
  • Disposition codes with written definitions and closure method field
  • Investigation steps, queries and analyst notes included, with secret scanning applied
  • Incident linkage and escalation outcome, with the confirmation basis
  • ATT&CK technique IDs and how they were assigned
  • Tokenization scheme for clients, hosts, IPs, users and hashes, consistent across records
  • Exclusion of malware binaries, exploit code and raw PCAP
  • Collection window, client mix by sector, and alert volume per source product
  • Allowed uses: training, evaluation, or both, and whether notes may be used for SFT

Rights: MSSP client telemetry is usually contractually confidential

The hardest question in MSSP alert data is whether the provider may license data derived from client telemetry at all. Managed security contracts typically treat client logs and alerts as the client's confidential information, so rights review has to confirm that the provider's agreements permit licensing derived, tokenized alert records for AI use. FTC staff have warned AI companies that they can face liability if they break commitments not to use customer data for undisclosed purposes such as model training [7], so check what the provider promised its clients.

In-house SOCs avoid the multi-client problem but raise employee monitoring and internal policy questions. Either way, your counsel will want the chain of permissions documented; the in-house counsel guide to AI data deals covers that review, and the SourceX insight on compliance checks for cybersecurity and MSSP firms describes the supplier-side view.

Evaluating a triage agent: count missed true positives separately

Accuracy alone hides the failure that matters most in a SOC: a confirmed incident closed as benign. Because true positives are a small share of most alert queues, an agent that closes everything can post high accuracy. Report false-negative rate on incident-confirmed alerts separately from false-positive reduction, and break both out by ATT&CK tactic and source product.

Hold out evaluation data by time and by client, not by random split, so the agent cannot memorize a host's history. The same principle appears in outcome-labeled evaluation data, and the adjacent MSP RMM alert-to-remediation records page covers IT-operations alerts rather than security detections. For broader context on industry data, see the industry-specific operational data hub and the SourceX cybersecurity services buyer page.

How SourceX handles alert triage data requests

SourceX sources operational datasets from US companies on request; it does not hold alert data in stock, and a request does not guarantee a match. Buyers describe the alerts, labels and fields they need, and SourceX looks for US businesses that hold that data, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded, and delivery runs through private, access-controlled workflows only after an executed agreement. You can describe your triage dataset requirements to SourceX.

Request SOC alert triage data for training or evaluation

SourceX finds US companies that hold the operational data you describe, assesses data and licensing permissions, and manages the license and ongoing purchases; nothing is contracted until a supplier agrees. Start with the record fields, disposition codes and allowed uses you need at sourcex.si/buyers.

Frequently asked questions

Can I train a triage agent only on public SOC alert datasets?

You can prototype on them, but public sets such as SALAD are derived from lab network datasets [2], so their labels reflect injected attacks rather than analyst judgment on production alerts. Expect a distribution gap when the agent meets real vendor rules and benign admin activity.

Should benign true positives count as negatives?

Only if your task is "needs escalation or not." For detection tuning or suppression tasks, keep benign true positives as a separate class from false positives, because they call for different actions.

Are analyst notes safe to use for fine-tuning?

They are valuable but often contain hostnames, usernames, IPs and pasted credentials. Require tokenization and secret scanning on free text, and confirm the license allows the notes, not just the structured fields, to be used for training.

Sources

  1. arXiv, "AI-Driven Security Alert Screening and Alert Fatigue Mitigation in Security Operations Centers: A Survey" (2026). https://arxiv.org/pdf/2605.08316
  2. Hugging Face (dataset card by Nutthakorn Chalaemwongwan, KMITL), "SALAD: SOC Alert Labeled Analysis Dataset". https://huggingface.co/datasets/nutthakorn7/SALAD-SOC
  3. arXiv, "CORTEX: Collaborative LLM Agents for High-Stakes Alert Triage" (2025). https://arxiv.org/pdf/2510.00311
  4. GitHub, "SecAlertBench". https://github.com/Dxsssu/SecAlertBench
  5. arXiv, "Using AI/ML to Find and Remediate Enterprise Secrets in Code & Document Sharing Platforms" (2024). https://arxiv.org/html/2401.01754v1
  6. Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  7. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data