Tables, time series and transactional data
Production Log Data for Log Anomaly Detection and LLM Log Analysis
Quick answer
Training data for log anomaly detection should be real application, system and infrastructure logs with labels tied to actual incidents, not only the public HDFS, BGL, Thunderbird and Spirit sets most papers reuse. Buyers should license logs from modern microservice and cloud stacks that span several software versions, carry incident-window and root-cause labels with documented provenance, and arrive scrubbed of secrets, IP addresses and customer identifiers, sampled by agreed time windows and services.
By SourceX Editorial · Updated
Why public log benchmarks no longer cover production behavior
Public log collections are valuable for comparability, but they are narrow and old relative to the systems most teams now monitor. Loghub, the standard public collection, offers 19 real-world log datasets from distributed systems, supercomputers and operating systems, and only a minority carry anomaly labels, such as HDFS block-level labels [1]. A 2026 benchmark of LLM-based log anomaly detection still evaluates on four of those sets (HDFS, BGL, Thunderbird and Spirit) and notes that earlier LLM studies tested on limited data [2]. LogEval, a benchmark suite for LLM log analysis, also draws on BGL and Thunderbird for parsing and anomaly detection tasks [3].
That reuse creates three concrete problems. First, template diversity is low: a parser such as Drain or an LLM parser that scores well on supercomputer RAS logs has rarely seen Kubernetes events, Envoy access logs, JVM stack traces or structured JSON emitted by OpenTelemetry SDKs. Second, the label definitions are dataset-specific (a session, a block, a node-time window), so results do not transfer cleanly to request-level or trace-level anomalies. Third, heavily published sets raise contamination risk when you evaluate large language models that may have seen them during pretraining.
Which log types matter for each model task
Each task needs a different log shape, so name the task before you name the data. The table below maps common AIOps and security ML objectives to the records and labels a supplier would need to provide.
| Task | Useful log sources | Labels or annotations to request | Common failure if missing |
|---|---|---|---|
| Log parsing / template extraction | Application logs across several releases, syslog, web server and proxy access logs | Ground-truth event templates and parameter spans per line | Parser overfits to one release's message formats |
| Sequence anomaly detection | Session- or trace-correlated logs with request IDs, container and pod IDs | Anomaly tag per session, trace or time window | Labels cannot be joined to sequences |
| Incident detection and triage | Service logs aligned with alert and paging history | Incident start/end windows, severity, affected service | Model learns alert noise instead of real incidents |
| Root-cause localization | Multi-service logs plus deployment and config change events | Root-cause service or change, linked to the postmortem | No ground truth for "which component caused it" |
| LLM log-analyst SFT and evaluation | Log excerpts paired with on-call notes, tickets or postmortems | Expert explanation, next action, cited log lines | Answers that sound right but cite nothing |
For root-cause work, pair logs with metrics and traces as covered in telemetry aligned with incident labels; for network operations centers, see NOC alarm and incident logs. Industrial controllers produce a different record type, covered in PLC, SCADA and DCS alarm and event logs.
How to specify labels and label provenance
Labels are only as useful as the evidence behind them, so require the supplier to state how each label was produced. In production logs, the strongest labels usually come from incident records: a ticket or page opened at a known time, a postmortem that names the faulty service, and a resolution timestamp. Weaker labels come from threshold alerts or heuristic rules, which encode the old monitoring system's blind spots.
Ask for a label provenance field on every labeled unit (for example incident_ticket, postmortem, alert_rule, analyst_review) and for the inter-annotator process if humans relabeled anything. Label noise matters even in curated benchmarks: an audit of widely used ML test sets estimated an average label error rate of at least 3.3%, enough to reorder model rankings [5]. Incident-derived labels also tend to have fuzzy boundaries, so agree on how pre-incident degradation and post-fix recovery windows are marked.
Linked records often live outside the log store. SourceX lists related record types such as ITSM ticket datasets, incident postmortems and IT service tickets, which can supply incident-aligned ground truth when matched to the same systems and time ranges.
Scrubbing secrets and personal data from logs before delivery
Logs routinely contain credentials and personal data, so scrubbing must be designed and verified, not assumed. Typical leaks include bearer tokens and API keys in request headers or URLs, database connection strings in startup lines, session cookies, internal and client IP addresses, email addresses, usernames and customer or account IDs embedded in error messages. Research on enterprise secrets shows credentials spread across shared code and documents and calls for dedicated detection and remediation [4].
A defensible process combines pattern- and entropy-based secret scanners (tools in the class of gitleaks or TruffleHog), PII detection for names, emails and phone numbers, and consistent pseudonymization so that the same user or host maps to the same token across a session. Consistency matters for anomaly detection, because replacing every IP with a constant destroys the signal a model needs. See scanning for credentials and secrets in ticket, chat and log data for scanner selection and residual-risk checks.
Through SourceX, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked; no method is perfect. Buyers should still run their own scanners on arrival and treat secret-scanning as part of protecting training data integrity, which NIST SP 800-218A addresses in its secure AI development practices [7].
Sampling, volume and format: what to agree before extraction
Logs are large and repetitive, so define the slice before anyone exports data. A month of debug-level logs from a busy service can be mostly heartbeat and health-check lines; a useful sample instead covers chosen services, chosen time windows around and between incidents, and enough normal baseline to estimate false positive rates.
Illustrative example: invented to show structure; it does not describe an available dataset.
log_data_request:
task: sequence anomaly detection + LLM triage evaluation
stack: Kubernetes microservices, Java and Go services, Envoy ingress
sources: [application_json, envoy_access, k8s_events, deploy_events]
versions: at least 3 release lines per service
sampling:
incident_windows: 2h before to 1h after each labeled incident
baseline_windows: random 1h windows with no open incident
log_levels: INFO and above; DEBUG only inside incident windows
record_fields: [timestamp_utc, service, host_pseudonym, pod_id,
trace_id, level, message_raw, template_id]
labels: [incident_id, window_start, window_end, severity,
root_cause_service, label_source]
scrubbing: secrets scan + PII pseudonymization, consistent per session
format: Parquet partitioned by date and service; templates as CSV
holdout: one service and one month withheld for evaluation
Keep raw messages alongside any parsed templates; parsing errors are themselves training signal. Hold out whole services or time periods, not random lines, to avoid leakage between train and test splits.
Licensing log data for AI: what to check
License terms for log data should be explicit because logs are rarely created with model training in mind. Public dataset licenses are often missing or wrong: an audit of more than 1,800 datasets found license omission above 70% and error rates above 50% on popular hosting sites [6]. For commercial logs, confirm that the supplying company owns the systems that produced them, that customer contracts and privacy notices permit the use, and that the license names the records, permitted uses (training, evaluation, or both), term and delivery method.
SourceX sources operational datasets from US companies on request, rights-reviews each dataset for ownership and consents, and delivers under a license that defines records, uses, term and delivery. Data is not held in stock, and a request does not guarantee a match; every release is approved by the supplying company. Delivery runs through private, access-controlled workflows after an executed agreement, and SourceX does not train models. You can describe your target stack, labels and volume on the SourceX buyers page.
For the broader category, see the tabular, time-series and transactional data buyer's guide, the AI data hub, and SourceX's page on training data for IT operations agents.
Request labeled production logs for anomaly detection
SourceX looks for US businesses that hold the logs you describe and manages the commercial process from assessment of data and licensing permissions through an agreed license and delivery. Nothing is contracted until a supplier agrees, and terms are agreed per deal. Describe the systems, labels and volume you need at sourcex.si/buyers.
Frequently asked questions
Can I just use Loghub for production anomaly detection?
Loghub is the right baseline for comparability with published work, but most of its labeled sets come from older HDFS and supercomputer systems [1][2]. Use it as one benchmark and add licensed logs from your target stack for training and held-out evaluation.
Do unlabeled production logs still have value?
Yes. Unlabeled logs support template mining, self-supervised pretraining and normal-only anomaly models, which flag deviations from baseline behavior. Labels become essential for evaluation and for LLM triage tasks that must name a cause.
How should LLM log-analysis evaluation sets be built?
Pair real log excerpts with incident tickets or postmortems so each question has a verifiable answer, and keep them private to limit contamination. The guide to regression suites for production LLM applications covers how to keep such sets stable over time.
Sources
- arXiv (Zhu et al., LogPAI), "Loghub: A Large Collection of System Log Datasets for AI-driven Log Analytics" (2020). https://arxiv.org/pdf/2008.06448
- arXiv, "LLM-Enhanced Log Anomaly Detection: A Comprehensive Benchmark of Large Language Models for Automated System Diagnostics" (2026). https://arxiv.org/pdf/2604.12218
- arXiv, "LogEval: A Comprehensive Benchmark Suite for Large Language Models In Log Analysis" (2024). https://arxiv.org/pdf/2407.01896
- arXiv, "Using AI/ML to Find and Remediate Enterprise Secrets in Code & Document Sharing Platforms" (2024). https://arxiv.org/html/2401.01754v1
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- National Institute of Standards and Technology, "Secure Software Development Practices for Generative AI and Dual-Use Foundation Models: An SSDF Community Profile (NIST SP 800-218A)" (2024). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-218A.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.