Industry-specific operational data
Geotechnical Boring Logs and Reports as AI Training Data
Quick answer
Geotechnical data for machine learning means four linked record types from subsurface investigations: boring logs (depth intervals, soil descriptions, SPT blow counts, groundwater), laboratory test results, cone penetration test (CPT) soundings and the engineering report that interprets them. Public sets are thin and scattered, so most usable volume sits in engineering firms' project archives. Buyers should specify the exchange format, how legacy PDFs were digitized, coordinate precision, and whether client contracts allow the reports to be reused for model training.
By SourceX Editorial · Updated
What a usable geotechnical record set contains
A usable set links every observation to a hole, a depth and a site, so a model can learn soil behavior in context rather than from isolated numbers. The hole (borehole or sounding) is the parent; samples, test intervals and lab specimens are children keyed by hole ID and top and bottom depth. Without that hierarchy, an SPT N-value or a plasticity index cannot be tied to the stratum it describes.
Expect these components in an engineering firm's archive:
- Boring logs: hole ID, location, ground elevation, drilling method, hammer type, depth intervals, visual descriptions using the Unified Soil Classification System (ASTM D2487 for lab-based classification, ASTM D2488 for visual-manual description), SPT blow counts per increment, sample type and recovery, and groundwater observations with date.
- Laboratory results: moisture content, Atterberg limits, grain-size distributions, unit weight, consolidation curves, triaxial or direct shear results, each keyed to a sample depth.
- CPT and CPTu soundings: cone resistance, sleeve friction and pore pressure at fine depth increments, usually exported as delimited text from the rig's acquisition software.
- Reports: site description, subsurface profile narrative, design parameters, foundation and earthwork recommendations, and appendices that reproduce the logs and lab sheets.
The report is where report-drafting and retrieval use cases get their value: it pairs raw evidence with an engineer's interpretation. That pairing is rare in public data and is the main reason buyers approach firms directly rather than relying on agency portals.
Why public geotechnical data rarely suffices
Public data covers the wrong sites and lacks interpretation, which is why researchers keep turning to synthetic or borrowed data. A 2025 conference paper built GeoSyn, a generator of synthetic geotechnical cross-sections, specifically because real labeled subsurface data was too scarce to train a conditional GAN for CPT interpretation [4]. Another study augmented geotechnical datasets with resource drilling data and named small datasets and the subjectivity of geotechnical logging as core limits [5].
Some state departments of transportation publish boring logs for highway structures, and these are useful for corridor geology. They skew toward bridges and embankments, though, and rarely include the full report or the private-sector building, industrial and residential sites where most commercial tools will be deployed. Mixed provenance is also a licensing hazard: an audit of more than 1,800 text datasets found license information frequently missing or wrong on hosting sites [6], so "found online" is not the same as "cleared for training".
Exchange formats: DIGGS, AGS and proprietary logging exports
Ask for data in a documented exchange format first and accept proprietary exports only with a field map. In the US, DIGGS (Data Interchange for Geotechnical and Geoenvironmental Specialists) is the reference schema: an XML standard built on Geography Markup Language, stewarded by the ASCE Geo-Institute, with the schema, code lists, dictionaries and validation tools developed openly [1]. In the UK and many other markets, the AGS format, created in 1991 by the Association of Geotechnical and Geoenvironmental Specialists, is a software-independent, comma-delimited text format organized into groups such as hole location, samples and in situ tests [2].
Many US consultancies log in commercial boring-log or geotechnical database software and can export to DIGGS, AGS, CSV or a vendor database. Each route has distinct failure modes:
- DIGGS XML: strong structure and units, but schema version drift and optional elements mean two "DIGGS files" may populate different fields. Validate against the published schema [1].
- AGS: easy to parse, but comma and quote errors break automated loading, and key fields and parent-child links must be intact [3]. Regional localizations add or alter groups.
- Proprietary exports: often the richest source (custom fields, project metadata), but codes and units are firm-specific and need a data dictionary from the supplier.
The British Geological Survey's submission checks for AGS are a reasonable acceptance baseline for any format: clean delimiters, coordinates for every borehole, key fields present in every group, and every record tied to a parent hole or sample without duplicates [3].
Digitizing legacy logs without poisoning labels
Treat digitized legacy logs as a separate tier with its own quality metrics, because extraction errors become training labels. Much pre-2000s material, and plenty of later work, exists only as scanned PDF logs and hand-annotated field sheets. Graphic logs mix strata symbols, depth scales, blow-count columns and free-text descriptions, which defeats generic OCR.
Ask suppliers or their processors how they handled:
- depth-scale alignment (are interval boundaries read from the graphic column or the text?),
- SPT increment parsing (6-inch increments versus reported N, refusal notation such as "50/3"),
- symbol-to-USCS mapping and how ambiguous dual symbols (for example SM-SC) were coded,
- handwritten field notes, which often differ from the final typed log.
Request the source image alongside the extracted record, a per-field confidence or review flag, and a double-keyed sample to estimate error. If log digitization is itself your product, the scanned originals are the training data; see our page on licensing scanned forms and handwritten documents.
Coordinates, site identity and confidentiality
Geology needs location, but precise coordinates plus a report can identify a client and a project. Exact borehole coordinates let a model learn regional stratigraphy and join to geologic maps; they also reveal who owns the site and what was built there. Agree the trade-off before data is prepared, not after.
Common options include rounding coordinates to a grid cell, snapping to a regional geologic unit or county, or keeping exact coordinates while removing client names, project numbers, addresses and engineer names from reports. Elevation, groundwater depth and nearby-structure descriptions can also re-identify a site, so check them too. For buyers combining subsurface data with other location layers, our page on licensing geospatial and location data covers precision and use terms.
Reuse rights: client contracts, reliance and professional seals
The firm that drilled the hole does not automatically hold the right to license the report for model training. Geotechnical investigations are performed under client agreements that often assign ownership of instruments of service, restrict use to the named project, or require consent for disclosure. Reports frequently carry reliance language limiting who may rely on them and for what purpose.
For a buyer, that means diligence on each tranche: who owns the data, what the client agreement says, whether reports were submitted to a public agency, and whether sealed engineering judgments may be used to train a drafting model. Datasheets-style documentation of collection process, composition and intended uses [7] gives counsel a record to review. SourceX rights-reviews each dataset for ownership and consents, and delivers under a license that defines the records, uses, term and delivery; every release is approved by the supplying company.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Matching data shape to the model you are building
Each application needs a different slice, so specify the target task before sizing a request. The table below maps common geotechnical AI use cases to the minimum record linkage they need.
| Use case | Minimum data | Key quality risk |
|---|---|---|
| Soil property prediction (e.g., strength or compressibility from index tests) | Lab results keyed to sample depth, USCS class, SPT or CPT at the same interval | Mismatched depths between field and lab records |
| CPT-based soil behavior or stratigraphy | CPTu traces plus an adjacent boring log for ground truth | Distance between paired CPT and boring; unrecorded cone calibration |
| Log digitization and extraction | Scanned logs with verified structured transcription | Label errors from OCR; firm-specific templates |
| Report drafting (SFT) and retrieval (RAG) | Full reports with appendices and the underlying structured logs | Client identifiers in text; boilerplate dominating content |
| Evaluation sets | Held-out projects from different geologic settings and firms | Same-site leakage between train and test |
Split by project and region, not by row, or nearby borings from one site will leak into your test set. For mining-style drillhole data with assays rather than engineering properties, see drill hole logs and assay data; for the review comments engineers leave on design deliverables, see design QA/QC review comments and markups.
A request specification for geotechnical data
A precise request describes the data, not the companies that hold it, and states acceptance tests up front. Use the template below as a starting point.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: geotechnical_investigation_records
use_case: soil property prediction + report drafting SFT
geography: US, mixed coastal plain and glacial till settings
record_types:
boring_logs: {required: true, fields: [hole_id, lat_lon_rounded, ground_elev, method, hammer_type,
depth_top, depth_bottom, uscs, description, spt_increments, recovery, groundwater]}
lab_tests: {required: true, keyed_to: [hole_id, sample_depth],
tests: [moisture, atterberg, grain_size, unit_weight, consolidation]}
cpt: {required: false, fields: [depth, qc, fs, u2], paired_boring_max_distance: "agree with supplier"}
reports: {required: true, include_appendices: true}
formats_accepted: [DIGGS XML, AGS, CSV with data dictionary]
legacy_scans: {accepted: true, require_source_image: true, require_review_flag: true}
coordinates: rounded to agreed grid; exact values withheld
redaction: client, project number, address, engineer names removed from reports
acceptance_tests:
- every child record resolves to a parent hole
- no duplicate hole_id within a project
- units declared per field
- sample of 50 digitized logs double-checked against images
split_policy: by project and region
Personal details such as names, emails and phone numbers that appear in reports are removed or replaced before delivery through SourceX, the method is recorded and a sample is checked; no method is perfect, so keep your own scan in acceptance testing. For broader guidance on tabular test data, see our tabular and time-series data guide.
How SourceX fits a geotechnical data request
SourceX sources operational datasets, including engineering records and documents, from US companies on request; nothing is held in stock and a request does not guarantee a match. The process runs Find, Assess (the data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. SourceX does not publish prices and does not train models. You can see how this works for subsurface and design records on engineering consultancy buyer guidance, browse related categories in the industry-specific operational data hub, or start from the AI data hub. To describe what you need, use the SourceX buyer intake.
Request geotechnical boring log and report data
If your team needs boring logs, lab results, CPT soundings or engineering reports, describe the record types, formats and use case and SourceX will look for US businesses that hold that data. Every dataset is rights-reviewed and delivered through private, access-controlled workflows after an executed agreement and supplier approval. Describe your geotechnical data request.
Sources
- ASCE Geo-Institute, "DIGGS (Data Interchange for Geotechnical and Geoenvironmental Specialists)". https://geoinstitute.org/special-projects/diggs
- The National Archives (UK), "AGS Data Format (PRONOM fmt/1649)". https://www.nationalarchives.gov.uk/PRONOM/fmt/1649
- British Geological Survey, National Geoscience Data Centre, "AGS data format". https://www.bgs.ac.uk/ngdc/?p=518
- EYGEC 2025, "GeoSyn: Synthetic Geotechnical Cross-Sections for Machine Learning Applications" (2025). https://eygec2025.uniri.hr/paper/geosyn-synthetic-geotechnical-cross-sections-for-machine-learning-applications/
- Academia.edu, "Augmenting Geotechnical Datasets With Resource Drilling Data Using Machine Learning". https://www.academia.edu/143019094/Augmenting_Geotechnical_Datasets_With_Resource_Drilling_Data_Using_Machine_Learning
- Longpre et al., arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- 7. Gebru et al., arXiv, "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.