Industry-specific operational data
Well Files for AI: Completion Reports, Workover Histories and Engineering Notes
Quick answer
Well file data for AI comes in two layers. Regulators publish permits, completion reports, plugging records and production volumes, much of it as scanned images. Operators privately hold the rest: workover procedures, failure analyses, artificial-lift changes, wellbore diagrams and engineers' notes that explain why a well was touched. Extraction models can train on the public layer, while production-engineering copilots and intervention-recommendation agents need the private layer, linked by API number to production outcomes.
By SourceX Editorial · Updated
What a complete well file contains
A complete well file is the operator's lifetime record of one wellbore, keyed by its API number, and only some of it was ever filed with a regulator. Treat it as a heterogeneous document archive, not a table. For document AI and retrieval, the useful unit is "every document for one well, in date order, with type labels."
Typical contents, roughly from spud to plug and abandonment:
- Regulatory filings: permit to drill, directional surveys, completion reports, test and potential filings, plugging reports.
- Completion records: perforation intervals, stimulation summaries (stage counts, proppant and fluid volumes), tubing and casing tallies, initial potential tests.
- Intervention history: workover AFEs, rig daily summaries for each job, slickline and coiled-tubing tickets, pump and rod changes, scale and paraffin treatments.
- Engineering work product: failure analyses, artificial-lift design calculations, decline-curve reviews, recompletion proposals, emails and handwritten margin notes.
- Schematics and logs: wellbore diagrams, cement bond logs, and LAS or raster well logs.
Industry writing describes exactly this content problem: key header attributes such as coordinates, water depth and completion date are often buried in scanned reports or handwritten notes [1], and engineers can face hundreds of pages of handwritten logs, emails and manual forms for a single well [2]. That density is what makes well files good training and evaluation material, and what makes them expensive to use. Teams sourcing the wider category of mixed-format archives should also read the owner page on enterprise document datasets.
Public regulatory records versus operator-held files
Public records give you labeled forms at scale, while operator files give you the reasoning, the interventions and the outcomes. Decide which your model needs before you start sourcing.
State oil and gas commissions, such as the Railroad Commission of Texas, publish imaged well records searchable by API number, lease, operator and county, with newer completion filings often submitted electronically and older ones available only as scans or microfilm. Expect inconsistent keys: state code prefixes dropped or kept in API numbers, oil leases that cover several wells, and log images in TIFF. Each of those details becomes a data-engineering task: key normalization, lease-to-well joins and image handling. Confirm the current coverage and reuse terms on each commission's own records pages before planning around them.
Offshore federal wells follow a different format. Operators on the US Outer Continental Shelf file an End of Operations Report on Form BSEE-0125 under 30 CFR 250.744, which records well status and depths, producing zones, perforated intervals and abandonment details in a fixed layout. That fixed layout makes it a clean target for key-value extraction benchmarks; check the current form version and BSEE's public-release rules for proprietary data before building on it.
What regulators do not hold is usually what buyers want most:
| Content | Usually in public filings | Usually only in operator files |
|---|---|---|
| Permit, location, surface and bottom-hole coordinates | Yes | Also in operator files |
| Completion intervals and initial potential | Yes | Fuller detail (stage-level treatment data) |
| Monthly production volumes | Yes, often at lease level | Well-level allocation, test data |
| Workover procedures and job tickets | Rarely | Yes |
| Failure analyses and root cause | No | Yes |
| Artificial-lift design and change history | Partially, if filed | Yes |
| Engineer emails and recommendations | No | Yes |
Lease-level production is a common trap. If a model learns intervention outcomes from lease totals, a workover on one well is confounded with every other well on the lease. Ask for well-level allocated or tested volumes wherever an outcome label matters.
Use cases and the data each one needs
Each well-file application needs a different slice of the archive, so write the request around the model rather than the document pile.
- Well-record extraction (document AI): page images plus ground-truth field values for headers, perforations, casing strings and plugs. Public completion and EOR-style forms suit this well because their fields are defined by the form.
- Document classification and file assembly: whole well files with a type label per document. Commercial tools already classify well files by extracting metadata automatically [4], so a buyer's eval set needs hard cases: misfiled pages, mixed-well scans and undated memos.
- RAG for production engineers: text-searchable files with reliable well, field and date metadata, plus the engineering notes that answer "what did we try last time."
- Workover and intervention recommendation agents: intervention events joined to production before and after each job, with cost and failure cause where recorded. Time-aligned production and notes overlap with paired time-series and text data.
- Acquisition due-diligence copilots: complete files for packages of wells, including title-free technical summaries. This is where rights restrictions are tightest (see below).
Daily drilling reports are a separate source with their own structure; for activity codes, time logs and NPT narratives see daily drilling reports for AI. Well logs in LAS format and seismic are usually licensed through specialist geoscience data brokers and are better scoped separately; parsing LAS files is a well-documented pipeline problem in its own right [3].
Scan quality, OCR and labeling expectations
Most legacy well files are scanned, partly handwritten and inconsistent across operators and decades, so specify document types, image quality and OCR status in the request. A vague request for "well files" returns whatever was easiest to export.
Points to fix up front:
- Image format and resolution: TIFF or PDF, color or bilevel, minimum DPI; whether original microfilm conversions are included.
- OCR layer: none, vendor OCR, or corrected text, and which engine produced it. Handwritten pages need their share stated separately.
- Document-type taxonomy: a fixed label set (for example, permit, completion, workover summary, failure analysis, schematic, correspondence) applied per page or per document.
- Keys: API number at 10, 12 or 14 digits, with sidetrack and completion suffixes resolved; lease and field identifiers as secondary keys.
- Ground truth: for extraction evals, double-keyed field values with an adjudication log. ISO/IEC 5259-4 provides a process framework for data quality in ML training and evaluation data, including labelling [5].
Scanned handwritten material has its own licensing and preparation profile; the owner page on scanned forms and handwritten documents covers it in general terms.
Rights, NDAs and personal data in well files
Operator well files carry rights and privacy issues that public filings do not, and three of them recur: data-room NDAs, joint-venture ownership and personal data in land and royalty records.
- Acquisition data rooms: files received during a divestiture process are typically under a confidentiality agreement restricting use to evaluating the transaction. A company that bought wells may hold full files, but files it only reviewed in a data room are usually not its to license.
- Partner and operator rights: non-operating working-interest owners receive reports under joint operating agreements; confirm who controls the technical data before relying on a non-operator's copy.
- Personal data: royalty-owner names and addresses, division orders, landowner correspondence and lease records identify individuals. Exclude land and royalty folders unless the use case truly needs them. Engineering notes also name field staff and contractors, and small-field context can re-identify people; see indirect identifiers in business text.
- Safety and incident content: failure analyses can reference injuries or spills. Those records follow the patterns in HSE incident and near-miss reports.
Ask any supplier to state the provenance of each document class: originally generated by the operator, received from a partner, or obtained from a regulator. Public records are often free or cheap to retrieve but still need their reuse terms checked against your intended use.
Specifying a well file request
A well file request should define the wells, the document classes, the linkage to outcomes and the delivery format, so a supplier can tell quickly whether it holds a match.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example specification |
|---|---|
| Basin and well type | Onshore Permian horizontal oil wells, rod pump or gas lift |
| Vintage | Spud 2010 or later; interventions 2015 to present |
| Document classes | Completion reports, workover summaries and tickets, failure analyses, lift change records, wellbore diagrams, engineer notes |
| Exclusions | Land, title, royalty and division-order files; LAS logs; seismic |
| Keys | 14-digit API number; operator internal well ID mapped to API |
| Outcome linkage | Well-level daily or monthly production 180 days before and after each intervention |
| Format | Searchable PDF or TIFF plus OCR text in JSONL with page-level document-type labels |
| Volume | Stated as wells and intervention events, not pages |
| Use | Fine-tuning an extraction model and a held-out eval set for an intervention-recommendation agent |
An illustrative record for one intervention event, after linkage:
{
"api14": "42XXXXXXXX0000",
"event_type": "workover",
"event_start": "2021-03-04",
"job_summary": "Pulled rods and tubing; found rod part at 4,212 ft; replaced 18 rods and pump.",
"failure_cause_coded": "rod_part_wear",
"lift_before": "rod_pump",
"lift_after": "rod_pump",
"source_docs": ["workover_summary_p1-3.pdf", "eng_note_2021-03-09.pdf"],
"oil_bpd_90d_before": 41.2,
"oil_bpd_90d_after": 63.5
}
Assessing a sample before you license
A sample of 20 to 50 complete well files, chosen by the supplier at random from the stated scope, tells you more than any data dictionary. Run your extraction pipeline on it and check:
- share of pages that are handwritten or below your OCR quality threshold;
- whether every document resolves to one API number, and how many are mixed-well scans;
- how many interventions have a recorded cause and matching production data;
- whether redaction of personal details left engineering content readable;
- date coverage gaps, such as a missing decade after an operator change.
How SourceX approaches well file requests
SourceX sources operational datasets from US companies on request, including engineering records and documents, and manages the commercial process through licensing agreements and ongoing purchases. Nothing is held in stock, and a request does not guarantee a match. Buyers describe the data they need, SourceX looks for US businesses that hold it, and every release is approved by the supplying company.
Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval; if your spec is ready, you can describe your well file requirements to SourceX. For other operational sources by sector, start at the industry data hub.
Source well file data for your AI project
If you need completion reports, workover histories or engineering notes linked to production, describe the wells, document classes and use in a buyer request. SourceX will assess whether a US operator holds matching data and what its licensing permits before anything is agreed. Submit your well file data request.
Sources
- Journal of Petroleum Technology (SPE), "AI/ML Offers New Solutions for Data Challenges in the Energy Industry". https://jpt.spe.org/ai-ml-offers-new-solutions-for-data-challenges-in-the-energy-industry
- LTIMindtree, "AI-Powered Well Data Management for Instant Engineering Insights" (2025). https://www.ltm.com/uploads/povs/2025/11/AI-Powered-Well-Data-Management-for-Instant-Engineering-Insights-PDF.pdf
- Databricks Community technical blog, "Processing LAS well log files and other semi-structured data". https://community.databricks.com/t5/technical-blog/processing-las-well-log-files-and-other-semi-structured-data/ba-p/133583
- American Oil & Gas Reporter, "Software optimizes well data management". https://www.aogr.com/web-exclusives/exclusive-story/software-optimizes-well-data-management
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality process framework" (2024). https://www.iso.org/standard/81093.html
- BSEE via LII / Legal Information Institute, "30 CFR § 250.744 - What are the end of operation reporting requirements?". https://www.law.cornell.edu/cfr/text/30/250.744
- Railroad Commission of Texas, "Oil and Gas Well Records - Online". https://www.rrc.texas.gov/oil-and-gas/research-statistics/obtaining-commission-records/oil-and-gas-well-records-online
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.