Industry-specific operational data
Construction Schedule Data for AI: Primavera P6 and MS Project Files with Updates
Quick answer
A useful construction schedule dataset is not a pile of baseline plans. It is a series of native CPM files per project, usually Primavera P6 XER or XML and Microsoft Project MPP or XML, running from the approved baseline through monthly updates to completion. The training value is in the deltas between updates: logic edits, duration changes, actual dates and float erosion. Buy update series with intact relationships and calendars, screen them against DCMA-style checks, and join them to delay and change records before labeling.
By SourceX Editorial · Updated
Why update series beat single schedules
The update series is the asset, because a single schedule shows intent while consecutive updates show how intent failed. For delay prediction, the target variable is usually the movement of forecast finish or a key milestone between data dates, which you can only compute from two or more updates of the same project. For schedule generation, a baseline paired with the as-built final update tells a model which sequencing choices survived contact with the site.
Public research corpora show the scale gap. A Mendeley dataset published in September 2025 offers construction project information aimed at duration estimation [1], and a Cambridge thesis on hybrid machine learning for schedule risk worked with roughly 300 schedules in Python and Primavera formats [2]. One schedule-AI vendor says, as of October 2026, that its training corpus holds over 750,000 real schedules [3]; treat that as a vendor claim, but it shows why commercial teams look beyond open data.
Single-file collections also hide a common failure mode: survivorship. Projects that finished cleanly are more likely to have tidy archives, while disputed jobs have missing or contested updates. If your sample skews toward clean archives, a delay model will underpredict slippage on exactly the projects that matter.
What a P6 XER file actually contains
An XER export is a tab-delimited text file holding a set of P6 tables, and Oracle documents how each table and column maps to P6 fields. The relationship table, TASKPRED, holds one row per logic link with task_id, pred_task_id, pred_type and lag_hr_cnt, so lag is stored in hours rather than days. Activities live in TASK, the work breakdown in PROJWBS, working time in CALENDAR, and resources in RSRC and TASKRSRC.
For modeling, the fields that matter most are activity identity (task_code, task_name), status and dates (planned, early, actual and remaining), total float, constraint type and date, and the project data date. Calendars are not optional: a 10-day duration on a 4x10 calendar and a 7-day calendar produce different finish dates, and dropping CALENDAR rows silently corrupts any float or duration feature. Open-source parsers such as xerparser load XER tables into Python objects, which makes bulk feature extraction practical.
MS Project files bring different traps. MPP is a binary format tied to the authoring version, while the XML export is easier to parse but may lose custom fields or baseline sets the scheduler used. Ask suppliers which baseline slot (Baseline, Baseline1 and so on) holds the contract baseline, because projects frequently re-baseline without renaming.
Screening schedule quality before you train
Screen every file against a DCMA 14-point style assessment before it enters a training set, because poor logic produces float and critical-path values that teach a model the wrong lessons. Commonly cited targets include zero leads (negative lags), lags on no more than about 5% of relationships, and no more than 5% of incomplete activities with total float above 44 working days. Thresholds and denominators vary across practitioner guides, so record which variant you applied.
The other checks cover missing predecessors or successors, hard constraints, relationship types (finish-to-start share), negative float, high durations, invalid dates (actuals after the data date, forecasts before it), resources, missed tasks, critical path test, critical path length index and baseline execution index. For training data, a failing schedule is not automatically useless. A schedule-quality agent needs labeled bad schedules, so keep failures and tag them rather than filtering them out.
| Check | What breaks for modeling | Keep or exclude |
|---|---|---|
| Open ends (missing logic) | Float inflates; critical path is wrong | Keep for quality agents; exclude from float-based delay labels |
| Leads and long lags | Sequencing hidden in lags; float unreliable | Keep, tag at relationship level |
| Hard constraints | Dates forced, not calculated | Keep, record cstr_type as a feature |
| Invalid dates vs data date | Status not trustworthy | Exclude the update; keep the project if others are clean |
| Missing calendars | Durations and float uncomputable | Exclude until calendars are supplied |
| Re-baseline mid-project | Baseline variance meaningless | Keep; flag the new baseline update |
Labeling delay causes from project records
Schedules record that dates moved, not why, so cause labels must come from a join to other project records. The usual candidates are change orders and change order logs, RFIs, daily reports and weather logs, time impact analyses, notices of delay and the narrative that accompanies each monthly update. Submittal review cycles are another strong signal for procurement-driven slippage, which our guide to construction submittal review data covers in depth.
Join keys are the hard part. Activity IDs often change between updates when schedulers renumber or split activities, so build a crosswalk keyed on task_code plus WBS path, and track activity additions and deletions as events in their own right. Dates in daily logs are calendar dates while P6 stores durations in hours against a calendar, so normalize before you compare.
Narratives are worth sourcing even when they are thin. A two-paragraph monthly narrative that names the driving path and cites a pending change gives you weak supervision for cause labels at low cost. Pair schedules with construction cost estimate and bid data when you need cost-loaded forecasts, and with construction progress photo datasets when you want visual evidence for percent-complete claims.
Delivery manifest for a schedule update series
A clear manifest is the simplest way to make a schedule purchase auditable and reproducible. It should list every file per project with its data date, format, authoring version and quality results, so you can detect gaps in the series before training.
Illustrative example: invented to show structure; it does not describe an available dataset.
project_ref: PRJ-0193 # pseudonymous, no owner or site name
sector: healthcare_vertical_build
contract_type: CMAR
calendar_set: [5x8_std, 6x10_concrete, holiday_2024]
updates:
- file: PRJ-0193_BL00.xer
role: contract_baseline
data_date: 2024-02-01
p6_version: "23.12"
activities: 2140
relationships: 3388
dcma: {leads_pct: 0.0, lags_pct: 3.1, high_float_pct: 4.2, open_ends: 7}
- file: PRJ-0193_UP01.xer
role: monthly_update
data_date: 2024-03-01
forecast_finish_delta_days: 3
narrative: PRJ-0193_UP01_narrative.pdf
- file: PRJ-0193_UP07.xer
role: rebaseline
data_date: 2024-09-01
note: "owner-approved recovery schedule"
joins:
change_orders: PRJ-0193_co_log.csv # co_id, issued, days_requested, days_granted
daily_logs: PRJ-0193_dailies.parquet # date, weather_code, crew_count, delay_note
scrubbing:
resources: role_codes_only # names replaced with role codes
cost_loading: removed
Check the manifest for gaps: a missing month between UP03 and UP05 breaks any delta feature for that interval. Also confirm the final update reaches substantial completion, otherwise your labels are censored.
Privacy, confidentiality and owner restrictions
Schedules carry less personal data than most operational records, but they are not clean by default. Resource dictionaries often contain named superintendents, foremen and subcontractor crews, and activity names sometimes embed people or tenant names. Research on process event logs shows that combining timestamps with resource attributes creates measurable re-identification risk [4], so replace named resources with role codes and check activity names and notebook topics for free-text identifiers.
Confidentiality is the larger issue. Owner contracts, especially on public works, healthcare, data centers and defense or utility facilities, can restrict sharing schedules, and some schedules reveal security-relevant sequencing or site layout. Screen for third-party confidential terms using the approach in our third-party confidential information screening guide, and for defense-adjacent work review export-controlled technical data checks. Cost-loaded schedules can expose subcontractor pricing; strip cost loading unless you have specific rights to it.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
How this fits with other construction data
Schedules answer "when and in what order", while other construction records answer "what" and "how much". Specifications and project manuals define scope, and our guide to construction specifications data covers that layer; estimates and bids carry quantities and costs. For a wider view of project records, see the construction project datasets page, the construction management buyer page and project plans and status reports.
If you are comparing industries, the industry-specific operational data hub and the main AI data buyer guides show how the same update-series logic applies to other workflow records. When you are ready to scope a custom collection, a data collection statement of work helps fix formats, series completeness and quality thresholds up front.
Request template for schedule data
A precise request gets better matches, because suppliers need to know which formats, series lengths and joins you actually need. Describe the data, not specific companies. When you are ready, you can describe your schedule data needs to SourceX.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Use case: delay forecasting at the milestone level, plus a schedule-quality agent.
- Formats: P6 XER (version stated) or P6 XML; MS Project XML accepted with baseline slot documented.
- Series: contract baseline plus at least monthly updates through substantial completion; re-baselines flagged.
- Required tables: TASK, TASKPRED, CALENDAR, PROJWBS, PROJECT; resources as role codes.
- Joins wanted: change order log, update narratives, daily logs with weather.
- Quality: DCMA-style results per update; failing files kept and tagged.
- Exclusions: named individuals, cost loading, facilities with owner sharing restrictions.
Source construction schedule data with SourceX
SourceX sources operational datasets, including engineering records and project documents, from US companies on request; it does not hold schedules in stock, and a request does not guarantee a match. Every release is approved by the supplying company, rights-reviewed and delivered under a license that defines records, uses, term and delivery. Describe the schedule data you need.
Sources
- Mendeley Data, "Construction Schedule data set" (2025). https://data.mendeley.com/datasets/9g5ntn3n4f
- University of Cambridge Apollo repository, "A hybrid machine learning method for construction schedule risk analysis". https://www.repository.cam.ac.uk/handle/1810/307046
- nPlan, "Why AI couldn't build construction schedules before, and why it can now". https://www.nplan.io/blog-posts/why-ai-couldnt-build-construction-schedules-before-and-why-it-can-now
- arXiv (CAiSE 2020), "Quantifying the Re-identification Risk of Event Logs for Process Mining" (2020). https://arxiv.org/pdf/2003.10707
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.