Industry-specific operational data
Resident Maintenance Requests for AI: From Request Text to Completed Work Order
Quick answer
Useful property maintenance request data for AI links what a resident actually wrote, and any photos, to what happened next: the category and priority assigned, whether it was treated as an emergency, permission to enter, the technician or vendor dispatched, notes, parts, cost, completion time and whether the same unit reported the problem again. Request text alone trains a classifier. The full chain trains triage, troubleshooting chat and dispatch agents, and lets you evaluate them against real outcomes.
By SourceX Editorial · Updated
What a complete maintenance request record contains
A training-grade record joins the resident's intake message to the work order lifecycle, not just to a category label. Intake agents typically collect contact details, an issue description and photos [2]. For model training you want the parts those templates skip: the decisions staff made and how the job ended.
Property management systems such as Yardi Voyager, RealPage, AppFolio, Entrata and Buildium, and CMMS tools layered on them, export work orders with different field names, status codes and category trees. Ask every supplier for a data dictionary and the category taxonomy as it existed during the export window, because categories are often renamed or merged mid-year. A "Plumbing - Leak" code in one portfolio may be split into "Leak - Active" and "Leak - Stain" in another.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"request_id": "WO-0000418",
"submitted_at": "2025-01-14T06:52:00-05:00",
"channel": "resident_portal",
"request_text": "No heat since last night, thermostat says 58. Baby in the apartment.",
"photos": ["ph_01.jpg"],
"resident_category": "Heating/AC",
"staff_category": "HVAC - No heat",
"priority": "emergency",
"emergency_rule": "no_heat_below_threshold",
"permission_to_enter": true,
"pets_present": "unknown",
"assigned_to": {"type": "in_house_tech", "id": "TECH_07"},
"dispatched_at": "2025-01-14T07:20:00-05:00",
"tech_notes": "Ignitor failed on furnace. Replaced. Cycled 3x OK.",
"parts": [{"desc": "hot surface ignitor", "qty": 1}],
"labor_minutes": 55,
"completed_at": "2025-01-14T09:05:00-05:00",
"resident_rating": 5,
"repeat_request_30d": false,
"property_class": "garden_150_units",
"region": "US-Northeast"
}
Unit numbers, resident names and phone numbers are absent by design; the record keeps only pseudonymous IDs and coarse property attributes.
Emergency definitions drive triage labels
The most valuable label in this data is whether the emergency decision was correct, because a missed "active leak" or "gas smell" costs far more than a false alarm. Each operator writes its own emergency policy, usually covering no heat, no water, active leaks, gas odor, sewage backup, lockouts, fire or smoke alarms and inoperable entry doors. Several of those overlap with state habitability and repair statutes, which differ by state, so each record should carry the property's state.
Ask suppliers for the written emergency policy and its version dates, then check whether priority labels follow it. Useful derived labels include:
- Under-triage: routine priority, but the technician notes show an active leak or a gas shutoff.
- Over-triage: emergency dispatch for a cosmetic issue, often visible as a short after-hours call with "no issue found."
- Wrong trade: a "plumbing" ticket closed by an electrician or reassigned within hours.
- Safety-language misses: text containing "smell gas," "sparking" or "water coming through ceiling" that was not escalated.
Where a request mentions gas odor or carbon monoxide, the correct intake behavior is usually to direct the resident to leave and call the utility or 911, not to troubleshoot. Keep those conversations in an evaluation set with expected refusals.
Troubleshooting chat and dispatch agent data
Resident chat agents need the back-and-forth, not just the final ticket. Look for request threads where staff asked clarifying questions ("Is water dripping or pooling?", "Is the breaker tripped?") and records where a self-help step closed the ticket without a visit, such as a GFCI reset, a garbage disposal reset key or a thermostat battery swap. Those are the positive examples for deflection; the reopened ones are the negatives.
Dispatch agents need the resources and constraints behind each assignment: in-house technician skills, vendor trade and coverage, after-hours rates, access windows, permission-to-enter flags and pet notes. Vendor maintenance AI commonly relies on response delays, repeat requests, vendor completion rates and tenant communication logs [1]. For evaluation, a tool-and-policy setup like tau-bench, which tests agents against simulated users, programmatic APIs, realistic databases and policy documents [3], adapts well: replay historical requests with the operator's emergency policy as the policy document and the work order system as the tool.
The closest neighboring data types are covered elsewhere: hotel guest requests have a different urgency and messaging profile (see hotel guest service request data), and general field service maintenance records cover technician work orders without resident intake text.
Outcome labels from repeat requests and completion data
Repeat requests on the same unit and asset are the cheapest reliable outcome label in this domain. A second "no heat" ticket within 30 days on the same unit suggests the first fix failed or the diagnosis was wrong. Join on unit and asset IDs before de-identification, then keep only the derived flags and pseudonymous keys.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Label | How to derive it | Common failure mode |
|---|---|---|
| Fix held | No same-category request on the unit within N days | Resident moved out, so silence is not success |
| First-time fix | One visit, no "parts on order" status | Status codes used inconsistently across sites |
| Triage correct | Priority matches policy given tech notes | Policy changed mid-period |
| Resident satisfaction | Post-completion survey rating | Low response rates skew toward complaints |
| Time to complete | Completed minus submitted timestamps | Time zones and batch closures at month end |
Check completion timestamps for bulk closures, where staff mark dozens of tickets complete at the end of a shift or month. Those records corrupt any time-to-resolution label. Our training data quality guide covers broader coverage and leakage checks.
Privacy: residents, units and photos
Resident maintenance data is personal data even after names are removed, because a unit number plus a timestamp identifies a household. Request text routinely contains names of children, medical details ("I use oxygen"), immigration or employment hints, and access codes. Photos show people, pets, mail with names, medications and the interior of a home.
Ask suppliers how they remove or replace names, emails, phone numbers, unit numbers, gate and lockbox codes and street addresses, and how they handle faces and documents in photos. Ask for the method in writing and a sample check result. Our de-identified AI training data guide covers methods and residual risk, and the guide to captions from work records covers pairing technician notes with photos.
Fair housing and response-time bias
Response-time and priority labels can encode unequal treatment across properties, and a model trained on them will reproduce it. Fair Housing Act disparate-impact theory, including HUD's discriminatory-effects regulation at 24 CFR 100.500, can reach housing practices that disadvantage a protected group without any intent; confirm the current status of that regulation with counsel as of October 2026.
Practically, compare median time to dispatch and emergency rates by property, region and building age before training. If one property class consistently waits longer for the same category, do not use raw response times as a target. NIST's AI RMF MEASURE function gives a framework for documenting those checks [4].
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Buyer request checklist for maintenance request data
A clear request names the agent you are building, the decisions it must make and the labels you need.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Use case: intake classification, emergency triage, troubleshooting chat, dispatch, or evaluation only.
- Record scope: residential multifamily, single-family rental, student or senior housing; US regions; date range covering at least one heating and one cooling season.
- Fields: original request text, photos, resident and staff category, priority, emergency rule, permission to enter, assignment, tech notes, parts, cost, timestamps, survey, repeat flags.
- Documents: data dictionary, category taxonomy history, emergency policy versions, after-hours procedure.
- Privacy: de-identification method, photo handling, sample check.
- Rights: confirmation the operator can license resident-submitted content for AI training, and the allowed uses.
Use our guide to requesting a training data sample to test field completeness before committing, and see the industries hub for related operational data such as property claim estimates and adjuster reports.
How SourceX sources resident maintenance request data
SourceX sources operational datasets, such as support histories and engineering records, from US companies on request; nothing is held in stock and a request does not guarantee a match. Buyers describe the data, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with personal details removed or replaced before delivery. Property teams can also see the property management buyer page, the maintenance dispatch workflow and customer support ticket datasets, or describe the data on the buyers page.
License resident maintenance request data for your triage agent
SourceX sources operational records from US companies, manages licensing and ongoing purchases, and serves AI teams wherever they are based. Nothing is contracted until a supplier agrees, and terms are agreed per deal. Describe the maintenance request data you need.
Frequently asked questions
Is request text enough to train a maintenance triage model?
Text alone trains a category classifier, but triage needs outcome labels. Without technician notes and repeat-request flags you cannot tell whether the original priority was right, so the model learns staff habits rather than correct decisions.
Why ask for the category taxonomy history?
Operators rename, split and merge categories over time. Without a dated taxonomy, the same issue carries different labels in different months, which looks like label noise during training.
Should synthetic maintenance requests replace real ones?
Synthetic requests help cover rare emergencies, but they miss how residents actually describe problems: misspellings, mixed languages, vague symptoms and photos taken in poor light. Use real records as the base and the evaluation set.
Sources
- Oxmaint, "AI-Powered Tenant Experience Optimization Through Maintenance Data". https://oxmaint.com/industries/property-management/ai-tenant-experience-optimization-maintenance-data
- Jotform, "Property Maintenance Request AI Agent". https://jotform.com/agent-templates/property-maintenance-request-ai-agent
- Sierra Research, "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.