Software companies
What the RealPage settlement means for SaaS vendors training AI on pooled data
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
The RealPage settlement matters to SaaS vendors because the case treats software fed with competitors' nonpublic data as a possible channel for price coordination. Any vendor that pools customer data to train models or make recommendations should check data age, aggregation level and which features touch pricing, and confirm the final terms on the court docket.
Key takeaways
- The RealPage case concerns pooled nonpublic data from competing landlords feeding pricing software, not AI training in general.
- Settlement terms bind the parties, but they show which data practices enforcers consider risky.
- Risk rises when pooled data is recent, granular and used to recommend prices, wages or bids to competitors.
- Licensing de-identified operational records such as support tickets is a different question from pooling competitors' pricing data.
- Confirm the entered judgment's text and court approval status before relying on any summary, including this one.
What is the RealPage case about?#
The RealPage case is a Justice Department antitrust action alleging that RealPage's revenue management software used nonpublic, competitively sensitive data from competing landlords to generate rent recommendations that could align pricing across a market. RealPage disputed the allegations, and the company and the government later reached a proposed settlement.
For software vendors outside property management, the case matters less for its specific terms than for its theory. A model trained on pooled customer data can become the channel through which competitors share information they could never lawfully exchange directly, and that theory reaches any vertical SaaS product whose customers compete with each other.
This analysis does not restate the settlement's terms. Settlements can change between proposal and final judgment, so read the filed documents on the court docket and confirm approval status before relying on any summary. This is general information, not legal advice.
Which settlement terms should vendors read closely?#
When you read the filed judgment, look at how it treats the dimensions below. They are the levers a decree on pooled data tends to pull, and they map directly onto choices your own product team makes every quarter.
| Dimension to check | What to look for in the final text | Why it matters for your product |
|---|---|---|
| Data age | Any minimum age for nonpublic customer data used in models | Older data is less useful for coordinating current prices |
| Aggregation level | The geographic or portfolio granularity allowed in model inputs and outputs | Fine-grained data can reveal an individual competitor's position |
| Feature scope | Which features are covered: pricing recommendations, benchmarks, analytics | Restrictions may reach features beyond the core pricing tool |
| User controls | Whether customers must make their own pricing decisions | Defaults that push acceptance can look like coordination |
| Communications | Limits on sharing competitor information in meetings, user groups or support | Information exchange can happen outside the software too |
| Compliance and monitoring | Monitors, audits, reporting duties and the decree's duration | Signals what documentation enforcers expect a vendor to keep |
Why pooled customer data is the real issue for vertical SaaS#
Pooled customer data is the issue because vertical SaaS vendors often serve many of the competitors in one niche. A property management platform, a freight rate tool or a contractor estimating product can end up holding current, nonpublic pricing, capacity and cost data from businesses that compete for the same customers.
Training a model on that pool is not unlawful in itself. Antitrust risk tends to rise when the model's outputs steer competitors' prices, wages or bids, when inputs are recent and granular enough to reveal a rival's position, and when customers are nudged to follow recommendations rather than set their own terms. Each factor is assessed case by case with antitrust counsel under the Sherman Act and state laws.
How different uses of pooled data compare#
Different uses of pooled data carry different sensitivity. The tiers below are a general way to triage your own products before a legal review; they are not a safe harbor and do not replace counsel's assessment.
| Use of pooled customer data | Example | General sensitivity | What to check |
|---|---|---|---|
| Price, rate or wage recommendations | Suggested rents, freight rates or bid prices | Higher | Data age, granularity and user control over final decisions |
| Benchmarks shown to customers | Market occupancy or average job value by region | Moderate | Aggregation level, contributor counts and publication lag |
| Internal product analytics | Feature adoption and error trends | Lower | Whether outputs ever reach customers |
| Fraud and security detection | Shared patterns of payment fraud | Lower | Whether competitively sensitive fields are used |
| Models or datasets licensed to outside AI developers | A dataset built from pooled customer records | Depends on content | Customer rights, pricing fields and re-identification risk |
| Licensing your own operational records | De-identified support tickets and engineering issues | Generally lower | Customer details and confidential pricing removed |
Implications checklist for vendors pooling customer data#
The checklist below helps a vendor find exposure before a customer, competitor or regulator asks. Most items are product and data decisions, not legal drafting, so the product lead and head of data should own the first pass.
- Map every feature that uses data from more than one customer and note what it outputs.
- Flag features whose outputs influence price, wages, capacity or bids.
- Record the age and granularity of pooled inputs for each flagged feature.
- Confirm customers make their own final decisions and that defaults do not push acceptance.
- Review customer contracts for rights to aggregate and use data across accounts.
- Train support and customer success teams not to pass along competitor-specific information.
- Keep a written rationale for each design choice, reviewed by antitrust counsel.
Questions customers and acquirers may start asking#
Customer questions about pooled data tend to arrive through procurement and security questionnaires before they arrive from regulators. Larger customers ask whether their data trains shared models, whether competitors benefit from it and whether they can opt out of pooling, and a vague answer can stall a renewal.
Acquirers ask the same questions in diligence, framed as risk. A buyer of a vertical SaaS company will want the feature map from the checklist above, the contract basis for aggregation and any counsel memos, because an antitrust problem in a pricing feature can follow the product into the buyer's portfolio.
Prepare one consistent answer set and keep it current. When product, sales and legal describe pooling the same way, customers are less likely to assume the worst.
Illustrative: a self-storage software vendor reviews its benchmarks#
Illustrative: a fictional management software vendor for independent self-storage operators offers a market report showing each operator nearby occupancy and street rates, drawn from other customers on the platform. The product team also wanted to train a model that suggests unit prices.
After reading coverage of the RealPage case, the CEO paused the pricing model and asked counsel to review the market report. The vendor moved the report to older, wider-area aggregates with minimum contributor counts, left the pricing model off the roadmap pending advice and documented each decision.
Separately, the company kept preparing a licensing package of its own support tickets and engineering history, with operators' names and rates removed, because those records did not depend on pooling competitors' pricing data.
Does this change licensing data to AI developers?#
Licensing records to AI developers raises a different question from pooling competitors' data inside a pricing product. A package of de-identified support conversations, issue histories or code reviews teaches a model how work gets done; it does not hand competitors each other's current prices.
The overlap appears when a dataset would carry recent, identifiable pricing or capacity data from competing customers. In the SourceX five-step transaction, the Rights step screens for customer-owned and competitively sensitive fields, Preparation removes them, and the SourceX Evidence Packet records the permitted use that the license then restricts.
Nothing is shared during the initial SourceX fit check, which collects metadata rather than files, and large datasets stay in the seller's own storage or ship on encrypted drives. A vendor can therefore scope a package of its own records without moving any pooled customer data while an antitrust review is still open.
Frequently asked questions
Is the RealPage settlement final?
Check the court docket. Proposed consent decrees in Justice Department civil antitrust cases generally go through a public comment process under the Tunney Act and court review before entry, and terms can change along the way. Rely on the entered judgment, not on early summaries.
Does the settlement bind other software vendors?
No. A consent decree binds the parties to it. It still matters to other vendors because it shows which practices enforcers challenged and which remedies they accepted, and private plaintiffs and state enforcers may draw on the same theory.
Are there state or local rules on algorithmic pricing?
Some states and cities have proposed or adopted rules aimed at algorithmic rent-setting or at shared nonpublic data in pricing tools. Requirements vary and are changing, so check the jurisdictions where your customers operate and revisit the review when new rules pass.
Can we still offer market benchmarks to customers?
Many vendors do, but design matters. Older data, wider aggregation, enough contributors that no single competitor can be inferred, and outputs that inform rather than direct pricing are the usual starting points. Have antitrust counsel review the specific design before launch.
Who inside the company should own this review?
The CEO or general counsel should sponsor it, with the product lead and head of data doing most of the mapping and antitrust counsel reviewing the flagged features. A documented review matters as much as its conclusions if questions come later.
Related resources
- QuestionDo AI labs buy code?
- InsightWhat permitted uses should a code license allow: training, evaluation or RL environments?
- InsightVendor AI training vs licensing your own data: who captures the value?
- SolutionWhat is AI evaluation data?
- SolutionHow AI developers source data
- IndustryBPO & contact centers data
See if your company qualifies
A short company assessment. No data uploads are needed.