Fine-tuning and post-training data
Compliance-reviewed communications as policy-alignment data
Quick answer
Compliance-reviewed communications are drafts of customer-facing material (ads, emails, scripts, disclosures, product pages) paired with the reviewer's decision, the edits required and the rule or policy cited. They teach a model to apply a domain policy, not just recite it: rejected-then-revised drafts become preference pairs, approved finals become supervised targets, and cited rules become rationales. Buy them with jurisdiction, rule-version dates and reviewer identity pseudonymized, and scope confidential internal policy text before release.
By SourceX Editorial · Updated
This page sits inside the fine-tuning and post-training data guide. For the sector view, see financial services LLM fine-tuning data; for general SFT sourcing, see how to source supervised fine-tuning data.
Why review decisions beat policy documents as training signal
Review decisions carry the judgment that policy text leaves implicit, which is what a policy-adherence model actually needs. The Alignment Studio authors argue that fine-tuning on policy documents gives a model the vocabulary of a policy but not the ability to judge whether a specific response complies [5]. PAM's authors note that policy-aligned behavior usually requires curating custom datasets and retraining, which is costly and still does not guarantee robust compliance [6].
A compliance review log is that curated dataset, produced as a by-product of regulated work. Each record says, in effect, "this sentence breaks rule X for reason Y; change it to Z." That is the shape of a critique-and-revise example, and it already exists in many firms that must approve communications before use.
Volume matters less than precision here. LIMA showed that 1,000 carefully curated prompt-response pairs could produce strong alignment behavior, while also noting that such curation is labor-intensive [8]. A few thousand reviewed drafts with clean decision labels and rule citations can outweigh a much larger set of unlabeled marketing copy.
Where these records come from
The richest sources are industries where approval before publication is mandatory and the approval trail must be retained. In US broker-dealers, FINRA Rule 2210 requires an appropriately qualified registered principal to approve each retail communication before the earlier of its use or filing with FINRA, and firms must keep records that include who approved it and when [1]. FINRA Rule 4511 ties those records to SEA Rule 17a-4 preservation [4], so the draft, approval and final often survive in archiving systems for years.
As of October 2026, FINRA has requested comment in Regulatory Notice 26-14 on replacing the blanket pre-use approval requirement with written procedures each firm designs to decide which retail communications need pre-use approval [2]. If adopted, review coverage will vary more by firm, so ask suppliers which communication categories were reviewed pre-use and in which period.
Registered investment advisers review advertisements against the SEC Marketing Rule, Rule 206(4)-1, which uses principles-based general prohibitions and sets conditions on testimonials, endorsements, third-party ratings and performance presentation [3]. Other common sources:
- Pharmaceutical and medical-device promotional review (often called MLR, for medical, legal and regulatory review) of claims, fair-balance language and references.
- Insurance advertising and agent-script review against state advertising rules and carrier guidelines.
- Lending and collections scripts reviewed for disclosures and prohibited phrasing.
- Internal communications policies: brand claims, sustainability claims, and employee social-media pre-clearance.
SourceX's compliance records category describes this family of operational records from the supply side.
Anatomy of a usable review record
A usable record links the exact draft version reviewed, a categorical decision, span-level comments, the cited rule with its version, and the revised version that was finally approved. Records missing the draft-to-final linkage are mostly useless for preference tuning, because you cannot form a chosen-versus-rejected pair.
Review platforms typically store each round as a separate version with comments anchored to text spans. Exports from archiving or workflow tools often flatten that history into a PDF of the final and a comment log, so request versioned text, not rendered documents.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"review_id": "rv_000184",
"channel": "email_campaign",
"audience": "retail",
"jurisdiction": "US",
"rule_set": [{"rule": "FINRA 2210(d)(1)", "version_date": "2025-01-01"}],
"internal_policy_refs": ["POL-MKT-07 v3"],
"round": 1,
"draft_text": "Our fund delivers guaranteed income with no downside.",
"decision": "rejected_with_changes",
"comments": [
{"span": [18, 35], "issue": "promissory_claim", "note": "Income is not guaranteed; remove."},
{"span": [41, 52], "issue": "unbalanced_risk", "note": "Add principal-loss risk language."}
],
"final_text": "The fund seeks to provide income. Investing involves risk, including possible loss of principal.",
"final_decision": "approved",
"reviewer_role": "registered_principal",
"reviewer_id": "pseud_r_12",
"reviewed_at": "2025-03-14"
}
Field notes: span offsets must refer to draft_text exactly, issue should come from a controlled taxonomy rather than free text, and rule_set.version_date is what lets you exclude examples decided under superseded rules.
Turning decisions into SFT, preference and eval sets
One review record can yield up to three training artifacts, and the decision label determines which ones are valid. The supervised fine-tuning target is the approved final; the preference pair is rejected draft versus approved revision; the rationale is the comment set plus cited rule.
| Record outcome | SFT use | Preference use (DPO, reward model) | Eval use | Main failure mode |
|---|---|---|---|---|
| Approved first time | Target text for "write compliant copy" | Chosen only; needs a synthetic or sampled rejected pair | Positive control | Over-represents easy, low-risk copy |
| Rejected, then revised and approved | Final as target; draft plus comments as critique task | Draft rejected, final chosen | Strong held-out items | Edits that changed meaning beyond the cited issue |
| Rejected, abandoned | Critique task only | Rejected only | Negative control | No ground-truth compliant version |
| Approved with conditions (for example, "use only with disclosure X") | Target plus condition as context | Weak; condition is outside the text | Context-sensitivity tests | Condition lost in export |
| Escalated or overturned | Exclude from training | Exclude, or treat as low-confidence | Disagreement set | Inconsistent labels poison the reward signal |
Direct Preference Optimization fits a policy directly to such chosen-versus-rejected pairs without training a separate reward model [7], which makes draft-final pairs immediately usable. Format the records as multi-turn critique-and-revise conversations if you want the model to explain its edits; the chat fine-tuning data format page covers role structure and loss masking.
Keep a slice of general instruction data in the mix, because a model tuned only on compliance edits tends to hedge everything. The data mixtures page explains how to size that slice.
Rule versions, jurisdiction and reviewer drift
Rules and house policies change, so every example needs the rule version and jurisdiction it was judged under, or the model will learn contradictions. A disclosure that was optional in one year can be required the next, and a pending FINRA change to pre-use approval [2] is a live reminder that review practice itself shifts.
Three drift sources show up in real review logs:
- Rule drift. Amended rules or new internal policy versions. Filter or tag by
version_date. - Reviewer drift. Different principals apply different thresholds. Measure per-reviewer rejection rates and inter-reviewer agreement on duplicated or near-duplicate drafts.
- Channel drift. Social posts, emails and long-form brochures face different rules (FINRA treats correspondence and retail communications differently [1]). Keep
channelandaudienceas conditioning fields.
Hold out whole time periods, not random rows, when you build evals. Near-duplicate campaign copy across rows will otherwise leak into your test split and inflate scores.
Confidentiality, personal data and rights to release
These records mix three kinds of sensitive content: confidential internal policy, personal data in drafts and comments, and third-party material quoted in the communications. Internal policy manuals and reviewer notes may be trade secrets the supplying company will not release in full, so negotiate a release scope (for example, rule citations and issue labels without full policy text).
Drafts often contain customer names, account numbers and personal examples used for personalization tests. In financial services, nonpublic personal information received from a financial institution carries reuse and redisclosure limits under the GLBA Privacy Rule [9]; in pharma or health-plan communications, any protected health information needs HIPAA de-identification by Safe Harbor or Expert Determination [11]. Reviewer names are personal data too, so require pseudonymous reviewer_id values that still allow agreement analysis.
Treat alignment data as a security asset as well. NIST SP 800-218A extends secure development practices to the integrity of training, fine-tuning and alignment data [10], which supports asking for hashes, versioned manifests and a record of who prepared the extract. For license scope, see fine-tuning-only data licenses; for counsel's review path, see reviewing an AI data license as in-house counsel.
Buyer checklist before you sign
Use this checklist to test whether a compliance-review dataset will actually train policy adherence.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Draft and final are both present as text, linked by a stable review ID and round number.
- Decision labels use a fixed vocabulary (approved, approved with conditions, rejected with changes, rejected, escalated).
- Comments are anchored to spans, and issue types map to a documented taxonomy.
- Each record cites a rule or policy reference with a version or effective date.
- Jurisdiction, channel and audience fields are populated for at least the records you plan to train on.
- Reviewer identity is pseudonymized but consistent, so agreement can be measured.
- A sample of 50 to 100 records has been checked by your own compliance staff for label accuracy.
- Personal data removal method is documented, and the supplier's sample check results are available.
- Release scope for internal policy text and third-party quoted material is written into the license.
- The license names fine-tuning and preference tuning as permitted uses, plus any eval use.
For a broader acceptance process, see how to evaluate a fine-tuning dataset before you buy it and the data provenance guide.
How SourceX approaches compliance-review data requests
SourceX sources operational datasets, including finance and legal workflows and documents, from US companies on request; nothing is held in stock and a request does not guarantee a match. You describe the review records you need, and SourceX looks for US businesses that hold them; each release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and no method is perfect.
Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and delivery runs through private, access-controlled workflows only after an executed agreement. Teams scoping domain-specific fine-tuning can describe their dataset requirements to SourceX.
Requesting policy compliance fine-tuning data
SourceX manages the commercial process for operational datasets from US companies, from Find and Assess through Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Each dataset is delivered under a license that defines records, uses, term and delivery, with terms agreed per deal. Tell SourceX what compliance-review records you need.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Frequently asked questions
Can I use approved-only communications without the rejected drafts?
Yes, for SFT, but you lose the contrast that teaches the model what to avoid. Approved-only sets also skew toward low-risk copy, so pair them with synthetic or sampled negatives and audit those negatives with your own reviewers.
Do I need the full internal policy manual?
Usually not. Rule citations, issue labels and span-level comments carry most of the signal; you can supply your own policy text at inference time through retrieval if the model must quote it.
How do I handle disagreement between reviewers?
Keep escalated and overturned records out of preference training and use them as a disagreement eval set. Persistent per-reviewer differences are a reason to condition on reviewer group or to train only on records with second-level sign-off.
Sources
- FINRA, "2210. Communications with the Public". https://www.finra.org/rules-guidance/rulebooks/finra-rules/2210
- FINRA, "Regulatory Notice 26-14" (2026). https://www.finra.org/rules-guidance/notices/26-14
- U.S. Securities and Exchange Commission, "Investment Adviser Marketing (Small Entity Compliance Guide)". https://www.sec.gov/resources-small-businesses/small-business-compliance-guides/investment-adviser-marketing
- FINRA, "4511. General Requirements". https://www.finra.org/finramanual/rules/r4511
- arXiv, "Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations" (2024). https://arxiv.org/pdf/2403.09704
- arXiv, "PAM: Training Policy-Aligned Moderation Filters at Scale" (2025). https://arxiv.org/pdf/2505.19766
- arXiv, "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
- arXiv, "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- Federal Trade Commission, "How To Comply with the Privacy of Consumer Financial Information Rule of the Gramm-Leach-Bliley Act". https://www.ftc.gov/business-guidance/resources/how-comply-privacy-consumer-financial-information-rule-gramm-leach-bliley-act
- National Institute of Standards and Technology, "Secure Software Development Practices for Generative AI and Dual-Use Foundation Models (NIST SP 800-218A)" (2024). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-218A.pdf
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.