Skip to content

Fine-tuning and post-training data

Draft-to-final document pairs for rewriting and editing models

Quick answer

A text revision dataset for training rewriting and editing models pairs a professional's draft with the version that was actually approved, plus the reason for the change: a review comment, a style instruction or an edit-intent label. Public corpora such as IteraTeR and CASIMIR come from Wikipedia and scientific papers [1][4]. Teams building business writing assistants usually need licensed version histories from real companies, cleaned of formatting churn, de-identified, and aligned so that every pair carries its instruction.

By SourceX Editorial · Updated

What public text revision datasets cover, and where they stop

Public revision data is useful for research but skews heavily toward encyclopedic and academic prose. IteraTeR collected about 31K document revisions from Wikipedia, arXiv and Wikinews and annotated edit intentions such as fluency, clarity, coherence and style [1][2]. CASIMIR aligns multiple revised versions of 15,646 OpenReview papers at sentence level with revision intent [4], and ParaRev pairs scientific paragraph revisions with the instruction that motivated them [5].

A 2026 survey of edit-intention work organizes the field by granularity and label scheme and confirms the same source mix: Wikipedia, academic writing and student essays [3]. None of these reflect how a claims adjuster tightens a denial letter, how a support lead rewrites a macro, or how counsel redlines an indemnity clause. If your assistant ships into enterprise documents, the domain gap is the main reason to source licensed data rather than stack more public pairs.

ParaRev's design points to the target shape for instruction-tuned editors: each revision is paired with the natural-language instruction that motivated it [5]. The lesson for buyers is that the instruction column matters as much as the before and after text.

Where draft-to-final pairs live inside companies

Draft-to-final pairs exist wherever a business keeps version history and review comments on the same document. The practical sources are:

  • Word documents with tracked changes. In DOCX (Office Open XML), insertions and deletions are stored as w:ins and w:del elements with w:author and w:date attributes, and comments sit in comments.xml anchored by range markers. One file can yield draft, final and the reviewer's rationale.
  • Collaborative editors. Google Docs revision history, SharePoint and OneDrive version history, and Confluence page versions hold intermediate states, though suggestion and comment metadata may need a separate export.
  • Docs-as-code repositories. Markdown or reStructuredText in Git gives clean diffs, commit messages and pull-request review threads that act as edit instructions.
  • Workflow systems. Contract lifecycle tools, proposal software, and support macro or knowledge-base editors often store a submitted draft and a published version with an approver note.

Domain owners for adjacent data already exist: contract markups are covered in contract redline datasets, proposal writing in RFP responses and proposals, and correction-as-feedback in human feedback datasets with QA scores and corrections. For legal work product specifically, see legal LLM fine-tuning data.

Building a usable pair: instruction, endpoints and intermediate versions

A usable pair needs three things: the source text, the revised text, and an instruction or rationale that a model can condition on. Without the instruction, you are training a model to imitate an unspecified house style, which is hard to steer at inference time. ParaRev attaches explicit revision instructions to its pairs for this reason [5].

Where instructions are missing, there are three fallbacks, in order of quality. First, recover them from review comments anchored to the changed span. Second, have domain experts write short instructions after the fact. Third, assign intent labels from a taxonomy such as IteraTeR's [1], accepting that a label like "clarity" is weaker supervision than a sentence of reviewer guidance.

Intermediate versions are worth requesting even if you only train on endpoints today. Iterative revision data shows that documents improve over several rounds with different intents at each step [1][4]. Keeping version 1, 2 and final lets you build multi-step editing tasks, sample single-intent pairs, and check whether a "final" was later reverted.

Separating substantive edits from formatting churn

Most raw version diffs are noise, so filtering is the main cleaning cost. Typical churn includes style and font changes, field-code refreshes, autonumbering, smart-quote conversion, whitespace and line-wrap differences, template updates and boilerplate swaps such as footers or dates. A pipeline that diffs extracted text rather than XML removes much of it, but you still need rules for trivial token edits.

Practical filters: normalize Unicode and whitespace before diffing; compute a token-level edit ratio and drop pairs below a floor (pure typo fixes) or above a ceiling (full rewrites that no longer align); and segment by paragraph or clause so one document yields several aligned pairs. Templates create many near-identical drafts, so run near-duplicate detection across pairs, since duplicated training text increases memorized output [8]. Our guide to MinHash and LSH near-duplicate detection covers thresholds.

Small, clean sets can be enough: LIMA fine-tuned on 1,000 curated examples, while noting that curation is labor-intensive [7]. Budget for expert review of a filtered sample rather than for raw volume.

Illustrative record schema for an edit pair

The schema below is one way to package a pair so it works for SFT and can be reused for preference training.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "pair_id": "doc-0192_v2-v3_para-07",
  "document_type": "customer_notice_letter",
  "source_version": "v2",
  "target_version": "v3",
  "source_text": "We regret to inform you that your request cannot be processed at this time due to missing items.",
  "target_text": "We could not process your request because two documents are missing: a signed form and proof of address.",
  "instruction": "Be specific about what is missing; avoid passive voice.",
  "instruction_origin": "review_comment",
  "edit_intents": ["clarity", "specificity"],
  "edit_ratio": 0.62,
  "reviewer_role": "compliance_reviewer",
  "reviewer_id": "R-14 (pseudonymized)",
  "deidentification": {"method": "entity replacement", "sample_checked": true},
  "final_status": "approved_and_sent"
}

Fields worth insisting on: instruction_origin (comment, expert-written or inferred label), reviewer_role rather than name, and final_status, which separates an approved final from an abandoned branch.

Using human edits as preference data

A draft and its approved final can be framed as a rejected and chosen response, which fits Direct Preference Optimization's pairwise objective [6]. This is attractive because the preference signal came from a real reviewer, not a rating task. See preference datasets for DPO for formatting.

Three caveats apply. The draft was written by a person, not your model, so these are off-policy pairs; read on-policy vs off-policy preference data before relying on them alone. A final is not always better on every axis, since some edits reflect a deal concession or a manager's taste. And if reviewer comments explain the change, consider training a critique model too, as covered in critique and revision data for generative verifiers.

Privacy and rights in revision metadata

Revision files carry more personal data than the visible text. Tracked changes record w:author names and timestamps, comments include reviewer names and initials, and document properties can hold the creator, last editor and company. Strip or pseudonymize all of these alongside names, emails, phone numbers and account numbers in the body; our guide to PII redaction for LLM training data covers detection recall.

Deleted text is a specific failure mode. A w:del element (its w:delText) can hold a client name or figure that was removed from the final for confidentiality, so de-identification has to run on both sides of the pair. Health documents need HIPAA de-identification through Safe Harbor or Expert Determination [10].

Rights matter as much as privacy. An audit of public dataset hosting found license information omitted in over 70% of cases and wrong in over 50% [9], so confirm that the supplying company owns the documents, that client confidentiality terms permit the use, and that the license names editing-model training as an allowed use.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Buyer checklist for a text revision dataset request

Illustrative example: invented to show structure; it does not describe an available dataset.

Question to settleWhy it mattersAcceptable evidence
Which document types and roles produced the edits?Domain and seniority of reviewers set the target styleDistribution by document type and reviewer role
Are instructions or comments attached to each pair?Instruction-conditioned SFT needs them [5]Share of pairs with comment, expert or inferred instruction
Are intermediate versions kept?Enables multi-step and single-intent pairs [1]Versions per document, timestamps
How was formatting churn removed?Raw diffs are mostly noiseFilter rules, edit-ratio bounds, sample before and after
Were near-duplicates removed?Template drafts inflate counts [8]Dedup method and threshold
How were author and comment metadata handled?Reviewer identities leak through XMLDe-identification method, checked sample
Does the license permit editing-model training?Licensing gaps are common [9]Allowed uses, term, delivery terms in writing

Before purchase, score a held-out sample against your own rubric, as described in how to evaluate a fine-tuning dataset before buying, and browse related tasks in the fine-tuning and post-training data hub.

How SourceX handles requests for draft-to-final revision data

SourceX sources operational datasets from US companies on request, including documents and the finance, legal, support and sales workflows where draft-and-review history accumulates. Data is not held in stock and a request does not guarantee a match; you describe the data, such as document types, edit instructions and version depth, and SourceX looks for businesses that hold it, with every release approved by the supplying company. You can describe the revision data you need at any stage.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement.

Request draft-to-final revision data

If you are training a rewriting or editing model and need professional version histories with review comments, tell SourceX the document types, instruction coverage and version depth you require. SourceX assesses data and licensing permissions with candidate suppliers, and nothing is contracted until a supplier agrees. Start a buyer request.

Sources

  1. arXiv (Du et al.), "Understanding Iterative Revision from Human-Written Text" (2022). https://arxiv.org/pdf/2203.03802
  2. Grammarly Engineering, "Introducing IteraTeR" (2022). https://www.grammarly.com/blog/engineering/introducing-iterater
  3. arXiv, "Making Revisions Understandable: A Survey of Edit Intentions, Methods, and Applications" (2026). https://arxiv.org/pdf/2609.01610
  4. arXiv, "CASIMIR: A Corpus of Scientific Articles enhanced with Multiple Author-Integrated Revisions" (2024). https://arxiv.org/abs/2403.00241
  5. arXiv, "ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction" (2025). https://arxiv.org/html/2501.05222v1
  6. arXiv (Rafailov et al.), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  7. arXiv (Zhou et al.), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  8. arXiv (Lee et al.), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
  9. arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  10. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data