Skip to content

Industry-specific operational data

Product returns and RMA reason data for retail AI

Quick answer

Product returns data for machine learning means order-linked RMA records: the customer-selected reason code, the free-text comment, the warehouse inspection result, the disposition (restock, refurbish, liquidate, destroy) and the refund or exchange outcome, tied to SKU attributes and the return policy in force. Public review corpora cannot substitute, because they hold no return outcomes. Buyers should specify both labels (stated and inspected), a fraud confirmation status and policy versions before sourcing from retailers or returns operators.

By SourceX Editorial · Updated

What a usable returns record contains

A usable returns record joins four systems: the order management system (OMS), the RMA or returns portal, the warehouse management system (WMS) where inspection happens, and the payments or refund ledger. Records taken from only the returns portal give you a stated reason with no ground truth. Records taken from only the WMS give you a condition grade with no customer intent. The value for classification and fraud work sits in the join.

Fields to request, at minimum:

  • Linkage: tokenized order ID, line ID, RMA ID and customer token, stable across orders so repeat-return behavior is visible.
  • Product attributes: SKU, category path, size, color, brand tier, price band and whether the item was on promotion.
  • Customer-stated reason: the reason code exactly as shown in the portal dropdown, plus the code list version.
  • Free-text comment: the raw comment, after personal details are scrubbed.
  • Timing: order date, delivery date, RMA creation date, carrier receipt date and days to return.
  • Inspection: condition grade, inspector notes, tags-attached flag, item-match flag (right SKU, right serial).
  • Disposition and money: restock, refurbish, liquidate, return to vendor or destroy; refund type (original tender, store credit, exchange) and any restocking or return-shipping fee.

For catalog normalization upstream of this, see product categorization and taxonomy-mapping data; returns models degrade quickly when category paths drift between seasons.

Why stated reasons are noisy labels

Customer-stated reasons are weak labels because customers often pick whichever option produces free return shipping. "Item defective" or "not as described" can carry no fee while "changed my mind" or "no longer needed" can, so the distribution of stated reasons partly reflects the fee schedule rather than the product. A classifier trained only on stated reasons learns the policy, not the problem.

Inspection outcomes provide a second, independent label. When the stated reason says "defective" and inspection says "new, tags attached, fully functional," the disagreement is useful in two ways: it is a cleaner training target for root-cause models, and the gap itself is a feature for return-abuse scoring. Ask suppliers what share of returns were physically inspected, because many low-value items are refunded without inspection ("returnless refunds"), and those rows have no second label at all.

The general method for testing whether operational outcome fields can serve as ground truth is covered in verifying outcome labels in operational records.

Return-fraud labels need confirmation status

Return-fraud training data is only as good as its confirmation workflow, so a suspicion flag alone is not a label. Wardrobing (worn then returned), empty-box or swapped-item returns, receipt fraud, and serial-number swaps are each confirmed in different ways: inspector photos, serial mismatch logs, carrier weight discrepancies, or a loss-prevention case closed with a finding.

Request a fraud field with explicit states, such as suspected, under review, confirmed, cleared and unresolved, plus the confirming evidence type and the date of the decision. Without the "cleared" state, your negatives are contaminated by uninvestigated cases. Treat refund reversals and chargebacks as a related but separate outcome; the chargeback representment case data page covers that dispute trail.

Public proxies and their limits

Public datasets cover product sentiment, not returns transactions. Large product-review corpora are the usual stand-in [1][2]. They help with fit and quality language ("runs small," "seams split"), but they contain no RMA reasons, inspection results or dispositions, and some large ones carry non-commercial license terms that rule out production training [1].

License metadata on hosted datasets is also unreliable: one audit of more than 1,800 text datasets found licenses omitted in over 70% of cases and miscategorized in over 50% on popular hosting sites [3]. If a public set is used for pretraining-style warm-up, record its license terms and keep it separable from licensed operational data in your lineage.

Policy versions and time drift

Return-reason distributions shift whenever the policy changes, so every record needs the policy version in force on its order date. Shortening a return window, adding a mail-back fee or moving to box-free drop-off all change who returns and which reason they pick. A model trained across an undocumented policy change will read a fee change as a product-quality shift.

Ask suppliers for a dated policy changelog (window length, fee by reason, eligible categories, holiday extensions) and a reason-code changelog when dropdown options were added, merged or renamed. The same problem in other industries is described in code-set revisions in multi-year operational datasets.

Privacy and rights checks for returns data

Returns records carry personal data in obvious and less obvious places, so de-identification has to cover structured fields, comments and attachments. Names, shipping addresses, emails, phone numbers and payment details must be removed or replaced; free-text comments often contain addresses, order numbers and phone numbers typed by the customer. Customer-submitted photos of damaged items can show faces, home interiors or labels, and are usually best excluded or handled as a separate, reviewed stream.

The supplying retailer's original data can be governed by state consumer-privacy laws, depending on the retailer's size and where its customers live. Under the CCPA, information counts as deidentified only if it cannot reasonably be linked to a consumer and the business meets conditions that include contractually obligating recipients [4]. FTC staff have also warned (January 2024) that companies may face liability if they use customer data for model training contrary to their privacy commitments [5], so ask what the supplier's privacy notice said when the data was collected.

Returns record specification for a sourcing request

A written field specification shortens supplier assessment, and the example below shows one returns line as delivered in JSON Lines, one UTF-8 object per line [6].

Illustrative example: invented to show structure; it does not describe an available dataset.

{"rma_id":"rma_7f3a","order_token":"ord_91c2","line_id":2,"customer_token":"cus_44be","sku":"SKU-118204","category_path":"Apparel>Women>Dresses","size":"M","color":"navy","promo_flag":true,"order_date":"2025-11-28","delivered_date":"2025-12-02","rma_created":"2025-12-19","days_to_return":17,"policy_version":"2025-10-v3","reason_code":"DEFECTIVE","reason_code_list":"rc-v7","comment":"[SCRUBBED] stitching came apart after one wear","inspection_grade":"A_NEW","tags_attached":true,"item_match":true,"disposition":"RESTOCK","refund_type":"ORIGINAL_TENDER","return_fee_usd":0.00,"fraud_status":"CLEARED","fraud_evidence":null}

Buyer checklist before signing:

CheckWhat to ask forFailure mode if missing
Join coverageShare of RMAs with OMS, WMS and refund records attachedStated reasons with no ground truth
Inspection coverageShare of returns physically inspected, by categoryReturnless refunds mislabeled as verified
Fraud statesSuspected, confirmed and cleared, with evidence typeContaminated negatives
Policy logDated window, fee and reason-code changesPolicy effects learned as product effects
ScrubbingMethod used on comments and attachmentsAddresses and phones leaking via free text
License fieldsAllowed uses, term, record scopeData usable for analytics but not training

How SourceX handles returns data requests

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases; it holds no inventory, and a request does not guarantee a match. You describe the data you need, such as the specification above, and SourceX looks for US businesses that hold it; every release is approved by the supplying company. The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. You can describe your returns data requirement on the buyers page.

Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. For broader context, see the industry-specific operational data guide, the AI data hub, buyers in e-commerce, multi-brand retail buyers, customer support ticket datasets for service conversations, and what AI companies build with retailer data.

Sourcing returns and RMA data for your models

SourceX sources returns records from US companies on request, rights-reviews each dataset and delivers it under a license that defines records, uses, term and delivery. Prices are not published; terms are agreed per deal. Start a returns data request at sourcex.si/buyers.

Sources

  1. HyperAI, "Amazon Reviews dataset". https://hyper.ai/en/datasets/5481
  2. Cleanlab (Hugging Face), "amazon-reviews README". https://huggingface.co/datasets/Cleanlab/amazon-reviews/blob/main/README.md
  3. Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  4. California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  5. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  6. jsonlines.org, "JSON Lines". https://jsonlines.org/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data