Image data
Item Condition Photos for Resale and Returns Grading Models
Quick answer
A usable item condition grading dataset is not a folder of "used product" photos. It is multi-angle images of returned or pre-owned items, each tied to a grade from a documented taxonomy, defect labels with locations, the disposition the operator actually chose (restock, refurbish, liquidate, recycle) and, ideally, the refund or resale outcome. Public sets rarely carry these labels, so most teams license operational records from returns processors, refurbishers or resale marketplaces and verify grader agreement before training.
By SourceX Editorial · Updated
Why condition grading needs its own label scheme
Condition grading is a separate task from catalog attribute extraction because the label describes this unit's wear, not the product's design. A catalog model learns "navy wool crew-neck"; a grading model has to learn "pilling on cuffs, small stain near hem, grade B, route to resale at a discount." Training on product catalog photo datasets teaches identity and attributes, while condition data teaches deviation from new.
Condition labels do show up in public data, but thinly. One public garment dataset on Hugging Face includes a "visible defects" attribute alongside other clothing fields [1]. That is useful for prototyping a VLM that describes defects, yet it rarely includes a calibrated grade, a disposition or any link to what happened to the item afterward.
Most production grading teams need three label layers at once:
- Grade: an ordinal class such as A/B/C/D, or like-new/good/fair/poor, defined by written criteria per category.
- Defects: typed findings (scratch, dent, stain, tear, missing part, pilling, screen burn-in, seal broken) with a bounding box or polygon and a severity.
- Outcome: the disposition decision and, where available, the realized resale price band, refund decision or customer dispute.
For pure surface-anomaly work on new goods, industrial defect image datasets and rare defect coverage strategies are closer fits than condition grading.
Where condition-labeled photos come from
The strongest condition photos come from operational grading stations, not from crowdsourced uploads. Each source type has a typical structure and a typical weakness.
| Source | What you usually get | Common gaps |
|---|---|---|
| Returns processing center (3PL or retailer) | Fixed-station photos, SKU, return reason code, grade, disposition | Grades tuned to a single retailer's policy; few angles for small items |
| Electronics refurbisher | Cosmetic grade, functional test results, part replacements | Cosmetic photos may be taken only for failures |
| Resale or recommerce marketplace | Seller photos, platform-assigned condition, sale price, return/dispute outcome | Inconsistent lighting and backgrounds; seller-declared condition is noisy |
| Rental or subscription operator | Repeated photos of the same unit across cycles | Narrow catalog; wear accumulates on the same items |
| Customer-uploaded return evidence | Phone photos submitted with a return request, refund decision | Faces, addresses, packing slips and home interiors in frame |
Customer-uploaded images look attractive because they carry real refund decisions, but they need the heaviest privacy work. Fixed-station images give cleaner training signal and easier redaction. Many teams combine both: station photos for the grade model and customer photos for a "claimed damage vs. verified damage" model.
Specifying the grade taxonomy and grader agreement
The grade taxonomy is the core of the dataset, so ask for the written grading rubric before you look at a single image. Without it, a "B" from one site and a "B" from another are different labels, and your model will learn the disagreement.
Ask the supplier for the rubric version in effect for each record, because rubrics change after policy updates or new categories. Ask whether grades were assigned by a single grader, by two graders with adjudication, or by a grader with a supervisor audit sample. An ordinal scale should be measured with an agreement statistic that respects order, such as weighted kappa or Krippendorff's alpha for ordinal data; the right metric depends on the label type and how many graders scored each item [3].
Plan for label noise even with good rubrics. A widely cited audit of benchmark test sets estimated an average label error rate of at least 3.3% across ten popular datasets [2]. Condition grades, which depend on lighting and judgment, are unlikely to be cleaner, so reserve budget for a re-grading pass on your evaluation split and run a confident-learning style error scan on the training split.
Linking photos to disposition and refund outcomes
Disposition and outcome fields turn a classifier into a decision model, so they are worth asking for even if they complicate licensing. A grade predicts what a human inspector would say; a disposition label predicts what the business actually did; a resale price band or dispute flag tells you whether that decision held up.
Useful outcome fields include the disposition code, the date of disposition, the channel (own store, outlet, liquidation lot, parts harvest), the realized price band rather than exact price, and whether the item was later returned again. Ask how the supplier handles overrides: if a supervisor changed a grade after a customer complaint, you want both the original and final grade, because overrides are your best hard examples.
Watch for leakage. If the return reason code ("arrived damaged") is visible in a filename or packing slip, the model can learn the text rather than the defect. Split by SKU family and by grading site, not just randomly, so evaluation reflects new products and new stations.
Capture conditions, formats and metadata
Capture metadata decides whether the model generalizes beyond one grading bench. Ask for camera or station ID, lighting setup, number and names of standard angles (front, back, sides, close-up of defect, serial plate), and image resolution. Consistent station photography is good for accuracy; a share of phone-quality images helps if you will grade seller uploads later.
For labels, COCO JSON is a practical default for defect boxes and polygons and is supported by common annotation tools such as CVAT [9]. For training at scale, ask whether the supplier can package images and per-item JSON as WebDataset tar shards, where files sharing a basename form one sample [10], so all angles of one item stay together. Decide in advance which EXIF fields to strip; GPS and device serials rarely help grading, as covered in our guide to EXIF metadata in image training data.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"item_id": "RT-000184",
"category": "consumer_electronics/headphones",
"sku_family": "over_ear_wireless",
"images": ["RT-000184_front.jpg", "RT-000184_left.jpg", "RT-000184_defect_01.jpg"],
"station_id": "grading_bench_07",
"rubric_version": "electronics_cosmetic_v3.2",
"grade_initial": "B",
"grade_final": "C",
"graders": 2,
"adjudicated": true,
"defects": [
{"type": "scratch", "severity": "minor", "image": "RT-000184_left.jpg", "bbox": [412, 220, 96, 18]},
{"type": "ear_pad_wear", "severity": "moderate", "image": "RT-000184_defect_01.jpg", "bbox": [130, 88, 240, 210]}
],
"functional_test": "pass",
"return_reason_code": "no_longer_needed",
"disposition": "refurbish_then_resell",
"resale_price_band": "40-55pct_of_new",
"returned_again": false,
"redaction": {"faces": 0, "text_regions": 1, "method": "blur_v2"}
}
Privacy and rights checks for return photos
Return photos carry more personal data than catalog shots, so privacy review is a gating step, not a cleanup task. Customer uploads often show faces, hands, kitchens, shipping labels with names and addresses, order numbers and screens displaying account details. Research on a large image training set found personal information such as faces and identity documents that filtering did not fully remove [5].
Faces are the highest-risk element. If any faces could be processed for face geometry, Illinois BIPA section 15 sets notice, written-release and retention requirements for biometric identifiers [6], so most buyers require faces to be blurred or the image dropped. Treat redaction as a risk assessment: the ICO frames identifiability around what a motivated intruder could reasonably do [7], which matters when a distinctive home interior or a handwritten note is in frame. Generative uses add another reason to scrub: diffusion models have been shown to regenerate training images, including photos of people and logos [8].
Rights questions sit alongside privacy. Confirm that the supplier owns or controls the photos (seller uploads may be governed by marketplace terms), that its customer terms allow the transfer, and whether brand logos on products need any treatment; see model and property releases for AI training images and image licensing terms for AI training. Our privacy guide for de-identified training data covers redaction verification in more depth.
Request checklist for condition grading image data
A precise request filters out unsuitable data early and makes supplier assessment faster.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Categories and volume: product categories, SKU families, target item count per grade per category.
- Grade scheme: rubric document, version history, ordinal levels, how "unsellable" is labeled.
- Grader process: single vs. double grading, adjudication, audit sample size, agreement statistic reported.
- Defect labels: defect type list, severity scale, geometry (box, polygon or none).
- Outcomes: disposition codes, price band, re-return flag, override history.
- Capture: station vs. phone, angles per item, resolution, lighting notes.
- Privacy: faces, labels and documents handled by blur, crop or exclusion; method recorded; sample checked.
- Documentation: a datasheet covering collection, labeling and known gaps [4].
- Use and delivery: allowed uses (classifier training, VLM fine-tuning, evaluation), term, format and delivery path.
If you are weighing whether to buy labeled data or label raw station photos yourself, compare costs in pre-labeled vs. raw image datasets, and size the request with how many images a computer vision model needs.
How SourceX approaches condition photo requests
SourceX sources operational datasets from US companies on request; it does not hold condition photos in stock, and a request does not guarantee a match. You describe the data you need, such as graded return photos with dispositions, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows after an executed agreement. You can start by describing your grading use case on the SourceX buyer page, and browse related categories under images and inspection photos, product catalogs and descriptions and e-commerce buyers. For the wider image category, see the image datasets hub or the AI data hub.
Source condition grading image data for your model
SourceX finds US companies whose operational records match your described data, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Nothing is contracted until a supplier agrees, and SourceX does not train models. Describe your grade taxonomy, categories and outcome fields at sourcex.si/buyers.
Sources
- Hugging Face (Denali-AI), "Denali-AI/train-35k". https://huggingface.co/datasets/Denali-AI/train-35k
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv, "Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation" (2026). https://arxiv.org/pdf/2603.06865
- Gebru et al., "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010
- arXiv, "A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset" (2025). https://arxiv.org/pdf/2506.17185
- Illinois General Assembly, "740 ILCS 14/15 (Biometric Information Privacy Act)". http://www.ilga.gov/legislation/ilcs/fulltext.asp?DocName=074000140K15
- Information Commissioner's Office, "How do we ensure anonymisation is effective?" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
- Carlini et al., "Extracting Training Data from Diffusion Models" (2023). https://www.usenix.org/conference/usenixsecurity23/presentation/carlini
- CVAT documentation, "COCO format". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
- WebDataset project, "webdataset (GitHub repository)". https://github.com/webdataset/webdataset
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.