Skip to content

Image data

Agricultural Image Datasets for Crop, Weed and Field Models

Quick answer

An agricultural image dataset for weed detection or spot spraying is only useful if its labels match your target species at the growth stages your sprayer sees, and its metadata records the camera height, lighting, soil and season behind every frame. Public sets such as CropAndWeed and Weed25 are strong starting points for research, but species coverage, regional fit and license terms vary widely, so commercial teams often need to supplement them with field imagery licensed from operators who already capture it.

By SourceX Editorial · Updated

What public weed and crop datasets actually cover

Public agricultural datasets cover a narrow and uneven slice of the species, regions and conditions a commercial sprayer meets. CropAndWeed, published at WACV 2023, labels 74 crop and weed species with roughly 112,000 annotated instances across more than 8,000 images, and it records environmental conditions and recording parameters per image [1]. Weed25 offers 14,035 images, but of only 25 weed species [2]. Smaller sets built for robotic control, such as the annotated food-crop and weed collection published in Data in Brief, are useful for prototyping but rarely span multiple seasons or geographies [5].

The failure modes are predictable. A model trained on European row crops can miss Palmer amaranth or waterhemp in US soybeans, confuse cotyledon-stage broadleaf weeds with crop seedlings, or degrade under low-angle sun and wet residue that never appeared in training. Before you pick a set, map its species list and capture conditions against your deployment fields; the image cluster guide covers this coverage-first approach for computer vision generally.

Labels and metadata that separate usable field data from photos

Usable agricultural training data carries species-level labels plus the field context needed to explain errors. For spot spraying, the label set should distinguish crop from weed at instance or pixel level, name species using a consistent taxonomy (EPPO codes or scientific names, not local common names), and record growth stage on a standard scale such as BBCH. Bounding boxes are often enough for detection; segmentation masks matter when nozzle actuation depends on canopy area or when weeds overlap crop rows.

Metadata is where most public and vendor sets fall short. Ask for capture platform (boom-mounted, ATV, drone, handheld), camera height and angle, ground sampling distance, sensor and lens, illumination (natural, shaded hood, strobe), soil type and moisture, residue cover, crop row spacing, date and region at a coarse level. CropAndWeed shows that per-image environmental and recording annotations are feasible at scale [1]. For lens and lighting variation specifically, see camera, lens and lighting diversity in image datasets.

Licensing farm images for commercial training

A commercial license for agricultural imagery has to cover training, fine-tuning and deployment in a product you sell, and many public weed datasets do not. One drone weed dataset on Hugging Face, for example, is released under CC BY-NC-SA 4.0, which excludes commercial use and adds share-alike obligations [3]. Other research sets are available only by request to the owner, with terms set case by case [4]. Read the dataset card's license field first, since that is where the Hub records it [6], but treat it as a claim rather than proof.

License metadata on hosting sites is often missing or wrong: in an audit of text datasets, the Data Provenance Initiative reported license omission above 70% and error rates above 50% on popular dataset hosts [7], so apply the same check to image sets. For farm imagery captured by growers, agronomists or equipment operators, confirm who owns the images (the farm, the service provider or the equipment maker), whether existing data-sharing agreements allow licensing for AI training, and whether any images show workers or neighboring properties. The guide to commercially trainable image and caption datasets explains how to triage open licenses.

Privacy and location risk in field imagery

Field images look impersonal, but their embedded metadata and occasional bystanders can identify farms and people. GPS coordinates in EXIF tags pinpoint individual fields, and capture timestamps reveal operating patterns; an audit of a large web-scraped image dataset found non-empty EXIF tags on geolocation, timestamps and individuals [8]. Decide which fields to coarsen (for example, region and month instead of coordinates and timestamps) before delivery, as described in EXIF metadata in image training data.

Handheld scouting photos and drone passes over farmyards can capture faces, license plates and buildings. Agree on whether such frames are excluded, blurred or kept under releases; the guide to face data consent and anonymization covers the options.

Specification template for an agricultural image request

A precise specification shortens sourcing more than any other step, because suppliers can check their archives against concrete fields. Use the template below as a starting point and remove fields that do not affect your model.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy it matters
TaskInstance segmentation, crop vs. weed, for spot sprayingSets label geometry and QA rules
Target cropsSoybean, corn, cottonDefines negative class and row geometry
Target weedsPalmer amaranth, waterhemp, morningglory, grasses (grouped)Drives species taxonomy and class balance
Growth stagesBBCH 10-19 for weeds; crop V1-V6Early stages are the hardest and most valuable
Region and seasonUS Midwest and Mid-South; April-July; at least two seasonsCaptures regional biotypes and year-to-year variation
Capture platformBoom-mounted RGB, 0.5-1.0 m height, nadirMust match sprayer camera geometry
IlluminationNatural plus shaded hood; include dawn, overcast, harsh sunPrevents lighting-driven false positives
Field conditionsBare soil, heavy residue, wet soil, cover-crop interseedingCommon background failure modes
Metadata per imageSensor, height, GSD, timestamp (month), region (state), soil classEnables error slicing and stratified splits
VolumePilot of a few thousand frames; scale after QAValidates fit before committing
FormatJPEG or PNG plus COCO JSON masks; CSV metadataFits existing training pipelines
License scopeCommercial training and deployment in sprayer productExcludes NC-only sources
PrivacyGPS coarsened to state; frames with people removedLowers location and personal-data exposure

If you plan to annotate raw frames yourself, compare costs with pre-labeled vs. raw image datasets, and size the pilot with how many images you need to train a computer vision model.

Evaluating a sample before you commit

A sample review should test label accuracy, coverage and metadata completeness against your own field conditions, not the supplier's summary. Request a stratified sample across species, stages and conditions, then have an agronomist re-label a random subset to estimate species-level error rates, especially between look-alike pigweeds and among grass weeds. Check for near-duplicate frames from continuous boom capture, which inflate counts and leak across train/test splits unless you split by field and date.

Run your current model on the sample and slice errors by metadata. If false positives cluster under one lighting condition or soil type, that tells you where additional data is worth paying for. The training data quality assessment guide describes coverage and contamination checks in more depth.

Where licensed operational imagery fits

Agricultural service providers, crop consultants and equipment operators often hold large archives of field photos tied to scouting reports, application records or inspection notes. That operational context can supply labels and metadata that public sets lack. SourceX sources operational datasets from US companies on request, including documents and new recordings of hands-on work, and manages the commercial process from licensing through ongoing purchases; nothing is held in stock, and a request does not guarantee a match. Buyers describe the data they need on the SourceX buyer page, and SourceX looks for US businesses that hold it, with every release approved by the supplying company.

Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. SourceX does not source scraped web content or generic photos. For broader context, see what AI companies build with agriculture business data, the overview of licensed images and inspection photos, and robotics training data for embodied AI.

Source field images for your crop or weed model

If public datasets do not cover your species, stages or regions, describe the imagery and metadata you need and SourceX will look for US businesses that hold it. Pricing and allowed uses are agreed in a license per deal, and nothing is contracted until a supplier agrees. Start a buyer request at SourceX.

Sources

  1. CVF Open Access (WACV 2023), Steininger et al., "The CropAndWeed Dataset: A Multi-Modal Learning Approach for Efficient Crop and Weed Manipulation" (2023). https://openaccess.thecvf.com/content/WACV2023/html/Steininger_The_CropAndWeed_Dataset_A_Multi-Modal_Learning_Approach_for_Efficient_Crop_WACV_2023_paper.html
  2. PubMed Central (Frontiers in Plant Science), "Weed25: A deep learning dataset for weed identification" (2022). https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9748680/
  3. Hugging Face (Mobiusi), "Weed-Detection-Dataset README". https://huggingface.co/datasets/Mobiusi/Weed-Detection-Dataset/blob/main/README.md
  4. arXiv (2103.01415), "A Survey of Deep Learning Techniques for Weed Detection from Images" (2021). https://arxiv.org/pdf/2103.01415
  5. PubMed Central (Data in Brief), "Dataset of annotated food crops and weed images for robotic computer vision control" (2020). https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7305380/
  6. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
  7. arXiv (2310.16787), Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  8. arXiv (2506.17185), "A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Training Dataset" (2025). https://arxiv.org/pdf/2506.17185

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data