Video data
Can You Train Commercial Models on Open Driving Video Datasets? A License Check
Quick answer
Usually not without a separate agreement, because the most-cited public driving datasets are published for non-commercial research: the Waymo Open Dataset license covers non-commercial purposes only [3], Argoverse 2 ships under CC BY-NC-SA 4.0 [4], and KITTI under CC BY-NC-SA 3.0 [5]. Training a production perception model, running internal benchmarking that supports a product, or shipping weights derived from these sets generally falls outside those grants. Read each dataset's current terms, check clauses on models and derived labels, and license commercial data where the answer is no.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Which popular driving datasets allow commercial training?
As of October 2026, the large public autonomous driving datasets that procurement teams most often encounter restrict use to non-commercial purposes, and the license text, not the download page, decides. The table below summarizes what the primary terms say; verify the version in force on the date you download, because terms are revised.
| Dataset | Published license or agreement | Commercial training? | What to check next |
|---|---|---|---|
| Waymo Open Dataset (Perception, Motion) | Waymo Dataset License Agreement for Non-Commercial Use [3] | No, under the standard agreement | Definition of "Non-commercial Purposes"; attribution on derivative works; limits on sharing extracts |
| Argoverse 2 (Sensor, Lidar, Motion Forecasting) | CC BY-NC-SA 4.0 for data; code under a separate license [4] | No | ShareAlike on derived annotations; separation of code from data |
| KITTI and KITTI-360 | CC BY-NC-SA 3.0 [5] | No | Account terms; derivatives such as relabels inherit NC-SA |
| nuScenes and nuImages | Dataset-specific terms of use accepted at registration | Commonly described as non-commercial; read the current terms | Whether a separate commercial license is offered and on what scope |
| BDD100K | Toolkit code and dataset governed by different licenses | Read the dataset license, not the code license | Whether images and labels carry the same terms |
Two traps recur. First, the permissive license on a GitHub devkit (MIT, BSD-3-Clause, Apache-2.0) says nothing about the sensor data it loads. Second, hub labels are unreliable: the Data Provenance Initiative found license information on popular hosting sites frequently omitted or recorded incorrectly [1], so a "cc-by-4.0" tag on a mirror of a driving set is not evidence of anything.
What does "non-commercial" actually exclude for an ADAS or autonomy team?
Non-commercial clauses exclude far more than selling the data; they typically exclude any use primarily directed toward commercial advantage. Waymo's agreement, for example, defines non-commercial narrowly and excludes uses intended for or directed toward commercial advantage or monetary compensation [3]. The Creative Commons NC licenses used by Argoverse 2 and KITTI [4][5] define NonCommercial with similar "not primarily intended for or directed towards commercial advantage" language.
For a company building perception, prediction or planning stacks, that reaches activities engineers often treat as harmless:
- Pre-training a backbone (BEV encoder, occupancy network, video foundation model) that later feeds a production model.
- Using a public validation split as an internal regression benchmark that gates releases.
- Fine-tuning a vendor model on the data during a proof of concept for a paying customer.
- Generating pseudo-labels or auto-labeling models trained on the data and applying them to your own fleet logs.
Academic collaborators at universities may operate under the research grant, but once outputs flow into a company's model registry, the company's use is the one that counts. Statutory exceptions rarely rescue this: the UK text and data analysis exception in CDPA s29A covers non-commercial research only [6].
Which clauses decide whether trained models and weights are clean?
The clauses that matter most are those on derivative works, model outputs, redistribution and attribution, because they determine whether a model trained on the data can ship. Licenses for driving data vary in whether they treat trained weights as derivatives, and many are silent, which leaves the question to copyright and contract law rather than the license text.
Read for these specific provisions:
- Derivative definition. Does "derivative work" or "modification" include models, embeddings or features? Waymo's agreement requires derivative works to carry attribution to the dataset and the license [3], so check whether your team's trained artifacts fall within that term.
- ShareAlike. Under CC BY-NC-SA, adapted material you share must use the same license [4][5]. A relabeled KITTI subset you publish or hand to a supplier inherits NC-SA.
- Redistribution. Most agreements allow only limited extracts for illustration in publications [3]. Copying raw camera or lidar frames into a vendor's annotation tool can be a redistribution.
- Termination and survival. If the license ends, do obligations to delete data and derived labels survive?
- Vehicle operation. Some driving licenses have, in earlier versions, restricted use in operating a vehicle; confirm against the current text [3].
US case law is not settled in a way that makes non-commercial terms safe to ignore. On 29 September 2026 the Third Circuit affirmed that copying for a competing non-generative AI tool was not fair use in Thomson Reuters v. Ross [7]. Do not plan around a predicted fair-use outcome for perception training.
Where do hidden rights layers sit inside a driving dataset?
A driving dataset is a bundle of separately owned layers, and each can carry its own terms. Research on public datasets found that sets compiled from multiple sources often carry mixed licenses that a single top-level label hides [2]. For driving data, map at least these layers:
- Raw sensor data: camera video, lidar sweeps, radar, IMU and GNSS logs, owned by the collecting organization.
- Annotations: 3D cuboids, 2D boxes, panoptic masks, lane topology and tracking IDs, sometimes produced by a third-party labeling vendor under its own terms.
- HD maps and map-derived features: vector maps, drivable area rasters and lane graphs, which can come from a mapping provider.
- People and plates: faces and license plates visible in frames. Check whether blurring was applied to every frame and modality, and which privacy regimes applied at capture.
- Derived benchmarks: leaderboards and challenge splits with separate rules on test-server submissions.
Our guide to rights layers in a video clip covers footage, people and on-screen content in more depth, and the open data license compatibility matrix shows how CC, ODC and custom terms combine.
A license audit worksheet for public driving datasets
Run one row per dataset version, and keep the evidence, not just the conclusion. Structured documentation in the spirit of Data Cards, recording upstream sources, collection and annotation methods and intended use, makes the audit repeatable [8].
Illustrative example: invented to show structure; it does not describe an available dataset.
dataset_license_audit:
dataset: "ExampleDrive Open v2.1"
downloaded_on: 2026-09-14
source_of_terms: "license PDF from project site, SHA-256 recorded"
hub_label_seen: "cc-by-4.0 (mirror)" # do not rely on this
governing_license: "custom non-commercial agreement, rev. 2025-03"
layers:
camera_video: { owner: "dataset publisher", terms: "NC" }
lidar: { owner: "dataset publisher", terms: "NC" }
3d_annotations: { owner: "labeling vendor", terms: "assigned to publisher" }
hd_map: { owner: "map provider", terms: "separate EULA, not reviewed" }
intended_internal_use: "backbone pre-training for production ADAS model"
clause_findings:
commercial_use: "prohibited"
trained_models: "silent; counsel review required"
redistribution: "extracts for publications only"
sharealike: "n/a"
commercial_license_available: "ask publisher"
decision: "exclude from production training; allow in sandboxed research only"
reviewer: "legal + ML lead"
Pair the worksheet with a technical control: tag every training run with the dataset versions it consumed, so a model trained on an NC set cannot be promoted to production by accident. The open dataset license audit guide covers the general method across modalities, and the AI training data due diligence checklist lists the documents to request from any supplier.
What are the options when the license says no?
When a public driving dataset's terms exclude commercial training, the practical options are to negotiate a commercial license with the publisher, restrict the data to quarantined research, or source licensed fleet footage directly. Each carries different costs and diligence.
| Option | When it fits | What to verify |
|---|---|---|
| Commercial license from the publisher | You need that exact sensor suite, geography or benchmark comparability | Scope covers training and weights; fee basis; term; deletion on expiry |
| Sandboxed research only | Exploratory work with no path to product | Access controls and run tagging that prevent leakage into production |
| Licensed real-world footage from data holders | Production training at scale, specific ODDs or edge cases | Chain of title, consents, de-identification method, written allowed uses |
| Synthetic or simulated scenes | Rare events and controllable variation | Simulator and asset licenses; domain gap to real sensors |
The real cost of "free" datasets explains why audit, relabeling and legal review often outweigh the zero list price, and licensed vs synthetic vs scraped training data compares the three routes. For real footage, see dashcam video datasets from fleet footage and, for in-cab cameras, driver monitoring video, which raises biometric consent questions that exterior footage does not.
How SourceX fits when you need licensed driving footage
SourceX sources operational datasets from US companies on request and manages the commercial process, including the licensing agreement and ongoing purchases; it holds no stock, categories are not inventory, and a request does not guarantee a match. It looks for US businesses that hold the data you describe, such as new recordings of hands-on work, and every release is approved by the supplying company. It does not source scraped web content or generic CCTV.
Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which maps directly onto the audit worksheet above.
You can describe the driving or vehicle data you need and SourceX will search for holders; nothing is contracted until a supplier agrees. Browse more in the video data hub or the full AI data guides.
Need driving data you can train commercial models on?
SourceX finds US companies that hold the operational data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is transacted. Prices are not published and terms are agreed per deal, and SourceX serves AI teams wherever they are based. Start a buyer request at sourcex.si/buyers.
Sources
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/pdf/2310.16787.pdf
- arXiv, "Can I use this publicly available dataset to build commercial AI software? A case study on publicly available image datasets" (2021). https://arxiv.org/abs/2111.02374v4
- Waymo, "Waymo Open Dataset License Agreement" (2025). https://waymo.com/open/licensing
- arXiv (Wilson et al.), "Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting" (2023). https://arxiv.org/pdf/2301.00493
- Karlsruhe Institute of Technology and Toyota Technological Institute (cvlibs.net), "The KITTI Vision Benchmark Suite". https://www.cvlibs.net/datasets/kitti/
- UK Intellectual Property Office (GOV.UK), "Copyright, Designs and Patents Act 1988 - Consolidated (section 29A)". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (3d Cir.)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
- arXiv (Pushkarna, Zaldivar, Kjartansson), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.