Agent, workflow and domain-reasoning data
AI agent training data: the workflow, action and decision data agents learn from
Quick answer
AI agent training data is the record of real work that an agent learns to repeat or is graded against. It comes in six kinds: screens with labeled elements, GUI and tool-call trajectories, process event logs, decisions with their rationale, errors with their fixes, and seed state for practice environments. Public benchmarks cover websites and desktop apps; enterprise systems of record, approval waits and exceptions are the gap. Filling it usually means rebuilding trajectories from business logs and licensing them carefully.
By SourceX Editorial · Updated
Definitions live in the glossary entries for agent trajectory, computer-use data and tool-use data; why AI agents need business data covers the motivation. For data types outside agent work, start from the AI data sourcing guide for buyers.
Six kinds of agent data, sorted by what they teach
Agent data is easiest to scope by the skill it teaches: grounding, action, process, judgment, recovery or environment. Each kind comes from different systems in a different record unit, and most programs need several.
| Kind | What it teaches | Where it originates | Record unit | Go deeper |
|---|---|---|---|---|
| Grounding | Locating on-screen elements in the target apps | Screenshots paired with accessibility trees, DOM snapshots or element boxes | One screen with labeled elements | GUI grounding data from enterprise applications |
| Action | The next click, keystroke or tool call toward a goal | Screen recordings with input events, browser action logs, API call logs | Step: observation, action, result | Computer-use step records, API call logs as tool-use data |
| Process | Step order across systems, handoffs and waiting time | Event logs in XES [1] or OCEL 2.0 [2], ticket and case histories, field-level audit trails | Case or object-centric event log | Process mining event logs, ticket histories as trajectories |
| Judgment | Which option to take under policy, and why | Approvals, rejections, underwriting and claims decisions with reason codes and notes | Inputs at decision time, choice, rationale, later outcome | Decision records with rationale, approval and rejection records |
| Recovery | Noticing an error and repairing it | Reopened cases, reversals, credit memos, exception queues | Failure plus corrective steps | Exception handling records, rework and reversals |
| Environment | A realistic world to act in and a test of success | Seed records, state snapshots, app configuration, user-simulator data | Initial state plus success criteria | Seed data for agent sandboxes, user simulator data |
A tool-calling agent behind an API can skip grounding; a computer-use agent for claims intake may need all six. SourceX's capability pages list business records for computer-use agents, tool use and function calling, long-horizon task agents, agentic workflow planning, IT operations agents and email and calendar agents.
Train, build or test: three jobs for the same records
The same business records can train a policy, furnish the environment and grader that policy practices in, or evaluate it on held-out tasks. Each job needs different fields, so fix the job before writing the specification.
Training a policy. Supervised fine-tuning on demonstrations, and reinforcement learning (RL) warm-starts, need step-level observation-action pairs plus a stated goal. Mind2Web is a public example: open-ended tasks paired with crowdsourced action sequences on real websites [3]. Ask for coverage across task variants, not hundreds of repeats of one path; see data for post-training teams.
Building environments and graders. RL and repeatable evaluation need a world the agent can act in and an automatic check of the result. τ-bench builds that world from databases and APIs, domain policy documents and simulated user scenarios, grades the database state at the end of an episode, and reports pass^k, the probability of succeeding on all k repeated trials of a task [4]. OSWorld supports task setup and execution-based evaluation on real operating systems [5].
Business data supplies the realistic parts: seed records, app configuration, the policy in force and the end state that counts as done. See configuration data for enterprise app replicas and the RL environment entry.
Evaluating on held-out tasks. Test tasks must stay out of every training set, including the model vendor's. In a February 2026 post, OpenAI said SWE-bench Verified had become increasingly contaminated and that it had stopped reporting scores on it [6]. SWE-Bench Pro keeps a held-out set private and adds a commercial set from startup codebases whose code is not released [7]: private company data used to keep part of an agent benchmark out of public training corpora. Agent evaluation task suites and human baseline data cover design.
Assign each case to training, environment or test by case ID and time period before anyone uses it. Splitting at the step level leaks, because steps from one case land on both sides.
What public agent benchmarks cover, and what enterprise work adds
Public agent benchmarks cover consumer websites, self-hosted web apps, desktop operating systems and simulated service desks, but say little about configured back-office systems, approval waits measured in days, or the exceptions enterprise agents meet. Use them as baselines and task templates, and source data for the gaps.
| Benchmark | What it covers | What enterprise buyers still have to source |
|---|---|---|
| Mind2Web [3] | Over 2,000 tasks from 137 real websites in 31 domains, with crowdsourced action sequences | Logged-in internal applications; tasks spanning several sessions |
| WebArena [8] | Self-hosted sites for e-commerce, forums, collaborative software development and content management | ERP, claims, ITSM and other systems of record; role-based permissions |
| OSWorld [5] | 369 tasks in the 2024 release, with real desktop and web apps, file I/O and multi-app workflows, in an environment supporting Ubuntu, Windows and macOS | Customized enterprise clients, virtual desktops, terminal applications |
| τ-bench [4] | Simulated retail and airline domains with APIs, written policies and simulated users | Real policy versions over time; the real mix of exceptions |
| Spider 2.0 [9] | 632 data-workflow problems on databases from real applications, often over 1,000 columns, on systems such as BigQuery and Snowflake | The business request behind a query and the action taken after it |
Benchmarks are revised, so scores from different releases are not comparable, and public tasks leak into training data, as the SWE-bench Verified case shows [6]. Computer-use evaluation tasks with verifiable end states covers building your own, and legacy desktop and terminal interaction data covers interfaces no benchmark models.
Why business records rarely arrive as clean trajectories
Business systems log state changes, not the steps a person took to cause them. A trajectory has to be reconstructed by joining audit trails, event logs, tickets and messages on shared object IDs and ordering them by time.
An ERP audit trail shows that an invoice's payment block changed at 14:07 and who changed it, not the screens checked first, the value copied from an email or the abandoned attempt. Clock skew between systems reorders steps, and scheduled jobs and RPA bots write under service accounts that look like people. Field-level audit trails, cross-system workflow records and RPA bot logs cover each source.
Process-mining formats help. IEEE 1849-2023 (XES) defines an XML format for event logs and event streams and supersedes the 2016 edition [1]. OCEL 2.0 links events to several objects, records changes to object attributes over time and qualifies object-to-object relationships, with SQLite, XML and JSON exchange formats [2]. That object-centric model fits back-office work, where one order-to-cash case touches an order, deliveries, invoices and payments; see procure-to-pay and order-to-cash records.
Require every step to state its evidence: observed in an audit trail, observed on screen, or inferred. Inferred steps help process supervision but are weak action labels. SourceX's workflow task histories page describes records rebuilt from system audit logs.
Illustrative example: invented to show structure; it does not describe an available dataset.
task:
task_id: ap-exc-0447
family: invoice_price_variance_resolution
goal: "Release an invoice blocked by a purchase-order price variance"
systems: [erp_accounts_payable, erp_purchasing, email, itsm_queue]
outcome: {status: posted_after_po_change, label_source: erp_document_status, reopened_within_30d: false}
timing: {active_minutes: 38, wall_clock_hours: 52} # most of the elapsed time is waiting for the buyer
steps: # steps 1-2 (invoice receipt and automatic match) omitted
- {n: 3, ts: "2025-03-12T14:07:55Z", actor_role: ap_specialist, system: erp_accounts_payable,
action: {type: field_update, object: invoice_4471, field: payment_block, from: none, to: price_variance},
evidence: audit_trail}
- {n: 4, ts: "2025-03-12T14:11:20Z", actor_role: ap_specialist, system: email,
action: {type: send_message, to_role: purchasing_buyer, template: price_query},
evidence: mail_log}
- {n: 5, ts: "2025-03-14T09:40:02Z", actor_role: purchasing_buyer, system: erp_purchasing,
action: {type: field_update, object: po_88213_line_2, field: net_price, from: 41.20, to: 43.05},
rationale: "Supplier increase accepted under contract amendment", evidence: audit_trail}
- {n: 6, ts: "2025-03-14T09:52:10Z", actor_role: ap_specialist, system: erp_accounts_payable,
action: {type: release_block, object: invoice_4471},
evidence: inferred_from_status_change} # no screen capture exists for this step
privacy: {people: role_pseudonyms_consistent_across_systems, supplier_names: tokenized}
Rights and privacy: what a screen or log reveals
Agent data captures more parties than most datasets: customers in the records, employees doing the work, vendors whose software is on screen and, in deployment logs, your own users. Each needs a rights answer before delivery.
| What the data captures | The question to settle | Go deeper |
|---|---|---|
| Customers or patients | Identifiers sit in pixels, window titles and autocomplete lists, not only in fields. Health data needs HIPAA de-identification: Safe Harbor removal of 18 identifiers, or an expert's determination that identification risk is very small [10] | PII redaction for screen recordings and trajectories |
| Employees doing the work | Screen, keystroke and call capture at work raises notice and consent questions that vary by state; California prohibits recording confidential communications without all parties' consent [11]. Keystroke logs can also capture passwords | Collecting computer-use demonstrations at work |
| Vendors' software on screen | Screenshots and environment replicas reproduce third-party interfaces whose license terms may restrict reuse | Replicating third-party software, third-party content in screen recordings |
| Users of your deployed agent | FTC staff warned in 2024 that model-as-a-service companies may be liable if they break promises not to use customer data for training [12] | Training on agent deployment logs |
Documentation duties follow some agents. From 1 January 2027, Colorado's SB26-189, signed in May 2026, requires developers of automated decision-making technology that materially influences consequential decisions, such as lending, insurance, employment or health care, to give deployers documentation that includes the categories of training data [13]. As of October 2026, the Attorney General's implementing rules are an interim draft [14].
In the EU, providers of general-purpose AI models must publish a sufficiently detailed summary of the content used for training under Article 53(1)(d) of the AI Act [15]. The compliance hub and license terms for agent data go further.
For the data it sources, SourceX runs a rights review that checks the business owns or may share the records and that required consents are in place. Names, emails, phone numbers and account numbers are removed or replaced before delivery and the method is recorded, though no de-identification method is perfect; health records must meet HIPAA Safe Harbor or Expert Determination before they are considered for a license. This section is general information, not legal advice.
Scoping your first agent data request
A workable request names the task family, the systems, the observation and action format, how success is labeled, and every use the license must cover.
- Task family and goal: for example, "release an invoice blocked by a price variance," with in-scope and out-of-scope variants.
- Systems and surfaces: application, version, customization level, and whether it runs in a browser, desktop client, virtual desktop or terminal.
- Observations: pixels, accessibility tree or DOM, API request and response payloads, or event log only.
- Actions: clicks and keystrokes, tool calls with arguments, or business-level actions such as approve or reassign.
- Outcome label and its source: document status, reopen flag or QA review; see task success labels.
- Horizon: steps per task, active versus wall-clock time, and handoffs; see long-horizon task records.
- Exceptions: a minimum share of cases off the happy path.
- Privacy: pseudonyms consistent across steps and systems, so the agent can follow one customer through a case.
- License uses: training, environment replication, derived tasks and benchmark publication, plus split rules.
- Delivery: JSONL with one episode or step per line (UTF-8, one JSON value per line [16]), OCEL or XES for event logs, and a datasheet covering motivation, composition and collection process [17].
Writing an agent data specification expands each item, evaluating an agent data sample covers acceptance, and cost drivers for trajectory data explains pricing variables. For the sourcing route, compare licensed, commissioned or synthetic trajectories.
SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases. The kinds of data it sources include support and sales histories, engineering records, documents, finance and legal workflows, and new recordings of hands-on work. Datasets are sourced on request rather than held in stock, so a request does not guarantee a match, and every release is approved by the supplying company.
SourceX does not train AI models. To start, describe the workflow records your agent needs.
Agent data guides by interface and decision type
Choose the guide that matches how your agent touches software and what it decides.
- Screens and browsers: browser action logs, screen recordings to action-labeled trajectories, task mining data, spreadsheet task trajectories, contact-center desktop activity with transcripts.
- Tools and APIs: MCP tool-use data, policy-following service agent data.
- Back-office procedures: SOP-to-execution pairs, document-to-system entry pairs, runbook execution records.
- Regulated judgment: underwriting decision rationale, claims adjudication decisions, agent-to-human handoff data.
- Adjacent hubs: coding-agent data sits in the code datasets hub, including developer session recordings; evaluation design sits in the evaluation hub. Enterprise AI agent training data for computer use summarizes record types for that use case.
Mistakes that waste an agent data budget
Each of these gaps is visible in the specification before signing and costs more to fix after delivery.
- Screen video with no input events, task boundaries or outcomes. It is raw material, not a trajectory, and inferring actions from pixels afterwards is harder than logging input events at capture.
- "Closed" taken as success. Auto-close rules, reopened cases and later reversals make status fields poor labels unless checked against what happened next.
- Exceptions filtered out for clean demonstrations. The rework and edge cases removed are the ones agents fail on in production.
- A training-only license. Building an environment replica, deriving tasks or publishing a benchmark may fall outside a grant that names only training, so name each use.
Building agents that need real work histories?
Describe the task family, the systems involved, the outcome labels you need and the uses the license must cover. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Specify your workflow dataset.
Guides in this section
- Agent Data License Terms: Environments, Replay, BenchmarksThe license grants agent builders need beyond training: environment construction, trajectory replay, derived synthetic tasks and benchmark publication.
- Agent training data requirements: a specification templateWhat an agent training data specification must state: task inventory, observation and action fields, coverage cells, outcome labels, privacy and delivery.
- Agent-to-Human Handoff Data: Escalation Records for AIWhat agent-to-human handoff and escalation data must contain: triggers, context summaries, receiver re-asks, outcomes and missed escalations.
- API call logs for tool-use training: a sourcing guideHow to source API request/response logs and integration run histories for function-calling training: bodies, intent links, secrets, errors and rights.
- Approval Workflow Data for AI Agents: Decisions and ReasonsWhat approval workflow data for AI agents must hold: decision types, reasons, delegation-of-authority versions, signal checks, privacy and a request list.
- Claims adjudication training data for AI agentsWhat claims adjudication data for AI agents must hold: coverage decisions, cited policy provisions, adjuster reasoning and appeal outcomes per claim.
- Collecting computer-use demonstrations from employeesHow to commission screen and keystroke capture from employees for agent training: monitoring-notice laws, consent, recorder scope and vendor terms.
- Computer-use agent trajectory data format, field by fieldWhich observation, action and episode fields a computer-use trajectory dataset needs for training and replay, with an illustrative step record.
- Cross-System Workflow Records for AI Agents: Linking GuideHow to source and link ERP, CRM, ITSM and email records into one task timeline for agent training and evaluation: keys, pseudonyms, clocks, coverage.
- Decision Rationale Data for AI Agents: What to SpecifySpecify decision rationale data for AI agents: inputs as known at decision time, reason codes and notes, linked outcomes and decider consistency checks.
- Document-to-System Entry Pairs for Back-Office AgentsDocument-to-system entry pairs for back-office agents: what posted records add, how pairs link, corrections, non-entries, privacy and a request template.
- Exception Handling Data for AI Agents: What to SourceWhat exception handling data for AI agents must hold: queue triggers, cross-system investigation steps, resolution codes and matched clean cases.
- GUI grounding dataset sourcing for enterprise app screensWhat to specify in a GUI grounding dataset from enterprise software: element boxes, roles, referring expressions, coverage, redaction and rights.
- Human demonstrations vs synthetic agent trajectoriesCompare licensed workflow records, commissioned human demonstrations and synthetic agent trajectories on cost, realism, rights and coverage.
- Policy-Following Agent Training Data from Service RecordsHow to source policy-following agent training data: versioned policies, service conversations and back-office actions linked to account state.
- Process Mining Event Log Datasets for AI AgentsWhat a process mining event log dataset needs for agent training and evaluation: source systems, fields, XES or OCEL format, leak-free splits and privacy.
- Procure-to-Pay and Order-to-Cash Event Logs for AgentsWhat a procure-to-pay or order-to-cash event log dataset must hold for finance agents: document flow, match exceptions, resolutions, privacy and format.
- Seed Data for Agent Sandbox Environments: Real vs SyntheticSourcing seed data for agent sandbox environments: time-consistent snapshots, intact joins, preserved messiness, consistent masking and license scope.
- SOP execution pairs for agent training: a sourcing guideHow to source SOP execution pairs for agents: version-matched procedures, linked cases, adherence and deviation labels, sample checks and rights questions.
- Ticket history to agent trajectory: a conversion methodHow to convert support and ITSM ticket histories into step-level agent trajectories: event mapping, waits, outcome labels, missing steps and pseudonyms.
- Underwriting Rationale Data for AI: Referrals and ExceptionsWhat underwriting rationale data for AI holds: notes, referral and exception records linked to evidence and outcomes, plus proxy and privacy checks.
- Agent Trajectory Data Cost: What Drives the PriceWhat drives agent trajectory data cost: capture method, action labeling, screen de-identification, environment rights and refresh, plus a quote worksheet.
- Audit Workpaper Data for AI Agents: Judgment and ReviewHow to source audit workpaper data for AI agents: risk assessments, sampling decisions, exceptions and reviewer notes, with confidentiality checks.
- Browser Action Logs for Web Agents: What to SourceWhich browser logs to source for web agent training: DOM snapshots, selectors, network calls and outcomes, and why analytics clickstream falls short.
- Commercial Credit Memo Data for AI Lending AgentsHow to source commercial loan credit memos, financial spreads, risk ratings and committee decisions to train and evaluate credit analysis agents.
- Configuration Data for Enterprise App Replicas for AgentsWhich tenant configuration artifacts to source for enterprise app replicas: custom fields, picklists, workflow rules, permissions and layouts for agents.
- Contact-Center Desktop Activity Paired With Call TranscriptsHow to source aligned call transcripts and agent desktop actions in CRM and billing systems for training and evaluating voice and chat agents that act.
- Error Recovery Data for AI Agents: Rework and ReversalsHow to source human error-recovery records for AI agents: reversals, reopened tickets, credit notes and re-shipments linked into detect-and-fix sequences.
- Evaluating an Agent Trajectory Data Sample Before PurchaseHow to request and test an agent trajectory sample before buying volume: random draws, field coverage, harness replay, small fine-tunes and sample terms.
- Field-Level Audit Trails as Agent Action Training DataHow to rebuild agent action sequences from field-level change logs (SAP CDHDR/CDPOS, Salesforce field history, CDC): grouping, actor filtering and limits.
- HR Service Case Data for Employee-Service AI AgentsWhat HR case data an employee-service agent needs: case types, policy versions, HRIS actions and outcomes, plus the exclusions and privacy checks to set.
- Human Baselines for Agent Evaluation: Time, Cost and QualityHow to source and structure human time-on-task, cost and quality baselines so agent evaluation results compare fairly with real human performance.
- Human Override and Correction Logs for AI TrainingSourcing human override and correction logs from rules engines, extraction and model scores: fields, selection bias, preference pairs and rights.
- Legacy Terminal and Desktop Interaction Data for AgentsHow to source interaction data from 3270/5250 terminals, thick clients and remote desktops to train and evaluate computer-use agents on no-API systems.
- Long-Horizon Task Records: Measuring and Specifying HorizonHow to measure long-horizon agent task data: human completion time, action count, waits, handoffs and context growth, plus a request spec you can reuse.
- MCP Tool-Use Datasets: Traces, Schemas and Eval TasksWhat an MCP tool-use dataset must contain: tool catalogs, tools/list and tools/call traces, selection labels, error cases and injection-aware eval tasks.
- OCEL 2.0 Object-Centric Event Logs for Agent DataWhat OCEL 2.0 object-centric event logs capture, when they beat flat XES logs for orders, items and deliveries, and what to ask suppliers before buying.
- Rights Issues in Cloning SaaS Apps for Agent TrainingCopyright, terms of service, trademark, access and tenant-data questions to clear before replicating commercial software for agent training environments.
- RPA Bot Definitions and Run Logs as Agent Training DataHow to source RPA workflow definitions, bot run histories and exceptions handed to people as data for training and benchmarking AI agents.
- Runbook Execution Data for SRE and IT Operations AgentsHow to source runbooks linked to alerts, executed commands, automation runs and outcomes for training and evaluating SRE and IT operations agents.
- Screen Recordings to Action-Labeled Agent TrajectoriesThree ways to label actions in screen recordings for agent training: concurrent event capture, human annotation and inverse dynamics pseudo-labels.
- Spreadsheet Agent Training Data: Edit-History TrajectoriesHow to source spreadsheet edit histories for agent SFT and evaluation: version snapshots, cell-level diffs, task requests, scrubbing and formula grading.
- Task Mining Data for AI Agents: Desktop Logs as TrajectoriesCan task-mining captures of clicks, app switches and sampled screenshots become agent trajectories? Fidelity limits, notice scope and licensing checks.
- Task Success Labels for Agent Trajectories from RecordsHow to derive success and failure labels for agent trajectories from business records: proxy outcomes, observation windows, partial credit, expert audits.
- Tool-Call Error Recovery Data: Real Failures for AgentsHow to source and label real tool-call failures (validation, auth, rate limits, timeouts) and the recoveries that followed, for agent training and eval.
- User Simulator Training Data for Agent RL and EvaluationWhat real requester-side interaction data you need to build and calibrate simulated users for agent RL and evaluation, with a field schema and checks.
- Who Owns AI Agent Logs? Training Rights and Contract TermsWho owns AI agent interaction logs and human review data, and when an agent vendor may train on them: content vs usage data, consent and key clauses.
Sources
- IEEE Standards Association, "IEEE 1849-2023 - IEEE Standard for eXtensible Event Stream (XES) for Achieving Interoperability in Event Logs and Event Streams" (2023). https://standards.ieee.org/ieee/1849/10907
- OCEL standard authors (ocel-standard.org), arXiv:2403.01975, "OCEL (Object-Centric Event Log) 2.0 Specification" (2023). https://arxiv.org/pdf/2403.01975
- Deng, Su et al., The Ohio State University (arXiv:2306.06070; NeurIPS 2023), "Mind2Web: Towards a Generalist Agent for the Web" (2023). https://arxiv.org/abs/2306.06070v1
- Yao et al., Sierra (arXiv:2406.12045), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- Xie et al., XLANG Lab, HKU and collaborators (arXiv:2404.07972 v2; NeurIPS 2024), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- Deng et al., Scale AI (arXiv:2509.16941), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
- Zhou, Xu et al., Carnegie Mellon University (arXiv:2307.13854 v4), "WebArena: A Realistic Web Environment for Building Autonomous Agents" (2024). https://arxiv.org/abs/2307.13854v4
- Lei et al., XLANG Lab, HKU and collaborators (arXiv:2411.07763; ICLR 2025), "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows" (2024). https://www.arxiv.org/pdf/2411.07763
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- California Legislature (California Legislative Information), "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN§ionNum=632
- Federal Trade Commission, Office of Technology (Tech@FTC staff blog), "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
- Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
- Colorado Attorney General, "Colorado Automated Decision-Making Technology & Chatbot Safety Rulemaking" (2026). https://coag.gov/ai/
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- Gebru et al. (arXiv:1803.09010; Communications of the ACM 2021), "Datasheets for Datasets" (2021). https://arxiv.org/pdf/1803.09010
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.