Schemas, packaging and delivery
Data Dictionary Template for Licensed Dataset Deliveries
Quick answer
A data dictionary for a licensed dataset delivery should document every table and every field: name, type, nullability, units, allowed values or code list, source system and table, transformation applied, redaction status and an example value. At table level it should state grain, primary key, row count, date range and refresh behavior. Require it as machine-readable CSV or JSON plus a rendered human version, versioned with each delivery, so your pipeline can validate the export rather than reverse-engineer it.
By SourceX Editorial · Updated
Why a dictionary matters more for exported business-system data
Exports from CRM, ticketing, ERP and billing systems are full of meaning that lives outside the data, so a dictionary is the only reliable way to recover it. A column called status with values 3, 6 and 7 is useless for supervised fine-tuning or evaluation labels until someone tells you that 6 means "Resolved" and 7 means "Closed" in this tenant's configuration. Admin-customized picklists, renamed labels and retired values are routine in long-running systems, and none of them survive a flat CSV export.
Source-system structure also hides content in unexpected places. In ServiceNow, for example, incident comments and work notes are not stored on the incident row; they are written as entries in the sys_journal_field table and joined back by element and record ID [8]. A supplier who exports only the incident table delivers ticket metadata with no conversation text, and a dictionary that maps each delivered field to its source table makes that gap visible before you train on it.
A dataset card or datasheet answers different questions. Datasheets for Datasets frames dataset-level documentation around motivation, composition, collection and recommended uses [1], and Hugging Face dataset cards carry license, language and size in a YAML header [3]. The data dictionary sits one level down: it is the field-level contract your loaders, validators and labelers depend on. For the dataset-level layer, see our guide to dataset cards for licensed enterprise data.
The table-level entries every delivery needs
Each table in a delivery needs a header block that states what one row represents and how the table relates to the others. Without an explicit grain statement, the most common failure is double counting: an "orders" file that is actually one row per order line, or a "tickets" file that repeats the ticket for every status change.
Require these table-level entries:
- table_name and description: the delivered name and a one-sentence plain-language meaning.
- grain: "one row = one ticket comment" or "one row = one invoice line," stated in those words.
- primary_key: the column or columns that are unique per row, with a statement that uniqueness was tested.
- foreign_keys: each join to another delivered table, including cardinality (one-to-many, many-to-many).
- source_system and source_objects: for example, Salesforce
CaseplusCaseComment, or NetSuite transaction lines. - row_count and date_range: counted at export time, so you can reconcile against the delivery manifest and checksums.
- refresh_behavior: full snapshot, append-only increment, or upsert with deletes; see incremental deliveries vs full refreshes.
- time_zone_policy: whether timestamps are UTC, source-local, or mixed by field.
Primary keys deserve extra attention when records come from several systems. Salesforce, for instance, uses External ID fields to match incoming records during upsert [7], so a supplier's External_Ticket_ID__c may be the only stable link back to a ticketing system. Document which identifier is authoritative and how it persists across deliveries, as covered in stable record IDs and join keys.
Field-level columns: the template
The core of the dictionary is one row per field, with columns that a pipeline can read and a reviewer can audit. The columns below extend the familiar name-type-description-example layout with the provenance and redaction detail that licensed data needs.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Column | What it records | Example value |
|---|---|---|
| table_name | Delivered table the field belongs to | ticket_comments |
| field_name | Exact delivered column name, case-sensitive | comment_body |
| description | Plain-language meaning, not a restated name | Text of a customer-visible reply or internal note |
| data_type | Logical type in the delivered format | string, int64, decimal(12,2), timestamp[us, UTC] |
| nullable | Whether nulls occur, and what null means | true; null = note deleted at source |
| unit_or_format | Units, currency, or string format | USD; ISO 8601; seconds |
| allowed_values | Inline enum or a pointer to a code list | codelist:ticket_status |
| source_system | System of record | ServiceNow |
| source_object | Source table and column | sys_journal_field.value |
| transformation | Every change between source and delivery | HTML stripped; signatures removed; trimmed |
| redaction_status | none, masked, pseudonymized, removed, generalized | pseudonymized (names, emails replaced with tokens) |
| pii_category | Category of personal data before treatment | contact details, free text |
| example_value | A synthetic or already-redacted sample | Hi [PERSON_1], the patch is attached. |
| first_delivery_version | Version in which the field first appeared | v1.0 |
| notes | Known quirks, sparsity, quality caveats | Empty before 2021 migration |
Types should be stated in the delivered format's terms, not the source database's. Parquet carries logical types for decimals, timestamps and dates in its own type system [4], while JSON Lines only guarantees that each line is valid JSON in UTF-8 [5], so a JSONL delivery must spell out how timestamps and decimals are encoded as strings or numbers. The trade-offs are covered in Parquet vs JSONL for licensed training data.
Code lists for status, priority and reason codes
Every enumerated field needs a separate code list that maps each code to its human-readable meaning and records when the code was in use. Status, priority, category, disposition and reason codes are where label quality for SFT and eval sets is won or lost, because a model trained on raw codes learns the tenant's configuration rather than the task.
Ask for these code-list columns: codelist_name, code, label, definition, valid_from, valid_to, is_terminal (for workflow states), replaced_by, and source (local configuration or an external standard). Validity dates matter because picklists change. A "Pending Vendor" status added in 2023 cannot appear on a 2020 ticket, and if it does, your join or the export is wrong.
Externally governed code sets need their own version reference. Healthcare billing exports, for example, use X12 Claim Adjustment Reason Codes, which the code committee adds, modifies and retires over time and publishes with status and dates per code [6]. The dictionary should state which list version the supplier mapped against, and treat any HIPAA-covered content under a documented de-identification method such as Safe Harbor or Expert Determination [9].
Illustrative example: invented to show structure; it does not describe an available dataset.
codelist_name,code,label,definition,valid_from,valid_to,is_terminal,replaced_by,source
ticket_status,1,New,Created and not yet triaged,2018-01-01,,false,,local_config
ticket_status,6,Resolved,Fix provided; awaiting customer confirmation,2018-01-01,,false,,local_config
ticket_status,7,Closed,Closed after confirmation or auto-close timer,2018-01-01,,true,,local_config
ticket_status,9,Pending Vendor,Waiting on third-party supplier,2023-03-15,,false,,local_config
close_reason,DUP,Duplicate,Merged into another ticket,2018-01-01,2022-06-30,true,MERGED,local_config
Source-system field mapping and transformation logging
The mapping columns exist so you can trace every delivered value to a source object and every difference to a named transformation. This is what lets you distinguish a genuine pattern in the data from an artifact of the export job.
Require transformations to be written as discrete, ordered steps rather than prose, for example: 1. decode HTML entities; 2. strip quoted reply chains; 3. replace emails with [EMAIL_n] tokens; 4. truncate at 32,000 characters. Flag lossy steps explicitly, because truncation, deduplication and language filtering all change the distribution you will train or evaluate on. When records come from several systems, the mapping should name the join path, which ties into packaging linked records from multiple business systems.
Redaction status belongs at field level, not just in a cover letter. Structured fields such as account_number may be removed outright while free-text fields are pseudonymized with consistent tokens, and the dictionary should say which, using what method, so your privacy reviewers can assess residual risk. No automated method catches every identifier in free text, so treat redaction_status as a statement of method, not a guarantee. Our buyer's guide to de-identified data covers how to evaluate those methods.
Machine-readable format and versioning across deliveries
Ship the dictionary in two forms: a machine-readable file that your pipeline consumes and a rendered version for reviewers. CSV (one file for fields, one for tables, one for code lists) is the lowest-friction choice; JSON lets you nest code lists and table metadata in one document and validate it with a schema, as described in using JSON Schema to validate dataset deliveries.
If your team already uses ML documentation standards, map the dictionary into them rather than replacing it. Croissant-RAI extends the MLCommons Croissant vocabulary to make responsible-AI documentation machine-readable and reusable [2], and because a Hugging Face dataset card is the repository README with a YAML metadata header [3], the card can link to the dictionary files shipped alongside the data.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"dictionary_version": "2.1.0",
"dataset_version": "2026-09-30",
"tables": [{
"table_name": "ticket_comments",
"grain": "one row = one comment or work note on a ticket",
"primary_key": ["comment_id"],
"foreign_keys": [{"field": "ticket_id", "references": "tickets.ticket_id", "cardinality": "many-to-one"}],
"refresh_behavior": "append_only_with_tombstones",
"time_zone_policy": "all timestamps UTC"
}],
"changes_since_previous": [
{"type": "added_field", "table": "tickets", "field": "sla_breached", "reason": "new source field"},
{"type": "added_code", "codelist": "ticket_status", "code": "9"}
]
}
Version the dictionary with the dataset, not independently of it. Each delivery should carry a dictionary_version, and a change log should list added, removed, renamed and retyped fields plus added or retired codes. Renames and type changes are breaking changes for your loaders; agree in advance how they are announced, using schema evolution for recurring deliveries and data contracts as the mechanism.
Acceptance checks to run against the dictionary
A dictionary is only useful if you test the delivery against it on arrival. Automate these checks before any data reaches a training or evaluation job.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Every delivered column appears in the dictionary, and every dictionary field appears in the delivery.
- Declared types parse without coercion errors; decimals keep declared precision.
- Primary keys are unique; foreign keys resolve at the declared cardinality.
- Every enumerated value appears in its code list, and its timestamp falls within
valid_fromandvalid_to. - Null rates match the
nullabledeclaration and the documented meaning of null. - Row counts and date ranges match table-level entries and the manifest.
- Fields marked
removedare absent; fields markedpseudonymizedcontain tokens, not raw values, on a sample. - The change log explains every schema difference from the previous delivery.
Write failures back to the supplier against field names in the dictionary, which keeps the conversation precise. The broader specification that the dictionary plugs into is covered in our technical delivery specification template, and file-format expectations are answered in what file formats AI buyers accept. For manifest structure, see the sample manifest, and for the full cluster, the schemas, packaging and delivery hub.
How SourceX handles documentation for operational datasets
SourceX sources operational datasets from US companies, such as support and sales histories, engineering records and finance and legal workflows, and manages the commercial process through licensing and ongoing purchases. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and personal details such as names, emails, phones and account numbers are removed or replaced before delivery with the method recorded. If you are scoping a request, describe the fields and code lists you need on the SourceX buyer page.
Request operational data with field-level documentation
SourceX sources datasets on request rather than from stock, and every release is approved by the supplying company under a license that defines records, uses, term and delivery. Describe the data you need, including the dictionary you expect alongside it, and SourceX will look for US businesses that hold it; a request does not guarantee a match. Start a data request for your AI team.
Sources
- Gebru et al. (arXiv), "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010
- Jain et al., MLCommons Croissant RAI task force (arXiv), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
- The Apache Software Foundation, "Apache Parquet Documentation". https://parquet.apache.org/docs/file-format/types/logicaltypes/
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- X12, "Claim Adjustment Reason Codes". https://x12.org/codes/claim-adjustment-reason-codes
- Salesforce, "upsert() (SOAP API Developer Guide)". https://developer.salesforce.com/docs/atlas.en-us.api.meta/object_ref/sforce_api_calls_upsert.htm
- ServiceNow Community, "How to get all comments of incident using REST API" (2015). https://www.servicenow.com/community/developer-forum/how-to-get-all-comments-of-incident-using-rest-api/m-p/1343519
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.