Leadership and readiness
Data profile one-pager: how to describe your records without sharing files
By SourceX Editorial · Updated
Short answer
A data profile one-pager describes company records with metadata only: which systems hold them, the date range, approximate volumes, the fields and links between records, and what will be excluded. It lets a buyer or partner judge fit without seeing a single file. The rule: if writing a line would require an export, leave it out.
Key takeaways
- A data profile describes records with metadata only and never includes sample rows or screenshots of real records.
- Every line can be filled from admin reports, settings pages, migration notes and system owners.
- Accessible history, not years in business, is the date coverage that matters.
- Exclusions and known restrictions belong on the first page, before anyone asks.
- A specific profile maps easily onto dataset documentation standards such as the Data Provenance Standards.
What is a data profile one-pager for?#
A data profile one-pager is a single page that describes a company's records well enough for a buyer, partner or board to judge fit, without a single record changing hands. It answers the question every conversation starts with, what do you have, in a form that can be shared early and safely.
The page sits between a vague claim and a sample. A vague claim tells nobody anything, while a sample discloses real records before rights and privacy have been reviewed. A profile names systems, time spans, record types and limits, and stops there.
The same page gets reused: in the fit check, in the internal approval request, in a board update and as the starting point for a full data inventory.
The template, line by line#
The template has ten lines, each answerable from system knowledge and admin screens. Keep every entry to a phrase or two; the profile works because it is short. The example entries below describe a fictional commercial plumbing contractor.
| Line | What to write | Example entry |
|---|---|---|
| Company context | Industry, status (operating, acquired or wound down), length of operating history | Commercial plumbing contractor, operating, family owned |
| Systems | Each system of record and what it holds | Job management system for estimates, jobs and invoices; shared service inbox |
| Record families | The kinds of records, in plain words | Service calls, estimates, job notes, callbacks, warranty claims |
| Date coverage | Earliest and latest accessible records, plus known gaps | Accessible since the current system went live; older jobs on paper only |
| Volume | Approximate size from admin reports, as an order of magnitude | Tens of thousands of completed jobs, per the system's job report |
| Fields and links | Key fields and how records connect | Each job links to its estimate, invoice, technician notes and any callback |
| Language and format | Main language, structured versus free text, attachments | English; structured fields plus free-text notes; photos attached |
| Exclusions | What the company will not license | Customer payment details, employee HR records, site photos |
| Known restrictions | Contracts, notices or vendor terms that may limit use | Some commercial service agreements include confidentiality clauses |
| Contact and approver | Who answers questions and who would sign | COO for questions; president as signer |
How to fill it in without exporting anything#
Every line in the profile can be completed from places that show information about records rather than the records themselves. If a line seems to need an export, write unknown and leave it for the inventory stage.
- Admin dashboards and built-in reports for counts by year, status or type.
- Settings pages that list custom fields, ticket forms, job types and workflow stages.
- Plan and billing pages that show storage use and account age.
- Migration notes, statements of work and old project plans that record when systems changed.
- The IT asset register or the list of active subscriptions.
- A short conversation with each system owner about what the records contain and where history breaks.
Vague descriptions versus useful ones#
Useful descriptions name the record, the system and the outcome it captures. Vague descriptions make a reader guess, and guesses tend to be discounted.
| Vague | Useful |
|---|---|
| Lots of customer data | Help desk tickets with full agent replies, categories and resolution codes |
| All our email | Shared support inbox only; personal mailboxes excluded |
| Years of history | Accessible from the current CRM go-live; earlier records in a read-only archive |
| CRM data | Opportunity records with stage history, loss reasons and linked activities |
| Clean data | Cause, fix and parts fields required before a job can close |
| Engineering records | Jira issues linked to pull requests and release notes by issue key |
How the profile lines up with dataset documentation standards#
A specific profile already answers the questions that formal dataset documentation asks, which saves rework later. The Data & Trust Alliance's Data Provenance Standards group dataset metadata into Source, Provenance and Use, and describe that metadata as necessary for proper dataset selection for AI model training.
The Use group in that standard includes items such as confidentiality classification, where consent documentation is held, privacy-enhancing technologies applied, license to use and intended data use. The exclusions and known restrictions lines on your one-pager are early versions of those entries.
Machine learning teams also use approaches such as Datasheets for Datasets, published by Gebru and colleagues in Communications of the ACM in December 2021, and MLCommons' Croissant, a metadata format built on schema.org's Dataset vocabulary. Suppliers do not need to produce either, but a plain, specific profile translates into them easily.
Illustrative: an engineering firm drafts its first profile#
Illustrative: a fictional structural engineering firm keeps project setup, staffing and time in Deltek Vantagepoint, RFIs and submittals in Procore, markups in Bluebeam and internal QA/QC checklists in SharePoint. The managing principal wanted to know whether the firm had records worth discussing, without sending anything that might belong to clients.
The COO drafted the one-pager from Procore's project list and RFI log counts, Deltek's project reports and a conversation with the IT manager. It described RFI and submittal records with responses, reviewer comments and approval status, and listed drawings, models and calculation packages as exclusions because client agreements governed them. The date coverage line noted that older projects sat on an archived file server, accessible but unindexed.
The fit check that followed focused on the RFI and submittal history alone. Because drawings were excluded on the first page, the rights review never had to open a client deliverable.
Mistakes that make a profile misleading#
The most damaging profile mistake is overstating accessible history, for example counting years in business instead of years of records that can still be exported. A reader who later finds a migration gap will question every other line.
Other common mistakes are attaching a few real rows to show the format, pasting screenshots that include customer names, leaving exclusions blank because none have been decided and quoting volumes from memory instead of admin reports. Each one either discloses too much or promises too much.
How SourceX uses a data profile#
The SourceX fit check collects metadata, not files, so a data profile is exactly the kind of input it starts from. The profile feeds the Supply step of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery.
The SourceX Enterprise Data Value Framework, a SourceX-developed methodology rather than an industry standard, rates records on drivers such as human-generated signal, domain expertise, recency, rights and AI utility, and lines such as fields and links, date coverage and known restrictions give early evidence for them. If a package proceeds, its SourceX Evidence Packet later records provenance, licensing rights and permitted use in full.
Frequently asked questions
Can we include screenshots of the system?
Only of configuration screens that show field names, forms or workflow stages without record content. A screenshot of a ticket list or a CRM account view shows customer names and real details, which is the disclosure the profile is meant to avoid. When in doubt, describe the screen in words instead.
Who should write the profile?
The COO or an operations leader is usually best placed, with input from each system owner and a check of volumes by IT. The CEO or the person who would sign should read it before it goes out, since it describes what the company might license and what it will not.
Should the profile go out under an NDA?
A well-written profile is low in sensitivity because it contains no records, but it still describes systems and operations, so label it confidential. Whether to sign an NDA before sharing it depends on the counterparty and the conversation, and counsel can advise on the specific case.
How detailed should the field list be?
List field names and types, such as category, priority, resolution code or technician notes, and say whether each is required, optional or free text. Never include example values; a single real value in a field description can disclose more than the rest of the page.
Do we need one profile per company or one per system?
Write one profile per likely package. A software company may need one for support history and another for engineering records, while a smaller contractor might fit everything on a single page. Splitting by package keeps the exclusions and restrictions clear for each.
Sources
- The Data & Trust Alliance's Data Provenance Standards define dataset metadata in three groups, Source, Provenance and Use, needed to enable proper dataset selection for AI model training; the Use group includes confidentiality classification, consent documentation location, privacy-enhancing technologies applied, license to use and intended data use. Source
- Datasheets for Datasets by Timnit Gebru and colleagues was published in Communications of the ACM, vol. 64, no. 12 (December 2021). Source
- MLCommons' Croissant format is a metadata standard for machine-learning datasets built on schema.org's Dataset vocabulary. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.