Skip to content

Consulting and recruiting

Consulting firm data inventory template

By SourceX Editorial · Updated

Short answer

A consulting firm data inventory lists every system that holds firm records, with typical records, an owner, a date range and a rights status for each, then adds client-content, personal-data and export columns. Start with the CRM, PSA, shared drives, email, Slack or Teams, the proposal library and playbooks. Record metadata only, so it can circulate without exposing client work.

Key takeaways

  • One row per system and record family is enough for a first version; detail comes later.
  • The rights status column decides what the firm can do with each source, so mark unknowns honestly.
  • Shared drives and email usually mix client deliverables with internal records and need splitting.
  • Project reviews, proposals with outcomes and staffing plans are often the most useful firm-owned rows.
  • An inventory collects descriptions, not files; nothing is exported to build it.

What a consulting firm data inventory is for#

A consulting firm data inventory is a register of where the firm's records live, who is accountable for them and what the firm is allowed to do with them. It supports decisions that come up sooner than most partners expect: retiring a PSA, choosing an AI assistant, answering a client security questionnaire, preparing for an acquirer's diligence or reviewing whether records could be licensed.

It is not a records retention schedule and not a data map of every field. It sits one level up, describing each source in a sentence or two so that a managing partner can see the whole estate on one page.

Firms that skip it tend to discover their records in a hurry, during a system shutdown or a diligence request, when the people who know where things are may already be busy or gone.

The template: starting rows#

The starting rows below cover the systems almost every consulting firm runs. Copy them, replace the example system names with yours, and add rows for anything firm-specific, such as a survey platform or a client portal.

Keep all six columns even when some cells stay blank at first. A blank owner or rights cell is itself a finding, and it is easier to spot in a consistent layout.

The template: starting rows
SourceTypical recordsOwnerDate rangeRights statusNotes
CRM such as Salesforce or HubSpotAccounts, contacts, opportunities, win and loss reasonsHead of business developmentFrom CRM go-liveFirm-owned; contains client contact detailsCheck opportunity notes for client confidential detail
PSA such as Kantata, Certinia or BigTimeProjects, tasks, time entries with narratives, budgets, change ordersCOO or financeSince adoption, plus exports from older systemsFirm-owned records about client workConfirm narrative fields export in full
Shared drives such as SharePoint, Google Drive or BoxWorking files, research, deliverablesPractice leadsVaries by folderMixed; deliverables are often client-ownedSeparate client folders from internal ones
Email in Microsoft 365 or Google WorkspaceClient and internal correspondenceITPer retention policyMixed, with heavy personal dataUsually last on any reuse list
Slack or TeamsInternal discussion and project channelsIT and practice leadsPer retention settingsFirm-owned internal discussionTreat channels shared with clients separately
Proposal libraryProposals, SOWs, pricing rationale, outcomesBusiness developmentSince the library beganFirm-owned; client RFP material may be confidentialLink each proposal to its CRM outcome
Playbooks and methods in Confluence or NotionFrameworks, checklists, training materialKnowledge leadSince adoptionFirm-ownedFlag third-party licensed content
Project reviews and retrospectivesLessons learned, quality reviewsPractice leadsVariesFirm-owned; may name client staffOften the most useful rows for AI
Staffing and resource plansAssignments, skills, utilization forecastsResource managerSince adoptionFirm-owned; employee personal dataLeave compensation detail out of any reuse

Column definitions and example entries#

Column definitions keep entries comparable across practices. The columns below extend the starting six for a fuller version; add them once the first pass is complete.

Keep values short and consistent. Drop-down lists for client content, personal data and rights status make the sheet sortable, which matters once the inventory holds more rows than fit on one screen.

Column definitions and example entries
ColumnWhat to enterExample entry
Record familyThe type of record, not the file formatTime entries with narratives
VolumeAn approximate count or storage size, as knownRough count from an admin report
Client contentNone, some or mostlySome
Personal dataThe kinds presentEmployee names, client contact details
Export routeNative export, API, admin report or vendor requestAPI, because reports truncate narratives
RestrictionsKnown contract, vendor or policy limitsSome client MSAs prohibit reuse of working files
ProvenanceHow records were created and changedEntered by consultants, approved by managers
Permitted useWhat the firm may do with the records todayInternal use; external use unreviewed

How to fill it in without moving any files#

Filling in the inventory takes interviews and admin consoles, not exports. Keep client files closed throughout; the point is to describe sources, not to read them.

  • Start from the vendor list and card statements to find every system the firm pays for.
  • Ask each system owner short questions and record answers as stated, marking guesses as unknown.
  • Read date ranges and approximate counts from admin consoles and standard reports.
  • Flag every source that mixes client deliverables with internal records.
  • Ask finance which systems were retired and where their exports went.
  • Review the draft with a partner who knows the firm's client contract terms.

Rights status: the column that decides what you can do#

Rights status decides whether a row can feed an internal assistant, appear in a product, or be considered for licensing. Use four values: firm-owned, mixed, client-owned and unknown. Most consulting MSAs let the firm keep its methods and general know-how while assigning deliverables to the client and restricting client confidential information to the engagement.

That pattern makes playbooks, staffing plans and internal reviews likely firm-owned, and final reports likely client-owned. Drives and email are usually mixed. Unknown is an acceptable answer on a first pass; a wrong confident answer is not. This is general information, not legal advice, so confirm contested rows with counsel.

For column design, an external reference helps. The Data and Trust Alliance's Data Provenance Standards group dataset metadata into three families, Source, Provenance and Use, and say this metadata is needed to enable proper dataset selection for AI model training. Their Use group includes elements such as confidentiality classification, consent documentation location, license to use and intended data use, which line up with the rights status, restrictions and permitted use columns above.

Illustrative: an operations consultancy builds its first inventory#

Illustrative: a fictional strategy and operations firm built its first inventory ahead of a PSA change. The COO found project reviews in three places: drive folders, Confluence pages and email attachments. The proposal library sat in SharePoint, while win and loss outcomes lived in HubSpot with no shared ID.

The firm consolidated reviews into Confluence, added the HubSpot opportunity ID to each proposal file, and marked every drive folder as client or internal. The exercise also surfaced a retired PSA whose exports sat on a former finance manager's drive; the firm moved them into a controlled archive and gave them their own row.

The finished inventory guided the choice of an internal assistant, limited to firm-owned rows, and gave the firm the metadata it needed for a licensing fit check without exporting a single file.

How SourceX uses an inventory like this#

SourceX starts every assessment with metadata of exactly this kind: systems, years of history, record families and known restrictions. The SourceX Enterprise Data Value Framework uses it to judge which rows might interest AI developers, and the rights column feeds the Rights step of the SourceX five-step transaction.

If a license proceeds, the rows that were included become the core of a SourceX Evidence Packet, which records provenance, licensing rights, permitted use, the privacy record and release authorization. A firm that keeps its inventory current has most of that record before the conversation starts.

Frequently asked questions

How detailed should the first version be?

A first version at the level of system and record family is enough, typically one row per source with the six starting columns. Add volume, export route and provenance once the rows are agreed. Going field by field on the first pass slows the work without changing any decision.

Who should maintain the inventory?

The COO or an operations partner should own it, with each system owner responsible for their rows. Update it whenever a system is added, retired or migrated, and review it with the partners on a regular schedule so rights and restrictions stay current.

Should former employees' mailboxes and drives be included?

Yes, as their own rows. Firms often keep leavers' mailboxes and drives for continuity or legal reasons, and those archives can hold years of client correspondence and personal data. Record who controls them, the retention rule that applies and whether any legal hold covers them.

Is a data inventory the same as a retention schedule?

No. A retention schedule says how long each record type is kept and why. An inventory says what exists, where it lives, who owns it and what the firm may do with it. Each should reference the other, and gaps between them usually reveal records kept longer than policy allows.

Can we share the inventory with an acquirer or partner?

Usually yes, under a confidentiality agreement, because it holds descriptions rather than records. Remove client names from the notes column if they appear, and share the version that matches what the firm can actually support with evidence. Date the version you share, so later questions refer to the same document.

Sources

  • The Data and Trust Alliance's Data Provenance Standards (version 1.0.0 specification) define dataset metadata in three groups, Source, Provenance and Use, and say this metadata is needed to enable proper dataset selection for AI model training. Source
  • The Use group of the Data Provenance Standards includes elements for confidentiality classification, consent documentation location, license to use and intended data use. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify