Consulting and recruiting
Consulting firm data inventory template
By SourceX Editorial · Updated
Short answer
A consulting firm data inventory lists every system that holds firm records, with typical records, an owner, a date range and a rights status for each, then adds client-content, personal-data and export columns. Start with the CRM, PSA, shared drives, email, Slack or Teams, the proposal library and playbooks. Record metadata only, so it can circulate without exposing client work.
Key takeaways
- One row per system and record family is enough for a first version; detail comes later.
- The rights status column decides what the firm can do with each source, so mark unknowns honestly.
- Shared drives and email usually mix client deliverables with internal records and need splitting.
- Project reviews, proposals with outcomes and staffing plans are often the most useful firm-owned rows.
- An inventory collects descriptions, not files; nothing is exported to build it.
What a consulting firm data inventory is for#
A consulting firm data inventory is a register of where the firm's records live, who is accountable for them and what the firm is allowed to do with them. It supports decisions that come up sooner than most partners expect: retiring a PSA, choosing an AI assistant, answering a client security questionnaire, preparing for an acquirer's diligence or reviewing whether records could be licensed.
It is not a records retention schedule and not a data map of every field. It sits one level up, describing each source in a sentence or two so that a managing partner can see the whole estate on one page.
Firms that skip it tend to discover their records in a hurry, during a system shutdown or a diligence request, when the people who know where things are may already be busy or gone.
The template: starting rows#
The starting rows below cover the systems almost every consulting firm runs. Copy them, replace the example system names with yours, and add rows for anything firm-specific, such as a survey platform or a client portal.
Keep all six columns even when some cells stay blank at first. A blank owner or rights cell is itself a finding, and it is easier to spot in a consistent layout.
| Source | Typical records | Owner | Date range | Rights status | Notes |
|---|---|---|---|---|---|
| CRM such as Salesforce or HubSpot | Accounts, contacts, opportunities, win and loss reasons | Head of business development | From CRM go-live | Firm-owned; contains client contact details | Check opportunity notes for client confidential detail |
| PSA such as Kantata, Certinia or BigTime | Projects, tasks, time entries with narratives, budgets, change orders | COO or finance | Since adoption, plus exports from older systems | Firm-owned records about client work | Confirm narrative fields export in full |
| Shared drives such as SharePoint, Google Drive or Box | Working files, research, deliverables | Practice leads | Varies by folder | Mixed; deliverables are often client-owned | Separate client folders from internal ones |
| Email in Microsoft 365 or Google Workspace | Client and internal correspondence | IT | Per retention policy | Mixed, with heavy personal data | Usually last on any reuse list |
| Slack or Teams | Internal discussion and project channels | IT and practice leads | Per retention settings | Firm-owned internal discussion | Treat channels shared with clients separately |
| Proposal library | Proposals, SOWs, pricing rationale, outcomes | Business development | Since the library began | Firm-owned; client RFP material may be confidential | Link each proposal to its CRM outcome |
| Playbooks and methods in Confluence or Notion | Frameworks, checklists, training material | Knowledge lead | Since adoption | Firm-owned | Flag third-party licensed content |
| Project reviews and retrospectives | Lessons learned, quality reviews | Practice leads | Varies | Firm-owned; may name client staff | Often the most useful rows for AI |
| Staffing and resource plans | Assignments, skills, utilization forecasts | Resource manager | Since adoption | Firm-owned; employee personal data | Leave compensation detail out of any reuse |
Column definitions and example entries#
Column definitions keep entries comparable across practices. The columns below extend the starting six for a fuller version; add them once the first pass is complete.
Keep values short and consistent. Drop-down lists for client content, personal data and rights status make the sheet sortable, which matters once the inventory holds more rows than fit on one screen.
| Column | What to enter | Example entry |
|---|---|---|
| Record family | The type of record, not the file format | Time entries with narratives |
| Volume | An approximate count or storage size, as known | Rough count from an admin report |
| Client content | None, some or mostly | Some |
| Personal data | The kinds present | Employee names, client contact details |
| Export route | Native export, API, admin report or vendor request | API, because reports truncate narratives |
| Restrictions | Known contract, vendor or policy limits | Some client MSAs prohibit reuse of working files |
| Provenance | How records were created and changed | Entered by consultants, approved by managers |
| Permitted use | What the firm may do with the records today | Internal use; external use unreviewed |
How to fill it in without moving any files#
Filling in the inventory takes interviews and admin consoles, not exports. Keep client files closed throughout; the point is to describe sources, not to read them.
- Start from the vendor list and card statements to find every system the firm pays for.
- Ask each system owner short questions and record answers as stated, marking guesses as unknown.
- Read date ranges and approximate counts from admin consoles and standard reports.
- Flag every source that mixes client deliverables with internal records.
- Ask finance which systems were retired and where their exports went.
- Review the draft with a partner who knows the firm's client contract terms.
Rights status: the column that decides what you can do#
Rights status decides whether a row can feed an internal assistant, appear in a product, or be considered for licensing. Use four values: firm-owned, mixed, client-owned and unknown. Most consulting MSAs let the firm keep its methods and general know-how while assigning deliverables to the client and restricting client confidential information to the engagement.
That pattern makes playbooks, staffing plans and internal reviews likely firm-owned, and final reports likely client-owned. Drives and email are usually mixed. Unknown is an acceptable answer on a first pass; a wrong confident answer is not. This is general information, not legal advice, so confirm contested rows with counsel.
For column design, an external reference helps. The Data and Trust Alliance's Data Provenance Standards group dataset metadata into three families, Source, Provenance and Use, and say this metadata is needed to enable proper dataset selection for AI model training. Their Use group includes elements such as confidentiality classification, consent documentation location, license to use and intended data use, which line up with the rights status, restrictions and permitted use columns above.
Illustrative: an operations consultancy builds its first inventory#
Illustrative: a fictional strategy and operations firm built its first inventory ahead of a PSA change. The COO found project reviews in three places: drive folders, Confluence pages and email attachments. The proposal library sat in SharePoint, while win and loss outcomes lived in HubSpot with no shared ID.
The firm consolidated reviews into Confluence, added the HubSpot opportunity ID to each proposal file, and marked every drive folder as client or internal. The exercise also surfaced a retired PSA whose exports sat on a former finance manager's drive; the firm moved them into a controlled archive and gave them their own row.
The finished inventory guided the choice of an internal assistant, limited to firm-owned rows, and gave the firm the metadata it needed for a licensing fit check without exporting a single file.
How SourceX uses an inventory like this#
SourceX starts every assessment with metadata of exactly this kind: systems, years of history, record families and known restrictions. The SourceX Enterprise Data Value Framework uses it to judge which rows might interest AI developers, and the rights column feeds the Rights step of the SourceX five-step transaction.
If a license proceeds, the rows that were included become the core of a SourceX Evidence Packet, which records provenance, licensing rights, permitted use, the privacy record and release authorization. A firm that keeps its inventory current has most of that record before the conversation starts.
Frequently asked questions
How detailed should the first version be?
A first version at the level of system and record family is enough, typically one row per source with the six starting columns. Add volume, export route and provenance once the rows are agreed. Going field by field on the first pass slows the work without changing any decision.
Who should maintain the inventory?
The COO or an operations partner should own it, with each system owner responsible for their rows. Update it whenever a system is added, retired or migrated, and review it with the partners on a regular schedule so rights and restrictions stay current.
Should former employees' mailboxes and drives be included?
Yes, as their own rows. Firms often keep leavers' mailboxes and drives for continuity or legal reasons, and those archives can hold years of client correspondence and personal data. Record who controls them, the retention rule that applies and whether any legal hold covers them.
Is a data inventory the same as a retention schedule?
No. A retention schedule says how long each record type is kept and why. An inventory says what exists, where it lives, who owns it and what the firm may do with it. Each should reference the other, and gaps between them usually reveal records kept longer than policy allows.
Can we share the inventory with an acquirer or partner?
Usually yes, under a confidentiality agreement, because it holds descriptions rather than records. Remove client names from the notes column if they appear, and share the version that matches what the firm can actually support with evidence. Date the version you share, so later questions refer to the same document.
Sources
- The Data and Trust Alliance's Data Provenance Standards (version 1.0.0 specification) define dataset metadata in three groups, Source, Provenance and Use, and say this metadata is needed to enable proper dataset selection for AI model training. Source
- The Use group of the Data Provenance Standards includes elements for confidentiality classification, consent documentation location, license to use and intended data use. Source
Related resources
- QuestionCan CRM data be licensed?
- QuestionDo AI labs buy CRM data?
- InsightHow do I de-identify sales call recordings for AI training?
- InsightPurpose limitation: can records collected for one purpose be licensed for AI?
- InsightWhich state privacy laws exempt B2B contact data?
- SolutionData partnerships between businesses and AI developers
See if your company qualifies
A short company assessment. No data uploads are needed.