Consulting and recruiting
Knowledge management at consulting firms: from shared drives to AI-ready records
By SourceX Editorial · Updated
Short answer
Knowledge management at consulting firms means capturing the firm's own methods, proposals and project lessons so people and AI tools can find and reuse them. Most firms start from client-named shared drives. The first rule: separate firm-owned knowledge from client material before indexing anything, then tag each record by subject, ownership and status.
Key takeaways
- Client-named folder structures make firm knowledge hard to find and easy to leak from one client to another.
- Sort firm methods, firm records about the work, client confidential material and client-owned deliverables before any AI tool indexes them.
- Ownership, confidentiality and status tags do more for AI retrieval than the choice of search product.
- Archive superseded versions rather than deleting them, unless a retention rule or client agreement requires deletion.
- A platform migration or a partner retirement is the natural moment to sort and tag, instead of copying every folder unchanged.
- An ownership-tagged knowledge base serves internal AI and doubles as the starting inventory for any licensing review.
What counts as knowledge at a consulting firm?#
Knowledge, for a consulting firm's KM program, is what the firm learns on engagements and is entitled to reuse: methodologies, proposal language, project lessons, benchmarks it is allowed to keep, and a map of who knows what. It is the firm's own know-how, not every client file the firm happens to store.
At many mid-size firms, KM amounts to a shared drive organized by client, a few partners who remember where the good material is, and a proposal library someone maintained for a while. That holds up while the firm is small. It breaks when new hires cannot find past work, when senior partners retire, or when an AI assistant is pointed at the drive and returns the wrong client's numbers.
The KM maturity table: where does your firm sit?#
The KM maturity table describes five stages, from scattered files to governed records, and what AI can realistically do at each one. Many firms sit between the second and third rows, often with one practice ahead of the rest.
The jump that matters most is from the second row to the third. Until firm-owned knowledge is separated from client material, every AI tool placed on top of the drive inherits the mix.
| Stage | What it looks like | What AI can do | Next move |
|---|---|---|---|
| Scattered | Material on personal drives and in email attachments | Very little beyond one person's own files | Move work into shared, firm-controlled storage |
| Client-folder drive | SharePoint, Google Drive or Box organized by client and engagement | Search that mixes client material with firm knowledge | Classify folders by ownership |
| Curated library | Methods, templates and proposal sections in Confluence, Notion or a SharePoint site with named owners | Reliable search and drafting over firm-owned content | Link items to CRM pursuits and PSA projects |
| Linked records | Library items tied to pursuit IDs, project codes, outcomes and reviews | Answers that cite similar engagements and how they ended | Add rights, confidentiality and retention tags |
| Governed | Classification, review cycles, retention and access rules enforced | Safe use across practices and for outside review | Maintain, audit and track reuse |
Why client-named folders defeat search and create risk#
Client-named folders defeat search because they answer whose work a file was, not what it was about. A consultant looking for prior pricing-strategy work for a distributor has to know which clients those were, open each engagement folder and guess which file was final.
The same structure creates confidentiality risk once AI tools arrive. Many enterprise search and assistant tools follow existing file permissions, so anything a user can open, the assistant can quote. If engagement folders are broadly shared across the firm, one client's data can surface in a draft for another. Restricting client folders to their engagement teams and building a separate, curated firm library addresses both problems at once.
Sort firm knowledge from client material first#
Sorting firm knowledge from client material is the first real KM task, and four ownership classes cover nearly everything. The engagement letter or MSA decides which class a document falls into, so treat the firm's standard terms as the default and flag engagements signed on a client's paper.
Borderline cases are common: a playbook refined during one client's engagement, or a benchmark built from several clients' data. Record who decided the classification and why, and send unclear items to whoever handles contracts. That note becomes useful evidence later.
| Ownership class | Typical contents | Default handling |
|---|---|---|
| Firm methods | Frameworks, playbooks, templates, training material | Curate, tag and reuse across the firm |
| Firm records about the work | Proposals, staffing plans, project reviews, close-out notes | Reuse after removing client identifiers |
| Client confidential | Data the client supplied, interview notes, internal client documents | Keep inside the engagement, out of any firm-wide index |
| Client-owned deliverables | Final reports and models assigned to the client under the MSA | Reference only as the contract permits; do not copy into the library |
What metadata makes a consulting record AI-ready?#
AI-ready metadata tells a retrieval system, and a human reviewer, what a record is, where it came from and how it may be used. A workable starting set fits on one row of a spreadsheet per item.
This mirrors how dataset documentation is developing outside consulting. The Data & Trust Alliance's Data Provenance Standards, an open specification, group dataset metadata into Source, Provenance and Use, and state that this metadata is needed to enable proper dataset selection for AI model training. Their Use group includes elements such as confidentiality classification, license to use and intended data use.
- Subject tags: service line, industry, problem type and engagement type.
- Source: originating system, folder path and engagement or pursuit code.
- Ownership class and confidentiality level, carried over from the sorting step.
- Status: current, superseded or draft, with the date it was last reviewed.
- Owner: the partner or practice lead who answers questions about the item.
- Restrictions: NDA coverage, return-or-destroy duties and the retention date.
A cleanup sequence that does not stall billable work#
A KM cleanup that protects billable time runs one practice at a time, with operations staff doing most of the sorting and partners deciding only the borderline cases. Trying to fix the whole drive at once is the usual reason these projects stall.
The other common mistakes are predictable: copying every folder unchanged to a new platform, letting IT decide what counts as firm knowledge, connecting an assistant to the whole tenant before permissions are fixed, and deleting old versions to make the library look tidy. Archived versions show how a pricing model or diagnostic changed across engagements, so keep them unless a retention schedule or client agreement says otherwise.
- Step 1: choose one practice with an engaged partner sponsor.
- Step 2: list the systems that hold its knowledge, such as SharePoint or Drive sites, Teams channels, Confluence or Notion spaces, the CRM and the PSA tool.
- Step 3: classify folders into the four ownership classes, starting with the most-used ones.
- Step 4: copy firm methods and de-identified project reviews into a curated library with named owners.
- Step 5: label superseded versions and move them to an archive instead of deleting them.
- Step 6: add pursuit IDs and project codes so library items link back to outcomes.
- Step 7: connect the AI assistant to the curated library only, then widen access as tagging matures.
Illustrative: a consultancy uses a drive migration to rebuild its library#
Illustrative: a fictional operations consultancy of about 110 people, serving utilities and transport operators, is retiring an old Box account and moving to SharePoint. Two founding partners plan to step back within a few years, and much of the firm's method lives in their folders and inboxes. IT proposes copying every client folder to the new tenant as it stands.
The managing partner stops the straight copy. The operations lead classifies the most-used folders into the four ownership classes before anything moves. Client confidential folders migrate as restricted engagement sites, and final deliverables are referenced from the engagement record rather than copied into the library. The founders set aside a few hours a month to confirm which of their playbooks are current, and each playbook gets a named successor owner.
When the curated library opens, every item carries subject tags, an ownership class, a status and its PSA project code, and archived versions stay retrievable. The firm connects its assistant to the library only. It also holds a record-family inventory, with dates and owners, that a later licensing review could start from without reopening the old drive.
From KM program to licensing candidate: the SourceX view#
A well-run KM program produces the same inventory a licensing review needs, which is why SourceX treats KM cleanup as a practical first step rather than a separate project. The ownership classes map directly onto the Rights step of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery.
Preparation still needs care. Client names, figures and individuals must come out of the firm's records about the work, and automated tools help without finishing the job: Presidio, an open-source de-identification SDK, warns in its own documentation that there is no guarantee it will find all sensitive information. Preparation therefore pairs automated detection with human review, the result is recorded in the SourceX Evidence Packet, and the firm approves every step. Any license grants defined use of a prepared copy; the records are licensed, not sold, and the firm keeps ownership.
Frequently asked questions
Do we need a new KM platform to become AI-ready?
Usually not at first. Many firms already run SharePoint, Google Drive, Confluence or Notion, and the gap is classification and tagging rather than software. A platform bought before sorting tends to relocate the same mix of client and firm material. Choose or change platforms after the first practice is sorted, when you know which tags and permissions you need, and name a KM or practice operations lead to own the routine.
Should we roll out an AI assistant before the cleanup is finished?
Only against a curated collection. Connecting an assistant to the whole file store before ownership and permissions are sorted is how client material ends up in the wrong draft. A small, clean library with clear owners gives better answers than a large, mixed drive, and it can grow practice by practice.
What happens to firm knowledge when consultants leave?
Their files often sit in personal OneDrive or Google Drive folders and mailboxes that are removed after departure under the firm's account policies. Add a leaver step that moves engagement material and useful working files into firm-controlled locations and captures a short handover note on active pursuits and methods.
Do return-or-destroy clauses affect the knowledge base?
They can. Some engagement letters and NDAs require the firm to return or destroy client confidential information when work ends. Firm methods and general know-how usually fall outside those clauses, but copies of client documents in the library may not, so check the terms before curating anything that came from a client.
Is a consulting knowledge base valuable to AI developers on its own?
Method pages alone are less useful than records of real work. AI developers tend to value material that links a problem to the approach taken, the review comments and the result. A knowledge base tied to pursuits, projects and outcomes is a stronger candidate than a standalone library of templates.
Sources
- The Data & Trust Alliance's Data Provenance Standards (version 1.0.0 specification) define dataset metadata in three groups: Source, Provenance and Use, and say this metadata is needed to enable proper dataset selection for AI model training. Source
- The Use group of the Data Provenance Standards includes elements for confidentiality classification, consent documentation location, privacy-enhancing technologies applied, processing and storage geographies, license to use, intended data use, and copyright, patent and trademark status. Source
- Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee that Presidio will find all sensitive information, and additional systems and protections should be employed. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.