AI uses for records
Information governance for AI: what mid-sized companies need
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Information governance for AI is the set of owners, rules and records that decide which company information may feed internal AI tools and which may be licensed outside. A mid-sized company needs six parts: an inventory, sensitivity tiers, a rights register, use rules, retention controls and a named release approver. Build the inventory first.
Key takeaways
- AI governance at a mid-sized company works best as a few named owners and short documents, not a new department.
- One inventory and one classification scheme can serve both internal AI tools and any external data license.
- Rights come from contracts, notices and vendor terms, so a rights register belongs next to the data inventory.
- No record should leave the company without a named approver and a written account of what was released.
- Retention schedules and legal holds should be checked before a system migration deletes old records.
What does information governance for AI cover?#
Information governance for AI covers how a company decides which of its records may be used by AI, by whom, for what purpose and with what protections. It points in two directions at once: inward, to staff using AI assistants and vendor AI features on company data, and outward, to any decision to license records to an AI developer.
Traditional information governance already handled retention, legal holds and access. AI adds new questions. Can a support transcript be pasted into an assistant? Does the CRM vendor train on your records? Can a decade of project files be licensed, and who signs? A framework answers these once, instead of case by case in a hallway.
Why mid-sized companies need a lighter version#
Mid-sized companies need a lighter version of AI governance because they rarely have a records manager, a privacy office and a data team to staff a heavy program. The general counsel, or outside counsel, often covers privacy, contracts and employment at once, and IT may be a small team managing dozens of SaaS tools adopted one department at a time.
The practical answer is a short set of owners and documents, roughly one page per component. Each component needs a named owner, a written rule and a place where decisions are recorded. Governance that the people already in the building cannot maintain tends to exist only on paper.
The six components of the framework#
The six components are inventory, classification, rights, use rules, retention and release. Each one answers a question about internal AI use and a related question about external licensing, which is why one framework can serve both.
Two components are often underestimated. Retention now reaches AI prompts, outputs and connector logs: where an AI tool stores them, they may be subject to legal holds and discovery like other business records, so counsel should know where they live and for how long. Release also covers more than licensing: connecting a vendor AI feature to a system can amount to a release in practice if the vendor's terms let it use the data.
| Component | Internal AI question | External licensing question | Usual owner |
|---|---|---|---|
| Inventory | Which systems and record families do staff and vendor tools touch? | Which record families exist, for how many years, and where? | COO or IT lead |
| Classification | Which tiers may go into which AI tools? | Which tiers are excluded or need preparation? | General counsel with IT |
| Rights | Do vendor terms let the vendor train on our records? | Do customer contracts, notices and NDAs allow licensing? | General counsel |
| Use rules | What may staff paste, upload or connect? | What permitted uses would a license allow? | General counsel with the CEO |
| Retention | How long are AI outputs and logs kept? | Are records preserved until they have been assessed? | Controller or records owner |
| Release | Who approves connecting a new AI tool to a system? | Who signs before any record leaves the company? | CEO or authorized signer |
How should records be classified for AI use?#
Records should be classified for AI use by sensitivity and by who controls them, because a record can be harmless in content and still belong to a client. Five tiers cover most mid-sized companies.
| Tier | Examples | Internal AI tools | External licensing |
|---|---|---|---|
| Internal operational | Support tickets, job notes, order exceptions, project schedules | Approved tools with access controls | Candidate after privacy preparation |
| Confidential business | Pricing, margins, contracts, board materials | Restricted to approved tools and roles | Usually excluded or narrowly limited |
| Personal | Customer contact details, employee names inside records | Minimized and governed by notices | Removed or de-identified before release |
| Restricted personal | HR files, health details, government identifiers, payment data | Generally kept out of AI tools | Excluded |
| Third-party controlled | Client deliverables, customer code, licensed content | As the controlling contract allows | Excluded unless the owner agrees |
What external licensing adds to the framework#
External licensing adds a documentation duty: a company that releases records needs a written account of where they came from, what rights apply and what was done to them. Careful buyers may ask for that account, and keeping it helps the company answer questions that come up later.
One published reference point is the Data & Trust Alliance's Data Provenance Standards. The specification sorts dataset metadata into three groups, Source, Provenance and Use, and states that this metadata is needed to choose datasets properly for training AI models. Within the Use group sit items a general counsel will recognize: a confidentiality classification, the location of consent documentation, any privacy-enhancing technologies applied, where the data may be processed and stored, the license to use, the intended use, and the status of copyrights, patents and trademarks.
Most of those elements come straight from a well-kept inventory, classification scheme and rights register, so a company that maintains them for internal AI has done much of the groundwork for a license. Counsel decides, deal by deal, which privacy, consumer protection or AI laws may apply; the framework here is general information rather than legal advice.
Illustrative: an engineering firm sets up governance before a CRM change#
Illustrative: a fictional mid-sized civil engineering firm planned to retire its old CRM and an archive server while staff were already using AI assistants to draft proposals. The general counsel, who also handled privacy part-time, was asked two questions at once: could the archive be deleted, and could proposals be pasted into assistants?
She started with an inventory: Deltek for projects and time, the CRM, the file server holding drawings and RFIs, and Microsoft 365 email. Each system was tagged with tiers. Drawings and calculations were marked third-party controlled where client contracts assigned ownership to the client, and proposals were marked confidential business.
The firm paused deletion of the archive until it had been assessed, approved one AI assistant for internal use with rules that kept client deliverables out, and named the managing principal as release approver. When leadership later asked whether internal RFI and project review records could be licensed, the inventory and rights register were already in place.
Where to start: a sequenced checklist#
Starting with the inventory is the right sequence because every other component depends on knowing which systems exist and what they hold. A small team can run these steps without buying new software.
- List every system that holds records, including SaaS tools adopted by a single team.
- Record for each system its record families, approximate years of history, owner and export route.
- Tag each system with the sensitivity tiers it contains.
- Pull the AI and data-use clauses from vendor terms, customer contracts and employee notices into a rights register.
- Write short use rules for internal AI tools and circulate them with worked examples.
- Check retention schedules and legal holds, and pause any deletion tied to a planned migration.
- Name the release approver and define the written record each release decision must produce.
How SourceX fits into a governance program#
SourceX fits at the release component, where a company decides whether to license records outside. Its fit check uses metadata only, so the inventory and classification work described here is what the first conversation draws on, and no files change hands at that stage.
For any package that proceeds, the SourceX five-step transaction runs Supply, Rights, Preparation, Approval and Delivery. The result is documented in a SourceX Evidence Packet covering provenance, licensing rights, permitted use, the privacy record and release authorization, and that packet can be filed with the company's own governance records. SourceX's dataset rights are set out in the signed supplier agreement, and records are licensed rather than sold, so the company keeps ownership.
Frequently asked questions
Do we need a separate AI governance policy?
Not always. Many mid-sized companies add an AI section to existing acceptable use, data classification and vendor management policies. A standalone policy helps when AI use is widespread or when the company plans to license records, because it gives staff, auditors and counterparties one document to read.
Who should own AI governance at a mid-sized company?
One executive should be accountable, often the general counsel, COO or CFO, with a named owner for each component. The inventory owner is usually closest to the systems, while rights and classification decisions sit with counsel. Release decisions belong with whoever can sign for the company.
How do vendor AI features fit in?
Vendor AI features are part of internal use. Review each vendor's terms for whether your records may be used to train its models, whether that can be switched off, and how outputs are stored. Record the answer in the rights register so the question is not re-asked by every team.
Does governance slow down AI adoption?
Light governance usually speeds it up. Staff stop guessing what is allowed, IT can approve tools against written rules, and leadership can say yes to a licensing conversation without first reconstructing what the company holds and what its contracts permit.
What should we do with old archives during a migration?
Treat them as unassessed records rather than clutter. Check retention schedules and legal holds first, then decide whether to preserve, export or delete. Old help desk, CRM and ERP archives often hold years of linked history that is hard to recreate once the subscription ends.
Sources
- The Data & Trust Alliance's Data Provenance Standards (version 1.0.0 specification) define dataset metadata in three groups: Source, Provenance and Use, and say this metadata is needed to enable proper dataset selection for AI model training. Source
- The Use group of the Data Provenance Standards includes elements for confidentiality classification, consent documentation location, privacy-enhancing technologies applied, processing and storage geographies, license to use, intended data use, and copyright, patent and trademark status. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.