Software companies
Knowledge base cleanup before AI: fix, archive or delete old articles?
By SourceX Editorial · Updated
Short answer
Knowledge base cleanup for AI means deciding, article by article, whether to fix, archive or delete before an AI support agent starts quoting your help center. The safest default is archive, not delete: remove outdated articles from what the bot can retrieve, but keep them and their revision history, which records how your product and answers changed.
Key takeaways
- An AI support agent repeats whatever it retrieves, so stale and conflicting articles do more harm than they did with human agents.
- Audit first: owner, last update, product version, visibility and how often agents attach each article to tickets.
- Archive outdated articles out of AI retrieval rather than deleting them, and keep their revision history.
- Delete only empty stubs, test pages and content that should never have existed, after checking retention rules.
Why an AI support agent makes old articles a bigger problem#
An AI support agent makes outdated knowledge base articles more dangerous because it retrieves and repeats them with the same confidence as current ones. A human agent who sees an old screenshot knows to double-check; a retrieval system has no such instinct unless you give it one through metadata.
Three problems usually show up first. Conflicting articles that describe the same task two ways produce inconsistent answers. Articles about retired plans or features send customers down dead ends. Internal notes published by mistake leak workarounds or details that were never meant for customers.
Start with a KB audit, not a rewrite#
A KB audit is a spreadsheet of every article with enough facts about each to make a decision, and it comes before any rewriting. Most help desks let you export article lists, either directly or through an API; check your vendor's documentation for what the export includes, especially revision history.
The ticket link column is the most telling. An article that agents attach often is load-bearing even if its public view count is modest, and it deserves a fix before anything else on the list.
- Article ID, title, URL and category.
- Visibility: public, signed-in customers only, or internal agents only.
- Owner and last updated date.
- Product area and the feature or version the article describes.
- Views, and how often agents link the article in tickets.
- Known duplicates or overlapping articles.
- Whether revision history exists and can be exported.
The decision table: fix, archive or delete#
Each article's state points to one of three decisions, and the table below covers the common cases. When in doubt between archive and delete, archive: an archived article can be deleted later, but a deleted one cannot be brought back.
Record the decision and the reason in the audit sheet itself. That log becomes the explanation, months later, for why a customer can no longer find an article they remember.
| Article state | Decision | Reason |
|---|---|---|
| Accurate and frequently used | Keep, with a light fix | Load-bearing content needs only current screenshots and terms |
| Accurate but rarely viewed or linked | Keep, or merge into a broader article | Low views alone are not a reason to delete; check internal use first |
| Correct steps, outdated UI names or images | Fix | Small edits restore accuracy |
| Feature has changed | Fix the current version; archive the old one | Customers on older versions may still need it |
| Retired product or plan | Archive from public; keep internally | Support still meets legacy accounts |
| Duplicate of a better article | Redirect and archive | One answer per task reduces conflicts |
| Wrong or unsafe instructions | Unpublish now; archive with a note | Removes the risk but keeps the record of what was published |
| Contains personal data or credentials | Redact, or delete after a retention check | Content that should not exist is the main case for deletion; rotate any exposed credential |
| Empty stub or test page | Delete | No history worth keeping |
Why archive old versions instead of deleting them#
Archiving old knowledge base versions keeps a record that deletion destroys for good: how your answers changed as the product changed. Revision history shows which steps were rewritten after which release, and the tickets that cited each version show whether a rewrite reduced repeat contacts.
That history has practical uses. Support still meets customers on older versions, disputes sometimes turn on what your documentation said at a given time, and AI developers value records of knowledge maintenance, the work of keeping answers accurate over time. In some help desks, deleting an article also deletes its revision history, so check how yours behaves before you delete anything.
Archiving does not override retention policy. If your policy calls for deleting certain content after a set period, or a legal hold requires keeping it, those rules still apply.
Keep AI retrieval separate from the archive#
AI retrieval and the archive should be two different sets, because what the bot may quote is narrower than what the company should keep. Configure the AI agent to draw only from current, published sources, and label articles by audience and product version so retired material stays out.
Agent-assist tools that suggest replies to your own staff can reasonably see more than a customer-facing bot. Draw that line explicitly in the configuration rather than relying on article visibility settings alone.
| Content source | Include in AI retrieval? | Keep in archive? |
|---|---|---|
| Current public articles | Yes | Yes, with revision history |
| Internal agent-only articles | Only for agent-assist tools, not customer-facing bots | Yes |
| Macros and saved replies | Rarely; they lack context | Yes |
| Archived versions | No | Yes |
| Community forum posts | With care; accuracy varies | Depends on the forum terms |
| Release notes | Yes, for version questions | Yes |
Cleanup mistakes that cause trouble later#
Cleanup mistakes tend to surface months after the project ends, when a customer, an auditor or a buyer asks a question the help center can no longer answer. Most come from treating the work as a one-time purge instead of a documented change with a reason recorded for each article.
- Bulk-deleting every article with low recent views, including load-bearing internal ones.
- Overwriting articles in place for a new product version, so customers on the old version lose their instructions.
- Changing URLs without redirects, which breaks links inside old tickets and macros.
- Letting the AI agent index internal notes because their visibility flag was never set.
- Migrating to a new help desk with current articles only and leaving revisions behind.
Illustrative: an HR software company prepares its help center#
Illustrative: a fictional HR software company is about to switch on an AI agent in its help center. The help center covers two product generations, internal runbooks live in Confluence, and agents rely on a large library of macros.
The support operations lead exports every article with views, ticket link counts and revision history. The team fixes the most-linked articles first, moves first-generation articles into an internal legacy category, merges duplicates with redirects and deletes test pages. The AI agent is pointed only at current public articles and release notes.
When the company later moves to a new help desk, it exports the full archive, revision history included, before the old account closes. Later still, when leadership asks whether its support records could be licensed, the archive and its ticket links are intact.
How SourceX looks at knowledge base history#
SourceX looks at knowledge base history as part of a support package, not as a standalone set of articles. In the SourceX Enterprise Data Value Framework, domain expertise and human-generated signal raise value while preparation cost and privacy burden reduce it, so articles with revision history and links to the tickets that used them, showing how answers were maintained over time, tend to say more than a snapshot of today's help center.
Nothing is shared during the initial fit check, only metadata such as the help desk in use, the period covered and whether revisions exist. In Preparation, customer names, internal URLs and screenshots showing account details are removed, and the SourceX Evidence Packet records which versions were included.
Frequently asked questions
Should we let AI rewrite our old articles?
AI can draft rewrites, but a person who knows the product should review each one before it is published. Mark which articles were AI-drafted and when. That label helps your own team, and it matters later if the content is ever licensed, since buyers often want to know which text people wrote.
How often should we re-audit the knowledge base?
Tie the audit to your release cycle rather than a calendar. Any release that changes a workflow should trigger a check of the articles that describe it. A lighter full review on a regular schedule catches drift between releases.
What about community forum posts?
Forum posts are written by customers and sometimes staff, under the community's own terms. They can help an AI agent with edge cases, but accuracy varies and posts often contain personal details. Review the forum terms before reusing posts for AI or keeping them in an archive you might license.
Does deleting old articles reduce legal risk?
Not automatically. Deletion under a documented retention policy is defensible; deleting content ad hoc, especially during a dispute, can create problems. If an article was wrong, archiving it with a note of when it was corrected is often the better record.
Can archived articles be part of a data license?
They can, if the company owns the content and preparation removes personal and confidential details. Archived versions are often more useful than current ones in a license because they show change over time. Third-party content, such as copied vendor documentation, is usually excluded.
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.