Skip to content

Manufacturing

Data quality problems in manufacturing ERP: the usual suspects

By SourceX Editorial · Updated

Short answer

The usual ERP data quality issues in manufacturing are orphaned orders and jobs, stale routings and BOMs, reused part numbers, catch-all reason codes, duplicate customers and history cut off at migrations. For AI use, broken links matter more than typos: a buyer can live with messy fields, but not a job that no longer connects to its order or NCR.

Key takeaways

  • Most ERP data problems fall into two kinds: wrong values, which can be documented, and broken links, which usually cannot be repaired later.
  • Overwritten dates and backflushed actuals hide the variance and judgment that make manufacturing history useful.
  • Data quality is the condition of records; quality-control data, such as inspections, CMM reports and defect images, is a type of record.
  • A licensing buyer needs long, linked, documented history and can accept known flaws that would stop an internal AI rollout.
  • Document problems in a known-issues note before cleaning anything, and keep crosswalks whenever records are merged.

What counts as an ERP data quality problem in a plant?#

An ERP data quality problem in a plant is any gap between what the records say and what actually happened on the floor, in purchasing or with the customer. Some gaps are wrong values, such as a routing that still lists a machine the plant sold. Others are missing links, such as a job that no longer points to the sales order it was built for.

The mess is rarely spread evenly. Master data such as parts, routings and bills of materials decays slowly as the plant changes around it. Transaction history decays in jumps, usually at a migration, an acquisition or a change in how a module was used.

The two kinds call for different responses. Wrong values can often be documented and lived with. Broken links usually cannot be rebuilt after the fact, so they set the limits of what any analysis or AI project can do with the history.

This is not only a licensing concern. In RSM's 2026 middle-market AI survey, data quality and availability issues were the most cited inhibitor to AI deployment, named by 34% of respondents, ahead of security and privacy concerns and legacy systems integration. The same ERP problems that slow an internal AI project also shape what a plant can license.

The ten usual suspects#

The ten problems below turn up again and again in mid-sized manufacturing ERPs, whatever the brand. Each one has a different effect depending on whether the data is used internally or licensed to an AI developer, so the last column focuses on AI use.

The ten usual suspects
ProblemWhat it looks likeEffect on AI use
Orphaned orders and jobsJobs with no sales order, shipments with no invoice, NCRs with no job numberBreaks the chain from request to outcome, the part buyers value most
Stale routingsStandard setup and run times never updated from actual laborPlans look precise but mislead any model that treats standards as reality
BOMs that differ from what was builtFloor substitutions never recorded; phantom assemblies left in placeMaterial history disagrees with the bill, weakening cost and quality links
Reused part numbersA retired part number assigned to a new, unrelated itemMerges two products' histories and corrupts failure patterns
Duplicate customers and vendorsOne customer entered under several names after acquisitions or typosSplits one relationship's history and inflates counts
Catch-all reason codesScrap, downtime or return reasons coded as other or miscellaneousRemoves the label a model would learn from; notes become the only clue
Backflushed labor and materialConsumption posted at standard when the job closesActuals mirror the plan, hiding the variance that carries signal
Overwritten datesPromise and due dates replaced on each change with no historyLoses every reschedule, the record of planner judgment
Migration cut-offsClosed jobs, old quotes or NCRs left behind at go-liveHistory looks short, and older years sit in a system nobody opens
Work kept outside the ERPExpedite lists, schedule boards and inspection logs in spreadsheetsKey decisions are missing from the system of record unless the files are kept

Which problems a buyer can live with, and which end the conversation#

Broken links and lost history matter most for AI use; cosmetic errors matter least. An AI developer licensing manufacturing records wants to see how a request turned into a decision and an outcome. Typos, odd descriptions and inconsistent capitalization rarely change that story.

Problems in the middle tier are survivable if they are written down. Catch-all codes, backflushing and stale standards each distort part of the picture, and a buyer can work around the distortion once it knows where it sits. Undisclosed distortion is worse than known distortion.

  • Usually disqualifying for the affected years: orphaned records across the main chain, reused part numbers with no way to separate histories, and history that survives only as printed reports.
  • Workable with a note: catch-all codes, backflushed actuals, stale standards and duplicate customers that can be mapped to one ID.
  • Mostly cosmetic: description typos, inconsistent abbreviations and unit labels that are at least consistent within each part.

Data quality is not the same as quality-control data#

Data quality describes the condition of records, while quality-control data is a type of record, and the two get confused in planning meetings. Quality-control data includes inspection results, NCRs, CAPAs, CMM reports, SPC charts and the images saved at inspection stations.

Much of that material lives outside the ERP, in a QMS, in CMM software, in SPC tools or in folders written by a vision inspection camera. It can be among the most distinctive records a plant holds. Model builders working on visual inspection need labeled examples of defects, and a defect library with disposition decisions attached is exactly that.

The caution is ownership. Many inspection images and CMM reports show a customer's part made to a customer's print, so they need a rights review before anyone treats them as licensable.

What AI-ready means for a licensing buyer versus internal use#

AI-ready means something different for a licensing buyer than for an internal project. An internal project needs current, correct data because its output drives decisions in your plant. A licensing buyer needs long, linked, well-documented history with outcomes, and it can accept known flaws that would stop an internal rollout.

What AI-ready means for a licensing buyer versus internal use
QuestionInternal AI projectLicensing buyer
How much history?Enough recent data to reflect today's products and processesSeveral years, because variety and outcomes matter as much as recency
How clean?Correct values, since output drives daily decisionsConsistent and documented; known errors described, not hidden
Real names?Yes, users need real customers and partsNo, identities replaced with consistent tokens
What documentation?Team knowledge is often enoughA data dictionary, field notes and a known-issues list
What rights?Internal use is usually covered by existing termsA right to license each record family, with customer carve-outs

How to check your ERP without starting a cleanup project#

The fastest check of ERP data quality is a short set of counts by year, run by whoever administers the system. The goal is to find where the chain breaks, not to fix it, because rewriting closed history to make it tidy can destroy the signal an analyst or a buyer would use.

Write the results into a one-page known-issues note. When duplicates are merged later, keep a crosswalk from old IDs to new ones so older transactions still connect.

  • Count jobs without a linked sales order or stock reason, by year.
  • Count NCRs and returns that carry no job, lot or serial number.
  • Compare standard and actual labor on a sample of routings at the busiest work centers.
  • List part numbers whose description changed completely at some point, a sign of reuse.
  • Show reason-code distributions for scrap and downtime, and how often the catch-all code appears.
  • Record each migration and go-live date, and what history came across.
  • Ask planners and quality staff which spreadsheets they keep beside the ERP.

Illustrative: a valve manufacturer finds where its chain breaks#

Illustrative: a fictional maker of industrial ball valves runs a mid-market ERP adopted after it outgrew an older system, plus a separate QMS for NCRs and CAPAs. Leadership assumes the history is sound because month-end closes on time.

The counts tell a different story. Jobs from before the go-live have no sales-order links, because only open orders were converted. A block of retired valve-body part numbers was reassigned to a new product line. Scrap on two cells is coded almost entirely as other, though operators wrote useful notes in the comment field.

The COO decides not to clean anything yet. The team writes a known-issues note, builds a crosswalk for the reused part numbers using creation dates, and limits any outside review to the years after the go-live plus the QMS history, which links to lots throughout.

How SourceX looks at ERP data quality#

SourceX treats data cleanliness as one driver in the SourceX Enterprise Data Value Framework, weighed alongside uniqueness, domain expertise, human-generated signal, scale, recency, rights and AI utility. Messy but linked history can outrank tidy history that lost its outcomes.

The fit check collects metadata only: systems, years, record families and the known-issues note. If a package moves forward, the Preparation step of the SourceX five-step transaction handles consistent masking and documentation, and the manufacturer approves what is released.

Frequently asked questions

Should we clean ERP data before or after a migration?

Clean the master data the new system depends on, such as active parts, routings and customers, before cutover. Leave closed transaction history as it is, documented, and keep a full copy of the old database. Cleanup that rewrites closed history often breaks links the migration needs and erases evidence of how the business actually ran.

Does messy ERP data rule out licensing?

Not usually. Buyers expect operational records to be imperfect. What hurts is missing links between records, missing outcomes and problems nobody disclosed. A clear known-issues note and consistent identifiers often matter more than how tidy individual fields look.

Can AI tools clean our ERP data for us?

They can help find duplicates, suggest reason codes from free text and flag outliers, but they should not silently rewrite history. Treat each suggestion as a proposal that a planner or quality lead confirms, and keep original values alongside corrections so every change stays traceable.

Who should own the known-issues list?

The ERP administrator usually keeps it, with input from the controller on cost data, the quality manager on NCR and inspection records, and production planning on routings and schedules. A single owner stops the list from scattering across email threads and meeting notes.

Do spreadsheets kept outside the ERP count as records?

Yes, when they hold decisions the ERP does not, such as expedite lists, schedule boards or inspection logs. Inventory them with the system they supplement, note who keeps them and how far back they go, and check them for customer names before anyone shares them.

Sources

  • In RSM's 2026 middle-market survey, data quality and availability issues were the top inhibitor to AI deployment (34%), followed by security and privacy concerns (30%) and legacy systems integration (28%). Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify