Why a company's first AI initiative almost always fails on the data — not the model.
AI × ERP series · Part 1
This five-part series shows how ERP systems and AI grow together — from the data silo to the learning enterprise. Part 1 sets out the opening question across industries: why aren't ERP data alone enough?
The key answer in 60 seconds
A complete ERP delivers cleanly structured master and transaction data. But for robust industrial AI those data are too scarce, too uniform and too company-specific. Unlike a language model that learns from half the internet, industrial AI has no public ocean of data to draw on — the relevant data arise solely inside your own operation, and there they are limited.
That is why no AI initiative should begin with tool selection, but with a sober question: does our data foundation actually carry the load? Answer it honestly and you spare yourself the most expensive misinvestment. 33+ years of consulting experience, 500+ projects, 100% vendor-independent.
We hear this question in almost every first conversation about artificial intelligence. Managing directors have invested for years in an ERP system that maps orders, inventory, master data and financial postings cleanly. The natural expectation: if all the data already sit in the system, surely AI can be built from them. In practice, this is exactly where the first AI initiative breaks down — not because the model is too weak, but because the data foundation cannot bear the weight.
According to the Bitkom AI study 2026, Künstliche Intelligenz in Deutschland study report (n=604 companies), 36 percent of German companies now use AI actively — nearly double the 20 percent of a year earlier — with a further 47 percent examining a possible deployment.
The common error of thinking stems from experience with tools like ChatGPT. Such language models seem all-knowing because they were trained on vast, freely available volumes of text. Industrial AI works fundamentally differently: it is meant to make predictions about a specific machine, a specific supply process or a specific cost structure. For that there is no public ocean of data — the relevant data arise solely inside your own operation, and there they are naturally limited.
This scarcity is the core problem. A mid-sized manufacturer may have three years of machine data from a handful of plants — statistically too little and too uniform to train a dependable model. A scientific meta-review of data problems in industrial AI systems ranks data quality, missing context and inadequate validation as the dominant hurdles — not the choice of algorithm.
The ERP plays an important but bounded role in this logic. It supplies structured master and transaction data — items, orders, postings — at high consistency. What it does not supply are the sensor data, process contexts and quality assessments that industrial prediction models actually learn from. In our experience: the ERP is the bookkeeping of the business, not its sensory system. Anyone who wants to build AI on top of it must name that gap first.
When your own data are too scarce, there are three fundamental levers. Each is explored in depth in the following parts of this series.
Way 1
Many operations use only a fraction of the data they already generate. Sensor data, machine logs and inspection reports lie scattered alongside the ERP. Bringing them together enlarges the usable base before external sources become necessary.
Way 2
More data do not help if they are contradictory, incomplete or inconsistent. Master-data quality and data governance are the invisible foundation of any AI capability — the focus of Part 2, using wholesale and distribution as the example.
Way 3
When one company's data are never enough, the way out lies in shared use — without surrendering raw data or trade secrets. Data spaces are the framework for that. Parts 3 and 4 develop this.
The third way is the most demanding — and the one with the greatest leverage. The Fraunhofer Institute for Software and Systems Engineering ISST and the International Data Spaces Association describe data spaces as a framework for sovereign cooperation; the EU has designated manufacturing as one of the fourteen Common European Data Spaces.
Model drift describes the effect by which a once-trained model loses accuracy over time, because the underlying reality shifts — new materials, changed load profiles, different suppliers. In large, varied data sets, drift can be detected and corrected through retraining. In the small, homogeneous data volumes of a single mid-sized company, it often only surfaces once a wrong prediction has already done damage.
The example most often cited in the professional debate is predictive maintenance: a model that learns from the data of a few machines at one site generalises poorly to other operating conditions — and predicts failures least reliably exactly where they are most expensive. Boris Otto and colleagues at Fraunhofer ISST argue for precisely this reason that industrial AI depends on highly specific data that no single company holds in sufficient quantity on its own. The consequence is not a better model, but a broader, better-maintained and — where it makes sense — shared data foundation.
What we see across our projects
The pattern repeats across industries. A company arrives convinced the AI question is which model or which vendor. A structured data inventory — beyond the ERP, into machine data, quality records, the context that actually drives predictions — surfaces the uncomfortable finding: the usable data is thinner, more uniform and more scattered than anyone assumed. The initiative was about to be built on a foundation that could not carry it.
The projects that succeed are the ones that stop at that point and decide deliberately — enrich the internal data, raise its quality, or cooperate in an ecosystem — rather than pressing on into an expensive misinvestment. The inventory is not overhead before the real work; it is the cheapest insurance the programme will buy.
The parts that follow take each of those three paths in turn — starting, in Part 2, with the data-quality foundation that no model can repair on its own.
Before the first AI initiative comes no tool selection, but a sober stocktake: which data do we have, are they any good, and are they enough for our goal?
Our assessment: these three steps are not methodological overhead, but the cheapest insurance against an expensive misinvestment — and the entry point to the roadmap that Part 5 of this series brings together.
30 minutes of honest situational assessment — directly with Dr. Dreher
Does your data foundation support an AI initiative — or not? No sales pitch, no junior consultants.
If you are facing an AI initiative and are unsure whether your data foundation will carry it, begin with the question about the data — not about the model. You can find more on our approach to digitalisation, ERP and AI integration in our overview of consulting services.
Part 2 of the series: Data quality — the invisible foundation of AI-ready ERP systems (using wholesale and distribution as the example).
For the AI side of that assessment, see our AI-supported approach, SCOReX®.
|
|