EU Common Data Spaces Are Building the Data Layer Industrial AI Has Been Missing

· · Views: 2,402 · 6 min time to read

Europe’s industrial AI strategy increasingly rests on a resource that is harder to manufacture than GPUs: usable industrial data.

Factories already generate enormous volumes of sensor readings, machine logs, maintenance records, quality measurements and supply-chain information. Yet much of that information stays inside individual companies because competitors cannot simply pool commercially sensitive production data into a conventional cloud database.

The European Union’s Common European Data Spaces are attempting to solve that problem differently. Rather than creating one giant European database, the initiative combines shared infrastructure and governance so companies can exchange data while retaining control over how it is accessed and reused. The European Commission describes the spaces as providing a “secure and privacy-preserving infrastructure to pool, access, share, process, and use data” across sectors including manufacturing, energy, mobility and agriculture.

What began as a data-sharing policy is now moving closer to AI training infrastructure.

Europe Is Explicitly Connecting Data Spaces to AI Training

The direction became unusually clear in the EU’s latest manufacturing-data programme.

A 2026 Digital Europe Programme call says the manufacturing data space should “collect massive data from real industrial environments” to train or fine-tune generative AI models, while giving authorized AI developers access and coordinating the resulting datasets with Europe’s AI Factories.

The broader European Data Union Strategy makes the connection even more explicit. It identifies scaling access to high-quality data for AI as a priority and proposes Data Labs that link Common European Data Spaces with AI ecosystems so companies, SMEs and researchers can use sector-specific datasets for AI development.

Those Data Labs are intended to federate information from different sources, including the Common European Data Spaces, while providing data cleaning, enrichment, standardized formats, synthetic data and interoperability services connected to EuroHPC infrastructure.

That combination matters. Europe is effectively connecting three layers that AI developers typically have to assemble separately: industrial data, data-engineering services and high-performance compute.

Industrial AI Has a Data Fragmentation Problem

The scientific case for this infrastructure begins with a persistent weakness in industrial AI: individual factories often do not possess sufficiently diverse datasets.

An open-access study in Digital Communications and Networks found that smart-manufacturing AI suffers from “lack of high-quality data,” class imbalance and poor data diversity that can lead to inaccurate models. The researchers also noted that valuable industrial data remain fragmented across companies because of privacy, security and legal concerns.

This is especially problematic for tasks such as predictive maintenance.

A machine-learning model trained only on one factory’s equipment may see thousands of examples of normal operation but relatively few examples of rare failures. Another manufacturer may possess exactly those failure records, yet competitive and confidentiality concerns make conventional sharing difficult.

The same study proposed combining federated learning and data-space architecture so multiple organizations can contribute to collaborative predictive-maintenance models without requiring all sensitive information to be placed into one central repository. Its experiments demonstrated the potential of a multi-party architecture for sovereign data sharing and collaborative predictive diagnostics.

This is where European data spaces differ from ordinary datasets: the architecture can become part of the training system itself.

Manufacturing Data Spaces Are Already Moving Beyond Theory

A 2025 open-access survey examined 26 European manufacturing data-space applications and found that adoption was concentrated around real industrial use cases rather than generic data marketplaces. Predictive maintenance, machine monitoring and quality prediction appeared in more than 60% of the studied cases.

Machine data—particularly time-series information—was the most commonly reported type, while JSON was the most popular data format. The researchers found that machine data appeared in almost 70% of predictive-maintenance-related cases.

Those are precisely the kinds of continuous, operational datasets needed for industrial machine learning.

But the same research exposes how early the market remains. About 81% of the studied manufacturing data-space applications were still in the implementation stage, generally meaning pilots; just 11% were operational, and the researchers found no cases at the scaling stage within their sample.

For startups, that gap is important. Europe has considerable industrial data, but the infrastructure for making it reliably usable across organizations is still being built.

Interoperability May Matter More Than Data Volume

Pooling industrial data does not automatically create good training material.

Factories describe equipment differently. One company may represent temperature, tool wear or component identifiers through proprietary schemas while another uses different naming, units and metadata.

The EU-backed Data Spaces Support Centre therefore places data models, semantic interoperability, provenance and traceability among the core technical building blocks of a data space. Its architecture also includes identity management and machine-enforceable access and usage policies.

That semantic layer could become particularly valuable for AI developers. Training across 100 factories is less useful if engineers first have to manually interpret 100 incompatible representations of essentially the same machine state.

Recent Manufacturing-X research reached a similar conclusion. A 2026 study examining five advanced use cases inside a physical federated testbed found that, despite very different industrial applications, the fundamental technological barriers were surprisingly similar. Researchers proposed an architecture that reconciles operational-technology realities on factory floors with top-down IT systems.

Standardization is therefore not administrative overhead. It can determine whether industrial datasets are reusable enough to support scalable models.

Data Sovereignty Could Be Europe’s Competitive Feature

Europe’s approach also reflects why companies have historically hesitated to share industrial data.

Manufacturing information can reveal equipment performance, production volumes, component defects, supplier relationships and proprietary processes. Companies understandably resist handing those datasets to competitors—or even to AI vendors.

A 2026 open-access study of the EU-funded UNDERPIN and SM4RTENANCE projects identifies technical, business and regulatory barriers as obstacles to exploiting manufacturing data. Their data-space frameworks are designed around interoperability and data sovereignty so organizations can exchange information within controlled environments rather than surrender ownership of it.

The DSSC architecture follows the same principle: companies can define and enforce machine-readable policies governing who may access data and how it may be used, including enforcement during discovery, contract negotiation and actual data exchange.

For AI infrastructure companies, that creates an interesting product category between the model and the database: permission-aware training infrastructure.

The Opportunity Is Bigger Than Building a European Dataset

The most consequential outcome of the Common European Data Spaces may therefore not be a repository containing billions of industrial records.

It could be an infrastructure layer allowing European companies to train models across organizational boundaries without pretending those boundaries do not exist.

That opens opportunities for startups building connectors, semantic models, federated-learning systems, data-quality tooling, synthetic-data services, policy engines and industrial AI platforms capable of operating across multiple companies.

Europe has often been described as disadvantaged in AI because it lacks hyperscalers and foundation-model companies at the scale of the United States.

Industrial AI presents a different contest.

Europe already possesses something difficult to reproduce: dense manufacturing networks generating decades of specialized operational knowledge. If Common European Data Spaces can make that information interoperable and usable without forcing manufacturers to relinquish control, the EU may create an AI advantage from infrastructure that already exists inside its factories.

The strategic asset is not simply the data.

It is the system that finally makes fragmented industrial data trainable.

Share
f 𝕏 in
Copied