4  Architecture

The Open Music Observatory (OMO) is designed as a decentralised, federated, semantically interoperable knowledge infrastructure. Its architecture is the direct consequence of the policy, governance, and methodological requirements described in the Background chapter (see Chapter 2, Section 2.3, Section 2.4, and Section 2.5). The architecture is not an aesthetic or purely technical choice: it is the most sustainable and legally compliant way to realise the “open, scalable data-to-policy pipeline” mandated in the Grant Agreement and reflected in the Green Paper’s proposal for a federated European music data space. The architectural choice is also reinforced by the findings of the CITF First Project Report, which examined the requirements for trustworthy, lifecycle-aware copyright infrastructures in the AI era. CITF identifies the same structural needs—federated governance, interoperable identifiers, machine-readable rights metadata, and verifiable provenance—that underpin OMO’s design. This convergence shows that the Observatory’s dataspace architecture is not only technically justified but part of a broader European shift toward distributed, standards-based copyright and metadata infrastr

Traditional observatories created in the 1990s and 2000s were built on centralised databases, static data submissions, and annual reporting cycles. Such architectures are no longer viable in an ecosystem where:

These conditions make a federated dataspace (see Section 2.4) the only feasible approach. Within such a dataspace, Wikibase/Wikidata provides the most suitable technological foundation.

4.1 Knowledge Base for Humans and Machines

Knowledge representation describes how information about a given domain is structured so that humans and machines can interpret it in a consistent way. In the language of ISO standards, a knowledge representation identifies the concepts that matter, how these concepts relate to one another, and how they are recorded in a data model. In complex domains such as music production, rights management, or cultural statistics, knowledge representation acts as the bridge between raw data and meaningful analysis.

A knowledge base is the concrete form of this representation. It is a structured collection of facts, identifiers, rules, and relationships that model a specific domain and support reasoning over it. According to ISO definitions, a knowledge base may hold human-created mappings, machine-generated inferences, and rules that encode expertise from the various aspects of live music, recorded music, music education, cultural and music policies, music tech innovation, sustainability reporting and other areas.

A shared knowledge base is necessary because the data that describe music production are distributed across many actors and formats.

  • Live and recorded music producers maintain accounting ledgers and project-based cost codes.

  • Public authorities collect administrative data for tax rebates, audits, and cultural policy.

  • Statistical offices use NACE and CPA classifications, supply-use tables, and environmental accounts to represent the wider economy.

These systems describe overlapping parts of the world but in their own conceptual schemes and with their own definitions.

4.1.1 Shared Services: European Interoperability Framework

The European Interoperability Framework recognises that in such situations, achieving interoperability requires a shared space where meaning can be aligned without centralising all data. This is why we design MMAT as a data sharing space rather than a single repository. The data remain with the actors who produce or control them, but the shared knowledge base provides common identifiers, mappings, and semantic patterns that allow them to communicate.

A data sharing space solves a structural problem, which is explained further in the Chapter 51.

  • Live and recorded music productions use project-specific terminology and internal codes.
  • Accounting systems are arranged around financial categories.
  • National accounts rely on NACE activities and CPA products.
  • Environmental accounts follow the structure of air emissions accounts and input–output extensions.

None of these systems alone is sufficient to answer cross-cutting questions about economic impact, cultural policy, or sustainability. Without a shared data space, each actor must reconcile data manually, leading to inconsistencies, duplicated effort, and results that cannot be compared across projects or years.

A well-designed knowledge base captures the relationships among these systems and makes them reusable. It also makes it possible to combine administrative records with statistical classifications without violating legal boundaries around data protection or commercial confidentiality.

Interoperability is essential because screen production depends on many autonomous systems that must exchange information in a trustworthy and legally compliant manner. - Granting authorities need accurate and consistent metadata to evaluate public-funded projects and report aggregated results. - Producers need a simple way to reuse their accounting data for financial control, audit, and sustainability reporting. - Emerging sustainability requirements in Europe and the United States increasingly require value-chain information, including emissions embedded in purchased goods and services.

These requirements cannot be met if each system remains isolated. Semantic interoperability allows systems to describe their data in a shared language, even if each retains its own internal structures.

Technical interoperability ensures that these mappings can be exchanged across platforms.

Legal interoperability clarifies what can be shared and how data protection obligations are met. The European Interoperability Framework stresses that all three layers are necessary to enable cross-administration and cross-sector data flows.

The architecture of Open Music Observatory’s system, ReprexBase, adapts these principles for the music sector production. The music data sharing spaces of the OMO are not single databases. They form a a federated structure built around a common knowledge base.

The knowledge base contains the standard identifiers, mappings, and controlled vocabularies that describe the music domain in a machine-interpretable form.

These include mappings between chart-of-accounts categories, mandatory music-production cost codes, and national statistical classifications such as NACE and CPA. They also include alignments to supply-use and input–output tables, cultural satellite accounts, and environmental accounts. This allows a production ledger to be connected to national economic statistics through well-defined transformations. It also allows emissions accounting to be computed using harmonised environmental intensity data.

The architecture combines several layers.

At the foundational layer, Open Music Observatory’s central, supranational module maintains reference identifiers, domain vocabularies, and canonical mappings. This is where the authoritative concepts are defined, drawing on ISO terminology, national classifications, and legal definitions.

The semantic layer contains the patterns and equivalences that link different institutional models. It allows music cost codes to be interpreted as economic activities, or accounting categories to be reconciled with national accounts. Because different institutions organise their data according to different logics, this semantic layer is deliberately modular, following the approach we have used in other cultural sectors. It connects, rather than replaces, existing models.

At the technical layer, MMAT uses interoperable interfaces and automated pipelines to exchange data across systems. This is necessary for integrating ERP exports, digital audit files, statistical registers, and environmental accounting tools.

4.1.2 Fixing Metadata At Source; But Also Work with Legacy Metadata

This directly operationalises the Green Paper’s recommendation to “fix music data at the source” while enabling curative AI-assisted reconciliation of legacy metadata.

  1. Preventive workflows make data interoperable at the moment of creation by assigning identifiers, validating codes, and capturing provenance–this is what the European Parliament’s resolution calls for referring to the future: that future music assets should we well described.

  2. Curative workflows repair legacy data, reconcile inconsistent categories, and align historical records with current classifications.

In the music domain, both are necessary, because the copyright and neighbouring right protection terms necessitate the rights management and documentations up to about ~130 years (70 years after the passing of the last surviving composer.) This means that our systems must carry the data of the entire production catalogue of the 20th century and the first quarter of the 21st century.

Studio productions as well as live productions, particularly music festivals generate large volumes of financial and descriptive data in short time windows, and much of this data was not originally designed for reuse in national or environmental statistics. A knowledge base allows these datasets to be repaired and mapped in a consistent and transparent way.

Empirical findings from the Commission’s Study on Copyright and New Technologies confirm that persistent interoperability gaps, misaligned workflows, and incomplete, fragmented rights metadata continue to generate licensing inefficiencies and delayed remuneration in the music sector. Our Green Paper on AI, Data Governance, and Metadata Policies for Europe’s Music Ecosystem demonstrates how a federated data sharing space can connect public and private actors and enable cross-sector analysis while respecting legal and institutional boundaries2. It also shows that centralisation is unrealistic in a domain where autonomy, specialised workflows, and legal mandates differ across actors.

Reprexbase, and the core Open Music Europe module of our Observatory therefore provides a stable foundation for current and future needs.

  • It supports economic impact analysis by connecting production data to national accounts. It supports cultural policy by linking productions to cultural statistics and sectoral indicators.

  • Its diversity module helps to monitor the local content regulations of many countries (that require certain local, national or European production to be present in the broadcast stream) or KeyChange gender equity pledges.

  • It supports environmental analysis by integrating emissions accounting into the same data pipelines.It prepares the sector for future sustainability reporting obligations by establishing harmonised and reusable mappings.

In this sense, the knowledge base does not serve a single reporting purpose. It creates a shared language that enables producers, public bodies, and researchers to draw consistent and verifiable conclusions from heterogeneous data sources.

By grounding the architecture in open standards and the principles of the European Interoperability Framework, the Open Music Observatory’s ReprexBase system becomes a long-term infrastructure rather than a single project database. It creates a space where data can be shared meaningfully, legally, and reproducibly, and where new analytical tools can be integrated without restructuring everything from scratch. This ensures that the music (and also the audiovisual) sector can adapt to new policy requirements while retaining control over its own data and workflows.

4.2 Why Wikibase Is the Right Foundation for the Open Music Observatory

The European music ecosystem produces data in many incompatible formats:

  • repertoire and rights databases (CMOs, publishers, labels),
  • cultural-heritage and performing-arts collections,
  • company registries and economic statistics,
  • event metadata from festivals and venues,
  • community-maintained sources (Wikidata, folk archives, local heritage groups).

These sources use different identifiers, legal definitions, languages, and metadata schemas. A knowledge graph is therefore indispensable: no relational or document database can reconcile these sources while maintaining provenance, multilinguality, and entity-level linking.

NoteAlignment with CITF’s Copyright Infrastructure Model

The CITF First Project Report defines three layers that a modern copyright infrastructure must satisfy: a foundational identifier layer (authoritative PIDs and registries), a semantic layer (shared meaning and mapping across domains), and a technical layer (APIs, services, resolution, and provenance). Wikibase operationalises this same structure in practice. Its support for persistent identifiers, ontology alignment, multilingual semantics, and transparent versioning makes it fully compatible with the CITF model and positions the Observatory as a concrete, domain-specific implementation of that broader European framework (Partanen et al. 2025).

Wikibase is adopted because it has already proven its suitability in domains facing the same structural issues as the music ecosystem: fragmented identifiers, inconsistent authority control, multilingual metadata, cross-domain vocabularies, and parallel institutional workflows.

4.2.1 Proven in real-world scenarios highly similar to music

Wikibase is widely used by libraries, archives, museums, national cultural bodies, and open-science infrastructures. These institutions face the same challenges OMO addresses:

  • reconciliation of people, works, events, organisations, and places;
  • multilingual labels and aliases;
  • authority file alignment (ISNI, VIAF, ORCID, GND, BNF, corporate registers);
  • provenance tracking and version history;
  • SPARQL-based validation and constraint checking.

The GLAM-Wiki ecosystem, national knowledge graphs, and EU-funded linked-data projects have collectively demonstrated that Wikibase is an effective intermediary between:

  • authoritative PID systems;
  • domain ontologies (CIDOC-CRM, RiC-O, DDI, DCAT);
  • community-curated knowledge models.

This track record gives OMO a mature, future-proof, and standards-aligned foundation.

4.2.2 Already aligned with Europe’s digital knowledge infrastructure

Wikibase aligns with existing institutional practice across Europe. Many major knowledge centres and initiatives already use Wikibase/Wikidata:

  • national libraries and archives,
  • national cultural-heritage aggregators,
  • research infrastructures,
  • public-sector linked-data programmes,
  • the EU Knowledge Graph initiative.

Adopting Wikibase ensures that the Open Music Observatory fits directly into the European interoperability ecosystem (see Section 2.3).

Footnote: See Wikibase as an Infrastructure for Knowledge Graphs: the EU Knowledge Graph (Diefenbach, De Wilde, and Alipio 2021) and (2020 2020).

4.2.3 Demonstrated support for required OMO functionality

Everything OMO needs has already been demonstrated in production Wikibase environments:

  • authority control for creators, ensembles, organisations, venues;
  • multilingual and multiscript modelling for names, places, works;
  • cross-domain entity linking (work–recording–performance–rights–heritage);
  • event-based and entity-based models;
  • SPARQL validation, schema constraints, and automated reconciliation.

This means OMO does not invent an untested paradigm: the consortium integrates proven practices from:

  • national registries,
  • performing-arts knowledge graphs,
  • the Slovak pilot and Finno-Ugric metadata federations developed inside the project.

4.2.4 Fits EU policy preference for open-source and trustworthy AI

Wikibase is open-source, auditable, and non-proprietary.
It aligns with:

  • the EU’s preference for open-source digital public infrastructure,
  • FAIR and CARE principles,
  • trustworthy AI requirements (provenance, transparency, versioning),
  • cross-border interoperability mandates,
  • decentralised data governance models.

This makes it compatible with Europeana, EOSC/ECCCH, DCAT-AP, and the European Interoperability Framework. It is used in EU organisations, too. 3

4.2.5 The most widely used graph-editing interface in the world

Tens of thousands of data stewards, librarians, researchers, and citizen-scientists already know how to edit Wikibase/Wikidata.

This provides OMO with:

  • an immediate user base,
  • a ready-made contributor community,
  • institutional familiarity across Europe,
  • workflows already adopted in GLAM and research sectors.

No alternative open-source system has remotely this level of adoption.

4.2.6 A hybrid model that fits real institutional workflows

Wikibase uniquely accommodates:

  • spreadsheet-based workflows (Excel, CSV),
  • relational database exports,
  • statistical microdata reference linking,
  • complex semantic modelling,
  • API-based ingestion,
  • R and Python pipelines.

It is a practical compromise between triple stores, document databases, and relational systems — perfect for a music ecosystem where many partners still rely on basic tools. It has proven to be useful in music services4, and more generally on small- and large scale European knowledge institutions (national libraries, libraries, archives, museums5.)

4.3 How Wikibase Fits into the Open Music Europe Data-to-Policy Pipeline

The data-to-policy pipeline defined in the Grant Agreement and documented in the Background chapter (see Section 2.5) provides the methodological backbone of OME. Wikibase is the component that makes this pipeline operational.

4.3.1 Wikibase supports each stage of the pipeline

4.3.1.1 Data collection

  • imports from Excel, CSV, SQL, APIs, and legacy systems;
  • immediate linkage to persistent identifiers;
  • entity reconciliation as part of ingestion.

4.3.1.2 Validation and reconciliation

  • authority-control workflows for people, works, organisations, and places;
  • constraint-based quality checks;
  • SPARQL-driven validation;
  • alignment with external authority files.

4.3.1.3 Harmonisation and enrichment

  • multilingual labels and roles;
  • event-based and relationship-based modelling;
  • addition of contextual metadata by different institutions;
  • integration of domain vocabularies.

4.3.1.4 Activation for analysis

  • SPARQL endpoints for programmatic access;
  • JSON-LD, RDF dumps, and REST APIs;
  • R and Python pipelines use stable URIs for reproducibility.

4.3.1.5 Indicator construction

  • cross-domain indicators linking economic, cultural, rights, and heritage data;
  • entity-level referencing ensures indicators are traceable and verifiable.

4.3.1.6 Interpretation and contextualisation

  • experts review, annotate, and correct metadata through a human-readable interface;
  • provenance guarantees transparency.

4.3.1.7 Policy translation and observatory outputs

  • live, federated knowledge base powering the OMO front end;
  • entity profiles, metadata dashboards, and linked methodological documentation;
  • direct links from indicators to source entities and datasets.

4.3.1.8 Feedback loop

  • corrections flow back into the shared graph;
  • updated authority files update all downstream indicators;
  • the observatory improves over time.

4.4 Summary

Wikibase is adopted because it best fulfils the policy requirements, semantic needs, technical constraints, and governance expectations described in the Background chapter:

  • It is proven in domains identical to ours.
  • It is aligned with the EU’s open-source, dataspace, and interoperability agenda.
  • It is widely adopted by the very institutions we must interoperate with.
  • It supports the entire Open Music Europe data-to-policy pipeline.
  • It allows the Observatory to operate as a decentralised, federated, evolving knowledge infrastructure.

This architecture is therefore not an optional design preference but the only viable model for delivering a European Music Observatory described in the Grant Agreement, taking into consideration the developments in data governance and data interoperability since the creation of the feasibility study and awarding this grant.


  1. See in particular (Kung, Walshe, and Wenning 2023) and the Chapter 5 for more details.↩︎

  2. This green paper is the accompanying policy document of this technical description (Antal 2025).↩︎

  3. On official adoption: EU Knowledge Graph (Diefenbach, De Wilde, and Alipio 2021); SEMIC guidelines (SEMIC Support Centre 2023).↩︎

  4. On Belgian pilots: MetaBelgica (Stallmann et al. 2023) and Flemish performing arts enrichment (Magnus and Van D’huynslager 2021).↩︎

  5. See for example: On Wikidata/Wikibase in heritage: (Bianchini, Bargioni, and Pellizzari di San Girolamo 2021; Sardo and Bianchini 2022).↩︎