6 Federated Data Modules
This chapter introduces the four federated data modules that currently constitute the backbone of the Open Music Observatory.
Rather than treating “data sources” as a flat list, we follow the architectural logic established in Chapter 4 and the data-to-policy pipeline in Chapter 2: data is contributed, curated, harmonised, and activated inside federated modules, each with its own governance, provenance, and semantic profile.
The OpenMusE project was contractually expected to populate the Observatory’s four thematic pillars — Economy, Diversity, Society, and Innovation — through coordinated data collection, harmonisation, and processing. These contributions now materialise as a central Open Music Observatory module, complemented by three national or regional modules demonstrating how federation works in practice. Together, they show how the Observatory evolves as a network of interoperable, decentralised components rather than a single central repository.
The four federated modules presented in this chapter are:
- the Slovak Comprehensive Music Database (SKCMDb) — the first fully scaled national dataspace feeding the Observatory, accessible at https://hudobnadatabaza.sk/en/;
- the Hungarian Music Database (HU-MDb) — a replication and enhancement of the Slovak model, demonstrating cross-border portability and semantic continuity;
- the Finno-Ugric Data Sharing Space — an example of subsidiarity-based federation for regional, community, and low-depth archival collections, accessible at https://finnougric.net/en/;
- the Open Music Observatory Core Module — the central, pan-European module containing data created by WP1, WP2, WP3, and WP4 of the OpenMusE project, and the point of semantic linkage between national and regional modules; in its partially uploaded form available on https://openmusicobservatory.eu/.
Each module is described using the same structure:
1. Purpose and scope
2. Data inputs and contributing institutions
3. Metadata, identifiers, and semantic alignment
4. Governance and legal basis
5. Interoperability with OMO and other modules
6. Current status and next steps
In the sections that follow, we provide concise summaries of the existing modules and describe how they federate with the Observatory. The temporary landing page of the OMO can be reviewed on https://dataobservatory-eu.github.io/omo-landing-page/.
Once the DMP is updated, all datasets described in these modules will be linked to the DMP’s Data Summaries, vocabularies, legal bases, and provenance statements, and ingestion workflows will proceed accordingly.
6.1 Slovak Comprehensive Music Database (SKCMDb)
The Slovak Comprehensive Music Database (SKCMDb) is established to make Slovak music and music-related cultural assets more visible, discoverable, and interoperable across memory institutions, rights-management ecosystems, and public collections. It is available on https://hudobnadatabaza.sk/en/.
6.1.1 Purpose and scope
increase the precision of information about Slovak musical works, recordings, persons, and artefacts;
improve public accessibility to sheet music, sound recordings, books, and documents;
create a publicly accessible, linked database (“Slovak Summary Music Database”) as part of a broader Slovak Music Dataspace;
provide trusted, authoritative identifiers enabling cross-institutional coordination through VIAF, Wikidata, ISNI, library identifiers, and other persistent identifiers;
support digital curation, enrichment, and harmonisation of metadata across libraries, archives, collective management organisations, and publishers.
This module embodies the subsidiarity principle: national institutions retain control over their data while contributing to a shared knowledge infrastructure compatible with the Open Music Observatory.
The scope of the federated Slovak module covers:
musical works (compositions, arrangements)
recordings (audio carriers, CDs, digitised tapes, releases)
persons and organisations related to music (composers, performers, lyricists, pedagogues, musicologists, publishers, ensembles)
music-related documents (books, sheet music, archival holdings)
physical and digital artefacts in Slovak libraries and public collections
authoritative name control for Slovak creators and entities, with linkage to international authority files
cross-institutional identifiers, especially VIAF, ISNI, Wikidata Q-IDs, and internal catalogue identifiers
The scope explicitly includes both heritage-sector data (libraries, archives, SNL) and rights-management data (SOZA), forming a hybrid public-private data module.
6.1.2 Data inputs
Based on the Memorandum, the SKCMDb ingests the following categories of data:
From SOZA
- work registrations (musical works, identifiers, rightsholders) - personal data of creators and publishers (under collective-management legal basis)
additional metadata supporting identification, repertoire documentation, and cultural-value preservation
From Slovak Music Centre (MC)
documentation of professional music creativity and live music culture across genres
metadata from MC’s internal library, databases, and central information systems
catalogues of events, artists, performances, and publications
From Slovak National Library (SNL)
authoritative name authority files
VIAF IDs created for Slovak individuals and organisations not yet represented
internal library identifiers
catalogue records from SNL and associated Slovak libraries
From Music Fund (MF)
publishing metadata (sheet music, books, CDs)
catalogue and e-shop metadata (Musica Slovaca)
information on supported creators and institutions.
From Reprex
semantic models, schemas, reconciliation rules
enriched metadata with global PIDs (Wikidata, ISNI, ROR, etc.)
data linking workflows across institutions
These data inputs form a complete national module suitable for federation with regional, subnational, or thematic music datasets.
Additional libraries and memory institutions, labels, publishers, association may add: - internal identifiers - holdings metadata for music-related artefacts - distributed repertoire metadata - digitised and physical collection metadata
6.1.3 Metadata and semantic alignment
The Slovak Music Data Sharing Space adopts the same semantic modelling principles as the Open Music Observatory. Although it operates as a separate federated node with its own namespace, entity schemas, and property definitions, its conceptual model is fully aligned with OMO. All Slovak classes and properties that correspond to OMO concepts are connected through equality relations (e.g., owl:equivalentClass, owl:equivalentProperty, or the Wikidata’s “equivalent to” relation), ensuring semantic interoperability across federated modules.
This alignment allows the Slovak module to use:
- the same core ontology for persons, works, recordings, organisations, events;
- the same authority-crosswalk strategy linking VIAF, ISNI, ROR, and Wikidata;
- compatible entity schemas for music works, creators, and releases;
- the same provenance and versioning rules, enabling OPA and reproducible workflows.
At the same time, the Slovak dataspace includes a number of Slovakia-specific classes and properties—for example, detailed roles in Slovak musical traditions, local institutional identifiers, historical datasets, and legacy catalogue structures.
These are maintained locally but semantically bridged to OMO through mappings. This preserves local specificity while ensuring cross-border interoperability with other national modules and with the central OMO knowledge graph.
Crucially, the Slovak module is designed for high interoperability with major open and industry ecosystems, including Wikidata, MusicBrainz, Discogs, Europeana, the Cultural Heritage Cloud, and DDEX-based commercial metadata pipelines. The SKCMDb has already exchanged a significant volume of data with Wikidata in both directions, and we are now preparing a scaled-up, systematically designed exchange. Because Wikidata is deeply integrated with MusicBrainz and Discogs, these improvements will further increase the visibility and discoverability of Slovak music across the open-music ecosystem used by curators, streaming platforms, and many other downstream stakeholders.
We are preparing our first full round-trip proof-of-concept with the WikiProject Music for December 2025. To support this, the Reprex team actively participates in the Wikidata Ontology/Cleaning Task Force and the Wikidata Mereology Task Force, treating this round-tripping exercise as a pilot for broader interoperability improvement across the entire music metadata landscape.
We are planning our first round-trip proof of concept with project in December 2025. For improved interoperability, we are participating in the Wikidata Ontology Cleanup Task Force and its Mereology Working Group, and treat our round-tripping excersize as a broad interoperability improvement pilot.
The SKCMDb is also the first real-world implementation in the music domain that operationalises several requirements independently identified by both the CITF First Project Report (2025) and the Open Music Europe Green Paper (2025). Both documents diagnose the same structural obstacles—fragmented rights metadata, missing or unstable identifiers, incomplete provenance, and weak linkage between public cultural-heritage systems and private rights registries. The SKCMDb directly addresses these challenges by providing trusted identifiers, lifecycle-based provenance, cross-domain semantic mediation, and a federated governance model spanning national libraries, rights management, publishers, and public-sector institutions. This makes the Slovak module a practical demonstration of how a national music dataspace can support trustworthy AI, machine-readable copyright infrastructures, and cross-border cultural data interoperability.
6.1.4 Governance and legal basis
Governance is defined contractually in the Memorandum of Understanding and implemented through collaboration among five parties:
Slovak Music Centre (HC) — national documentation centre for professional music culture; maintains live-music databases and coordinates IAML and IAMIC activities.
Slovak National Library (SNL) — national authority for cataloguing and VIAF contributions; responsible for creating VIAF records for Slovak persons and organisations not yet represented.
SOZA — collective management organisation for musical works; provides authoritative work registrations, rightsholder data, and contributes identifiers.
Music Fund (MF) — public institution supporting music creation; contributes metadata from its publishing, catalogue, and Musica Slovaca activities.
Reprex B.V. — data and knowledge-management provider; responsible for semantic modelling, interoperability, and technical coordination.
The Memorandum establishes a multi-party governance structure, characterised by shared stewardship over data and identifiers and coordinated authoritative control (e.g., VIAF, national library authority files, SOZA work registrations). It is legally grounded cooperation under the missions of MC, SNL, SOZA, and MF and shows openness to additional Slovak libraries and stakeholders joining the dataspace.
This positions SKCMDb as a national federated node aligned with the European Interoperability Framework and ready for integration into the Open Music Observatory.
6.1.5 Interoperability and federation
6.1.6 Status and next steps
Due to difficulties with the Data Management Planning of the OpenMusE project, the prolonged grant agreement changes, and political changes in Slovakia, building the governance model took longer than we expected. The module is being populated with data.
- Trustworthy RMI chains — linking SOZA work registrations with SNL authority control, VIAF/ISNI identifiers, Wikidata entities, and OMO entities.
- Lifecycle-based provenance — each reconciliation step includes machine-readable provenance and versioning.
- Semantic interoperability — crosswalks between DDEX, MARC, EDM, RiC-O, DCTERMS, Wikidata patterns, and local Slovak schemas.
- Federated governance — each institution retains its own data; only identifiers and mappings are shared.
- AI readiness — stable reference objects for works, recordings, persons, and organisations suitable for trustworthy AI guardrails and fairness testing.
Together these elements turn the Slovak dataspace into a rights-aware, culturally inclusive, and technically interoperable national module fully compatible with the Open Music Observatory.
6.2 Hungarian Music Database (HUMDb)
The Hungarian Music Database (HuMDb) is the second national-level module of the Open Music Observatory’s federated dataspace architecture. Its purpose is to adapt and extend the principles tested in Slovakia to the specific institutional landscape, heritage depth, and rights environment of Hungary.
The conceptual basis of the Hungarian module follows the principles established in the Slovak Comprehensive Music Database (SKCMDb). As summarised in (Antal 2024), the SKCMDb demonstrates how a national music dataspace can provide trustworthy descriptions of all music connected to a territory or cultural community, without imposing ethnomusicological or legal definitions of identity. It links public memory institutions and private rights-management organisations through a shared semantic layer, enabling legal, organisational, semantic, and technical interoperability. This framework was presented as a model for cross-border replication and for future federation with Hungarian music data owners, showing how Hungarian institutions could join a common dataspace while retaining their own systems, authority structures, and cultural specificities. The Hungarian module applies the same principles but adapts them to Hungary’s heritage depth, institutional landscape, and contemporary music workflows, forming the second national node in the Open Music Observatory’s federated architecture.
At this early stage, two interoperable but distinct tracks are being developed in parallel: a heritage-focused knowledge graph in cooperation with the Hungarian Heritage House (HHH), and a broader music-sector knowledge base and data-sharing space developed with the House of Music Hungary and its library and pop music heritage collection (MZH).
we hope to include in this replication the members of the Independent Label Fair, i.e., Hungarian music microlabels that have a good working relationship with MZH.
These two tracks differ in governance readiness and legal clarity, but are semantically compatible, and together outline the full scope of the future Hungarian module.
6.2.1 Purpose and scope
Heritage-focused module (Hungarian Heritage House)
The cooperation with the Hungarian Heritage House focuses on integrating folklore, folk-music, and narrative-heritage collections into a multilingual, interoperable knowledge-graph environment. Using small pilot corpora—such as subsets of the Székely dance-music collections and the Hungarian Folk Tale Inventory—the pilot tests:
authority reconciliation and personal-name alignment,
geographic and settlement-level gazetteer integration,
thesaurus and ontology alignment,
AtoM (ICA-AtoM) round-trip export and re-ingestion.
These pilots demonstrate how legacy archival structures can be modernised and how Hungarian ethnomusicological and folklore datasets can be embedded into a wider European federation. Further documentation is provided in Enriching and Futureproofing the Databases of the Hungarian Heritage House (Antal and Zagyva 2025), which is available here in various formats.
6.2.1.1 Music-sector module (Magyar Zene Háza)
In parallel, the cooperation proposal with the House of Music Hungary (MZH) defines a richer, institution-wide music knowledge base that connects collections, studio recording workflows in line with the fixing music data at source principle1, events,and public-facing applications into a unified data-sharing space.
The scope includes:
harmonising library, archival, studio recording, and event workflows, and their data representation;
creating a modern collections database interoperable with open-source library software;
standardising legacy studio files and event data (KeleSys) through structured metadata, identifier strategies, and provenance capture;
establishing rights-aware workflows (ISRC, ISWC, UPC/EAN, performer roles, producer rights);
building a bilingual knowledge base capable of supporting AI-assisted interfaces, including chatbot prototypes.
The conceptual model links internal MZH systems (library, records, studio workflows, event management) to external authority systems such as VIAF, ISNI, ISRC, ISWC, Nemzeti Névtér. The knowledge base supports both internal processes (archival, rights management, acquisitions) and public-facing applications (web search, chatbot, event browsing).
6.2.2 Data inputs
Heritage module (HHH)
- folklore recordings and field notebooks - folk-tale inventory samples - dance-music metadata - archival catalogue exports - geographic and settlement-level metadata
Music-sector module (MZH)
- library and archival catalogue data - legacy collections requiring harmonisation - studio metadata and recording files - event metadata from KeleSys - collection and workflow data from web and EyeWall systems - initial rights metadata such as ISRC and ISWC
6.2.3 Metadata and semantic alignment
HuMDb uses the same mediation patterns as other OMO modules:
- mapping creators and organisations to VIAF, ISNI, Wikidata, and national authority files
- alignment of works and recordings with ISWC, ISRC, and DDEX categories
- mapping geographic entities through regional gazetteers, including Finno-Ugric materials
- thesaurus alignment for genres, traditions, and folk-taxonomies
- use of Wikibase schemas compatible with the OMO ontology layer
The heritage track follows semantic structures tested in Finno-Ugric pilots, while the MZH track follows contemporary rights and workflow patterns. Both converge within the same semantic layer.
6.2.4 Governance and legal basis
HuMDb governance is in formation. Current cooperation includes:
Hungarian Heritage House: stewardship of heritage and folklore collections; cooperation agreement based on the OpenMusE grant agreement between Reprex and the HHH.House of Music Hungary: management of contemporary collections, studio workflows, and event metadata. ; cooperation agreement based on the OpenMusE grant agreement between Reprex and the HHH.Reprex B.V.: semantic modelling, identifier strategy, mediation, and future-proofingweCan: chatbot integration.- optional technical coordination with KeleSys developers.
Each partner retains authority and control over its own datasets.
The emerging model follows the Slovak Memorandum of Understanding structure but adapted to Hungarian institutions.
6.2.5 Interoperability and federation
HuMDb is designed to interoperate with:
- the Slovak module, using shared schemas for persons, works, and events
- the Finno-Ugric Data Sharing Space, sharing heritage workflows and gazetteers
- Wikidata and VIAF/ISNI ecosystems
- open-source catalogue and archival systems including AtoM and Koha
- industry platforms aligned with DDEX for distribution and rights workflows
The Hungarian module introduces additional areas such as studio-file mediation and chatbot-ready knowledge-base services.
6.2.6 Status and next steps
- heritage datasets have been semantically lifted following the feasibility work (Antal and Zagyva 2025)
- the MZH cooperation document defines a two-month roadmap for studio files, events, and a public chatbot interface
- ingestion templates and minimal metadata rules are being prepared
- governance arrangements will be formalised after the first pilot integrations
- next step: establishing a unified namespace and SPARQL endpoint for federation with the Open Music Observatory.
6.3 Finno-Ugric Data Sharing Space
The Finno-Ugric Data Sharing Space (FUDSS) is the third federated module of the Open Music Observatory. The FUDSS can be accessed via https://finnougric.net/
It functions as a subsidiarity-based regional node, designed to support culturally endangered, minority, community, and heritage-rich datasets that lack the institutional structures available in larger countries.
This module demonstrates how the OMO federated architecture can scale to low-resource, multicultural, multilingual, and distributed memory environments, and how open-source semantic technologies allow small organisations to participate in European data spaces.
6.3.1 Purpose and scope
Cultural and linguistic preservation Providing a sustainable semantic infrastructure for Finno-Ugric musical traditions, including Estonian, Finnish, Sámi, Mari, Udmurt, Komi, and Livonian materials.
Linking heritage and contemporary music ecosystems Connecting archival field collections to rights-aware DDEX-compatible distribution workflows and multilingual knowledge graphs.
Demonstrating regional federation Implementing the principles described in the CITF First Project Report and the Open Music Europe Green Paper in small and distributed heritage environments.
The scope includes traditional songs, field recordings, contextual ethnographic metadata, contemporary reinterpretations, revival performances, notebook materials, archival finding aids, community-maintained materials, multilingual enrichment in Finno-Ugric languages and regional languages, and DDEX-ready metadata for selected recordings. It also includes digital twins for rights-aware distribution that separate non-commercial research uses from commercial streaming versions.
6.3.2 Data inputs
Data inputs come from four primary sources.
Latvian Archives of Folklore i.e., Garamantas.lv: - digitised field recordings and ethnographic notes - collector, performer, and informant metadata - archival structure (fonds, series, file, item) - settlement-level geographic data - Livonian and Latvian cross-border repertoires
University research datasets:
- ethnomusicology corpora - Finno-Ugric language and phonology datasets - structured vocabularies and contextual descriptions - annotations and transcriptions
Community organisations: - local archives of Sámi, Komi, Mari, Udmurt, Livonian, and other Finno-Ugric groups - recordings linked to cultural revitalisation - contextual information, translations, and performance metadata
Reprex and Unlabel: - semantic models and reconciliation rules - enriched metadata mapped to VIAF, ISNI, Wikidata, and geographic registers - DDEX catalogue-transfer metadata - digital-twin transformations
Selected distributors:
- ALOADED proof-of-concept for turning archival metadata into releasable DDEX messages
6.3.3 Metadata and semantic alignment
The Finno-Ugric module follows the same semantic principles used across the Open Music Observatory.
Key elements: - multilingual labels and scripts - reified contribution roles for collectors, performers, informants, translators - preservation of local vocabularies instead of normalisation - settlement alignment with national and European gazetteers - archival provenance aligned with Records in Contexts (RiC-O) - alignment to CIDOC CRM, DCTERMS, EDM - equivalence mappings to OMO ontology and Wikibase schemas
Industry identifiers are integrated through:
- ISWC, ISRC, and DDEX fragments defined in the OMO namespace - release metadata compatible with commercial distribution pipelines - multilingual contributor roles
This enables round-trip interoperability with archival systems, OMO, Wikidata, MusicBrainz, and commercial distributors where permitted.
6.3.4 Governance and legal basis
Governance follows a distributed subsidiarity model.
Each institution retains stewardship and decides which data to share. Rights metadata controls which versions are usable for research, public access, or distribution. Personal data follows GDPR balancing tests and cultural-heritage exemptions. Community organisations provide contextual data and approve sensitive uses. Reprex maintains the semantic layer and federated workflows.
Legal bases include public-domain status, non-commercial research licences, community authorisation, and explicitly granted commercial licences. The module implements digital-twin workflows separating versions for MIR research from versions licensed for streaming.
6.3.5 Interoperability and federation
6.3.6 Status and next steps
The Finno-Ugric Data Sharing Space demonstrates how minority and low-resource cultural communities can participate in a European data sharing space using a lightweight, federated, rights-aware approach. It connects archival heritage, community knowledge, research datasets, and modern music-industry workflows, showing how the Open Music Observatory architecture supports both cultural preservation and contemporary reuse.
6.4 Open Music Observatory Core Module
The Open Music Europe Core Module is the central, project-wide knowledge base that integrates the datasets, metadata structures, indicators, and workflows produced in WP1–WP4, in line with the Grant Agreement’s mandate to deliver a “360-degree intelligence” system for the European music ecosystem . Its function is not to act as a monolithic database, but to serve as the semantic, methodological, and interoperability backbone that connects the thematic, national, and regional modules of the Observatory.
6.4.1 Purpose and scope
The Core Module provides the shared foundations required by the data-to-policy pipeline defined in Annex 1 of the Grant Agreement. It ensures that all datasets curated and produced within the project:
follow the project’s shared DMP, licensing, and data-protection rules
use common indicator definitions, harmonisation schemas, and metadata standards
expose reproducible processing workflows (survey ingestion, statistical pipelines, streaming sampling, register harmonisation)
feed policy analysis through persistent identifiers, versioning, and provenance that fulfil the project’s obligations for transparency and open policy analysis under Horizon Europe rules.
It is the reference point against which national, regional, and thematic modules align their schemas, identifiers, and provenance models.
6.4.2 Data inputs (WP1–WP4 contributions)
The Core Module ingests curated outputs from:
WP1 policy and landscape mapping (conceptual definitions, indicator families, regulatory mappings)
WP2 identification of data gaps, methodological foundations, and pan-European variable definitions
WP3 data collection instruments, including survey pipelines, national statistical harmonisations, platform-data sampling, and register-based contributions
WP4 data processing tools, ontologies, entity schemas, and automated ingestion scripts.
These curated datasets correspond to the “backend datasets” referenced in the project summary: official statistics, survey participation data, rights-holder data, and streaming-service samples that power the OMO “living policy documents”
6.4.3 Metadata and semantic alignment
The Core Module maintains the canonical Wikibase ontology layer for the project. It defines:
shared classes for persons, organisations, works, recordings, events, economic indicators
crosswalks between Eurostat, national statistical offices, cultural heritage vocabularies, DDEX categories, VIAF/ISNI/Wikidata identifiers
provenance and versioning rules required for reproducible scientific workflows
standardised definitions for contested domain concepts (composer, lyricist, average income, microlabel, participation index), in line with the project’s standardisation tasks.
It acts as the semantic mediation layer between statistical, heritage, industry, and survey data—mirroring the federated architecture recommended in the CITF report.
6.4.4 Governance and legal basis
The Core Module is governed by the consortium as a whole under Annex 1 of the Grant Agreement, with SINUS as coordinator and Reprex as technical steward for semantic modelling, ingestion, and interoperability. It is also complemented by the Consortium Agreement and the elements defined in these agreements:
- The Data Management Plan of the consortium
- The OPA folders where the raw data and its contextual information, such as provenance information can be found.
All partners contribute datasets or workflows according to their WP obligations. The Module operationalises Horizon Europe requirements on open science, open-source software, transparent methodology, and reproducibility, using the licensing, privacy, and provenance conditions defined in the DMP.
6.4.5 Interoperability and federation
The Core Module is the linking hub for all other modules. It federates with:
the Slovak Comprehensive Music Database
the Hungarian Music Database (HHH + MZH tracks)
the Finno-Ugric Data Sharing Space
external infrastructures such as Wikidata, VIAF/ISNI, Europeana, the EU Open Data Portal, and the Cultural Heritage Cloud.
It does so by providing the shared namespace, identifier strategy, ontological patterns, and schema alignment that allow national and regional nodes to retain autonomy while contributing to a unified European knowledge infrastructure.
6.4.6 Status and next steps
The Core Module is operational and continuously expanded as WP3–WP4 pipelines mature. Next steps include:
finalising the unified indicator registry
completing automated ingestion for all WP datasets
publishing stable RDF/JSON-LD exports for long-term reproducibility
preparing the federation services that will onboard external stakeholders, as foreseen in the Grant Agreement and Amendment.
6.5 Summary and Integration
This modular, federated structure reflects the principles described in Chapter 2 and operationalises the architectural choices detailed in Chapter 4. The Observatory grows through interoperable contributions rather than central accumulation, aligning with the European data strategy, the Data Governance Act, and the European Interoperability Framework.
See the Fixing Music Data at the Source of our Open Music Europe Green Paper (Antal 2025) and the European Parliament’s resolution (European Parliament 2024). See Chapter 2.↩︎