Semantic Knowledge Production and the Review Algebra
Two Complementary Working Papers on Provenance-Aware Semantic Systems
Abstract
This document contains two closely related working papers in progress. The first develops a conceptual account of semantic knowledge production as a provenance-bearing process through which human and computational agents produce semantic assertions from evidence, existing knowledge, specialised interpretation, and inference. The second develops a Review Algebra for representing how candidate semantic assertions are generated, reviewed, finalised, and progressively stabilised. The two papers are developed together because the semantic-production problem motivates the requirements of the algebra, while the algebra provides a formal computational component through which those requirements can be implemented and tested.
The first working paper examines how semantic knowledge is produced through empirical observation, domain and linguistic interpretation, reconciliation with external knowledge, and inference. These activities are not epistemically interchangeable. They depend on different evidence, methods, competences, permissions, and conditions of applicability. Particular attention is given to the distinction between evidence as encountered in a computational or documentary environment, the objects represented through that evidence, and the semantic assertions produced about them. Provenance is therefore treated as integral to semantic knowledge production rather than as metadata added after knowledge has been produced.
The second working paper develops the Review Algebra as a narrower formal response to this problem. The algebra distinguishes three recurring operations: candidate generation, review, and finalisation. Review returns both a semantic value and an explicit review status, allowing corroboration, falsification, deferral, and alternative proposals to remain distinguishable. Finalisation determines which reviewed assertions are admitted to a subsequent stabilised semantic state. Candidate generation, review, and finalisation are provenance-bearing transformations, allowing semantic knowledge to develop without destructive modification of earlier states.
The Review Algebra extends ordinary relational and tidy data workflows rather than replacing them. Candidate assertions, reviewed values, review statuses, and provenance can be represented in normalised persistent form and projected into wide human-centred review surfaces. Persistent representations can in turn distinguish domain knowledge from provenance knowledge for exchange with RDF, graph, and other semantic infrastructures. Stabilised assertions may recursively support further candidate generation, review, and inference.
A running example based on multimodal Livonian cultural heritage connects the two working papers. It illustrates how digital and documentary evidence can support empirical observation, specialist interpretation, named-entity identification, multilingual and lexical knowledge, rights-related reasoning, and inference, while the Review Algebra provides a common formal representation for the resulting candidate, reviewed, and stabilised assertions.
The two working papers are deliberately retained in a single versioned document at this stage of development. Each has its own introduction and conclusion and is intended to become independently readable. Their joint development allows conceptual requirements, formal definitions, reference implementations, and empirical applications to be revised against one another before the papers are separated for publication.
This document contains two complementary working papers in progress:
Semantic Knowledge Production: Evidence, Competence and Provenance in Human–Computational Knowledge Systems develops the broader conceptual framework: how evidence, human and computational activities, domain and linguistic competence, inference, permissions, and limits of knowledge participate in the production of semantic assertions.
The Review Algebra: A Computational Algebra for Provenance-Aware Semantic Stabilisation develops a formal computational component of that framework: how candidate assertions are generated, reviewed, finalised, and transformed into progressively stabilised semantic states while preserving provenance.
The papers are currently maintained together deliberately. Semantic knowledge production provides the problem domain and requirements against which the Review Algebra is being designed, while implementation of the algebra provides a means of testing and refining those conceptual requirements. Each paper therefore has its own abstract or introductory framing and conclusion, and subsequent revisions are expected to separate them into independently publishable manuscripts.
The purpose of the present version remains developmental. Terminology, formal definitions, computational representations, reference implementations, and practical case studies are being refined together before comprehensive literature review, formal evaluation, and broader empirical validation.
The development of the work is preserved through a versioned Zenodo record:
- 7 July 2026 — earlier conceptual formulation: DOI 10.5281/zenodo.21250580
- 8 August 2026 — earlier revised formulation: DOI 10.5281/zenodo.21849208
- 9 August 2026 — earlier revised formulation: DOI 10.5281/zenodo.21862043
- 12 August 2026 — current revised formulation: DOI 10.5281/zenodo.21890589
The earlier version remains citable as a record of the development of the framework, while the present version represents its current working formulation.
Glossary
The following terms are used consistently across the related working papers, case studies, and reference implementations. Wherever an established standard provides a suitable concept, the terminology follows or builds upon that standard; framework-specific terms are introduced where the required distinction is not adequately captured by existing terminology. The glossary is not intended to replace definitions provided by ISO, PROV, RiC, CIDOC CRM, ODRL, or other external standards. Where those standards or models are used, their own formal definitions continue to apply.
Agent
A human or computational participant capable of performing or participating in a knowledge-producing activity. In the specific case of an AI agent, ISO/IEC 22989 defines an automated entity that senses and responds to its environment and takes actions to achieve its goals (ISO/IEC 2022). In semantic knowledge production, what an agent can responsibly establish additionally depends on the competence, method, permissions, and other conditions relevant to the particular activity.
Applicability
The conditions under which a question, activity, rule, method, or inference is relevant or warranted. Applicability may depend on previously stabilised assertions, entity classes, relationships, temporal or geographical conditions, scope, permissions, or other explicitly represented knowledge.
Assertion
A semantic statement concerning a subject, predicate, and value. In the Review Algebra, the elementary domain assertion is represented as \((s,p,v)\). Assertions may be candidates, reviewed assertions, or members of a stabilised semantic state.
Candidate assertion
An assertion made available for consideration but not thereby assumed to be corroborated or stabilised.
Example: depicted garment — instance of — Livonian folk dress may enter review as a candidate.
Candidate generation
The operation through which one or more semantic assertions become explicit candidates for subsequent review or finalisation. Candidate generation is an elementary operation of the Review Algebra; observation and inference are different mechanisms through which candidate generation may occur.
Competence
The capacity of an agent to responsibly perform a particular knowledge-producing activity or establish a particular kind of assertion. Competence is activity-specific rather than a general property of an agent.
Example: A textile researcher and a local speaker may competently describe the same coat in different conceptual vocabularies.
Conditions of applicability
The explicit conditions that must be satisfied before a rule, inference, question, method, or other knowledge-producing operation is applicable. Scope may contribute to these conditions but is not synonymous with them.
Deferred
A review status indicating that a candidate has been examined but cannot presently be corroborated or falsified. Deferral records the result of a completed review activity rather than an absent value.
Domain knowledge
Semantic knowledge about the entities, relationships, events, properties, classifications, or other subject matter of the domain under investigation. Domain knowledge is distinguished from provenance knowledge concerning how those assertions were produced.
Example: Hilda Grīva wears Livonian folk dress.
Evidence
An entity examined or otherwise used in a knowledge-producing activity in order to establish, challenge, or interpret semantic assertions. Evidence is not necessarily identical to the semantic object about which knowledge is produced.
Example: pict003.jpg is evidence used to establish what the depicted person is wearing.
Evidentiary frame
The relationship under which evidence is examined in a particular knowledge-producing activity, specifying what kind of semantic object the activity concerns.
Example: pict003.jpg is observed as evidence concerning a depicted garment.
Finalisation
The Review Algebra operation that applies an explicit policy to review results or other eligible candidates and determines which semantic assertions enter a subsequent stabilised semantic state. Finalisation is distinct from review: epistemic judgement and operational authority need not belong to the same agent or process.
Identification
A knowledge-producing activity that establishes or proposes which particular entity participates in an observation, representation, record, expression, or other semantic relationship. This builds on the standardised meaning of identify as referring without ambiguity and of an identifier as a value by which something is identified [@iso_39075_2024].
Example: The depicted person in pict003.jpg is identified as Hilda Grīva.
Identifier
In a register, an identifier is a linguistically independent sequence of characters capable of uniquely and permanently identifying that with which it is associated (ISO 2026). More generally, it is a “value by which something is identified.” (ISO/IEC 2024b)
Indeterminate
A state in which a requested property cannot presently have a determinate value because the event or condition that would determine it has not occurred. Indeterminacy is distinct from unknown knowledge. A living person, for example, does not presently have a death year that remains to be discovered.
Inference
ISO/IEC 22989 defines inference generally as reasoning by which conclusions are derived from known premises, where premises in AI may include facts, rules, models, features, or raw data (ISO/IEC 2022).
In the Review Algebra, inference is formally defined as the rule-governed generation of candidate assertions \(C_k^ = \bigcup_{r \in \mathcal{G}} g_r(S_k, R, \sigma)\) operating strictly over the stabilised semantic state \(S_k\)._Intermediate candidates generated within cycle \(k\) are explicitly barred from serving as premises for further inference during the same cycle. Example: Stabilised authorship and life dates may satisfy the conditions of a rule that generates a candidate copyright conclusion.
Interpretation
A knowledge-producing activity through which an agent applies linguistic, disciplinary, historical, legal, cultural, or other specialised competence to evidence or represented knowledge.
Example: A textile researcher interprets the depicted garment as Livonian folk dress.
Knowledge graph
A knowledge representation that uses a graph-structured data model to represent and operate on data (ISO 2023).
Knowledge-producing activity
A provenance-bearing human or computational activity through which semantic assertions are generated, examined, interpreted, reconciled, inferred, reviewed, or otherwise transformed. Different knowledge-producing activities may use different evidence, methods, agents, competences, rules, and permissions.
Knowledge representation The process or result of encoding and storing knowledge in a knowledge base [ISO/IEC 2382-28:1995; ISO (2023)].
Knowledge source source of information from which a knowledge base has been created for a specific kind of problem [ISO/IEC 2382-28:1995]. (ISO 2023)
Observation
Observation may be performed by human or computational agents. Computational performance of an activity does not by itself imply autonomy: ISO/IEC 22989 distinguishes automated operation under specified conditions from autonomy, in which a system can modify its intended domain of use or goal without external intervention, control, or oversight (ISO/IEC 2022).
Example: fscontext observes that pict003.jpg has MIME type image/jpeg.
Observe as
A practical expression for specifying the evidentiary frame under which an observation activity examines evidence: for example, observing a JPEG as a digital manifestation, as a photograph, or as evidence concerning a depicted garment or activity. It constrains what subjects and predicates are meaningful within the activity without becoming part of the elementary domain assertion.
Permission
A condition determining whether an agent or system is authorised to observe, use, disclose, infer, transform, or act upon particular knowledge or evidence. Permission is distinct from epistemic warrant: an assertion may be well supported while its use or disclosure remains restricted.
Provenance
Information concerning the activities, agents, evidence, methods, rules, models, competences, and other circumstances through which semantic assertions or semantic states were produced. Provenance is not merely an audit trail: where meaning or warrant depends on linguistic, disciplinary, evidentiary, or methodological context, provenance contributes to understanding the assertion itself.
Provenance knowledge
Knowledge describing how domain assertions and semantic states were generated, reviewed, inferred, finalised, or otherwise transformed. Provenance knowledge is represented separately from domain knowledge even where the two are closely connected.
Example: What Hilda Grīva wears was observed by a curator examining pict003.jpg.
Reconciliation
A knowledge-producing activity that proposes or establishes relationships between locally represented entities or expressions and entities or identifiers in external knowledge representations, authority systems, registers, or other namespaces. A namespace is understood as a classification scheme that permits a given name to identify multiple objects distinguished by context (ISO/IEC 2024b).
Review
The Review Algebra operation through which an existing candidate assertion, or one or more of its components, is examined and a semantic value together with an explicit review status is returned.
Example: A curator corroborates the candidate classification Livonian folk dress.
Review status
An explicit representation of what happened to a candidate during review, distinct from the semantic value returned by that review. Core statuses include corroborated, falsified, and deferred.
Scope
An explicit qualification of the semantic or applicability context within which an assertion is intended to hold. Scope may be retained where interpretation or validity depends on context and may contribute to determining whether a rule is applicable. Scope is distinct from the evidentiary frame of an observation activity and from provenance more generally.
Semantic stabilisation
The process through which incomplete, ambiguous, heterogeneous, contested, or partially conflicting semantic assertions become sufficiently stable for a specified purpose while preserving relevant uncertainty, provenance, competing interpretations, and opportunities for subsequent revision. Stabilisation is purpose-relative and does not imply permanent truth or universal agreement.
Semantic state
The set of semantic assertions available as represented knowledge at a particular stage of semantic production. A stabilised semantic state may provide input to further observation, interpretation, reconciliation, inference, candidate generation, review, or finalisation.
Taxonomy scheme of categories and subcategories that can be used to sort and otherwise organize itemized knowledge or information [ISO 25964-2:2013, definition 3.83 modified] (ISO 2013)
Unknown
A state in which relevant knowledge may exist or a property may have a determinate value, but that knowledge has not yet been established. Unknown knowledge may therefore justify further observation, investigation, reconciliation, interpretation, or review.
Unresolved dependency
A represented requirement for another assertion or condition that must be established before a candidate can be generated, an inference can proceed, or an operational decision can be made.
Example: A publication decision remains blocked until the recording date is established.
Review Algebra
The formal and computational framework developed in the companion paper for representing candidate generation, review, and finalisation as provenance-bearing transformations between semantic states. The Review Algebra deliberately does not prescribe the epistemology, evidence, competence, or domain models of the knowledge-producing activities that supply its candidates.
Semantic knowledge production
The broader process through which human and computational agents produce semantic knowledge by observing evidence, interpreting it under relevant competences, identifying and reconciling entities, applying explicit rules, reviewing candidate assertions, and progressively stabilising knowledge for particular purposes. The Review Algebra formalises a deliberately narrower part of this process.
Semantic knowledge production
Introduction
Semantic knowledge is produced through activities in which human and computational agents examine evidence, apply specialised competences, reconcile external knowledge, interpret linguistic expressions, or derive consequences from knowledge already represented. The resulting assertions do not emerge with a uniform epistemic status: they may be provisional, contested, corroborated, dependent on other assertions, or constrained by the competence and authority of the agents that produce them.
This paper examines these knowledge-producing activities as a computational and epistemic problem. Its purpose is not to provide a general epistemology of knowledge, but to identify recurring requirements that arise when heterogeneous human and computational agents contribute semantic assertions to a shared knowledge system. In particular, it asks what must be preserved if assertions are to remain connected to the evidence, activities, competences, rules, and purposes through which they were produced.
The central argument is that semantic knowledge production requires several distinctions that conventional metadata workflows can easily obscure. Evidence must be distinguished from the semantic objects about which it provides knowledge; observation must be distinguished from review and inference; identity must be distinguished from linguistic expression; and the production of an assertion must remain distinguishable from its subsequent corroboration, falsification, or use. Competence, permission, applicability, and unresolved dependencies must likewise be represented where they constrain what an agent can responsibly establish. At the same time, the current semantic state may condition which knowledge-producing activities become meaningful or applicable next.
These requirements motivate the development of a complementary Review Algebra, which formalises one part of this broader process: the generation of candidate assertions, their review, their finalisation into subsequent semantic states, and the provenance of those transformations. The algebra is developed separately as a formal component of the wider framework. The present paper instead concentrates on the semantic-production problem from which those formal requirements arise and on their application to cultural heritage and digital humanities.
The Livonian collection provides a running example. The activities considered below do not define a mandatory workflow. Rather, they illustrate different forms of scholarly and professional knowledge production in which evidence is examined, assertions are produced, specialised competences are applied, and the resulting knowledge may subsequently become the object of review or provide the basis for further knowledge-producing activities.
For exposition, the examples are organised into three broad families. The first concerns empirical observation and domain knowledge: agents examine evidence and produce assertions according to particular methods and domains of competence. The second concerns linguistic knowledge and interpretation: reference, multilingual expression, translation, and linguistic structures become objects of knowledge production in their own right. The third concerns the conditions and limits of semantic knowledge production: competence, permission, applicability rules, and unresolved dependencies constrain what particular agents may responsibly establish, use, or infer.
The boundaries between these families are deliberately permeable. Linguistic competence may be required to interpret empirical evidence; legal review may depend on historical or linguistic knowledge; and inference may operate over assertions produced by any of these activities. What unifies them is not a prescribed scholarly method but the requirement to preserve the relationship between semantic assertions and the provenance-bearing activities through which they are produced, examined, and progressively stabilised. This allows subsequent knowledge production to proceed from an intelligible semantic state while preserving how that state came to be established.
The Livonian collection provides a running example. The activities considered below do not define a mandatory workflow. They illustrate different forms of scholarly and professional knowledge production in which evidence is examined, assertions are produced, specialised competences are applied, and the resulting knowledge may subsequently become the object of review.
For exposition, the examples are organised into three broad families. The first concerns empirical observation and domain knowledge: agents examine evidence and produce assertions according to particular methods and domains of competence. The second concerns linguistic knowledge and interpretation: reference, multilingual expression, translation, and linguistic structures become objects of knowledge production in their own right. The third concerns the conditions and limits of knowledge production: competence, permission, applicability rules, and unresolved dependencies constrain what particular agents may responsibly establish, use, or infer.
The boundaries between these families are deliberately permeable. Linguistic competence may be required to interpret empirical evidence; legal review may depend on historical or linguistic knowledge; and inference may operate over assertions produced by any of these activities. What unifies them is not a prescribed scholarly method but the requirement to preserve the relationship between semantic assertions and the provenance-bearing activities through which they are produced, examined, and progressively stabilised.
The framework is being developed through an evolving ecosystem of open-source reference implementations rather than a single software package. The components address different parts of semantic knowledge production while remaining interoperable through common representations of assertions, review, provenance, and inference. Although several implementations are developed in the R statistical environment, the underlying concepts are language-independent and can be implemented in other programming environments.
fscontext supports empirical observation of digital resources and their contexts. It reconstructs evidence from local file systems, Git repositories, web archives, and other digital storage environments while preserving provenance about the resources and activities through which observations are made. The library is available through the Comprehensive R Archive Network (CRAN). Documentation and tutorials are available at https://fscontext.dataobservatory.eu/.
betwixt provides a reference application for semantic knowledge production and review. It organises candidate assertions into reviewable workspaces and preserves reviewed semantic values, review statuses, and the provenance-bearing activities through which semantic knowledge is progressively stabilised. The software is currently under active development and is not yet publicly released.
review provides a reference implementation of the Review Algebra. Building on relational algebra and the tidy data paradigm, it represents review operations and their provenance using vectorised data structures while maintaining compatibility with RDF and provenance standards. The library is currently in an early stage of development. Documentation is available at https://review.dataobservatory.eu/.
The ecosystem is complemented by libraries demonstrating how reviewed semantic knowledge can participate in broader analytical workflows. dataset provides interoperability with reproducible statistical workflows and research data management and has undergone peer review through both CRAN and rOpenSci. retroharmonize supports the retrospective harmonisation of statistical and business surveys across time, languages, and changing questionnaires, while regions supports historically consistent harmonisation of geographical entities and names.
Together these reference implementations illustrate the separation between the formal algebra and its applications. The Review Algebra provides a general representation of review and stabilisation, while applications such as betwixt use that representation within particular processes of semantic knowledge production.
Semantic production as a progressive programme
Semantic knowledge production is not a single transformation from evidence into a final representation. Assertions are produced under particular evidential, methodological, linguistic, and institutional conditions; they may subsequently be challenged, refined, connected to other assertions, or used to generate further questions.
Semantic knowledge production therefore begins from a problem of fallibility. Observations may be incomplete, interpretations may depend on historically or institutionally situated vocabularies, external knowledge may conflict, and inferences may depend on premises that are themselves revisable. A computational knowledge system must therefore be able to preserve useful assertions without treating its current representation as final.
Fallibility is compounded by semantic pluralism. Different agents may legitimately describe the same objects through different disciplinary, linguistic, institutional, or computational vocabularies. Semantic production cannot therefore be reduced to selecting a single representation and progressively eliminating alternatives. Candidate alternatives, disagreements, mappings, and reinterpretations must remain representable, while some assertions nevertheless become sufficiently stabilised for subsequent scholarly, organisational, or computational use.
Lakatos’s conception of a research programme provides a useful structural analogy. Scientific knowledge can progress while maintaining a body of commitments against which anomalies, alternatives, and auxiliary propositions are investigated (Lakatos, Worrall, and Zahar 1976). Applied to semantic knowledge production, this suggests a programme that maintains sufficiently stabilised semantic commitments while allowing them to be extended, challenged, qualified, or revised. The analogy is not exact: stabilised semantic assertions are not insulated from falsification. They are protected from uncontrolled modification, but not from criticism or revision.
A related constraint follows from Davidson’s account of radical interpretation: interpretation cannot in general presuppose independently settled meanings, beliefs, and intentions, because these are established interdependently through the interpretive process (Davidson 1973). Fallibility therefore implies that no semantic state should be treated as final. Semantic pluralism requires alternatives to remain representable. Criticism requires current commitments to remain challengeable. Stabilisation nevertheless allows some assertions to acquire operational authority. Progress occurs when a resulting semantic state changes what can meaningfully be investigated next.
In this limited sense, semantic knowledge production can be understood as a progressive research programme. In the systems architecture developed here, its stabilised and exploratory roles are separated between a Production Graph and a Candidate Graph. The Production Graph contains semantic commitments currently stabilised for production-level use. The Candidate Graph provides a space in which observations, interpretations, reconciliations, inferences, proposed corrections, and competing alternatives can be represented without immediately altering those commitments. The Production Graph can therefore be understood, within the limited Lakatosian analogy, as the protected semantic core of the programme: protected from uncontrolled change, yet deliberately exposed to governed challenge from candidate knowledge.
This epistemological structure can be translated into systems engineering as a governed progression between semantic states. From a current stabilised state \(S_k\), authorised candidate-generating activities produce reviewable candidates \(C_k^*\). Review activities, allocated according to relevant competence, authority, applicability, and the nature of the candidate, produce reviewed assertions and explicit review outcomes \(Q_k\). Finalisation determines which results acquire production-level authority in the subsequent state \(S_{k+1}\).
The resulting progression can be represented schematically as
\[S_k \longrightarrow \text{governed candidate generation} \longrightarrow C_k^* \longrightarrow \text{allocated review} \longrightarrow Q_k \longrightarrow \text{finalisation} \longrightarrow S_{k+1}.\]
This progression is not merely iterative. The resulting state \(S_{k+1}\) is not only the output of the preceding activities: it changes the represented knowledge from which subsequent semantic production proceeds and thereby conditions which questions, candidate-generating activities, and reviews become meaningful, applicable, or permissible next. Progress is therefore not measured simply by the accumulation of assertions, but by the capacity of increasingly stabilised knowledge to support further well-founded knowledge production.
Human governance concerns the Production Graph as a whole without requiring every assertion or computational operation to be individually reproduced or inspected by a human. Human review may operate directly on particular candidates or at review frontiers where the premises, rules, scope, competence, and conditions governing larger classes of computationally produced assertions can be assessed. The systems-engineering problem is therefore to support extensive and exploratory candidate generation while ensuring that the acquisition of production-level semantic authority remains human-governed, provenance-bearing, and open to subsequent criticism.
Taken together, these requirements imply three architectural invariants whose computational consequences are formalised more precisely in the companion Review Algebra. First, inference operates over stabilised knowledge: candidate-generating inference rules operate on the current stabilised semantic state \(S_k\), and candidates generated during the current production cycle cannot themselves become premises for further inference within that cycle. This prevents unreviewed assertions from recursively acquiring consequences before they have crossed an appropriate review and finalisation frontier.
Second, candidate generation remains logically separated from the stabilised production state. Human and computational agents may generate, compare, evaluate, and prune extensive candidate structures without thereby modifying the Production Graph. Implementations may enforce this separation through distinct graphs, namespaces, permissions, transactional boundaries, or sandboxed environments; the architectural requirement is that candidate generation alone cannot confer production-level semantic authority.
Third, changes to the stabilised semantic state occur through explicit, provenance-bearing transitions. A subsequent state \(S_{k+1}\) must remain traceable to the candidate-generating, review, and finalisation activities through which it was established. The progression is therefore auditable without requiring the underlying human judgements themselves to be deterministic.
Together, these invariants protect the distinction between exploratory semantic production and production-level semantic authority. They allow candidate generation to remain extensive, heterogeneous, and open to competing alternatives while ensuring that the Production Graph develops through governed and inspectable transitions. In this sense, the systems architecture implements the progressive character of the semantic programme: stabilised knowledge is protected without being made immutable, candidate knowledge can challenge it without silently modifying it, and each governed transition establishes the semantic state from which further knowledge production can proceed.
The same architecture also makes the Production Graph inspectable in the opposite direction. Because production-level assertions enter through explicit review, finalisation, or authorised derivation, their authority lineage can in principle be traced backwards through the provenance of those transformations. Human governance of the Production Graph therefore need not mean that every assertion has been directly inspected by a human. It requires instead that production-level assertions remain accountable to an appropriate governance frontier: either directly through human review or indirectly through authorised transformations whose premises, rules, scope, and conditions of application are themselves governed.
Three planes of semantic knowledge production
The architecture developed in this paper can be understood heuristically as operating across three related planes: a Domain Plane, a Provenance Plane, and a Normative Plane. The distinction is analytical rather than ontological. Each plane identifies a different role played by assertions and activities within a governed semantic system.
The Domain Plane concerns assertions about the objects, persons, activities, relationships, and other subject matter represented by the system. The Provenance Plane concerns how those assertions were produced: the observations, evidence, agents, methods, and review activities through which candidate claims are examined and stabilised. These two planes form the principal subject of the present paper and of the Review Algebra. The distinction is important because the review workflow is designed to preserve not only stabilised values but also the activities through which they acquired their warranted status. :contentReferenceoaicite:0
A third Normative Plane becomes relevant when stabilised knowledge is used to determine what an agent may, must, or must not do. It introduces questions of permission, prohibition, obligation, authority, purpose, and applicable policy. The Normative Plane is introduced here because it demonstrates an important consequence of semantic stabilisation: reviewed semantic knowledge can become an input to governed action. Its general formalisation, however, lies outside the scope of the present paper.
The distinction can be illustrated with two photographs and a hypothetical publication decision.
Photo A depicts a person. Successive knowledge-producing activities may establish that a person is present, identify that person, and subsequently establish the biographical assertion that the person is deceased. Photo B follows a shorter path: observation establishes that no person is depicted. These are developments in the Domain Plane. They concern what can be asserted about the photographs and their depicted subjects.
The production of these assertions simultaneously creates a history in the Provenance Plane. The system may record which photograph was inspected, which agent performed the review, what evidence was used, which assertions were proposed, and through which activities they acquired their stabilised status. Provenance therefore does not add another fact about the depicted world; it provides the warrant through which the authority of a domain assertion can be assessed.
Neither domain knowledge nor its provenance, however, is itself a permission to act. The assertion that a depicted person is deceased remains a domain assertion, however well supported it may be. A normative conclusion becomes available only when sufficiently warranted domain knowledge is evaluated under a policy applicable to a contemplated action.
Suppose, for example, that the workbench supports a prospective publication activity under a hypothetical policy concerned with permission from depicted persons. For Photo B, the relevant condition is inapplicable because no person is depicted. For Photo A, the condition is relevant because a person is depicted, but the stabilised assertion that the person is deceased may satisfy a policy rule under which permission from that depicted person on grounds of personality rights is not required. The two photographs can therefore reach the same practical outcome through different semantic paths.
This example is deliberately illustrative rather than a statement of law. Personality and privacy rights, post-mortem protections, archival restrictions, copyright, contractual obligations, sensitive-data rules, and other constraints differ between jurisdictions and circumstances. The example assumes only a hypothetical policy environment in which the established death of a depicted person removes the particular requirement to obtain that person’s permission for the contemplated publication. Other permissions, obligations, or prohibitions may remain.
The distinction can be expressed schematically as:
stabilised domain state + warranted provenance + applicable policy → normative conclusion
The important point is that these terms are not interchangeable. person1 status deceased does not mean publication permitted. The first is a claim about the world. Its provenance explains how that claim was established. A policy supplies the rule under which the stabilised claim becomes relevant to a contemplated action. Only their evaluation together can produce a normative conclusion.
The second figure places this distinction within the broader Review Algebra. It separates four operations that would otherwise be easy to conflate: reviewing domain claims, drawing domain inferences from stabilised states, recording the provenance that warrants those states, and applying policy to derive a normative decision.
Figure 2 makes two different forms of inference explicit. Domain inference remains within the Domain Plane: stabilised semantic states can support further classifications or relationships, such as assigning photographs to collections. Normative inference crosses into a different plane: stabilised domain states, together with their warrant and an applicable policy, are evaluated in relation to a contemplated action. The output is therefore not another description of the photograph but a decision concerning what may be done with it.
This distinction also explains why semantic production is progressive rather than merely accumulative. Establishing that a person is depicted makes person identification meaningful. Identification makes biographical reconciliation meaningful. Stabilising the person’s death may make a particular policy condition decidable. Each stabilised state can therefore change not only what the system knows, but also which subsequent questions are applicable and which operations can be warranted on the basis of that knowledge.
The contrast between the two photographs makes this especially visible. Their relevant normative outcomes may coincide, yet their semantic histories do not. For Photo B, permission from a depicted person is inapplicable because the relevant subject is absent. For Photo A, the question becomes applicable and is subsequently resolved through additional knowledge production. Treating both merely as publication allowed would erase precisely the semantic and provenance distinctions required to explain and audit the decision.
This separation is also consistent with the broader architecture developed here: semantic stabilisation reduces uncertainty without turning the system into an ontology or policy engine. The Review Algebra provides a lightweight mechanism for creating reviewable claims and documenting successive review activities and their provenance; interpretation and downstream inference remain separable operations.
The present paper therefore develops primarily the first two planes: the production and stabilisation of domain assertions and the provenance through which their authority can be assessed. The Normative Plane is retained as an explicit architectural frontier. A fuller treatment would need to address the representation and provenance of policies themselves, conflicts and precedence between policies, purpose-relative permissions and obligations, the authority of decision-making agents, and the conditions under which normative conclusions acquire sufficient warrant to govern action.
Empirical observation and domain knowledge
Empirical semantic knowledge production begins with evidence, but the evidence does not determine in advance which semantic assertions should be produced from it. An agent examines evidence through a provenance-bearing activity and produces or reviews semantic assertions. The activity is constrained by the method used, the competence of the agent, the relationship between the evidence examined and the semantic object about which knowledge is being produced, and the semantic state from which the activity proceeds.
Different agents may examine the same evidence and legitimately produce different kinds of knowledge. A technical agent may inspect the media type and embedded metadata of a digital file; a curator may identify the kind of documentary or cultural object made available through it; a dress historian may classify clothing visible in a photograph; and a musicologist may identify features of a recorded performance. Their activities differ not only in method and competence, but also in what the evidence is being examined as evidence of and which questions are meaningful at the current stage of semantic production.
The examples in this section therefore distinguish three related questions. First, what is the evidence and what kind of object does it make available for examination? Second, given the current semantic state, what candidate-generating or review activities are meaningful and applicable? Third, how can the resulting knowledge-producing activities and their relationships to evidence, agents, methods, generated assertions, and subsequent semantic states be represented through provenance?
Evidence and the object of observation
The term source is insufficiently precise for empirical semantic knowledge production. A source may mean a file, archival source, database, publication, source system, or provenance resource. Here, evidence denotes an entity that participates in a knowledge-producing activity because an agent can examine it in order to establish or challenge semantic assertions.
Evidence is not necessarily identical to the semantic object about which knowledge is produced. Consider a JPEG file available to a curator or computational system. The file is an observable entity in a computational environment, but it may also provide access to a documentary object, such as a historical photograph, which in turn may represent people, garments, places, objects, or activities.
Semantic knowledge production therefore does not assume that reviewed or labelled data constitute an epistemically privileged ground truth. ISO/IEC 22989 uses ground truth technically for the target-variable value attached to labelled input data and explicitly cautions that this does not imply correspondence with the real-world value of that variable (ISO/IEC 2022).
For documentary and archival evidence, Records in Contexts (RiC) provides a useful starting distinction between a Record Resource and its Instantiation. A historical photograph may be treated as a documentary record resource while a particular digital representation provides an instantiation through which that resource can be encountered computationally (Archives Expert Group on Archival Description 2023).
Record Resource
historical photograph
|
| has instantiation
v
Instantiation
photo003.jpg
This documentary distinction should not be confused with the further semantic movement from the documentary object towards what it represents. While examining the same accessible digital evidence, a knowledge-producing activity may concern different semantic objects:
| evidentiary frame | object under consideration | example assertion |
|---|---|---|
| digital manifestation | photo003.jpg |
media_type = image/jpeg |
| documentary object | photograph003 |
instance_of = photograph |
| represented entity | person1 |
instance_of = person |
| represented activity | activity1 |
instance_of = weaving activity |
These assertions may ultimately be connected, but they do not have the same subject. The MIME type belongs to the digital manifestation; the documentary classification concerns the photograph; and assertions concerning people or activities represented by the photograph concern still other semantic objects.
Also, the originator of data should not be conflated with the entity represented by those data. ISO/IEC 5259-4 explicitly distinguishes the data originator—the party that created the data and can have rights—from persons or legal entities mentioned, described, or otherwise associated with those data (ISO/IEC 2024a).
The distinction can therefore be summarised heuristically as
digital manifestation
|
| instantiates / makes available
v
documentary or representational object
|
| represents
v
represented subject matter
├── entity
└── activity
These are not proposed as additional components of the elementary semantic assertion, nor as a universal ontology of evidence. They identify recurring evidentiary frames within which empirical knowledge-producing activities can be designed.
The distinction is especially important for predicates whose interpretation depends strongly on their semantic subject. Allowing an activity to move implicitly between a digital manifestation, a documentary object, and represented subject matter can produce assertions that are individually plausible but semantically incorrect. An empirical activity should therefore make sufficiently clear what the available evidence is being examined as evidence of.
What it is meaningful to examine the evidence for may itself change as semantic knowledge becomes stabilised. Establishing that a digital resource makes a photograph available for examination makes questions about depicted subject matter applicable; establishing that the photograph depicts a person may make identification, biographical reconciliation, clothing classification, or other domain-specific observations applicable. The current semantic state therefore does not merely contain the results of earlier knowledge production: it helps determine the candidate-generating activities that become meaningful next.
Observation and review
Empirical semantic knowledge production involves two closely related but distinct activities: observation and review.
Observation examines evidence and produces explicit candidate semantic assertions about what an agent can establish from that evidence. Review begins from an assertion that has already been made explicit and determines what happens to that assertion when it is critically examined.
The distinction concerns the relationship between an activity and the current semantic state rather than whether the activity is performed by a human or computational agent. A filesystem observation performed by fscontext and ExifTool, an acoustic analysis performed by a music-information-retrieval system, and a visual observation performed by a human curator can all generate candidate assertions from evidence.
| activity | agent | evidence | result |
|---|---|---|---|
| observation | ExifTool | JPEG file | candidate assertions about embedded metadata |
| observation | MIR system | audio recording | candidate assertions about acoustic or musical features |
| observation | curator | photograph | candidate assertions about depicted entities or activities |
| observation | specialist | photograph | candidate assertions requiring specialist domain competence |
Observation can therefore be represented schematically as
\[E \xrightarrow{\mathcal{O}} C^{*},\]
where an observation activity \(\mathcal{O}\) examines evidence \(E\) and produces one or more candidate assertions \(C^{*}\).
Review has a different starting point. A candidate assertion is already available, whether it was generated through observation, inference, import, reconciliation, or another process. The reviewer examines that assertion using whatever evidence and knowledge are appropriate to the review activity.
| activity | input assertion | evidence | result |
|---|---|---|---|
| review | Hilda Grīva — clothing — Livonian dress |
photograph | corroborated |
| review | Hilda Grīva — clothing — Latvian dress |
photograph | falsified; alternative proposed |
| review | candidate identity for depicted person | photograph; authority record | corroborated |
| review | candidate life date | catalogue; scholarly source | corrected and corroborated |
The empirical act involved in observation and review may nevertheless be very similar. A curator may examine the same photograph using the same domain competence. What differs is the relationship of that activity to the existing explicit knowledge.
Suppose, for example, that the following candidate has already been generated:
| subject | predicate | candidate |
|---|---|---|
photograph1 |
depicts |
Hilda Grīva |
While examining the photograph, the curator may corroborate this assertion but also observe information that has not previously been represented:
| operation | subject | predicate | value | result |
|---|---|---|---|---|
| review | photograph1 |
depicts |
Hilda Grīva |
corroborated |
| observation | Hilda Grīva |
clothing |
Livonian folk dress |
new candidate |
| observation | Hilda Grīva |
participates_in |
weaving |
new candidate |
A single scholarly examination may therefore perform both functions. It may review existing assertions and generate new candidate assertions through observation. The operations remain distinguishable even when they arise from the same examination of the same evidence by the same agent.
This distinction is important for human-centred semantic knowledge production. A system that permits an expert only to corroborate, falsify, defer, or correct assertions already presented for review implicitly limits the expert to the candidate-generation competence of the preceding system. Allowing new observations preserves the expert’s ability to identify relevant knowledge that was absent from the existing representation and, consequently, to alter the subsequent direction of semantic production.
Human participation therefore cannot be reduced to checking machine-generated candidates. An expert may challenge a candidate, introduce an alternative, generate previously absent observations, or establish knowledge that changes which questions become applicable next. Human review and observation participate in the progression of the semantic programme rather than merely validating the output of preceding computational activities. Observation and review therefore answer complementary questions:
Observation: What can this agent establish from the available evidence?
Review: What happens to this existing assertion when this agent critically examines it?
Observation is consequently one important mechanism of candidate generation rather than an alternative to the Review Algebra. Inference, reconciliation, import, and computational proposal may generate candidates through other mechanisms. Review begins once a candidate assertion is available for critical examination.
| operation | principal input | empirical evidence | principal output |
|---|---|---|---|
| observation | evidence | constitutive | candidate assertion |
| inference | stabilised assertions and rule | not necessarily | candidate assertion |
| reconciliation or import | external representation | not necessarily | candidate assertion |
| review | candidate assertion | where required | reviewed value and review status |
| finalisation | reviewed assertion and policy | not necessarily | stabilised assertion |
These operations occupy different positions in the progression introduced above. Observation, inference, reconciliation, import, and other computational or scholarly activities can generate candidates, but their applicability may depend on the current stabilised semantic state. Review determines what happens when those candidates are critically examined, while finalisation governs which reviewed results acquire production-level authority. Candidate generation is therefore neither unrestricted graph expansion nor merely a preliminary step before review: it is itself part of the governed progression of semantic knowledge production.
The distinction is analytical rather than a requirement that scholarly work be fragmented into separate human tasks. One curatorial examination may encompass several observations and reviews, while their outputs remain individually represented and provenance-bearing. What matters for the semantic programme is that the resulting assertions and activities remain distinguishable so that their contribution to subsequent semantic states can be governed and reconstructed.
Provenance of knowledge-producing activities
The distinction between evidence, observation, and review clarifies the role of provenance. W3C PROV provides a general model for representing the activity through which an entity is used, an agent participates, and another entity is generated (Moreau and Missier 2013; Lebo et al. 2013).
A simple observation can be represented schematically as
A review has a related structure, but the activity also takes an existing candidate assertion as an input:
PROV therefore records how knowledge was produced or reviewed, but it does not by itself determine what the evidence represents or what the generated assertion means in a particular scholarly domain.
The relationship between an activity and its evidence may also need to be qualified. An activity does not merely use photo003.jpg; it may examine that evidence specifically with respect to its digital manifestation, its documentary character, a depicted entity, or a depicted activity. At the level of a human-centred interface, this can be expressed simply as an instruction such as:
Evidence: photo003.jpg
Observe as: depicted activity
Agent: curator
Method: visual observation
The phrase observe as identifies the evidentiary frame adopted by the activity. It does not become another component of the elementary semantic assertion. Rather, it characterises the relationship between the knowledge-producing activity and the evidence it uses. In a provenance representation, such a relationship can be modelled through a qualified use of the evidence and an appropriate controlled vocabulary for its evidentiary role.
This also explains why the evidentiary frame and conditions of applicability should normally be established when an activity is designed rather than changed independently for individual assertions. A curator asked to examine a photograph for depicted activities should be able to review existing assertions and add missing observations within that frame. Moving instead to technical properties of the JPEG or to authority information about a depicted person constitutes a different knowledge-producing activity, potentially requiring a different method or competence.
RiC, PROV, and domain-semantic models consequently have complementary responsibilities.
RiC can describe documentary evidence and distinguish record resources from the instantiations through which they are encountered and preserved.
PROV can describe the activities, agents, evidence, methods, and generated resources involved in producing and reviewing semantic knowledge.
Domain models such as CIDOC CRM can describe the heritage objects, people, places, events, activities, representations, and other subject matter about which assertions are produced.
The Review Algebra can represent how candidate assertions are generated, reviewed, finalised, and progressively incorporated into stabilised semantic states.
Reference applications such as
betwixtcan operationalise these distinctions through constrained observation and review activities appropriate to particular evidence, methods, and domains of competence.
The resulting architecture is therefore a division of responsibilities rather than an attempt to construct a single ontology for semantic knowledge production.
RiC
documentary evidence and instantiation
|
v
PROV
evidence --> knowledge-producing activity <-- agent
|
v
generated assertion
|
v
Review Algebra
candidate --> review --> finalisation
|
v
domain knowledge
CIDOC CRM / other domain models
The diagram should not be read as a strict processing pipeline. RiC, PROV, the Review Algebra and domain models describe different aspects of the same knowledge-producing process. A domain assertion may itself become evidence or input for subsequent activities, and stabilised knowledge may generate further candidates through inference or renewed observation.
This separation is important epistemically as well as computationally. The fact that an assertion was produced while examining a particular digital resource does not make the resource, the documentary object, and the subject matter represented by that object identical. Preserving these distinctions allows the provenance chain to answer not merely what is asserted?, but also what was examined, what was it examined as evidence of, who performed the activity, by what method, and how did the resulting assertion enter the current semantic state?
The examples below apply this architecture first to taxonomic classification and then to observation and identification. They illustrate how empirical evidence becomes explicit semantic knowledge and how stabilised knowledge can progressively determine which further questions, methods, and specialist competences are applicable.
Taxonomic review
The first application illustrates how taxonomic classification can generate candidate assertions and, once stabilised, determine which subsequent inferences and review questions are applicable.
Taxonomic review answers the question
What kind of object is under review?
This question is more subtle than identifying a file type. In the running example, the system already observes that object1, object2, and object3 are JPEG files, while object4 and object5 are WAV files. These are technical observations about digital resources. They may contribute to subsequent candidate generation, but they do not by themselves establish the semantic class of the documentary or cultural object made available through the file.
For example, an observation activity may examine object3 as a digital manifestation and establish its media type as image/jpeg. That technical observation may subsequently contribute to the generation of a different candidate assertion concerning the documentary object made available through the file:
\[(\texttt{object3},\texttt{instance\_of},\texttt{photograph}).\]
The classification may draw on the observed media type together with embedded metadata, catalogue information, collection context, or other available evidence. It remains a candidate rather than a logical consequence of the JPEG media type.
Likewise, technical observations concerning the WAV resources may contribute to candidate classifications of the corresponding documentary objects as sound recording.
| claim_id | subject_candidate | predicate_candidate | value_candidate | value_reviewed | value_status |
|---|---|---|---|---|---|
| claim1 | object1 |
instance_of |
photograph |
photograph |
corroborated |
| claim2 | object2 |
instance_of |
photograph |
photograph |
corroborated |
| claim3 | object3 |
instance_of |
photograph |
photograph |
corroborated |
| claim4 | object4 |
instance_of |
sound recording |
sound recording |
corroborated |
| claim5 | object5 |
instance_of |
sound recording |
sound recording |
corroborated |
In this example, the candidate classifications survive review and are eligible for finalisation. Their provenance records both how the candidates were generated and the review activity through which they were corroborated.
The distinction between candidate generation and review remains important. Technical observations may contribute to semantic classifications at scale, while review determines whether those candidate classifications are appropriate for the objects under investigation.
Taxonomic stabilisation then enables a different operation. Once
\[(\texttt{object3},\texttt{instance\_of},\texttt{photograph})\]
has entered the stabilised semantic state, object3 belongs to a class for which particular properties, relations, and review questions are applicable. A photograph may support assertions concerning its creator, date, technique, physical form, depicted agents, depicted places, or depicted activities. A sound recording may support assertions concerning a recorded performance, performers, musical works, composers, languages, instruments, or recording events.
This is an inference over class membership. The class photograph provides a domain of applicability for rules concerning photographs. Such a rule may be expressed schematically as
\[x\ \texttt{instance\_of}\ \texttt{photograph} \;\land\; \rho_{\text{photograph}}(p) \quad\Longrightarrow\quad p\ \text{is applicable to}\ x,\]
where \(\rho_{\text{photograph}}\) represents an explicit applicability rule. Class membership alone does not establish a value for property \(p\). It establishes only that candidate generation or review involving \(p\) may be applicable.
Taxonomic review therefore does more than add another descriptive property. Once finalised, a classification changes the domain over which subsequent candidate-generation and review operations are applicable. It allows the system to exclude questions that are meaningless for a particular class and to expose those for which further observation, inference, or review may be productive.
This is one of the mechanisms through which the review algebra can scale human semantic review: deterministic inference over stabilised classifications can eliminate entire classes of irrelevant questions before specialised human judgement is required.
In a practical implementation, taxonomic review can begin with resources captured from collection portals and organised into a common review workspace. Technical observations, catalogue metadata, collection context, and automated extraction can generate initial candidate classifications that are subsequently reviewed.
Suppose that a collection contains one hundred resources already stabilised as photographs. The classification photograph satisfies the applicability condition for candidate generation and review concerning what the photographs depict. Observation may then generate candidate assertions that ninety depict buildings, seven depict people, and three depict boats; review determines whether those candidates are corroborated.
These reviewed classifications can in turn determine the applicability of more specialised questions.
- Photographs depicting people may become eligible for person identification, clothing analysis, namespace review and, where relevant, personality-rights assessment.
- Photographs depicting boats may become eligible for vessel or vessel-type identification without requiring personality-rights review.
- Photographs depicting buildings may become eligible for architectural, geographical, or historical identification while excluding questions that are meaningful only for persons or performances.
The important mechanism is recursive. A reviewed and finalised assertion enters the stabilised semantic state. Explicit inference rules use that state to determine which properties, candidate-generation processes, and review questions are applicable next. Observation or other candidate-generation processes then produce the semantic values that those reviews examine.
The cycle therefore remains
\[C_k \xrightarrow{\mathcal{G}_{\gamma}} C_k^{*} \xrightarrow{\mathcal{R}} Q_k \xrightarrow{\mathcal{F}_{\phi}} C_{k+1}.\]
What changes between cycles is the knowledge available to determine the applicability of subsequent operations. Taxonomic stabilisation reduces semantic uncertainty while allowing deterministic inference to narrow the questions presented for further review. Human attention can consequently be directed towards increasingly specific questions rather than repeatedly applied to every possible property of every object.
Observation, interpretation, and identification
Once relevant semantic classes have been stabilised, empirical knowledge production can proceed from the type of documentary or representational object made available through a resource to what that object depicts, records, transcribes, or otherwise represents.
Observation and identification address the question
What or whom does this object depict or record, and what can be established about the represented agents, objects, or activities?
Consider again object3, the portrait photograph from the running example. Taxonomic review has already stabilised the documentary object as a photograph. A subsequent observation activity can now examine the photograph as evidence concerning its depicted subject matter.
This establishes an evidentiary frame for the activity. The photograph is no longer being examined primarily as a digital manifestation whose media type or embedded metadata is under observation. It is being observed as a representation of depicted entities and activities. The evidentiary frame does not itself assert what the photograph depicts; it specifies the kind of semantic object about which the observation activity is intended to generate candidate assertions.
Suppose examination of object3 indicates that four people are visible, that three faces appear recognisable, and that one of the depicted people may be Hilda Grīva. Observation can generate candidate assertions such as
| claim_id | evidence | observed as | subject | predicate | candidate |
|---|---|---|---|---|---|
| claim10 | object3 |
depicted subject matter | photograph | depicts | Hilda Grīva |
| claim11 | object3 |
depicted subject matter | photograph | depicts | unidentified person 1 |
| claim12 | object3 |
depicted subject matter | photograph | depicts | unidentified person 2 |
| claim13 | object3 |
depicted subject matter | photograph | number of depicted persons | 4 |
| claim14 | object3 |
depicted subject matter | photograph | number of recognisable faces | 3 |
These are candidate assertions generated through empirical examination of the photograph. They are not yet equivalent to reviewed or stabilised knowledge. In particular, observing that a face is recognisable does not establish the identity of the depicted person.
Different candidates may also require different kinds of knowledge-producing activity. Counting visible people can be treated as visual observation. Identifying a depicted person may require comparison with catalogue information, inscriptions, authority resources, other photographs, or specialist knowledge. Classifying visible clothing may require domain interpretation by a dress historian.
For example, a dress historian may examine the same photograph as evidence concerning the clothing of a depicted person. The historian observes visible features of the clothing and, using relevant domain knowledge, may interpret or classify the garment as Livonian folk dress:
| claim_id | evidence | observed as | subject | predicate | candidate |
|---|---|---|---|---|---|
| claim15 | object3 |
depicted clothing | Hilda Grīva | clothing | Livonian folk dress |
The distinction between observation and interpretation is useful here. Observation concerns what an agent can establish through examination of evidence. Interpretation becomes necessary where establishing the semantic assertion depends on domain or linguistic competence. The boundary is not always absolute: recognising a person, classifying a garment, or identifying an activity may combine empirical observation with specialist interpretation within the same scholarly examination. Provenance should therefore record the agent, evidence, method, supporting resources, and relevant competence rather than forcing every knowledge-producing activity into a single epistemic category.
Additional evidence or supporting resources may participate in these activities. A catalogue record may propose an identity; a dress guide may support garment classification; another photograph may provide comparative evidence; a computational model may generate candidate regions, faces, objects, or classifications. These resources form part of the provenance of the resulting assertions.
The resulting candidates can then enter review. For example, a review workspace presented to an appropriately competent reviewer might contain
| claim_id | evidence | observed as | subject | predicate | candidate | review_1 |
|---|---|---|---|---|---|---|
| claim13 | object3 |
depicted subject matter | photograph | number of depicted persons | 4 | 4 |
| claim14 | object3 |
depicted subject matter | photograph | number of recognisable faces | 3 | 3 |
| claim15 | object3 |
depicted clothing | Hilda Grīva | clothing | Livonian folk dress | Livonian folk dress |
The corresponding review decisions may record these candidates as corroborated. The reviewed semantic value and the review status remain distinct: Livonian folk dress is the semantic value returned by review, while corroborated records what happened to that candidate during the review activity.
Their provenance may also differ:
| provenance | claim13 | claim14 | claim15 |
|---|---|---|---|
| activity | visual observation and review | visual observation and review | dress interpretation and review |
| agent | reviewer1 | reviewer1 | dress historian |
| evidence | object3 |
object3 |
object3 |
| supporting resource | NA | NA | dress guide |
| observed as | depicted subject matter | depicted subject matter | depicted clothing |
| review status | corroborated | corroborated | corroborated |
| comment | four persons visible | three recognisable faces | classified as Livonian folk dress |
This example also illustrates that candidate generation and review may occur within the same scholarly examination. A curator presented with an existing candidate may review it. When the same curator observes something for which no candidate has yet been represented, the activity also generates a new candidate. The interface need not force these into separate human tasks, provided that their different semantic and provenance roles remain explicit.
Identification can then proceed selectively. The stabilised assertion that a depicted person is recognisable does not establish who that person is. It instead establishes that a more specific identification activity may be applicable. Candidate identities may be generated from catalogue metadata, authority resources, inscriptions, comparison with other photographs, computational recognition, or human knowledge and then subjected to appropriate review.
This progression should not be confused with logical inference. The assertion
\[(\texttt{depicted\ person\ 1},\texttt{recognisable},\texttt{TRUE})\]
does not logically imply a particular identity. It establishes that attempting identification may be applicable. The subsequent identity still has to be generated through an appropriate knowledge-producing activity and supported by its own provenance.
The same pattern applies to the audio resources in the running example. A recording may be examined as evidence concerning a represented performance. Listening or computational analysis may generate candidates concerning performers, musical works, language, instrumentation, or recording events. Specialist interpretation may subsequently be required to identify a musical form, linguistic expression, performer, instrument, or other domain-specific feature.
Observation and identification therefore do more than add descriptive metadata. They progressively establish explicit semantic relationships between evidence, documentary or representational objects, and the entities or activities represented through them.
Observation and identification are particularly suitable for allocating knowledge-producing and review activities according to competence.
Suppose a review package contains one hundred photographs that have already been stabilised as depicting people. A dress historian does not need to examine photographs of buildings or boats and does not need to answer every possible question about the photographs depicting people. Candidate generation and review can instead be restricted to questions for which the historian’s competence is relevant.
For photographs associated with traditional Livonian dress, the activity may ask the historian to assess
- the number of depicted people;
- whether particular individuals are visually recognisable;
- which depicted person is wearing the garment under examination;
- whether the clothing can be classified as traditional Livonian dress;
- whether it appears to be a historical reconstruction or costume;
- whether it appears to be contemporary everyday clothing; or
- whether its character cannot be determined from the available evidence.
These questions do not all require the same epistemic operation. Counting people is primarily observational. Assessing recognisability combines observation with judgement. Classifying clothing requires domain interpretation. Identifying a depicted person constitutes a further identification activity and may require additional evidence or specialist competence.
The historian need not identify every recognisable person. Nor must the historian determine copyright status, personality rights, or other questions outside the assigned competence and purpose of the activity. A judgement such as recognisable = TRUE is already useful semantic knowledge because it can make a subsequent identification activity applicable without prematurely asserting an identity.
For example:
| evidence | observed as | subject | predicate | reviewed value |
|---|---|---|---|---|
photo127 |
depicted person | depicted person 1 | recognisable | TRUE |
photo127 |
depicted clothing | depicted person 1 | clothing | Livonian dress |
photo341 |
depicted person | depicted person 1 | recognisable | FALSE |
photo341 |
depicted clothing | depicted person 1 | clothing | uncertain |
These stabilised assertions can subsequently contribute to further candidate generation. A recognisable depicted person may make an identification activity applicable. An identified person may make namespace reconciliation applicable. A particular dress classification may make more specialised historical or ethnographic interpretation applicable. Where the purpose of the workflow includes rights clearance, recognisability or identification may likewise determine whether particular legal questions become relevant.
The important distinction is between applicability and inference of a substantive value. Stabilised knowledge may establish that a further activity or rule is applicable without determining the answer that activity will produce. Where an explicit rule does derive a new candidate from represented knowledge, that derivation is an inference and should be recorded as such.
Empirical knowledge production therefore progressively restructures the remaining problem. Observation generates candidates from evidence; specialist interpretation generates candidates where domain competence is required; review examines explicit candidates; and stabilised results determine which subsequent knowledge-producing or review activities are applicable.
Human expertise is consequently not applied uniformly to every resource or every possible property. It can be allocated to those questions for which the relevant competence can reduce uncertainty.
Linguistic knowledge and interpretation
Language introduces a distinct but closely connected family of semantic problems. Identifying an entity, establishing how it is named or translated, and interpreting linguistic forms are related activities, but they do not establish the same kind of knowledge.
Linguistic competence may already be required during empirical observation. A reviewer identifying an inscription, interpreting a handwritten annotation, recognising a spoken language, or establishing what is said in a recording does not examine the empirical evidence independently of language. The assertions that such an observation can responsibly produce depend in part on the linguistic competence of the reviewing agent, just as the identification of a garment, musical form, or historical object may depend on other forms of domain competence.
The distinction made here is therefore not between linguistic and non-linguistic evidence. Rather, language itself becomes the object of semantic knowledge production. Names and titles raise questions of reference; multilingual expressions and translations raise questions about how the same entity or concept is expressed in different linguistic, historical, and communicative contexts; and words and other linguistic forms can themselves become objects of lexical and linguistic interpretation.
The examples below consequently move from reference, through expression and translation, to linguistic interpretation. Proper nouns provide an intuitive starting point for named entities. Multilingual names and translations demonstrate why an entity cannot always be assigned a single context-independent “good name”. Common nouns and other linguistic forms then lead to lexical concepts, senses, grammatical structures, and historically situated uses as objects of scholarly investigation in their own right.
As in empirical observation, these activities are competence-dependent. A native speaker, translator, historical linguist, lexicographer, curator, and computational language model may be able to examine different questions and may warrant different kinds of assertions. Making those competences and their limits explicit becomes important when the resulting linguistic knowledge is reviewed, combined, or used for further inference.
Named entity review
Once observations and interpretations have introduced people, places, organisations, works, events, and other potentially identifiable entities, a further problem is to establish which entities are being referred to.
Named entity review asks
Which entity is being referred to, and what can be established about its identity?
The most intuitive cases involve proper nouns, but the occurrence of a name and the identification of an entity must remain distinct. A photograph may generate an observation concerning a depicted person; an archival record may contain the string Riga; a sound recording may be catalogued under the title Min izāmō. Such observations provide evidence for identification, but the strings or observed features do not by themselves establish that two resources or assertions refer to the same person, place, or work.
Consider again the depicted person in object3. Observation may first establish assertions such as
\[(\texttt{depicted\ person\ 1},\texttt{recognisable},\texttt{TRUE}).\]
A subsequent identification activity may use the photograph together with catalogue metadata, inscriptions, other photographs, authority records, computational recognition, or human knowledge to generate the candidate identity
\[(\texttt{depicted\ person\ 1},\texttt{identified\ as},\texttt{person1}).\]
where person1 denotes the candidate entity subsequently represented as Hilda Grīva.
This distinction prevents the observation of depicted subject matter from being conflated with the identification of the depicted entity. The photograph remains evidence used by the identification activity, while the identity assertion is semantic knowledge generated from that evidence together with any additional resources and specialist knowledge used by the activity.
Once an identity has been sufficiently stabilised, the identified entity can itself become the subject of further assertions concerning names, life dates, identifiers, relationships, occupations, associated works, and other authority information.
Candidate assertions about the entity may be generated from local catalogue metadata, authority files, archival catalogues, scholarly publications, Wikidata, and other external knowledge representations. Reconciliation with such resources is itself a provenance-bearing knowledge-producing activity: it may generate identity assertions or further candidates concerning the identified entity. These candidates remain distinct from the external resources used to generate or review them.
For the running example, a stabilised representation may be presented in a convenient wide form as
| person | preferred_label | birth_name | latvian_name | birth_year | death_year |
|---|---|---|---|---|---|
| person1 | Hilda Grīva | Hilda Cerbach | NA | 1910 | 1984 |
| person2 | Kōrli Stalte | NA | Kārlis Stalte | 1870 | 1947 |
| person3 | Julgi Stalte | NA | NA | 1941 | NA |
This is a wide projection of individual semantic assertions rather than the persistent representation itself. The underlying review workspace may contain, for example,
| claim_id | subject_candidate | predicate_candidate | value_candidate | value_reviewed | value_status |
|---|---|---|---|---|---|
| claim20 | person1 |
birth_year |
1910 |
1910 |
corroborated |
| claim21 | person1 |
death_year |
1983 |
1984 |
falsified |
| claim22 | person2 |
birth_year |
1870 |
1870 |
corroborated |
| claim23 | person2 |
death_year |
1947 |
1947 |
corroborated |
claim21 illustrates why the reviewed value and review status remain distinct. The candidate 1983 is falsified, while 1984 is returned as an alternative semantic value. Depending on the finalisation policy, the alternative may enter the subsequent stabilised state or become a candidate for another review.
The provenance representation records how these semantic states were produced. For claim21, the candidate may have been generated by an import activity using a local catalogue, while the subsequent review activity uses additional scholarly and authority evidence:
| activity_id | activity | agent | used | generated |
|---|---|---|---|---|
| generation21 | catalogue import | import agent |
local catalogue |
candidate 1983 |
| review21 | named entity review | reviewer2 |
candidate 1983; Wikidata; Livonian monograph |
reviewed value 1984; status falsified |
The correction therefore does not overwrite the earlier assertion. The candidate, the activity through which it was generated, the evidence used in review, its falsification, and the proposed alternative remain distinguishable in the persistent representation.
Named entity review will often make use of a local authority namespace and external identifiers such as VIAF, ISNI, Wikidata, or GeoNames. In standard terminology, identification concerns unambiguous reference, while an identifier is the value through which something is identified (ISO/IEC 2024b). These infrastructures therefore support identification; they are not themselves the object of review. The scholarly problem is to establish that mentions, observations, descriptions, or other representations encountered in different resources or activities refer to the same entity.
This distinction applies beyond people. Places, organisations, musical works, songs, expeditions, and other identifiable entities can all be stabilised in this way. A recording carrying the title Min izāmō, for example, provides evidence for an identification problem: which musical work does this title in this recording refer to? Reconciliation with local or external knowledge representations may generate a candidate identity for the work, which can then be reviewed and stabilised independently of the particular recording and the linguistic form through which the work was encountered.
Named entity review therefore stabilises reference, not names as strings. A name, title, inscription, catalogue entry, or observed feature may provide evidence for identification, but the resulting entity identity must remain distinguishable from the linguistic expression through which the entity was encountered.
This distinction immediately creates the next problem. An identified entity does not necessarily have one uniquely correct human-readable name. Hilda Grīva may be represented by different personal names; a place may have names in several languages and historical periods; and a work may circulate under different titles. Once reference has been stabilised, the linguistic expressions used to refer to the entity can therefore become objects of review in their own right.
This leads to multilingual expression and translation review.
Multilingual expression and translation review
Once reference has been stabilised, the linguistic expressions used to refer to an entity can be examined independently of the entity itself. A person, place, work, or other identified entity may be associated with several names, labels, titles, historical forms, transliterations, or translations, none of which should be conflated with the entity to which they refer.
Multilingual expression and translation review asks
Which linguistic expressions refer to this entity, in which languages and contexts, and what relationships hold between those expressions?
Consider person1, identified in the preceding review as Hilda Grīva. Archival records, catalogues, publications, inscriptions, and external authority resources may contain different linguistic forms referring to the same person. Observation of these resources can establish that particular strings occur in the evidence. Reconciliation and linguistic interpretation can then generate candidate assertions concerning the relationship between those expressions and the identified entity.
For example, the evidence may support candidates such as
| claim_id | subject | predicate | candidate |
|---|---|---|---|
| claim30 | person1 |
preferred label | Hilda Grīva |
| claim31 | person1 |
birth name | Hilda Cerbach |
| claim32 | person2 |
Livonian name | Kōrli Stalte |
| claim33 | person2 |
Latvian name | Kārlis Stalte |
These assertions concern linguistic expressions associated with already identified entities. They should therefore be distinguished from the identity assertions through which the entities themselves were stabilised.
The distinction is particularly important in multilingual collections. Two different strings may refer to the same entity without being translations of one another. They may instead be language-specific names, historical forms, aliases, transliterations, orthographic variants, or names used under different social or institutional conditions.
Translation is therefore a more specific semantic relationship than multilingual reference. For an identified work, for example, the following assertions represent different questions:
| subject | predicate | value |
|---|---|---|
work1 |
Livonian title | Min izāmō |
work1 |
Latvian title | Mana tēvzeme |
title1 |
translation of | title2 |
The first two assertions associate linguistic expressions with the same identified work. The third makes an additional claim about the relationship between the expressions themselves. Shared reference does not by itself establish translation equivalence.
This distinction also separates observation from linguistic interpretation. An agent examining a catalogue, inscription, publication, or recording may observe that a particular expression occurs in the evidence. Establishing its language, referent, meaning, translation relationship, historical form, or preferred use may require linguistic or domain competence.
For example, observation of an archival catalogue may establish that the string Kōrli Stalte occurs in a particular record. If the record has already been associated with person2, this occurrence can contribute evidence for the candidate assertion
\[(\texttt{person2},\texttt{Livonian name},\texttt{Kōrli Stalte}).\]
A linguistic reviewer may then examine the candidate using the original evidence together with dictionaries, authority resources, scholarly publications, or other relevant material. Review determines whether the proposed relationship between the entity and the linguistic expression is corroborated, falsified, or deferred.
The provenance of these activities remains important. The occurrence of a string in a historical record, its identification as Livonian, its association with a particular person, and its selection as a preferred label may have been established through different activities by different agents using different evidence and methods. They should not be collapsed into a single assertion merely because they can ultimately be displayed together in an authority table.
A practical review surface may nevertheless present these assertions in wide form:
| person | preferred_label | birth_name | livonian_name | latvian_name |
|---|---|---|---|---|
| person1 | Hilda Grīva | Hilda Cerbach | NA | NA |
| person2 | Kōrli Stalte | NA | Kōrli Stalte | Kārlis Stalte |
| person3 | Julgi Stalte | NA | Julgi Stalte | NA |
Such a table is a convenient projection of stabilised assertions rather than a claim that each entity possesses one canonical name independently of language, historical context, or purpose.
The same principles apply to titles, place names, institutional names, lexical forms, inscriptions, and other linguistic expressions. Stabilising the identity of the referent makes it possible to review these expressions without confusing variation in language with variation in entity identity.
Multilingual expression review therefore stabilises relationships between identified entities and linguistic expressions, while translation review makes the stronger claim that particular expressions stand in a translation relationship. Both depend on evidence and linguistic interpretation, but neither should be reduced to named entity identification.
Once linguistic expressions themselves become objects of semantic analysis, a further distinction becomes necessary. An expression may not merely name an entity: it may carry lexical meaning that requires interpretation within a language and context. This leads from multilingual expression and translation review to lexical and semantic review.
Lexical and linguistic review
Named entity review stabilises reference, while multilingual expression and translation review establishes relationships between identified entities and the linguistic forms through which they are named. Linguistic knowledge, however, extends beyond names and reference. Common nouns, descriptive expressions, classifications, and other lexical forms introduce a further problem: their semantic interpretation may depend on language, community, disciplinary competence, historical context, and purpose.
Lexical and linguistic review asks
What does this linguistic expression mean in this context, according to which linguistic or disciplinary competence?
This question differs from asking which entity a proper noun refers to. An identified object does not necessarily have a single context-independent linguistic description. Different competent speakers or specialists may describe the same object using different lexical and conceptual systems.
Consider, for example, a coat depicted in an ethnographic photograph from an Udmurt village. A local Udmurt speaker may describe the garment using a vernacular term embedded in everyday language and local cultural practice. A textile researcher examining the same garment may classify it using a scholarly vocabulary concerned with construction, material, technique, historical type, or comparative morphology. The resulting descriptions need not be competing attempts to supply one correct name. They may constitute different, contextually warranted semantic interpretations of the same observed object.
The distinction can be represented schematically as
| evidence | observed as | agent / competence | linguistic interpretation |
|---|---|---|---|
| photograph | depicted garment | local Udmurt speaker | vernacular description |
| photograph | depicted garment | textile researcher | scholarly textile classification |
The evidence and represented object may therefore remain the same while the knowledge-producing activity, competence, and resulting linguistic assertion differ.
This makes provenance constitutive of linguistic semantic knowledge rather than merely an audit trail. Preserving only the resulting string or classification would lose information necessary to understand what kind of assertion has been made and under which conditions it is warranted. Provenance may therefore need to record the agent, relevant competence, evidence examined, language or linguistic community, disciplinary context, method, purpose, and supporting lexical or scholarly resources used in interpretation.
This relativity does not imply that linguistic interpretation is arbitrary. Assertions remain open to review against the evidence and according to the relevant linguistic or disciplinary competence. A vernacular description can be challenged as an inaccurate account of local usage; a scholarly classification can be challenged against the concepts and methods of textile research. What differs is the context within which corroboration is meaningful.
The same issue appears in lexicography and translation. A lexical item may have several senses, and the appropriate interpretation may depend on the linguistic context in which it occurs. Terms used by historical communities may not correspond exactly to categories in contemporary scholarly vocabularies. Translation may preserve some semantic distinctions while suppressing or introducing others. Mapping a vernacular term to a thesaurus concept is therefore itself a semantic assertion requiring appropriate evidence, competence, and provenance.
For example, suppose an Udmurt expression observed in a catalogue or elicited from a local speaker is proposed as referring to a particular type of garment. A lexical or terminological review may relate the expression to a scholarly concept:
| claim_id | subject | predicate | candidate |
|---|---|---|---|
| claim40 | udmurt_term_1 |
lexical meaning | local garment concept |
| claim41 | udmurt_term_1 |
mapped to | textile concept 17 |
These are different assertions. The first concerns the interpretation of a linguistic expression within a linguistic community; the second proposes a relationship between that interpretation and a scholarly conceptual system. Corroborating the first does not logically entail the second.
A practical knowledge system should consequently resist collapsing vernacular descriptions, translations, scholarly classifications, and thesaurus mappings into a single supposedly canonical label. They can instead coexist as explicit assertions whose relationships may themselves be observed, interpreted, proposed, reviewed, and, where justified by explicit rules, inferred.
This is also why linguistic competence matters to empirical review more generally. A curator examining a photograph may possess sufficient visual and domain competence to recognise a garment while lacking the linguistic competence required to interpret an inscription associated with it. A local speaker may possess linguistic and cultural competence unavailable to the curator while lacking the specialist vocabulary of textile history. Competence is therefore activity-specific rather than an undifferentiated property of the reviewer.
Lexical and linguistic review makes these differences explicit. The purpose is not to eliminate semantic plurality by selecting a universally correct vocabulary, but to produce reviewable knowledge about linguistic expressions, their interpretations, and the relationships between different linguistic and conceptual systems while preserving the provenance necessary to understand those assertions.
Conditions and limits of semantic knowledge production
Semantic knowledge production is constrained not only by the available evidence but also by the conditions under which particular human or computational agents are competent, permitted, or logically warranted to produce, review, disclose, infer, or act upon semantic assertions.
These conditions are not interchangeable.
Competence concerns what an agent can responsibly establish. A filesystem inspection tool may observe technical properties of a digital resource but cannot identify a depicted historical person. A local speaker may possess linguistic and cultural competence unavailable to a textile researcher, while the textile researcher may possess a specialist classificatory competence unavailable to the local speaker. As the preceding examples demonstrate, competence is therefore relative to a particular knowledge-producing activity rather than an undifferentiated property of an agent.
Permission concerns what an agent or system is authorised to observe, use, disclose, infer, or act upon. An assertion may be well supported by evidence while its disclosure or use remains restricted by law, contract, ethical obligations, institutional policy, access conditions, or the purpose for which the knowledge was produced. Epistemic warrant and permission must therefore remain distinct: can this be established? and may this be used or disclosed for this purpose? are different questions.
Logical applicability concerns what may be derived from knowledge already represented. Inference generates candidate assertions only where an explicit rule is applicable and its required premises are available. A rule that is not applicable does not produce a negative conclusion; nor does a missing premise establish that the conclusion is false. It establishes a limit on what can presently be inferred.
These conditions may produce several different forms of unresolved knowledge. An assertion may be unknown but potentially discoverable because appropriate evidence or competence has not yet been brought to the activity. A conclusion may be dependent on an unstabilised assertion that requires further observation, interpretation, identification, or review. A value may be presently indeterminate because the event or condition that would determine it has not yet occurred. An otherwise warranted assertion may also be known but not permitted for a particular use or disclosure.
These situations should not be collapsed into a common missing value. They describe different states of knowledge and imply different possible next actions.
| condition | what is missing or limited | possible consequence |
|---|---|---|
| insufficient evidence | empirical support | further observation may be required |
| insufficient competence | appropriate capacity to establish or interpret | another agent or specialist may be required |
| unstabilised dependency | required semantic premise | further candidate generation or review may be required |
| rule not applicable | conditions required for inference | no inference is warranted |
| presently indeterminate | determining event or condition | the value cannot yet be established |
| permission constraint | authority for use, disclosure, or action | knowledge may exist but cannot be used in the intended way |
Provenance is important to these limits for the same reason that it is important to empirical and linguistic knowledge production. Knowing an assertion without knowing the activity, evidence, competence, rule, or permission under which it was produced may be insufficient to determine whether that assertion can responsibly participate in another activity.
The provenance chain should therefore make it possible to ask not only
What is currently asserted?
but also
Who or what was competent to establish it, from which evidence, through which activity or interpretation, under which applicable rules, and for which subsequent uses is it permitted?
Making these conditions explicit turns limits of knowledge into part of the semantic representation rather than treating them as failures of knowledge production. A deferred review, an inapplicable inference rule, an unresolved dependency, or a permission restriction can itself provide actionable knowledge about what can—and cannot—happen next.
The following rights example illustrates this distinction particularly clearly. Rights review remains a form of semantic knowledge production: it establishes assertions about legal relationships, their applicability, relevant rights holders, or other legally significant properties of a resource. Like other semantic assertions, these claims require provenance concerning the evidence, rules, professional competence, and review activities through which they were established. A further step is required when such stabilised knowledge is used to determine whether an agent may, must, or must not perform an action. That step crosses into the Normative Plane and is not formalised here.
Rights review
Rights review provides an important application of the review algebra because it connects semantic stabilisation to decisions about how cultural heritage resources may be used.
Rights review asks
Which rights, legal relationships, or other legal considerations are relevant to an intended use, and what can presently be established about them?
Rights review is not a mandatory stage of the algebra. A research workflow concerned exclusively with linguistic or historical analysis may have no reason to perform it. In many digital humanities and cultural heritage applications, however, legal review becomes important when reviewed knowledge is intended to support publication, dissemination, reuse, digitisation, licensing, or another form of action.
Like the scholarly activities considered above, rights review is a provenance-bearing activity. A legal specialist, curator, or computational agent examines previously stabilised assertions together with relevant legislation, licences, contracts, institutional policies, ethical guidelines, or other evidence and produces further reviewable assertions.
Consider a sound recording for which previous review activities have already stabilised assertions concerning the recording, the performed work, its authors, its performers, and the recording event. These assertions may make particular legal questions applicable.
For example:
| claim_id | subject_candidate | predicate_candidate | value_candidate | value_reviewed | value_status |
|---|---|---|---|---|---|
| claim50 | recording1 |
copyright_relevant |
TRUE |
TRUE |
corroborated |
| claim51 | recording1 |
neighbouring_rights_relevant |
TRUE |
TRUE |
corroborated |
| claim52 | recording2 |
copyright_relevant |
TRUE |
TRUE |
corroborated |
| claim53 | recording2 |
neighbouring_rights_relevant |
TRUE |
TRUE |
corroborated |
These assertions do not establish that a particular right remains in force, belongs to a particular rights holder, or prevents a proposed use. They establish that a legal relationship or question is relevant and may therefore require further evidence, reasoning, or specialist review.
Rights review should therefore not be equated with normative decision-making. An assertion such as copyright applies, person1 is the rights holder, or neighbouring rights remain relevant belongs to the Domain Plane: it describes a legal relationship or legally significant state of affairs. The provenance of the legal or professional activity through which that assertion was established belongs to the Provenance Plane. Only when such stabilised and warranted knowledge is evaluated under an applicable policy for a contemplated action does the workflow cross into the Normative Plane.
Schematically:
Domain Plane:
copyright applies
Provenance Plane:legal reviewer established this using sources X and Y
Normative Plane:therefore agent A may publish recording R
The final step is deliberately outside the Review Algebra developed here. The algebra can produce and stabilise the assertions required by such a decision and preserve the provenance through which they acquired authority, without itself constituting a general policy or normative reasoning system.
The provenance of a rights review may therefore record both the semantic assertions and the legal resources used by the activity:
| activity_id | activity | agent | used | generated |
|---|---|---|---|---|
| rights50 | rights review | legal reviewer1 |
recording1; reviewed authorship and performance assertions; legal sources |
claim50; claim51 |
Candidate legal questions can themselves be generated from previously stabilised knowledge. Identification of an authored musical work embodied in a recording may make copyright-related questions applicable. Identification of a recorded performance may make neighbouring-rights questions applicable. A photograph depicting an identifiable person may make a different family of legal, ethical, or institutional-policy questions applicable.
Rights review therefore illustrates the recursive relationship between semantic stabilisation and specialised scholarly or professional activity. Earlier reviews determine which legal questions are worth asking; legal review produces further reviewed assertions; and those assertions may become inputs to subsequent inference and decision-making.
Where the intended outcome is machine-actionable access or reuse policy, reviewed rights assertions may provide warranted inputs to a policy representation such as ODRL (W3C 2017). The architectural distinction is important. Rights review establishes what can presently be asserted about the resource and its legal relationships; provenance records how and on what authority those assertions were established. ODRL can then express permissions, prohibitions, duties, and constraints applicable to particular actions or purposes. Rights review is therefore not itself an ODRL workflow or a normative decision procedure. It produces reviewed, provenance-bearing semantic knowledge that may subsequently be used by such a procedure.
Rights review also demonstrates how review effort can be allocated according to purpose and competence.
Suppose the objective is to make a collection of ethnographic photographs available for research. Previous taxonomic and observation reviews may already distinguish photographs depicting buildings, boats, landscapes, objects, recognisable people, and unrecognisable people.
A legal or ethical reviewer need not inspect every photograph independently. Previously stabilised assertions can determine which resources require review and which questions are relevant to the intended use.
A photograph depicting a building may require copyright or reproduction-rights analysis but no review concerning the identity of a depicted person. A photograph containing recognisable people may make additional questions concerning identification, personality rights, privacy, ethics, or institutional publication policy relevant, depending on the intended use and applicable legal framework.
The specialist performing the earlier observation activity does not need to answer these questions. Their reviewed assertions provide inputs from which appropriate legal review activities can be generated and allocated to agents with the relevant competence.
The same principle applies to sound recordings. Stabilised assertions concerning works, authors, performers, recording events, publication dates, and other relevant facts can determine which legal questions require specialist examination.
The result may subsequently be used operationally. Once the relevant rights assertions have been sufficiently stabilised for a particular purpose, they can become warranted inputs to an access or reuse decision and, where appropriate, to a machine-readable policy representation such as ODRL. The decision itself belongs to the Normative Plane rather than to rights review as defined here.
Rights review therefore does not require every reviewer to become a legal expert, nor does every resource require the same legal analysis. It uses previously stabilised semantic knowledge to construct a smaller, explicit, purpose-specific, and provenance-bearing legal review problem.
Inference and recursive candidate generation
Once sufficient semantic knowledge has been stabilised, explicit rules can generate further candidate assertions without requiring a new observation of the original empirical evidence.
Inference asks
What follows from the knowledge already represented under an applicable rule?
Inference is not a separate kind of review. Within the Review Algebra, it is a mechanism of candidate generation:
\[\mathcal{G}_{\gamma}(C_k) \rightarrow C_k^{\*},\]
where \(\gamma\) is an explicit inference rule operating on assertions in the current semantic state.
This distinguishes inference from observational candidate generation. An image classifier may generate a candidate by examining pixels, and a human observer may generate a candidate by examining a photograph. An inference rule instead operates on semantic knowledge already represented.
Inference depends both on the required premises and on the conditions of applicability of the rule. A rule may apply only to particular classes of entities, relationships, jurisdictions, time periods, analytical dimensions, or value ranges. Where assertions carry explicit scope, that scope may contribute to determining whether these conditions are satisfied. Scope is therefore one possible component of applicability, rather than a substitute for the rule’s explicit conditions.
Suppose the stabilised semantic state contains assertions concerning
- the identity of a musical work;
- its authors or composers;
- relevant life dates;
- the identity of performers;
- the date and circumstances of a recording; and
- legal relationships established as relevant during rights review.
A legal inference activity may combine these assertions with an explicitly identified legal rule:
\[{\text{stabilised assertions}} + {\text{applicable rule}} \xrightarrow{\mathcal{G}_{\gamma_{\text{legal}}}} {\text{candidate legal conclusion}}.\]
The provenance of the resulting candidate records the assertions used, the rule applied, and the activity through which the candidate was generated. Inference is therefore a provenance-bearing knowledge-producing activity.
An inferred candidate is not necessarily exempt from review. A deterministic rule may be sufficiently authoritative for a particular purpose to permit automatic finalisation. Another rule may produce a candidate that still requires legal, scholarly, or institutional review. The Review Algebra does not prescribe that policy: it preserves the distinction between candidate generation, review, and finalisation policy.
Inference also exposes an important property of knowledge representations: knowledge that follows implicitly from represented assertions need not always be materialised explicitly. A record-set assertion may efficiently represent knowledge that could, under an applicable inheritance rule, generate hundreds of member-level candidate assertions. Materialising those candidates makes them individually reviewable and exposes possible exceptions, but may also produce large numbers of assertions that add little value for the purpose at hand.
Inference therefore asks not only
What can be inferred?
but, operationally,
Which implicit knowledge is worth making explicit for the present purpose?
Explicit inference rules also expose dependencies and limits in the current semantic state. A candidate conclusion may fail to be generated because the conditions of applicability are not satisfied, because a required assertion is absent, or because a required assertion has not been sufficiently stabilised.
Suppose, for example, that a legal inference requires the date of a recording but the current semantic state does not contain a sufficiently stabilised recording date. The legal conclusion cannot yet be generated. The unsuccessful inference attempt nevertheless identifies an explicit dependency:
| subject | required assertion | status |
|---|---|---|
recording2 |
recording_date |
required |
That dependency can make a subsequent knowledge-producing or review activity applicable. If the recording date may be established from empirical evidence, an observation activity may be required; if a candidate date already exists, it may require review; if the date can be derived from other stabilised knowledge, another inference may generate it.
A different recursive case arises when inference can proceed but the resulting candidate does not survive subsequent examination. Rejection need not terminate the semantic process. The candidate, its returned semantic value, its review outcome, and the provenance of the review remain available to subsequent candidate-generation activities. A later activity may consequently refine the candidate, restrict its scope, introduce an additional condition, or generate an alternative assertion for review.
A precedent for this pattern can be found in computational mathematical discovery. Colton and Pease’s theorem-modification system deliberately moves beyond a conventional proof-or-failure model: when a conjecture cannot be proved, supporting and falsifying examples are used to generate modified conjectures that are subsequently subjected to further proof attempts (Colton and Pease 2005). Drawing on Lakatos’s account of mathematical development, their approach treats counterexamples not merely as reasons to discard a conjecture but as information from which more specialised or otherwise modified conjectures can be generated.
The Review Algebra abstracts a different part of this computational pattern. It does not prescribe how a rejected candidate should be repaired, nor does it assume that semantic review can be reduced to theorem proving or counterexample generation. Rather, by representing the semantic value returned by review separately from its review outcome and provenance, it permits subsequent candidate-generation activities to use that history. A rejected, deferred, or modified candidate may therefore remain computationally productive without acquiring production-level semantic authority.
The recursive structure can consequently be represented as
\[C_k \xrightarrow{\mathcal{G}_{\gamma}} C_k^{\*} \xrightarrow{\mathcal{R}_k} Q_k \xrightarrow{\mathcal{F}_{\phi}} S_{k+1}^{D} \xrightarrow{\mathcal{G}_{\gamma'}} C_{k+1}^{\*}.\]
The candidate-generation mechanism at each stage need not be inference. New evidence may generate candidates through observation; specialist knowledge may generate candidates through interpretation or identification; external knowledge representations may generate candidates through reconciliation; represented knowledge may generate candidates through inference; and previous review results may provide evidence or constraints for generating revised candidates.
Semantic knowledge production is therefore recursive without being epistemically uniform. Stabilised assertions make further activities and rules applicable; those activities generate new candidates through different knowledge-producing mechanisms; and the resulting candidates can themselves be reviewed and finalised. Review outcomes that do not result in finalisation may likewise alter the space of subsequent candidate generation without becoming part of the stabilised domain state.
Inference is particularly important because it can make dependencies explicit. An inference that cannot proceed because its premises or applicability conditions are not satisfied should not be confused with a negative semantic conclusion. It may instead reveal precisely which assertion, condition, or further activity is required before the intended conclusion can be generated. Conversely, where inference successfully generates a candidate that is subsequently rejected, the review result and its provenance may inform a new cycle of candidate generation. In both cases, an unsuccessful path can contribute information to the continuing semantic process without being admitted to the stabilised domain state.
Limits of current knowledge
Semantic knowledge production does not require every activity to terminate in a positive semantic assertion. Explicit representation of candidate generation, review, finalisation policy, and inference also makes it possible to distinguish different limits of current knowledge.
A candidate may be deferred because the reviewing agent cannot presently corroborate or falsify it. A required assertion may be unknown but potentially discoverable. A requested property may be indeterminate because the relevant event or condition has not occurred. An inference may remain blocked by an unresolved dependency on another assertion.
These situations should not be collapsed into a single missing value.
Consider the distinction between an historical person whose death year has not yet been established and a living person. In the first case, the value may be unknown and further investigation may be warranted. In the second case, the property has no determinate value at the present time: there is no death year to discover while the person remains alive. An empty table cell or NA cannot by itself express this distinction.
Likewise, a reviewer may defer a candidate death year because the available evidence is insufficient. Under a finalisation policy that provisionally retains deferred candidates, the candidate value may continue to appear in the current semantic state while its deferred status remains explicit in provenance. The semantic value alone therefore cannot represent the epistemic state of the assertion.
An inference may encounter a different limitation. A legal rule may be unable to generate a candidate rights conclusion because the recording date required by the rule has not been stabilised. Here the limitation is neither a legal judgement nor necessarily an unknown fact about the world. It is an explicit dependency in the inference process.
The semantic knowledge-production process can therefore distinguish at least four situations:
| state | interpretation | possible consequence |
|---|---|---|
deferred |
a candidate has been reviewed but cannot presently be corroborated or falsified | retain, escalate, or review again |
unknown |
relevant knowledge may exist but has not yet been established | investigate or generate further candidates |
indeterminate |
the requested property cannot presently have a determinate value | do not pursue until conditions change |
unresolved dependency |
another assertion is required before a conclusion can be generated | review or establish the dependency |
These distinctions are operationally important. An unknown assertion may justify further investigation. An indeterminate property should not automatically consume additional review effort. A deferred assertion records a completed review activity without pretending that the epistemic problem has been resolved. An unresolved dependency identifies a concrete requirement for another activity.
These limits of current knowledge can themselves inform candidate generation, activity allocation, review, inference, and finalisation policy. Semantic stabilisation does not require uncertainty or incompleteness to disappear; it requires them to become explicit enough that subsequent activities can respond to them appropriately.
From semantic production to governed action
The preceding sections stop at the production and stabilisation of semantic knowledge. They establish what is asserted in the Domain Plane and preserve the provenance-bearing activities through which those assertions acquired epistemic authority. Together, these can be represented as two related but distinct components of the stabilised semantic state: a domain state \(S_k^D\) and a provenance state \(S_k^P\).
These states provide a governed epistemic substrate for subsequent action, but they do not themselves determine what action should occur. An assertion that a person is deceased, that copyright remains relevant, or that a particular agent is a rights holder remains domain knowledge, however well supported it may be. Its provenance records how that assertion was established and provides information required to assess its authority. Neither component, by itself or in combination, constitutes permission, prohibition, or obligation.
This distinction becomes particularly important when semantic knowledge is produced for an operational purpose. As part of the Open Music Europe project, archival recordings of Livonian folk songs and other culturally significant works were prepared for dissemination on commercial streaming services. Semantic production may establish the identity of works, authors and performers, relevant dates, recording events, applicable legal relationships, and the provenance through which these assertions were reviewed. Inference may additionally expose unresolved semantic dependencies: for example, that a recording date or performer identification is still required before a particular legal conclusion can be generated.
None of these operations, however, is equivalent to deciding that the recording may be published.
Once stabilised knowledge is used to determine whether an agent may, must, or must not perform an action, an additional normative operation is required. Provisionally, such an operation can be represented as
\[(S_k^D, S_k^P, \Pi, a) \xrightarrow{\mathcal{I}_N} C_N^{*},\]
where \(\Pi\) denotes an applicable policy or normative context, \(a\) denotes a contemplated action, and \(C_N^{*}\) denotes one or more candidate normative conclusions concerning that action.
The present paper does not define \(\mathcal{I}_N\). The notation marks the boundary between the semantic-production architecture developed here and the further problem of governed action.
The distinction can be illustrated by the same operational objective without treating the objective itself as part of semantic stabilisation. A collection prepared for publication may contain recordings in different semantic states:
| recording | semantic state relevant to contemplated publication |
|---|---|
recording1 |
required domain assertions stabilised under the examined dependency structure |
recording2 |
unresolved dependency: recording date |
recording3 |
unresolved dependency: performer identification |
recording4 |
unresolved dependency: specialised legal assertion |
Such a representation can establish which semantic knowledge and dependencies are presently available. It does not itself establish that publication is permitted. That conclusion may additionally depend on the applicable policy, the authority and provenance of that policy, the contemplated action and purpose, and the criteria under which the available knowledge is considered sufficient for the decision.
The Normative Plane therefore introduces questions that are related to, but distinct from, those addressed by the Review Algebra. Policies themselves have provenance and authority. Different policies may apply to different agents, purposes, jurisdictions, resources, or periods and may conflict or establish different priorities. A normative conclusion may depend not only on the values of domain assertions but also on whether their provenance satisfies requirements concerning competence, evidence, review, recency, or institutional authority. The sufficiency of knowledge may likewise be relative to the contemplated action and to the risks associated with an erroneous decision.
This boundary also exposes a further Operations Research problem. Once an intended action and its normative requirements are represented, unresolved semantic dependencies may have different expected values for reaching a decision. With limited expert capacity, a future system could therefore allocate observation, identification, legal review, or other knowledge-producing activities according to their expected contribution to resolving the decision problem.
For example, establishing a missing recording date might resolve an applicability condition affecting many recordings, whereas reviewing a detailed musicological property might contribute valuable scholarly knowledge without affecting the immediate publication decision. The latter review remains legitimate and potentially important; it simply serves a different objective. Review allocation can therefore be understood as purpose-relative without reducing semantic knowledge production to a single operational goal.
Computational agents may also be able to explore such dependencies prospectively. Taking the stabilised state as their premise, they could generate and evaluate hypothetical candidate paths in a logically isolated environment, estimate which unresolved assertions have the greatest expected consequence for a contemplated objective, and identify productive review frontiers for human attention. Such look-ahead would remain exploratory: intermediate candidates generated within the search would not thereby acquire production-level semantic authority.
A fuller account of governed action would therefore require concepts not formalised in the present paper: the authority and provenance of policies; normative candidate generation and review; purpose-relative sufficiency; conflicting permissions, prohibitions, and obligations; risk-sensitive finalisation policy; and the allocation of scarce review resources under uncertainty. These questions also create an important human-computer interaction problem: how to present the relationship between stabilised knowledge, its authority lineage, unresolved dependencies, possible consequences, and contemplated actions without encouraging users to mistake a complete provenance path for guaranteed truth.
The Domain and Provenance Planes developed here make these questions tractable without attempting to answer them prematurely. Once stabilised domain knowledge can be distinguished from its authority-bearing provenance, a harder question can be stated explicitly:
Under what rules may that knowledge legitimately cause an action?
That question defines the boundary between the semantic knowledge-production framework developed in this paper and a prospective theory of governed action.
Conclusions
Semantic knowledge production cannot be reduced to the accumulation of descriptive metadata. Human and computational agents produce assertions through observation of evidence, specialist interpretation, identification and reconciliation, linguistic analysis, and inference under explicit rules. The resulting knowledge remains intelligible and reusable only if these activities and their limits remain distinguishable from the assertions they produce.
A first requirement is therefore to preserve the relationship between evidence and the semantic object about which knowledge is being produced. A digital manifestation, the documentary or representational object made available through it, and the entities or activities represented by that object may participate in the same knowledge-producing process without being identical. Documentary models such as RiC, provenance models such as PROV, and domain models such as CIDOC CRM consequently perform complementary functions rather than providing interchangeable descriptions of the same layer.
A second requirement concerns competence. The same evidence may support different assertions when examined by agents with different methods, linguistic communities, and domains of expertise. Technical observation, historical interpretation, linguistic analysis, identification, and legal assessment are therefore not interchangeable annotation tasks. Their outputs acquire epistemic meaning partly through the provenance of the activities, agents, and competences through which they were produced. Provenance is consequently not merely an audit trail: in some forms of semantic knowledge, it is necessary for understanding what kind of assertion has been made and under which conditions it is warranted.
Language makes this particularly visible. Stabilising the identity of an entity, establishing the expressions through which it is named or translated, and interpreting lexical or grammatical structures are related but distinct knowledge-producing activities. Preserving these distinctions allows multilingual and historically situated knowledge to be represented without assuming a single context-independent name, translation, or meaning.
Semantic knowledge production is also constrained by competence, permission, logical applicability, and the current limits of represented knowledge. A useful system must therefore be capable of representing not only positive assertions but also deferred judgements, unknown knowledge, indeterminate properties, permission constraints, and unresolved dependencies. These states can themselves determine which activities are useful, applicable, or permissible next.
The resulting progression is recursive. Stabilised assertions change the space of subsequent semantic work: classifications make new questions applicable, identifications enable authority reconciliation, linguistic interpretations enable further analysis, and explicit rules can derive new candidates or expose missing dependencies. The objective is therefore not necessarily to maximise the quantity of semantic description, but to produce and stabilise the knowledge required for a particular scholarly, institutional, or operational purpose. The separation of the Production Graph from exploratory candidate generation allows this progression to remain open to extensive computational search without allowing unreviewed results to acquire production-level authority.
The same architecture admits a complementary backward-looking view. Authority lineage asks how assertions already present in a stabilised state acquired production-level semantic authority. Directly reviewed assertions can be traced to the activities through which they were examined, while mechanically derived assertions can be traced through authorised rules and their stabilised premises to the governance frontiers on which they ultimately depend. This provides a more meaningful account of human governance than the proportion of assertions individually inspected by a person: the relevant question is whether assertions in the Production Graph remain connected through inspectable provenance-bearing transformations to appropriately governed premises and decisions.
These requirements motivate a formal account of review and stabilisation. The companion Review Algebra treats candidate generation, review, and finalisation as explicit provenance-bearing transformations over semantic assertions. Keeping that algebra distinct from the domain-specific models developed here allows the same formal machinery to support different forms of semantic knowledge production without prescribing what counts as evidence, competence, interpretation, or sufficient knowledge in a particular domain.
Not every revision has the same semantic consequence. ISO 19135 distinguishes clarifying or non-substantive changes, which have only minor impact on the use of information, from substantive changes that alter semantics or technical meaning (ISO 2026). A persistent implementation of the Review Algebra can preserve this distinction when determining whether a revised assertion represents clarification, replacement, or a new semantic state.
The architecture developed here deliberately stops at the boundary between semantic production and governed action. Stabilised domain knowledge and the provenance through which it acquired authority can provide the epistemic substrate for subsequent decisions, but they do not themselves determine whether an agent may, must, or must not act. Such a determination additionally requires an applicable policy or normative context, a contemplated action and purpose, and criteria for deciding whether the available knowledge and its warrant are sufficient. The normative operation through which these elements produce permissions, prohibitions, obligations, or other action-guiding conclusions is not formalised in the present paper.
This boundary nevertheless exposes a further research programme. Computational agents may be able to explore hypothetical semantic trajectories from a stabilised state, identify unresolved dependencies, and estimate which review activities would have the greatest expected value for a particular scholarly or operational objective. Such agent look-ahead and Operations Research allocation could help direct scarce human attention without allowing exploratory candidates to acquire production-level authority. Extending the architecture in this direction will require explicit treatment of policy authority and provenance, purpose-relative sufficiency, normative candidates, risk-sensitive finalisation policy, and the allocation of review effort under uncertainty.
Semantic knowledge production can consequently be understood as an interoperable progression in which evidence, human expertise, computational agents, domain models, and explicit rules contribute to progressively stabilised semantic states while retaining the provenance and limits necessary to understand how those states came to be known. The separation of stabilised knowledge from exploratory candidates allows this progression to combine computational breadth with human-governed semantic authority. Looking backward, authority lineage makes it possible to reconstruct how assertions acquired production-level authority; looking forward, stabilised knowledge determines which further questions, candidate-generating activities, and reviews become meaningful or applicable. Review and finalisation policy govern the passage between these states. The further question of when sufficiently warranted semantic knowledge may legitimately cause an action remains a distinct normative problem and a direction for future work.
The Review Algebra
Abstract
Semantic knowledge production generates assertions through heterogeneous human and computational activities. Such assertions may be provisional, contested, corroborated, falsified, deferred, or replaced, while their subsequent use may depend on explicit policies determining what is sufficiently stabilised for a particular purpose. This paper develops a Review Algebra for representing these transformations without conflating semantic assertions with the provenance-bearing activities through which they are generated, examined, and admitted to subsequent semantic states.
The algebra distinguishes three recurring operations: candidate generation, review, and finalisation. Candidate generation makes semantic assertions available for consideration. Review examines candidate assertions and returns semantic values together with explicit statuses describing what happened to the candidates. Finalisation applies a policy determining which reviewed assertions enter a subsequent stabilised semantic state.
Observation and inference are treated as distinct mechanisms of candidate generation rather than additional operators of the algebra. Observation generates candidates through the examination of evidence; inference generates candidates from represented knowledge under explicit rules and conditions of applicability. Other mechanisms may include reconciliation, import, computational analysis, or human proposal. Their epistemic differences remain represented through provenance while their outputs can participate in the same downstream review structure.
The algebra separates domain knowledge from provenance knowledge while extending ordinary relational and tidy data operations (Wickham 2014). Wide and long review workspaces provide computational and human-facing projections over persistent assertions and activities without becoming the canonical knowledge model. Stabilised assertions can subsequently participate in further candidate generation, making semantic stabilisation a recursive, provenance-bearing transformation between reviewable semantic states.
The Review Algebra is presented as a deliberately narrow computational component of the broader framework of semantic knowledge production developed in the preceding paper. It does not formalise the epistemology of observation, interpretation, or specialist judgement; it provides a common structure through which the assertions generated by heterogeneous human and computational activities can remain reviewable, provenance-bearing, and reusable.
Introduction
The preceding paper on semantic knowledge production (Section 1) examined the broader processes through which semantic assertions are produced. Human and computational agents may observe empirical evidence, reconcile entities with external knowledge systems, interpret linguistic or domain-specific material, assess legal or ethical conditions, or derive new candidates through inference. These activities differ in their evidence, methods, competences, permissions, and conditions of applicability.
The Review Algebra addresses a narrower computational problem. It does not attempt to formalise the epistemology of observation, interpretation, reconciliation, or specialist judgement. Instead, it asks how the semantic assertions produced through such heterogeneous activities can enter a common process of examination and progressive stabilisation without losing either their semantic histories or the provenance of the transformations between them.
The central question of this paper is therefore:
How can candidate semantic assertions be generated, reviewed, and progressively stabilised while preserving both their semantic states and the provenance of the transformations between them?
The algebra distinguishes three recurring operations: candidate generation, review, and finalisation. Candidate generation makes semantic assertions available for consideration. Review examines one or more components of a candidate assertion and returns a semantic value together with an explicit review status. Finalisation applies an explicit policy to determine which reviewed assertions, if any, are admitted to a subsequent stabilised semantic state.
Candidate generation is deliberately general. Observation generates candidates by examining evidence. Inference generates candidates from represented knowledge under an applicable rule. Reconciliation may generate candidates from external knowledge representations, while computational or human activities may propose classifications, identities, translations, relationships, or other assertions. These processes differ epistemically, but their outputs can participate in a common downstream review structure while their origins remain explicit in provenance.
This distinction allows the algebra to remain small. Observation, inference, reconciliation, and other knowledge-producing activities do not become additional elementary operators. They are different provenance-bearing mechanisms through which candidate assertions become available to the algebra.
The algebra consequently separates domain knowledge from provenance knowledge. Semantic assertions represent what is proposed, reviewed, or stabilised in the domain. Provenance represents the activities, agents, evidence, rules, methods, and policies through which those semantic states were produced. The Review Algebra connects these representations without replacing either the underlying domain ontology or a general provenance model such as W3C PROV.
Where an assertion requires an explicit scope, scope qualifies the semantic or applicability context within which the assertion is intended to hold. It is distinct from the evidentiary frame of an observation activity. Whether a photograph is examined as a digital manifestation, as a documentary object, or as evidence concerning a depicted entity or activity belongs to the provenance-bearing design of the knowledge-producing activity developed in the preceding paper.
A further distinction is made between review and finalisation. Review records what happened when a candidate was examined, including the semantic value returned and its review status. Finalisation determines the operational consequence of that result under an explicit policy. The authority to examine an assertion and the authority to admit it to a subsequent semantic state therefore need not be identical.
The resulting process is recursive. A stabilised semantic state can provide input to further candidate generation, particularly inference, while new empirical evidence may independently generate further candidates through observation. The resulting candidates can themselves be reviewed and finalised. Semantic stabilisation is therefore represented not as a terminal distinction between true and false statements but as a sequence of provenance-bearing transformations between reviewable semantic states.
The Review Algebra extends rather than replaces ordinary relational and tidy data operations. Candidate assertions, reviewed values, review statuses, and provenance can be represented in vectorised tabular structures, transformed between wide and long review workspaces, and persisted separately as domain assertions and provenance-bearing activities. Equivalent representations may be implemented using relational databases, RDF, property graphs, or other knowledge infrastructures.
The Livonian cultural heritage example introduced in the preceding paper continues here as a deliberately modest reference application. Its purpose is no longer to establish why empirical, linguistic, legal, or other scholarly activities constitute different forms of semantic knowledge production. Selected assertions instead provide concrete examples through which the operators and computational representations of the algebra can be developed.
The remainder of this paper defines the elementary objects and principles of the Review Algebra, formalises candidate generation, review, and finalisation, and examines the computational representations, persistence requirements, and implementation principles needed to realise these operations without losing semantic or provenance information.
These constitute the two planes formalised by the present algebra. The Domain Plane concerns the semantic assertions being produced and stabilised, while the Provenance Plane records the activities, agents, evidence, rules, and transformations through which those assertions acquire their reviewed status.
A third Normative Plane becomes relevant when stabilised semantic states and their provenance are used to determine whether an agent may, must, or must not perform an action. Although finalisation already introduces a limited policy-governed mechanism into the algebra, the present framework does not attempt a general formalisation of normative reasoning. Finalisation determines which reviewed assertions enter a subsequent stabilised semantic state; it does not by itself determine the permissibility, prohibition, or obligation of external actions on the basis of that state.
Running example
The running example illustrates candidate generation, review, finalisation, and the transformations between successive semantic states. The particular scholarly activities through which candidates arise are not themselves elements of the algebra. Observation, taxonomic classification, identification, linguistic interpretation, rights assessment, reconciliation, and inference may all generate or examine assertions, but their methods and conditions of validity belong to the broader processes of semantic knowledge production discussed in the preceding paper.
For the Review Algebra, the important common structure is that semantic assertions become available as candidates, are examined through provenance-bearing activities, and may subsequently be admitted to a stabilised semantic state. The same formal operations can therefore be illustrated using assertions concerning different semantic objects without prescribing a canonical sequence of domain-specific reviews.
Throughout the paper, we use a small collection of Livonian cultural heritage resources as a running example. Initially, the system contains observations concerning five digital resources.
| object | filename | observed file type |
|---|---|---|
| object1 | photo001.jpg | JPEG image |
| object2 | photo002.jpg | JPEG image |
| object3 | photo003.jpg | JPEG image |
| object4 | recording001.wav | WAV audio |
| object5 | recording002.wav | WAV audio |
These observations concern the digital resources as encountered in a computational environment. They provide technical knowledge that may itself be reviewed or used as evidence in subsequent knowledge-producing activities, but they do not establish what documentary or cultural objects the files instantiate or represent, what or whom those objects depict or record, or which semantic relationships are relevant to subsequent review.
For the benefit of the reader, successive knowledge-producing and review activities will eventually establish descriptions such as the following.
| object | semantic description established through knowledge production and review |
|---|---|
| object1 | Photograph depicting a building on the Livonian coast |
| object2 | Photograph depicting an unknown woman wearing Livonian folk costume |
| object3 | Portrait photograph depicting Hilda Grīva |
| object4 | Sound recording of a song composed by Kōrli Stalte and performed by Hilda Grīva |
| object5 | Sound recording of a traditional Livonian folk song performed by Julgi Stalte |
These descriptions are provided only so that the development of the running example can be followed. They are not initially available to the system as stabilised semantic knowledge.
The descriptions also deliberately combine assertions about different semantic objects. A technical assertion such as media_type = image/jpeg concerns a digital manifestation. The classification of a documentary object as a photograph concerns a different semantic subject. An assertion that the photograph depicts Hilda Grīva concerns a relationship between the photograph and a represented person, while subsequent assertions about Hilda Grīva concern that person rather than either the photograph or its digital manifestation.
These distinctions were developed in the preceding paper as part of the relationship between evidence, observation, and the object of knowledge production. The Review Algebra does not reproduce that model as part of the elementary assertion. Instead, it requires the semantic subject of an assertion to remain explicit while provenance records the activity, evidence, agent, method, and other conditions through which the assertion was generated or examined.
Where required, scope provides additional semantic or applicability context for an assertion. It does not identify what the evidence was observed as and does not substitute for the provenance relationship between evidence and a knowledge-producing activity.
As semantic knowledge accumulates, stabilised assertions may themselves become inputs to further candidate generation. An observation may generate a candidate from empirical evidence; reconciliation may propose an identity from an external authority; an applicable inference rule may generate a new candidate from already stabilised assertions. Regardless of origin, these candidates can enter the same review and finalisation operations while retaining distinct provenance.
The running example therefore develops recursively. Later assertions may depend on semantic states established earlier, while the provenance of each transformation remains explicit. Early activities may depend directly on multimodal evidence such as images, recordings, documents, or audiovisual material. Later activities may increasingly operate on already stabilised semantic assertions, allowing human attention to be concentrated on exceptions, ambiguities, conflicts, and questions requiring specialised judgement.
Literature Review
Colton and Pease’s theorem-modification system provides an instructive precedent for treating unsuccessful candidates as productive inputs to subsequent computational search rather than terminal failures (Colton and Pease 2005). Drawing on Lakatos (Lakatos, Worrall, and Zahar 1976), their system uses supporting examples and counterexamples to generate modified conjectures that can subsequently be tested. The Review Algebra addresses a more general and different problem: it does not prescribe how rejected candidates should be repaired, but records review outcomes and their provenance so that heterogeneous candidate-generation activities may use them in subsequent semantic production.
Nguyen, Wallace & Lease — decision-theoretic allocation of human review
Nguyen, Wallace, and Lease provide an empirical precedent for treating the allocation of heterogeneous human review as a decision-theoretic optimisation problem. Their active-learning system jointly selects which item should be labelled and whether it should be assigned to relatively inexpensive crowd workers or a costly domain expert, ranking actions by expected loss reduction relative to cost (Nguyen, Wallace, and Lease 2015). This corroborates a downstream consequence of the Review Algebra: once reviewable candidates, review states, and reviewer competences can be represented explicitly, the allocation of scarce human attention can itself become an Operations Research problem.
The boundary is important. Nguyen et al. address finite-pool classification and model crowd and expert labels principally through differences in cost and reliability, including a simplifying assumption of expert correctness. The Review Algebra does not equate heterogeneous semantic activities with noisy versions of a common labelling task: reviews may differ in evidence, competence, applicability, scope, method, and authority. Decision-theoretic review allocation is therefore a possible optimisation layer operating over states exposed by the algebra, not an elementary operator of the algebra itself.
A further implication concerns provenance. Nguyen et al. show that active selection creates sampling bias and use inverse-probability weighting to compensate for it. For provenance-aware semantic production, this suggests that an optimised review system may eventually need to record not only who reviewed an assertion and how, but also the allocation mechanism through which that assertion was selected for review.
Formalisation
The Review Algebra represents semantic stabilisation as transformations between reviewable semantic states. Candidate assertions are generated, examined through review activities, and admitted selectively to subsequent stabilised states through explicit finalisation policies.
The algebra is independent of any particular implementation. Its operations may be realised over tabular, relational, RDF, or graph representations, provided that assertion identity, semantic state, review outcome, and provenance can be preserved across the relevant transformations.
The formalisation maintains the distinction developed in the preceding paper between domain knowledge and provenance knowledge. Domain assertions represent what is proposed, reviewed, or stabilised. Provenance represents the activities, agents, evidence, rules, methods, and policies through which those semantic states are produced.
The Review Algebra does not attempt to formalise all knowledge-producing activities described in the preceding paper. Observation, reconciliation, interpretation, computational analysis, and inference may differ substantially in their epistemic basis. For the algebra, their common significance is that they may make semantic assertions available as candidates for review.
The algebra therefore defines three recurring operations:
| operation | function |
|---|---|
| Candidate generation | makes one or more semantic assertions available as candidates for review |
| Review | examines candidate assertions and returns semantic values together with explicit review statuses |
| Finalisation | applies a policy determining which reviewed semantic values enter a subsequent stabilised semantic state |
These operations have both semantic and provenance-bearing representations. The semantic representation records the candidate, reviewed, and stabilised assertions. The provenance representation records how those transformations occurred. The algebra connects these representations without replacing either the underlying domain ontology or a general provenance model such as W3C PROV.
Observation and inference are important mechanisms of candidate generation but are not additional elementary operators of the Review Algebra. Observation generates candidates through the examination of evidence. Inference generates candidates by applying an explicit rule to represented semantic knowledge under appropriate conditions of applicability. Other mechanisms may include reconciliation with external knowledge representations, import, statistical analysis, computational classification, or human proposal.
Core principles of the Review Algebra
The formalisation developed below follows four design principles.
First, the elementary reviewable object is a semantic assertion, represented in its simplest domain form as
\[(s,p,v),\]
where \(s\) denotes the subject, \(p\) the predicate, and \(v\) the asserted value.
Where required by the review problem, an assertion may additionally carry an explicit scope \(\sigma\):
\[(\sigma,s,p,v).\]
Scope qualifies the semantic or applicability context within which the assertion is intended to hold. It is not the evidentiary frame of an observation activity and does not identify what a resource was examined as. Those relationships belong to the provenance-bearing knowledge-producing activity described in the preceding paper.
Second, review acts on one or more components of a candidate assertion. A review need not be restricted to the asserted value \(v\): the subject, predicate, value, and, where applicable, scope may each be examined. For every component under review, the review records both the semantic value returned and the status of the review.
Third, the status of a review is distinct from the semantic value returned by it. A candidate may be corroborated, falsified, or deferred. A review may also produce an alternative semantic value. These are distinct pieces of information: the relationship between a candidate and a returned value does not in general determine whether the candidate was corroborated, falsified, or left unresolved.
This distinction follows the critical-rationalist interpretation of corroboration associated with Popper: a corroborated assertion is not established as permanently true but has survived a particular process of examination. Review therefore records the conditions under which an assertion has been corroborated rather than converting a candidate into an unquestionable fact.
The iterative character of the algebra is also compatible with Lakatos’s account of knowledge development through successive criticism, modification, and reconstruction. A falsified assertion may give rise to an alternative candidate; a deferred assertion may enter another review activity; and a previously corroborated assertion may later be challenged by new observations or evidence.
More specifically, the Review Algebra is indebted to the methodology developed by Imre Lakatos in Proofs and Refutations, which extends the critical tradition associated with Popper by examining how knowledge develops in the presence of anomalies and counterexamples. A counterexample does not necessarily require the abandonment of an entire conjecture or body of knowledge. It may instead reveal an exception, expose an inadequately specified domain, motivate a revised definition, or generate a new conjecture. This pragmatic treatment of incomplete and partially refuted knowledge has a close analogue in data and semantic workflows. We do not discard an otherwise useful tidy dataset because some observations contain missing or erroneous values, even though a complete and internally consistent table remains desirable. Likewise, the Review Algebra does not require an entire set of semantic assertions to be rejected because some cannot be corroborated or are shown to be erroneous. Problematic assertions can instead be isolated, challenged, revised, deferred, or replaced while the remaining semantic structure continues to be usable.
Semantic stabilisation is therefore represented as a sequence of reviewable transformations rather than as a terminal distinction between verified and falsified statements. Anomalies and refutations are productive elements of this process: they identify where existing semantic structures require qualification, correction, or further investigation without requiring otherwise useful knowledge to be discarded.
Fourth, review and finalisation are separate operations. Review records what an agent concluded about a candidate assertion, including the semantic value returned and its review status. Finalisation applies an explicit policy to these review results and determines which semantic values, if any, enter the subsequent stabilised semantic state. The authority to review and the authority to finalise may therefore be exercised by different agents or processes.
These principles are design commitments of the Review Algebra rather than requirements concerning how particular knowledge-producing activities must be implemented.
Elementary assertions
The elementary domain assertion of the Review Algebra is
\[(s,p,v),\]
where
- \(s\) denotes the subject,
- \(p\) denotes the predicate,
- \(v\) denotes the asserted value.
This representation corresponds naturally to an RDF triple, a statement in a property graph, or an equivalent representation in a relational or tabular data model.
Where the review problem requires additional semantic or applicability context, the review representation may retain an explicit scope \(\sigma\):
\[(\sigma,s,p,v).\]
Scope is therefore not assumed to be an intrinsic component of every domain assertion. It is an additional component of the review representation where the intended validity or interpretation of an assertion depends on an explicitly represented context.
A scoped semantic assertion may be projected onto its domain assertion,
\[(\sigma,s,p,v)\longrightarrow(s,p,v),\]
without requiring downstream knowledge representations to adopt the additional scope component.
The role of scope becomes particularly important when candidate-generation rules have restricted conditions of applicability. This relationship is considered after candidate generation has been defined.
Candidate generation
A semantic assertion must become available as a candidate before it can be reviewed. Candidate generation is the operation through which one or more semantic assertions become available for subsequent review.
Let \(\gamma\) denote a provenance-bearing candidate-generation process. At the most general level,
\[\mathcal{G}_{\gamma}\rightarrow C^{*},\]
where \(C^{*}\) denotes the set of candidate assertions produced by \(\gamma\).
The notation deliberately leaves the inputs to \(\gamma\) unspecified. Different candidate-generation mechanisms operate on different inputs. An observation may use empirical evidence; an inference may use stabilised assertions and an explicit rule; reconciliation may use an external knowledge representation; and a computational classifier may use a model together with a digital resource.
Candidate generation does not imply that the resulting assertions are corroborated or stabilised. It establishes only that they have become explicit candidates for consideration. Provenance records how, by whom or by what computational agent, and from which evidence, assertions, rules, models, or external resources they were produced.
Two particularly important mechanisms are observation and inference.
Observation as candidate generation
Observation is the empirical case developed in the preceding paper. An observation activity examines evidence and produces one or more candidate semantic assertions:
\[E \xrightarrow{\mathcal{O}} C^{*}.\]
For example, ExifTool may inspect a JPEG file and generate candidates concerning embedded technical metadata. A curator may examine a photograph and generate a candidate concerning a depicted person or activity. A music-information-retrieval system may analyse an audio recording and generate candidates concerning acoustic or musical features.
These activities differ in agent, method, competence, and evidentiary frame. The Review Algebra does not make them epistemically equivalent. It requires only that the assertions they generate can become explicit candidates while their different origins remain represented in provenance.
Observation is therefore one mechanism of candidate generation, not an additional elementary operation of the Review Algebra.
Inference as candidate generation
Inference is a second important mechanism of candidate generation. Whereas observation generates candidates through the examination of evidence, inference generates candidates from semantic knowledge that has already been represented.
Let \(S_k\) denote the stabilised semantic state available at stage \(k\), and let \(r\) denote an explicit inference rule. Where the conditions of applicability of \(r\) are satisfied by the knowledge represented in \(S_k\), inference may generate one or more new candidate assertions:
\[S_k \xrightarrow{\mathcal{I}_r} C^{*}.\]
The resulting assertions are candidates rather than automatically stabilised knowledge. Their provenance records the rule applied, the assertions or other resources used in the inference, the computational or human agent responsible for the activity, and the conditions under which the rule was considered applicable.
For example, suppose the stabilised semantic state contains assertions that a photograph depicts a particular person and that the person has been identified with an established authority record. An explicit rule may use these assertions to generate further candidates concerning names, relationships, dates, or other properties available from the reconciled authority. Similarly, a taxonomic classification may enable a rule that generates only those review questions applicable to objects of that class.
Inference therefore has the same downstream status as other mechanisms of candidate generation:
\[C_k \xrightarrow{\mathcal{I}_r} C^{*} \xrightarrow{\mathcal{R}_t} Q \xrightarrow{\mathcal{F}_\phi} C_{k+1}.\]
This does not imply that every inference requires human review. A finalisation policy may permit deterministic inferences produced under sufficiently constrained rules to enter a subsequent semantic state without additional human examination. The distinction remains important nevertheless: inference describes how a candidate was derived, while finalisation determines whether that result is authorised to enter the subsequent semantic state.
Observation and inference can consequently participate in the same algebra without being treated as epistemically equivalent:
| mechanism | principal input | basis of candidate generation | output |
|---|---|---|---|
| observation | evidence | examination by an agent using a method | candidate assertion |
| inference | represented semantic knowledge | application of an explicit rule | candidate assertion |
Other mechanisms, including reconciliation, import, computational classification, and human proposal, may likewise generate candidates. The Review Algebra does not require these mechanisms to share a common epistemology. It requires their outputs and provenance to be represented sufficiently explicitly for subsequent review and finalisation.
Scope and conditions of applicability
Scope qualifies the semantic or applicability context within which an assertion is intended to hold. It is distinct from the evidentiary frame of an observation activity and from the provenance of the activity through which an assertion was generated.
Where required, a scoped assertion is represented as
\[(\sigma,s,p,v),\]
where \(\sigma\) records contextual information necessary for interpreting or applying the assertion.
Scope may become particularly important when assertions participate in candidate-generation rules. A rule should be applied only where its conditions of applicability are satisfied. Scope can provide one source of information for determining those conditions, alongside the classes of the participating entities, explicitly represented relationships, temporal or geographical constraints, and other stabilised semantic knowledge.
For example, an assertion established within a particular collection, historical period, linguistic context, or analytical frame should not automatically be propagated beyond that context. Likewise, inheritance from a class to its members is warranted only where the relevant property is intended to be inherited and the conditions governing that inheritance are satisfied.
Scope therefore does not itself perform inference. Nor does it determine what evidence was examined or what an observation activity examined that evidence as. Its role is narrower: where an assertion requires explicit contextual qualification, scope preserves that qualification so that subsequent review, candidate generation, and inference do not silently generalise the assertion beyond the conditions under which it is intended to hold.
The distinction can be summarised as follows:
| concept | question answered |
|---|---|
| scope | Within what semantic or applicability context is this assertion intended to hold? |
| evidentiary frame | What is this activity examining the available evidence as evidence of? |
| inference rule | Under what explicit conditions may existing knowledge generate a new candidate? |
| provenance | Through which activity, agent, evidence, rule, or method was this assertion produced or examined? |
Keeping these concepts separate allows scope to remain an optional component of the semantic representation while observation, inference, and other knowledge-producing activities retain their richer conditions in provenance.
Review
Review is a provenance-bearing activity in which an agent examines one or more components of a candidate semantic assertion and returns, for each component under review, a semantic value together with an explicit review status.
For a reviewable component with candidate value \(x_c\), review may be written
\[\mathcal{R}(x_c)\rightarrow(x_r,q_x),\]
where
- \(x_c\) is the candidate presented for review,
- \(x_r\) is the semantic value returned by the review, and
- \(q_x\) is the review status.
For a complete domain assertion,
\[(s_c,p_c,v_c),\]
review may operate independently on any of its components and return
\[((s_r,q_s),(p_r,q_p),(v_r,q_v)).\]
Where scope is explicitly represented, \(\sigma\) may be reviewed in the same manner. Scope is not otherwise required as a component of every assertion.
The elementary review statuses are corroborated, falsified, and deferred.
A candidate is corroborated when it survives the review activity. This does not establish the candidate as permanently true; it records that the candidate has survived a particular examination performed by a particular agent under specified conditions.
A candidate is falsified when the review provides grounds for rejecting it. Falsification does not require the reviewer to know the correct alternative. The review may therefore return no alternative semantic value,
\[(x_c,\mathrm{NA},\texttt{falsified}),\]
or it may return an alternative,
\[(x_c,x_r,\texttt{falsified}),\qquad x_r\neq x_c.\]
The alternative value is semantic information produced during review. It is distinct from the falsification status and may itself become a candidate for subsequent review.
A candidate is deferred when the review does not currently establish whether it should be corroborated or falsified. Depending on the review design, the returned semantic value may retain the candidate,
\[(x_c,x_c,\texttt{deferred}),\]
or remain unassigned. The explicit status distinguishes deferral from corroboration, falsification, and missing information.
A review activity may also expose a component for which no candidate value is currently available. If an agent supplies a semantic value where
\[x_c=\mathrm{NA},\]
the activity has generated a new candidate:
\[(\mathrm{NA},x_r,\texttt{proposed}).\]
Here proposed does not describe what happened to a pre-existing candidate. It records that candidate generation occurred during the review activity. The resulting semantic value may subsequently be reviewed and finalised according to the applicable workflow.
A single review activity may examine more than one component of an assertion. For example, a reviewer may corroborate a candidate subject and predicate while falsifying the candidate value record and proposing photograph as an alternative. Component-level statuses preserve these judgements independently.
Review does not itself determine the subsequent stabilised semantic state. It records the semantic values returned by an examination and what happened to the candidates presented to it. Review is also distinct from verification and validation in standard AI terminology. Verification concerns objective evidence that specified requirements have been fulfilled, whereas validation concerns objective evidence that requirements for a specified intended use or application have been fulfilled (ISO/IEC 2022). A review activity in the present algebra can contribute evidence to either process, but corroborating a candidate does not by itself constitute verification or validation.
Finalisation is the separate operation that determines what, if anything, is carried forward. Finalisation should likewise not be conflated with contractual or engineering acceptance: acceptance can concern whether a specified result satisfies agreed acceptance conditions, whereas finalisation is the algebraic operation determining which semantic values enter a subsequent semantic state (ISO/IEC/IEEE 2017).
Review results can be represented conveniently in a wide claim table because candidate values, reviewed values, and review statuses are aligned within the same row. This workspace representation temporarily combines two kinds of knowledge: semantic assertions in the domain and provenance concerning the activities through which those assertions were examined.
As developed below under Dual representations, the persistent representation separates these dimensions again. Domain assertions are represented as semantic knowledge, while review activities, statuses, agents, evidence, methods, and other relevant information are represented as provenance knowledge. This separation is compatible with the established pattern of stand-off annotation, in which annotation is layered over primary data and serialised separately from the document containing those data (ISO/IEC 2024a). ### Finalisation
Finalisation applies an explicit policy to review results and determines which semantic values enter the subsequent stabilised semantic state.
Let a reviewed component be represented by
\[(x_c,x_r,q_x),\]
where \(x_c\) is the candidate value, \(x_r\) is the value returned by review, and \(q_x\) is its review status. A finalisation policy \(\phi\) determines the value, if any, that is carried forward:
\[\mathcal{F}_{\phi}(x_c,x_r,q_x)\rightarrow x_{k+1}.\]
Here, \(\phi\) denotes a finalisation policy internal to the Review Algebra: it determines how reviewed assertions are admitted to a subsequent stabilised semantic state. It should be distinguished from a broader normative policy environment \(\Pi\), under which stabilised knowledge may subsequently be evaluated in relation to contemplated actions.
Applying the finalisation policy is distinct from review. Review records the outcome of examining a candidate; finalisation determines the operational consequence of that outcome. The same review result may therefore produce different subsequent semantic states under different finalisation policies.
A simple policy might carry forward corroborated reviewed values, exclude falsified candidates, and leave deferred assertions unresolved. Another policy might retain a deferred candidate provisionally while preserving its deferred status in provenance. An alternative proposed during review may be admitted directly under one policy or returned as a candidate for further review under another.
Finalisation does not erase preceding semantic or epistemic states. Candidate assertions, reviewed values, review statuses, review activities, and finalisation policies remain available in the persistent semantic and provenance representations. Finalisation determines only which assertions are admitted to the next stabilised semantic state.
For a collection of review results \(Q_k\), finalisation can be written as
\[\mathcal{F}_{\phi}(Q_k)\rightarrow C_{k+1},\]
where \(C_{k+1}\) denotes the subsequent stabilised semantic state.
The complete review cycle can therefore be represented as
\[C_k \xrightarrow{\mathcal{G}_{\gamma}} C_k^{*} \xrightarrow{\mathcal{R}} Q_k \xrightarrow{\mathcal{F}_{\phi}} C_{k+1}.\]
The transformation is non-destructive. \(C_{k+1}\) represents the semantic state made available for subsequent use, while the candidate-generation, review, and finalisation activities remain represented in provenance.
The resulting state may itself provide input to further candidate generation. Stabilised assertions may, for example, satisfy the conditions of applicability of an inference rule and thereby generate new candidates. New empirical evidence may independently give rise to further observations and candidate assertions. These candidates can enter subsequent review and finalisation cycles.
Finalisation therefore closes one cycle of semantic stabilisation without terminating the broader process of semantic knowledge production.
Dual representations
The Review Algebra extends rather than replaces the relational and tidy algebras on which ordinary tabular data workflows are built. One of its motivations is precisely to retain their computational advantages: simple, deterministic, vectorised operations over explicitly structured observations. This is consistent with the algebraic treatment of provenance in relational databases, where richer annotations can be propagated through relational operations without replacing the underlying relational model (Green, Karvounarakis, and Tannen 2007). The Review Algebra extends this principle to a different problem: preserving review states and provenance-bearing transformations while semantic assertions remain available to ordinary relational and tidy operations.
Following Wickham’s formulation of tidy data, variables can be represented as columns, observations as rows, and different observational units in different tables. Once semantic review is represented explicitly, the same relational operations used for ordinary data transformation can also be applied to candidate assertions, reviewed values, review statuses, and provenance records. Filtering, joining, grouping, pivoting, and vectorised transformation therefore remain available rather than being replaced by a specialised semantic-processing environment.
The review algebra adds structure to this relational representation by distinguishing domain knowledge from provenance knowledge and by making the transformations between semantic states explicit. A review workspace may temporarily combine dimensions from both representations because this is convenient for a particular human or computational operation. The combined representation can subsequently be reshaped or disaggregated without losing the distinction between the underlying observational units.
Three representations are particularly useful.
A wide review representation aligns candidate values, reviewed values, and review statuses for rapid comparison and human interaction.
A long review representation makes review components and states explicit as observations and is convenient for filtering, grouping, joining, and vectorised computation.
A disaggregated representation separates domain assertions from the provenance-bearing activities through which those assertions were generated, reviewed, inferred, and finalised.
These are not competing data models. They are computationally equivalent projections of the same review process for different operations. Moving between them allows semantic review to retain the advantages of relational computation while preserving the distinction between what is asserted in the domain and how that assertion came to be established.
The identity of a semantic state should likewise be distinguished from its position in a review sequence. A semantic state is identified by the knowledge it represents, whereas the ordering of the activities through which that state was produced belongs to provenance. Meaningful state identifiers should therefore be preferred where their semantic role is known; generic identifiers such as review_1 or review_2 are implementation fallbacks rather than elements of the conceptual algebra.
Wide review representation
For human review, candidate values, reviewed values, and review statuses are conveniently aligned within the same row.
Wide review representation
For human review, candidate values, reviewed values, and review statuses are conveniently aligned within the same row.
| claim_id | subject_candidate | subject_reviewed | subject_status | predicate_candidate | predicate_reviewed | predicate_status | value_candidate | value_reviewed | value_status |
|---|---|---|---|---|---|---|---|---|---|
| claim1 | object1 |
object1 |
corroborated | instance_of |
instance_of |
corroborated | record |
photograph |
falsified |
| claim2 | object2 |
object2 |
corroborated | instance_of |
instance_of |
corroborated | photograph |
photograph |
corroborated |
This representation is deliberately redundant. Its purpose is comparative: the reviewer can see the candidate, the returned semantic value, and the review status together. HTML forms, spreadsheets, MediaWiki tables, and similar interfaces can therefore expose the dimensions required for a particular review without requiring the reviewer to navigate the persistent semantic and provenance representations separately.
Long review representation
The same workspace can be pivoted into a longer representation in which the components and review states become explicit observations.
| claim_id | component | candidate | reviewed | status |
|---|---|---|---|---|
| claim1 | subject | object1 |
object1 |
corroborated |
| claim1 | predicate | instance_of |
instance_of |
corroborated |
| claim1 | value | record |
photograph |
falsified |
| claim2 | subject | object2 |
object2 |
corroborated |
| claim2 | predicate | instance_of |
instance_of |
corroborated |
| claim2 | value | photograph |
photograph |
corroborated |
The long and wide forms contain the same review information. Their difference is organisational rather than epistemic. The wide representation is useful when components and states need to be compared side by side; the long representation is useful for filtering, grouping, composition, persistence, and vectorised computation.
Disaggregating domain and provenance knowledge
A different transformation occurs when the review workspace is disaggregated. Candidate and reviewed semantic values belong to the evolving domain representation, while review statuses and information about the activities that generated them belong to the provenance representation.
The example above can therefore be decomposed into an assertion representation such as
| assertion_id | claim_id | state | subject | predicate | value |
|---|---|---|---|---|---|
| a1 | claim1 | candidate | object1 |
instance_of |
record |
| a2 | claim1 | reviewed | object1 |
instance_of |
photograph |
| a3 | claim2 | candidate | object2 |
instance_of |
photograph |
| a4 | claim2 | reviewed | object2 |
instance_of |
photograph |
and a provenance representation such as
| activity_id | claim_id | activity | status | agent | used | generated |
|---|---|---|---|---|---|---|
| r1 | claim1 | review | falsified | curator1 |
a1 | a2 |
| r2 | claim2 | review | corroborated | curator1 |
a3 | a4 |
The two transformations should not be confused. Long and wide are alternative arrangements of a review workspace; domain and provenance are different kinds of knowledge. The former can be transformed through ordinary reshaping operations. The latter are normalised separately because their observations and cardinalities differ.
This structure is analogous to other block-structured data systems in which a combined representation is useful for computation while its constituent quadrants retain distinct interpretations. The review workspace similarly brings related dimensions together for a particular operation without implying that they constitute a single kind of observation.
The cardinalities of these representations need not coincide. A single observation or review activity may use or generate several semantic assertions, while a single assertion may participate in several successive or independent activities. Assertions and provenance activities are therefore different observational units and should be normalised separately.
The apparent alignment between them in a wide review workspace is a computational convenience rather than a property of the persistent representation.
Persistence
Wide and long representations are computational views of the review process. Persistence requires the distinct observational units exposed by disaggregation to retain stable identities and explicit relationships. In the terminology adopted here, persistent means continuing to exist until deliberately destroyed (ISO/IEC 2024b); this property should not be conflated with the more specialised concept of a persistent identifier, which provides permanent access to a digital object independently of its physical location or current ownership (ISO 2017).
At minimum, the persistent representation must distinguish semantic assertions, provenance-bearing activities, and the relationships through which assertions are used or generated by those activities.
| persistent object | purpose |
|---|---|
| assertion relation | records semantic assertions and their states |
| activity relation | records candidate-generation, observation, review, inference, and finalisation activities |
| assertion–activity relation | records how assertions are used or generated by activities |
The assertion relation preserves the semantic assertions participating in the process. Candidate assertions, reviewed alternatives, and assertions admitted to a stabilised semantic state retain their identities rather than being destructively overwritten. Their semantic content belongs to the underlying domain knowledge representation.
The activity relation records the provenance-bearing processes through which those assertions are produced or examined. Observation, inference, review, and finalisation may therefore be represented as activities with agents, timestamps, evidence, rules, software, policies, and other relevant provenance. Candidate generation may similarly be represented either as a general activity or through the more specific activity, such as observation or inference, by which the candidate was produced.
The assertion–activity relation connects these two kinds of information. An assertion may be used by an activity, generated by an activity, or participate in several successive activities. Conversely, a single activity may use or generate many assertions. This relation therefore preserves the many-to-many structure that is temporarily flattened in a review workspace.
This separation corresponds naturally to the dual representation introduced earlier. Domain assertions record what is asserted, while provenance records how those assertions came to be generated, examined, or admitted to a stabilised state. W3C PROV provides the general model for the latter: assertions and other evidence can be represented as entities, epistemic and computational operations as activities, and human or computational participants as agents (Gil et al. 2013; Lebo et al. 2013).
The persistent representation is therefore not tied to a particular physical storage model. The relations may be implemented as relational tables, RDF graphs, property graphs, or other structures capable of preserving the same identities and relationships. Wide and long review workspaces can be generated from this representation and returned to it without requiring the workspace itself to become the canonical data model.
Where a stabilised semantic object is admitted to a governed register, its registration is a distinct operation from semantic review or finalisation. ISO 19135 defines registration as assigning an unambiguous identifier to an approved register item and distinguishes the governed register from the information system on which it is maintained (ISO 2026). Semantic stabilisation, approval, registration, and technical persistence should therefore not be treated as synonymous operations.
Implementation principles
The review algebra is independent of a particular software implementation. Reference implementations should nevertheless preserve several computational properties that make the algebra useful in practical semantic workflows.
Reversibility principle. Workspace and persistent representations should be related through information-preserving transformations. A review workspace may expose only the dimensions required for a particular task, but its round-trip to the persistent representation must not destroy assertion identity, review state, or provenance.
Vectorisation principle. Review operations should be expressible over vectors of assertions rather than requiring assertion-by-assertion manipulation. This permits implementations to retain the computational advantages of relational databases, spreadsheets, and vectorised languages such as R and Python.
Interoperability principle. The algebra should not depend on a particular storage technology, ontology, or review interface. Equivalent semantic and provenance representations may therefore be persisted in relational databases, RDF, property graphs, or other interoperable representations.
Economy principle. Implementations should support purpose-driven review rather than indiscriminate metadata production. Candidate generation, inference, filtering, and prioritisation should make it possible to direct scarce human attention towards assertions whose review can reduce uncertainty relevant to the intended purpose.
Conclusion
This paper has developed a Review Algebra for representing how candidate semantic assertions are generated, examined, and progressively stabilised. The algebra addresses a specific computational problem within the broader process of semantic knowledge production: preserving transformations between semantic states while retaining the provenance-bearing activities through which those states were produced.
Three operations form the core of the algebra. Candidate generation makes semantic assertions available for consideration without implying that they are already corroborated. Review examines candidate assertions and returns semantic values together with explicit review statuses. Finalisation applies a policy determining which reviewed assertions enter a subsequent stabilised semantic state.
This deliberately small algebra accommodates heterogeneous processes of knowledge production without treating them as epistemically equivalent. Observation generates candidates through the examination of evidence; inference generates candidates from represented knowledge under explicit rules and conditions of applicability; reconciliation, import, computational analysis, and human proposal may generate candidates through still other processes. Their differences remain represented through provenance while the resulting assertions can participate in a common downstream review structure.
The distinction between reviewed value and review status is central to this structure. A candidate may be corroborated, falsified, or deferred independently of the semantic value returned by review. Falsification need not identify an alternative, while an alternative proposed during review constitutes newly generated semantic information rather than the falsification itself. Deferral likewise represents an explicit epistemic state rather than an absent value. Review can therefore preserve correction, disagreement, uncertainty, and incomplete knowledge without destructively overwriting earlier assertions.
Separating review from finalisation further distinguishes epistemic judgement from operational authority. A review activity records what happened when a candidate was examined; a finalisation policy determines the consequence of that result for the subsequent semantic state. Different policies may therefore operate over the same review results, and the authority to review an assertion need not imply the authority to finalise it.
The algebra is recursive. Stabilised assertions may satisfy the conditions of applicability of inference rules and thereby generate further candidates, while new empirical evidence may independently produce additional candidates through observation. Candidate generation, review, and finalisation can consequently recur over successive semantic states rather than forming a single linear workflow.
The Review Algebra also maintains a deliberate separation between domain knowledge and provenance knowledge. Assertions represent what is proposed, reviewed, or stabilised, while provenance represents the activities, agents, evidence, rules, methods, and policies through which those states were produced. Wide and long review workspaces provide useful computational and human-facing projections over these representations without becoming the canonical persistent knowledge model.
This separation allows the algebra to extend ordinary relational and tidy computation rather than replace it. Review operations can remain vectorised and compatible with filtering, joining, grouping, reshaping, and other familiar relational operations, while persistent representations can be implemented using relational databases, RDF, property graphs, or other interoperable infrastructures.
The boundary between the Review Algebra and the broader account of semantic knowledge production is itself an important result of the present work. Empirical observation, linguistic interpretation, legal assessment, and other specialist activities differ in evidence, competence, permission, method, and epistemic warrant. The Review Algebra should not attempt to formalise these differences as additional elementary operators. Its role is narrower: to provide a common computational structure through which the semantic assertions produced by heterogeneous activities can be generated as candidates, examined, finalised, and reused without erasing their provenance.
This division of responsibilities also clarifies the role of scope. Scope may qualify the semantic or applicability context within which an assertion is intended to hold, but it does not encode the evidentiary frame of an observation activity or substitute for provenance. Keeping these concepts separate allows the elementary semantic representation to remain small while richer conditions of knowledge production remain explicit in the activities through which assertions are generated and examined.
The reference implementation provides the next test of the algebra. Implementation can determine whether transformations between wide review workspaces, normalised assertion representations, and provenance records remain reversible and computationally tractable; whether component-level review semantics are sufficient in practice; and whether recursive candidate generation can be supported without introducing unnecessary complexity.
The Review Algebra should therefore be understood as a formal and computational component of a larger architecture for semantic knowledge production: narrow enough to be implemented and tested independently, but general enough to support heterogeneous human and computational processes whose semantic outputs must remain reviewable, provenance-bearing, and reusable.
Annex
Annex A: An Accounting-Oriented Review Workbench for Semantic Production
Purpose and status
This annex outlines a speculative human–computer interaction model for the Review Algebra and the wider semantic-production architecture developed in this paper. It is not part of the formal algebra. Rather, it considers how the underlying distinctions between stabilised knowledge, candidate knowledge, review scope, provenance, semantic authority, and possible downstream consequences could be projected into an intelligible review environment for human experts.
The proposal retains an intuition derived from input–output accounting: a governed semantic transition should be explainable in terms of what entered the review, what was actually examined, what decision was made, what acquired or lost production-level authority, and what consequences followed. The analogy is therefore one of accounting discipline, not a claim that semantic knowledge production obeys the mathematical structure of an economic input–output system.
Existing review interfaces can readily present candidate assertions, admissible values, contextual resources, and provenance metadata. Such interfaces are sufficient for recording individual review acts, but they provide limited information about the larger semantic state in which those acts occur. A reviewer may see what is being decided without seeing why the decision matters, what has already been governed, or which subsequent semantic activities the decision may enable.
A future review workbench could expose these relationships through a canonical four-quadrant projection:
┌───────────────────────────────────┬───────────────────────────────────┐
│ CURRENT / CANDIDATE STATE │ FORWARD CONSEQUENCES │
│ │ │
│ What am I considering? │ What would this decision enable? │
│ │ │
│ candidate assertion │ inferred candidates │
│ current production context │ applicable activities │
│ unresolved state │ dependencies resolved │
│ │ future review frontiers │
├───────────────────────────────────┼───────────────────────────────────┤
│ REVIEW SCOPE & EVIDENCE │ DECISION & AUTHORITY │
│ │ │
│ What did I actually examine? │ What am I deciding? │
│ │ │
│ evidence │ review outcome │
│ evidentiary frame │ semantic output │
│ review scope │ authority acquired or withdrawn │
│ competence / method │ transition towards S_(k+1) │
└───────────────────────────────────┴───────────────────────────────────┘
The quadrants do not constitute four matrices of the Review Algebra. They are four views over information maintained or derived by the semantic-production system. Their purpose is to make the state transition intelligible to the reviewer without requiring the reviewer to interact directly with the underlying graph, provenance model, or formal operators.
The accounting intuition
The proposed interface is organised around four questions:
What am I considering?
The reviewer needs to understand the candidate assertion and the relevant part of the current stabilised semantic state \(S_k\).What did I actually examine?
The system should preserve the evidentiary and semantic scope of the review activity: which evidence was examined, under which evidentiary frame, by which agent, using which competence or method, and which parts of the candidate were actually within the scope of judgement.What am I deciding?
Review outcome and semantic output must remain distinct. A candidate may be corroborated, falsified, or deferred; a review may also generate a new candidate assertion. The interface should therefore avoid representing a replacement value as though it were merely the disposition of the original candidate.What would this decision change or enable?
The reviewer may benefit from seeing the immediate semantic consequences of a possible decision: assertions that could become derivable, applicability conditions that would be satisfied, unresolved dependencies that would be removed, or further review activities that would become meaningful.
The accounting requirement is consequently stronger than a conventional progress indicator. A statement such as “59 of 400 triples reviewed” says little about the semantic coverage or consequences of human governance. A more informative projection might distinguish directly reviewed assertions, assertions admitted through authorised transformations, unresolved areas, deferred decisions, and exploratory candidates that remain outside the Production Graph.
The underlying invariant can be expressed without assuming numerical conservation:
No assertion should acquire or lose production-level semantic authority without an accountable provenance-bearing transition.
In this sense, semantic accounting requires explainable differences between successive production states. If \(S_k\) becomes \(S_{k+1}\), the system should be capable of accounting for assertions admitted, withdrawn, replaced, deprecated, or derived through authorised transformations. The relevant requirement is not that a numerical quantity remain constant, but that changes in semantic authority do not occur silently.
Architectural separation
The proposed workbench depends on a separation between four architectural concerns.
Review Algebra — formal transition model.
The Review Algebra represents candidate generation, review, finalisation, semantic states, review outcomes, semantic outputs, and the provenance-bearing transitions through which assertions may acquire production-level authority. The algebra should remain independent of any particular graphical interface or optimisation strategy.
Authority lineage — backward projection.
Authority lineage asks how assertions already present in the Production Graph acquired production-level semantic authority. An assertion may have been directly reviewed and finalised, or it may have been produced through an authorised transformation over stabilised premises. The backward projection reconstructs these relationships through provenance rather than assuming that every production assertion was individually inspected by a human.
Agent look-ahead — forward projection.
Agent look-ahead asks what could follow from the current stabilised state. Computational agents may explore hypothetical multi-step candidate paths in a logically isolated environment rooted in \(S_k\). Intermediate exploratory candidates may be used internally to search possible semantic trajectories, but they do not thereby become authoritative premises or acquire production-level status.
Review Workbench — HCI projection.
The workbench combines selected information from the current state, authority lineage, review activity, and agent look-ahead. Its purpose is not to expose the complete internal search or provenance graph. It should present enough information for a reviewer to understand the proposition under consideration, the evidence and scope of the requested judgement, the authority consequences of the decision, and selected downstream effects that make the decision relevant.
This separation allows the interaction model to evolve without making interface constructs such as quadrants, masks, matrices, spatial blocks, or progress indicators primitives of the Review Algebra.
Forward projection: agent look-ahead
The stabilised-state inference principle constrains the acquisition of semantic authority, not the computational depth of exploratory reasoning. Authoritative candidate-generating inference operates over the current stabilised state \(S_k\); unreviewed candidates generated within the current production cycle do not recursively become authoritative premises merely because they were generated by a computational agent.
This does not prevent an agent from exploring longer hypothetical paths internally. Starting from \(S_k\), an agent may construct speculative chains such as
\[S_k \longrightarrow c_1^* \longrightarrow c_2^* \longrightarrow c_3^* \longrightarrow \cdots\]
inside a logically isolated candidate environment. These paths represent possible semantic trajectories rather than successive production states. Their intermediate assertions may be used to evaluate what could become possible under different review decisions, but none acquires production-level authority through the exploratory computation itself.
An agent could evaluate and prune these paths using criteria such as:
- uncertainty reduction;
- expected information gain;
- structural consistency;
- resolution of unresolved dependencies;
- newly satisfied applicability conditions;
- newly enabled candidate-generating activities;
- expected review cost;
- relevance to an explicit scholarly, curatorial, organisational, or operational objective.
The workbench need not expose the resulting search tree. Its purpose would be to present selected consequences that help a reviewer understand the leverage of the decision currently requested.
For example, instead of merely asking whether a resource should be classified as a particular documentary object, the interface might indicate that corroborating the classification would satisfy several applicability conditions, resolve unresolved dependencies, enable further candidate generation, or make a small number of additional expert reviews meaningful.
The forward projection therefore asks:
What semantic possibilities become reachable if this candidate acquires production-level authority?
Review frontiers and allocation of human attention
Combining backward authority lineage with forward agent look-ahead changes the role of the review interface. The objective is not necessarily to expose the largest possible number of candidate assertions to human reviewers. It is to identify review frontiers at which appropriately competent human judgement has consequential value for the progression of semantic knowledge.
A review frontier may consist of a single candidate assertion, a small related group of assertions, an unresolved identity, an applicability condition, or another proposition whose resolution affects a larger region of possible semantic production. The workbench could therefore prioritise review tasks not only according to candidate uncertainty, but also according to their expected systemic leverage.
This creates a natural Operations Research problem. Given a stabilised state \(S_k\), a set of possible review activities \(\mathcal{R}\), limited human competence and review capacity, and an explicit objective, review allocation may be represented abstractly as seeking reviews with high expected semantic value relative to their cost:
\[\max_{r \in \mathcal{R}} \frac{\mathbb{E}[\text{semantic value produced or enabled by }r]} {\text{review cost}(r)},\]
subject to competence, permission, applicability, time, and other governance constraints.
The objective function remains an open research question. Depending on the application, semantic value might include resolved dependencies, reduction of uncertainty, newly applicable activities, authorised derivations enabled, contradictions resolved, improved coverage of a research question, or progress towards an institutional objective.
The interface need not expose this optimisation machinery directly. A reviewer might instead receive an intelligible explanation such as:
High-leverage review: resolving this relation would close three unresolved dependencies and make seven further assertions eligible for governed inference.
The computational system searches; the human reviewer judges within an appropriate domain of competence.
Spatial production awareness
The accounting projection may also support spatial representations of semantic governance. A conventional progress bar treats all reviewed assertions as equivalent units. A spatial interface could instead represent regions of the semantic state according to their governance and dependency structure.
For example, a workbench might visually distinguish:
- stabilised assertions established through direct review;
- stabilised assertions supported through authorised derivation;
- assertions currently under review;
- deferred or unresolved assertions;
- candidate assertions outside production authority;
- hypothetical downstream consequences exposed by agent look-ahead.
Such an interface could use blocks, regions, density, adjacency, or other visual encodings to make the structure of semantic production perceptible. A Tetris-like metaphor is one possible implementation: governed knowledge may appear as settled regions, unresolved dependencies as gaps, candidates as provisional elements, and downstream consequences as projected but not yet admitted structures.
The metaphor should not determine the underlying semantics. Its value would lie in helping a reviewer perceive that a local judgement participates in a larger semantic state transition.
A reviewer might therefore see not merely:
59 / 400 assertions reviewed
but a representation such as:
Current governed semantic region
Directly reviewed assertions 12
Supported through authorised transformations 47
Currently awaiting review 4
Deferred dependencies 3
Exploratory candidates 26
The precise counts and categories would depend on the application and on the authority-lineage model eventually adopted. Their purpose is to communicate semantic governance rather than merely task completion.
Pre-commitment impact preview
The four-quadrant projection is particularly useful before a review decision is finalised. The reviewer can be shown both the proposition under consideration and a bounded account of its predicted consequences.
Conceptually:
CURRENT / CANDIDATE
|
| evidence + production context
v
REVIEW FRONTIER
|
+---- if corroborated ----> projected consequences
|
+---- if falsified -------> projected consequences
|
+---- if deferred --------> unresolved dependencies remain
The displayed consequences remain predictions over candidate space. They must not be confused with assertions already present in the Production Graph.
This distinction makes it possible for computational agents to provide substantial assistance without silently determining semantic authority. The system can say what is likely to follow if this proposition is accepted while leaving the relevant review judgement to an appropriately competent human or other authorised governance mechanism.
Review outcome and semantic output
The HCI accounting model must preserve a distinction that is fundamental to the Review Algebra: review outcome is not semantic output.
If a reviewer determines that the candidate
\[(\text{birch tree},\ \text{measured height},\ 28)\]
is false and proposes the alternative value \(29\), the review activity has produced two conceptually different results:
- the candidate value \(28\) has been falsified; and
- a new candidate assertion concerning the value \(29\) has been generated.
The interface may present these actions together for convenience, but the underlying model should not encode \(29\) as though it were the review disposition of \(28\). The alternative assertion has its own semantic identity and provenance and may require its own review or finalisation conditions.
A future workbench might display the distinction as:
Candidate under review
Birch tree — measured height — 28
Review outcome
Falsified
Semantic output of review
New candidate: Birch tree — measured height — 29
Projected authority consequence
Original candidate excluded from finalisation
New candidate enters the appropriate review/finalisation path
This allows streamlined correction workflows without collapsing epistemic judgement and semantic production into the same operation.
From data-entry form to governance cockpit
The proposed workbench extends the role of the review interface from recording isolated decisions towards supporting awareness of semantic state transitions.
Its four quadrants answer complementary questions:
┌──────────────────────────────┬──────────────────────────────┐
│ What am I considering? │ What could this enable? │
│ │ │
│ semantic state │ semantic consequences │
├──────────────────────────────┼──────────────────────────────┤
│ What did I examine? │ What am I authorising? │
│ │ │
│ epistemic scope │ semantic authority │
└──────────────────────────────┴──────────────────────────────┘
Horizontally, the interface moves from epistemic input towards semantic consequence. Vertically, it connects the graph-level context with the provenance-bearing human review activity.
The resulting workbench can therefore provide two complementary projections around the current Production Graph:
Production Graph S_k
/ \
/ \
backward projection forward projection
authority lineage agent look-ahead
\ /
\ /
review frontier
|
review + finalisation
|
v
S_(k+1)
Backward projection explains why current production knowledge has authority. Forward projection explores what semantic trajectories may become possible next. Review and finalisation form the governed frontier through which selected possibilities may contribute to a subsequent stabilised state.
Research agenda
This annex deliberately leaves several elements open.
The first is the formal representation of review scope. A matrix or mask may prove useful for particular tabular implementations, but review scope may ultimately require a richer representation capable of expressing which assertions, components, evidence resources, temporal intervals, or graph regions were actually examined.
The second is the formalisation of authority lineage. Logical closure matrices provide one possible computational representation for restricted classes of inference, but a general authority-lineage model must also accommodate heterogeneous provenance-bearing activities, rule versions, competence constraints, applicability conditions, revisions, and non-monotonic changes.
The third is the measurement of semantic leverage. Information gain, uncertainty reduction, dependency resolution, graph connectivity, and review cost offer possible components of an Operations Research objective, but their relevance and weighting are domain-dependent and require empirical evaluation.
The fourth is the HCI representation itself. The proposed four-quadrant accounting view and spatial or Tetris-like metaphors should be treated as design hypotheses. Their usefulness must be evaluated with actual reviewers, including whether consequence previews improve judgement or instead create anchoring effects, automation bias, or excessive cognitive load.
Finally, the relationship between direct human review and indirect human governance requires further formalisation. Human reviewers are not epistemic oracles, and authority lineage should not be interpreted as proof of truth. Its purpose is to make the governance of production-level semantic commitments inspectable, attributable, and revisable.
The central hypothesis of this annex is therefore narrower than a claim that semantic production can be represented as an input–output economy. It is that a human-governed semantic-production system benefits from an accounting surface: an interface capable of showing what is being considered, what was actually examined, what authority a decision would confer or withdraw, and what semantic consequences may follow. The Review Algebra provides the state-transition machinery beneath that surface; authority lineage and agent look-ahead provide its backward- and forward-looking projections; and the review workbench makes those relationships available to human judgement.



