Identifier Management (CPP-005)

CPP-IdentifierCPP-005
CPP-LabelIdentifier Management
AuthorMikko Laukkanen, Juha Lehtonen
ContributorsBertrand Caron, Johan Kylander
EvaluatorsFelix Burger, Maria Benauer
Change historyComments
Version 1.0.0 - 2025-08-29Milestone version
Version 1.1.0 - 2026-04-01Migration to XML and different minor vocabulary corrections

1. Description of the CPP

Identifiers are assigned to Objects, Information packages and/or Metadata, and managed along to their life cycle.

Inputs and outputs

Input(s)
Data
Information package
Object
Metadata
Technical metadata
Descriptive metadata
Documentation/guidance
Identifier creation and management policy
Output(s)
Metadata
Identifier-enriched Information package, Object(s) or Metadata
Provenance metadata

Definition and scope

Identifier management is the process of creating and updating identifiers and assigning them to Objects, Metadata or Information packages. Identifiers are essential components of digital preservation systems, serving as stable, long-term references to Digital Objects that can remain valid even when the Objects themselves are moved, renamed, or migrated to new systems.

Identifiers must be managed throughout the entire life cycle, taking into account any changes to their associated Objects, Metadata or Information packages. It is important to consider that not all types of identifiers are globally unique, some are unique only within their own identifier system.

A Persistent Identifier (PID) system can be used to generate unequivocal [1] identifiers to ensure that Objects can be precisely identified worldwide. PIDs are machine-readable strings of characters that conform to a defined scheme. Through providing and updating the reference link in the Metadata, these identifiers prevent the fundamental problem of "link rot" and ensure reliable access to preserved Digital Objects over time. However, this requires continuous maintenance of the identifiers to keep the Metadata up-to-date. Depending on the use case, it may be useful to assign multiple identifiers from different systems to an entity. To be able to provide user-facing PIDs, a TDA must manage local identifiers which provide the minimal baseline for providing persistent access and control to the data.

Common examples of PIDs are Digital Object Identifier (DOI), Uniform Resource Name (URN), handles and Archival Resource Key (ARK). One advantage of using PIDs is that their Metadata can be used to not only provide information about the Object itself, but also about its status, access conditions, and storage location. Even Objects which are not publicly accessible or have been disposed, can be identified and described by a PID. PIDs can also be moved from one organisation’s administration to another.

All types of identifiers can be assigned to multiple levels of entities, creating a hierarchical identification structure that reflects the complex nature of digital collections and their preservation requirements. Identifiers are usually assigned on the level of a) Objects, b) collections and aggregations, and c) Information Packages (AIPs and DIPs). In addition, identifiers can be assigned to Metadata entities, collections of other related entities, and even institutions or persons.

Identifiers and their Metadata should be updated according to the entity’s life cycle. In particular, when an entity may be deleted, merged, split or become partially unavailable, its identifier should be preserved. Moreover, its Provenance metadata should be updated in order to provide proper detail of information to the end users about its initial entity as well as the relationships to potential new entities that were created from the initial one. When using PIDs, some changes (e.g. the creation of identical parallel copies of the data that create new internal identifiers for each copy) can be documented in the PID version Metadata without creating a new PID.

Identifiers in a TDA are created at specific strategic points throughout the preservation life cycle, with timing and methods varying based on institutional policies and system architectures. Identifiers are typically assigned during Ingest as part of the packaging process, ensuring that every preserved Object has a persistent reference from the moment it enters the system. However, some institutions create identifiers earlier in the workflow (e.g. during acquisition planning or transfer preparation). This is especially useful when using PIDs, since it allows for early referencing and tracking of Objects before they undergo preservation processing. Identifiers can also be created after the initial preservation processing is complete, particularly when the final preserved format and structure have been determined. Identifiers can be also assigned to services or Objects which are not stored in the TDA but only generated on the fly based on user requests.

Identifiers may reveal the hierarchical relationships in the identifier string (e.g. by using qualifiers [2]) or might hide them by creating a whole new string for components [3]. This CPP does not choose between these approaches. Similarly, it does not make any assumptions on the organisation in charge of managing identifiers and whether identifiers are managed by the TDA directly. Since Identifier Management is relatively resource-intensive and can also be performed outside the scope of digital long-term preservation, no assumptions are made here about the structure or organisation of this work area; instead, reference is made only to the entity “the identifier management service”.

[1]This term is preferred over the “unique” adjective applied traditionally to identifiers. Indeed, it suggests that an identifier must reference one and only one thing, while “unique” might suggest that the thing must be referenced by only one identifier

[2] For example, the identifier <id:c8b> will be assigned to a Representation, and <id:c8b/001> to its first component or file.

[3] In the previous example, the identifier <id:t5g> could then be assigned to the first component or file.

Process description

Trigger event(s)

Trigger EventCPP-identifier
Pre-ingest transfer preparationCPP-029 (Ingest)
Ingestion workflowCPP-029 (Ingest)
Creation of new Files or RepresentationsCPP-028 (Creation of Derivatives)
Replacement of corrupted FilesCPP-004 (Data Corruption Management)
Data exportCPP-006 (AIP Batch Export)
Data replicationCPP-011 (Replication)
Data migrationCPP-014 (File Migration)
Data normalisationCPP-026 (File Normalisation)
Metadata ingest and creationCPP-016 (Metadata Ingest and Management)
Data version updateCPP-021 (AIP Versioning)
Broken File needs a new identifierCPP-027 (File Repair)
Information package, File or Metadata is removed from the TDA holdingsCPP-017 (Disposal)

Step-by-step description

NoSupplierInputStepsOutputCustomer
 sequence
1 alternative
1.aA producer or a TDA has a need to reserve an identifier, (e.g. a PID, prior to the entity being added) Reservation of identifier prior to new entity assignment (step 2)
1.b Object New entity added or a need to assign an identifier to an existing entity (step 2)
Information package
Metadata
1.c Object Entity with an identifier has changed (step 4)
Information package
Metadata
1.d Object Entity is disposed (step 7)
Information package
Metadata
2 Identifier creation and management policy Create a new identifier according to the TDAs policy for identifier management Identifier
3 (new) Identifier Assign the new identifier to the entity and add it as a part of the entity’s Metadata MetadataCPP-004 (Data Corruption Management)
CPP-011 (Replication)
CPP-014 (File Migration)
CPP-016 (Metadata Ingest and Management)
CPP-021 (AIP Versioning)
CPP-025 (Enabling Access)
CPP-027 (File Repair)
CPP-028 (Creation of Derivatives)
CPP-029 (Ingest)
Object
Information package
Metadata
4 Identifier creation and management policy If the changed entity has a PID assigned to it: evaluate if a new PID is requiredNeed for new PID identified (go back to step 2)
No need for new PID identified (step 5)
5Update or add identifier relationships (e.g. hierarchical relations, sequential relations etc.) for the assigned entity Metadata
6Update Provenance metadata for the entity so that identifiers have a historyProvenance metadata
7CPP-017 (Disposal) Disposed entity If the entity is disposed: retain minimum metadata and the identifier for the disposed entity

Rationale(s) and worst case(s)

RationaleImpact of inaction or failure of the process
The rationale for implementing PIDs in TDAs stems from fundamental challenges in maintaining long-term access to digital objects and the core mission of preservation itself. Link rot as well as problems and challenges in:
  • Internal data management problems;
  • System migrations;
  • Format migrations;
  • Activity tracking;
  • Interoperability

2. Dependencies and relationships with other CPPs

Dependencies

CPP-IDCPP-TitleRelationship description
///

Other relations

RelationCPP-IDCPP-TitleRelationship description
Required byCPP-016Metadata Ingest and ManagementWhile ingesting into a TDA, the Metadata should be assigned an identifier. Also, the management functions of the Metadata may require replacing and/or updating identifiers.
Required byCPP-017DisposalWhen the life cycle of the Digital Object or File ends, the identifier should be updated to “retired” status.
Required byCPP-021AIP VersioningWhen an AIP gets a new version, the new AIP version must also be assigned a new identifier.
Required byCPP-024Enabling DiscoveryEnabling Discovery should make use of identifiers.
Required byCPP-025Enabling AccessAccessing Digital Object, File(s) or Metadata should be based on identifiers as parameters.
Required byCPP-029IngestThe ingestion workflow is responsible for assigning identifiers to various entities in TDA, such as Files and Metadata.
May be required byCPP-004Data Corruption ManagementIf a File is corrupted, it may need to be repaired or replaced. During this process, a new identifier may be created.
May be required byCPP-011ReplicationWhen a Digital Object or File is replicated, the replica may be assigned a new identifier.
May be required byCPP-013Object Management ReportingThe management and reporting should require that the data is identified with identifiers.
May be required byCPP-014File MigrationDuring File Migration, the migrated File may be assigned a new identifier.
May be required byCPP-019Data Quality AssessmentThe data quality assessment may include validating the identifiers and their linked resources.
May be required byCPP-026File NormalisationA normalised File format may be assigned with a new identifier.
May be required byCPP-027File RepairA repaired File may get a new identifier.
May be required byCPP-028Creation of DerivativesA derivative of a File may get its own identifier.

4. Reference implementations

Use cases

DOI given for a research dataset by TDA

Institutional background
InstitutionCSC, Finland (Digital Preservation Service for Research Data), FI
Hyperlinkhttps://wiki.eduuni.fi/x/9ZRYH

https://doi.org/10.23729/e399409c-ac6f-4893-a93f-9d0c10bf00bd
Description
Trigger eventSubmitting research data to the TDA
Problem statementResearch dataset does not have a DOI
Proposed solutionBefore submitting a dataset to TDA (DPS in Finland), the user describes the dataset via a description tool or via a metadata API. When a dataset has been submitted, TDA automatically creates a DataCite description including a new DOI, and eventually it creates a corresponding publicly available website for the dataset Metadata.

Publicly available documentation

InstitutionOrganisation typeLanguageHyperlink
TIB – Leibniz Information Centre for Science and Technology and University Library, DENational library
Non-commercial digital preservation service
Research infrastructure
Research performing organisation
English https://wiki.tib.eu/confluence/spaces/lza/pages/93608951/Metadata#Metadata-Identifyingmetadata
CSC – IT Center for Science Ltd., FINon-commercial digital preservation serviceEnglish https://urn.fi/urn:nbn:fi-fe2020100578094
(section 2.4.1.)
Archivematica, CADigital preservation systemEnglish https://www.archivematica.org/en/docs/archivematica-1.17/user-manual/transfer/transfer/#transfer-tab-microservices