Difference between revisions of "Archivematica 1.0 Micro-services"

From Archivematica
Jump to navigation Jump to search
 
(5 intermediate revisions by one other user not shown)
Line 1: Line 1:
 
[[Main Page]] > [[Documentation]] > [[Technical Architecture]] > [[Micro-services]] > Archivematica 1.0 Micro-services
 
[[Main Page]] > [[Documentation]] > [[Technical Architecture]] > [[Micro-services]] > Archivematica 1.0 Micro-services
  
A micro-service may consist of a number of discrete tasks, or jobs. In the Archivematica 1.o dashboard, micro-services are always shown, while jobs may be viewed by expanding the micro-service (i.e. by clicking on the grey background behind the micro-service name).
+
<blockquote style="background-color:orange;">
 +
'''Note: The documentation on this page is out of date. Please see the [https://www.archivematica.org/docs official documentation page] for the latest.'''
 +
</blockquote>
 +
 
 +
A micro-service may consist of a number of discrete tasks, or jobs. In the Archivematica 1.0 dashboard, micro-services are always shown, while jobs may be viewed by expanding the micro-service (i.e. by clicking on the grey background behind the micro-service name).
  
 
</br>
 
</br>
  
[[Image:NormPresAccess-10.png|600px|center|thumb|Arhivematica dashboard showing a micro-service and its jobs]]
+
[[Image:Micro-services1.png|600px|center|thumb|Archivematica dashboard showing a micro-service and its jobs]]
  
 
</br>
 
</br>
Line 153: Line 157:
 
</div>
 
</div>
 
| Verifies that the SIP conforms to the folder structure required for processing in Archivematica. The structure is as follows: ''/logs/'', ''/metadata/'', ''/metadata/submissionDocumentation/'', ''/objects/''.
 
| Verifies that the SIP conforms to the folder structure required for processing in Archivematica. The structure is as follows: ''/logs/'', ''/metadata/'', ''/metadata/submissionDocumentation/'', ''/objects/''.
 +
|-
 +
|  <div class="toccolours mw-collapsible mw-collapsed">'''Verify transfer compliance'''<div class="mw-collapsible-content">
 +
<br/>Verify mets_structmap.xml compliance
 +
</div>
 +
</div>
 +
| Verifies the METS from the transfer.
 
|-  
 
|-  
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Rename SIP directory with SIP UUID'''<div class="mw-collapsible-content">
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Rename SIP directory with SIP UUID'''<div class="mw-collapsible-content">
 
<br/>Rename SIP directory with SIP UUID
 
<br/>Rename SIP directory with SIP UUID
 +
<br/>Check if SIP is from Maildir Transfer
 
</div>
 
</div>
 
</div>
 
</div>
| Directly associates the SIP with its metadata by appending the SIP UUID to the SIP directory name.
+
| Directly associates the SIP with its metadata by appending the SIP UUID to the SIP directory name and checks if SIP is from Maildir transfer type to determine workflow.
 
|-  
 
|-  
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Include default SIP processingMCP.xml'''<div class="mw-collapsible-content">
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Include default SIP processingMCP.xml'''<div class="mw-collapsible-content">
Line 164: Line 175:
 
</div>
 
</div>
 
</div>
 
</div>
| Copies the processing config file added to the transfer in '''Include default Transfer processingMCP.xml''', above, to the SIP.
+
| Copies the processing configuration file added to the transfer in '''Include default Transfer processingMCP.xml''', above, to the SIP.
 
|-  
 
|-  
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Remove cache files'''<div class="mw-collapsible-content">
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Remove cache files'''<div class="mw-collapsible-content">
Line 170: Line 181:
 
</div>
 
</div>
 
</div>
 
</div>
| Removes any Thumbs.db files.
+
| Removes any thumbs.db files.
 +
|-
 +
|  <div class="toccolours mw-collapsible mw-collapsed">'''Clean up names'''<div class="mw-collapsible-content">
 +
<br/>Sanitize SIP name
 +
<br/>Set file permissions
 +
</div>
 +
</div>
 +
| Some file systems do not support unicode or other special characters in filenames. This micro-service removes prohibited characters and replaces them with dashes. Original filenames are preserved in the PREMIS metadata.
 
|-  
 
|-  
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
Line 176: Line 194:
 
<br/>Check for Service directory
 
<br/>Check for Service directory
 
<br/>Check for Access directory
 
<br/>Check for Access directory
 +
<br/>Set remove preservation and access normalized files to renormalize link.
 +
<br/>Grant normalization options for no pre-existing DIP
 +
<br/>Move to workFlowDecisions-createDip directory
 
<br/>Find options to normalize as
 
<br/>Find options to normalize as
 +
<br/>Set resume link after tool selected
 +
<br/>Move to select file ID tool
 +
<br/>Select pre-normalize file format identification commant
 +
<br/>Identify file format
 +
<br/>Resume after normalization file identification tool selected
 +
<br/>Normalize
 +
<br/>Move to processing directory
 
<br/>Create DIP directory
 
<br/>Create DIP directory
<br/>Normalize
+
<br/>Create thumbnails directory
 +
<br/>Normalize thumbnails
 
<br/>Normalize access
 
<br/>Normalize access
 
<br/>Normalize preservation
 
<br/>Normalize preservation
 +
<br/>Set file permissions
 
<br/>Remove files without linking information (failed normalization artifacts etc.)
 
<br/>Remove files without linking information (failed normalization artifacts etc.)
<br/>Create thumbnails directory
+
<br/>Move to approve normalization directory
<br/>Normalize thumbnails
 
<br/>Verify checksums generated on ingest
 
 
<br/>Approve normalization
 
<br/>Approve normalization
<br/>Check for manual normalized files
+
<br/>Load post approve normalization link
<br/>Move access files to DIP
+
<br/>Set resume link after handling any manually normalized files
<br/>Assign UUIDs to manually normalized preservation files
+
<br/>Move to processing directory
<br/>Assign checksums to manually normalized preservation files
+
<br/>Set file permissions
<br/>Run FITS on manually normalized preservation files
+
<br/>Check for manually normalized files
<br/>Relate manually normalized preservation files to the original files
+
<br/>Load finished with manual normalized link
 
</div>
 
</div>
 
</div>
 
</div>
Line 197: Line 225:
 
|-  
 
|-  
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Process submission documentation'''<div class="mw-collapsible-content">
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Process submission documentation'''<div class="mw-collapsible-content">
 +
<br/>Copy transfer submission documentation
 
<br/>Check for submission documentation
 
<br/>Check for submission documentation
<br/>Copy transfer submission documentation
+
<br/>Move submission documentation into objects directory
<br/>Move to processing directory
 
<br/>Set file permissions
 
 
<br/>Assign file UUIDs to submission documentation
 
<br/>Assign file UUIDs to submission documentation
 
<br/>Assign checksums and file sizes to submissionDocumentation
 
<br/>Assign checksums and file sizes to submissionDocumentation
<br/>Extract packages in submission documentation
 
 
<br/>Sanitize file and directory names in submission documentation
 
<br/>Sanitize file and directory names in submission documentation
 
<br/>Scan for viruses in submission documentation
 
<br/>Scan for viruses in submission documentation
 
<br/>Characterize and extract metadata on submission documentation
 
<br/>Characterize and extract metadata on submission documentation
<br/>Identify Files ByExtension
 
<br/>Normalize submission documentation to preservation format
 
 
</div>
 
</div>
 
</div>
 
</div>
 
| Processes any submission documentation included in the SIP and adds it to the ''/objects/'' directory.
 
| Processes any submission documentation included in the SIP and adds it to the ''/objects/'' directory.
 +
|-
 +
|  <div class="toccolours mw-collapsible mw-collapsed">'''Process metadata directory'''<div class="mw-collapsible-content">
 +
<br/>Set resume link after processing metadata directory
 +
<br/>Move metadata to objects directory
 +
<br/>Assign file UUIDs to metadata
 +
<br/>Assign checksums and file sized to metadata
 +
<br/>Sanitize file and directory names in metadata
 +
<br/>GScan for viruses in metadata
 +
<br/>Load finished with metadata processing link
 +
</div>
 +
</div>
 +
| Processes metadata.
 
|-  
 
|-  
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Prepare DIP'''<div class="mw-collapsible-content">
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Prepare DIP'''<div class="mw-collapsible-content">
Line 248: Line 284:
 
|-  
 
|-  
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Prepare AIP'''<div class="mw-collapsible-content">
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Prepare AIP'''<div class="mw-collapsible-content">
 +
<br/>Remove files without linking information (failed normalization artifacts, etc.)
 +
<br/>Verify checksums generated on ingest
 
<br/>Copy transfers metadata and logs
 
<br/>Copy transfers metadata and logs
<br/>Remove files without linking information (failed normalization artifacts etc.)
+
<br/>Remove empty manual normalization directories
<br/>Verify checksums generated on ingest
 
 
<br/>Generate METS.xml document
 
<br/>Generate METS.xml document
 
<br/>Index AIP contents
 
<br/>Index AIP contents
 
<br/>Prepare AIP   
 
<br/>Prepare AIP   
 +
<br/>Move to compressionAIPDecisions directory
 
<br/>Select compression algorithm
 
<br/>Select compression algorithm
 
<br/>Select compression level
 
<br/>Select compression level
 
<br/>Compress AIP  
 
<br/>Compress AIP  
 +
<br/>Create AIP pointer file
 
<br/>Set bag file permissions
 
<br/>Set bag file permissions
 
<br/>Removed bagged files
 
<br/>Removed bagged files
 
</div>
 
</div>
 
</div>
 
</div>
| Creates an AIP in Bagit format. Indexes the AIP, then losslessly compresses it.
+
| Creates an AIP in Bagit format. Creates the AIP pointer file. Indexes the AIP, then losslessly compresses it.
 
|-  
 
|-  
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Store AIP'''<div class="mw-collapsible-content">
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Store AIP'''<div class="mw-collapsible-content">
 
<br/>Move to the store AIP approval directory
 
<br/>Move to the store AIP approval directory
 
<br/>Store AIP     
 
<br/>Store AIP     
 +
<br/>Retrieve AIP Storage Locations
 
<br/>Store AIP location
 
<br/>Store AIP location
 +
<br/>Move to processing directory
 +
<br/>Verify AIP
 
<br/>Store the AIP
 
<br/>Store the AIP
 +
<br/>Index AIP
 +
<br/>Remove processing directory
 
</div>
 
</div>
 
</div>
 
</div>

Latest revision as of 14:13, 11 December 2019

Main Page > Documentation > Technical Architecture > Micro-services > Archivematica 1.0 Micro-services

Note: The documentation on this page is out of date. Please see the official documentation page for the latest.

A micro-service may consist of a number of discrete tasks, or jobs. In the Archivematica 1.0 dashboard, micro-services are always shown, while jobs may be viewed by expanding the micro-service (i.e. by clicking on the grey background behind the micro-service name).


Archivematica dashboard showing a micro-service and its jobs


The table below shows micro-services and jobs in Archivematica 1.0. Note that this is is only a list of micro-services; detailed user instructions are available in the user manual.

Transfer[edit]

Micro-service Description
Approve Transfer

Set file permissions

This is the approval step that moves the transfer into the Archivematica processing pipeline.
Verify transfer compliance


Set file permissions
Move to processing directory
Set transfer type: (Standard, Zipped bag, Unzipped bag, DSpace, Maildir)
Remove hidden files and directories
Remove unneeded files
Attempt restructure for compliance
Verify transfer compliance
Verify mets_structmap.xml compliance

Moves the transfer to a processing directory based on selected transfer type (standard, zipped bag, unzipped bag, DSPace export or maildir). Verifies that the transfer conforms to the folder structure required for processing in Archivematica and restructures if required. The structure is as follows: /logs/, /metadata/, /metadata/submissionDocumentation/, /objects/.
Rename with transfer UUID


Rename with transfer UUID

Directly associates the transfer with its metadata by appending the transfer UUID to the transfer directory name.
Include default Transfer processingMCP.xml


Include default Transfer processingMCP.xml

Adds a file named processingMCP.xml to the root of the transfer. This is a configurable xml file to pre-configure processing decisions. It can configure workflow options such as creating transfer backups, quarantining the transfer and selecting a SIP creation option.
Assign file UUIDs and checksums


Set file permissions
Assign file UUIDs to objects
Assign checksums and file sizes to objects

Assigns a unique universal identifier and sha-256 checksum to each file in the /objects/ directory and sets file permission to allow for continued processing.
Verify transfer checksums


Verify metadata directory checksums

Checks any checksum files that were placed in the /metadata/ folder of the transfer prior to moving the transfer into Archivematica.
Generate METS.xml document


Generate METS.xml document

Generates a basic METS file with a fileSec and structMap to record the presence of all objects in the /objects/ directory and their locations in any subdirectories. Designed to capture the original order of the transfer in the event the user chooses subsequently to delete, rename or move files or break the transfer into multiple SIPs. A copy of the METS file is automatically added to any SIP generated from the transfer.
Quarantine


Workflow decision - send transfer to quarantine
Move to quarantine
Remove from quarantine

Quarantine's the transfer for a set duration, to allow virus definitions to update, before virus scan.
Scan for viruses


Scan for viruses

Uses ClamAV to scan for viruses and other malware. If a virus is found, the transfer is automatically placed in /sharedDirectoryStructure/failed/ and all processing on the transfer is stopped.
Clean up names


Sanitize object's file and directory names
Sanitize Transfer name

Some file systems do not support unicode or other special characters in filenames. This micro-service removes prohibited characters and replaces them with dashes. Original filenames are preserved in the PREMIS metadata.
Identify file format


Move to select file ID tool
Select file format identification command
Determine which files to identify
Identify file format

Identifies formats of the objects in the transfer using either FIDO or file extension based on user choice. Format types are managed in the Format Policy Registry. This micro-service can be skipped and done in Ingest instead.
Extract packages


Extract contents from compressed archives

Extracts objects from any zipped files or other packages. Extracts attachments from maildir transfers.
Characterize and extract metadata


Characterize and extract metadata
Load labels from metadata/file_labels.csv
Set file permissions
Check for specialized processing

Characterizes and validates formats and extracts object metadata using the File Information Tool Set (FITS).
Complete transfer


Index transfer contents
Move to SIP creation directory for completed transfers

Indexes transfer contents, then marks the transfer as complete.
Create SIP from Transfer


Check transfer directory for objects
Load options to create SIPs
Create SIP(s)
Move to processing directory
Create SIP from transfer objects
Send transfer to backlog
Move to SIP creation directory for completed transfers

This is the approval step that moves the transfer to the SIP packaging micro-services (Ingest) if user chooses to Create single SIP and continue processing. User can also choose to Send transfer to backlog at this time.


Ingest[edit]

Micro-service Description
Verify SIP compliance


Set file permissions
Move to processing directory
Verify SIP compliance

Verifies that the SIP conforms to the folder structure required for processing in Archivematica. The structure is as follows: /logs/, /metadata/, /metadata/submissionDocumentation/, /objects/.
Verify transfer compliance


Verify mets_structmap.xml compliance

Verifies the METS from the transfer.
Rename SIP directory with SIP UUID


Rename SIP directory with SIP UUID
Check if SIP is from Maildir Transfer

Directly associates the SIP with its metadata by appending the SIP UUID to the SIP directory name and checks if SIP is from Maildir transfer type to determine workflow.
Include default SIP processingMCP.xml


Include default SIP processingMCP.xml

Copies the processing configuration file added to the transfer in Include default Transfer processingMCP.xml, above, to the SIP.
Remove cache files


Remove cache files

Removes any thumbs.db files.
Clean up names


Sanitize SIP name
Set file permissions

Some file systems do not support unicode or other special characters in filenames. This micro-service removes prohibited characters and replaces them with dashes. Original filenames are preserved in the PREMIS metadata.
Normalize


Identify manually normalized files
Check for Service directory
Check for Access directory
Set remove preservation and access normalized files to renormalize link.
Grant normalization options for no pre-existing DIP
Move to workFlowDecisions-createDip directory
Find options to normalize as
Set resume link after tool selected
Move to select file ID tool
Select pre-normalize file format identification commant
Identify file format
Resume after normalization file identification tool selected
Normalize
Move to processing directory
Create DIP directory
Create thumbnails directory
Normalize thumbnails
Normalize access
Normalize preservation
Set file permissions
Remove files without linking information (failed normalization artifacts etc.)
Move to approve normalization directory
Approve normalization
Load post approve normalization link
Set resume link after handling any manually normalized files
Move to processing directory
Set file permissions
Check for manually normalized files
Load finished with manual normalized link

Determines which normalization options are available for the SIP and presents them to the user as choices. Normalizes (i.e. generates preservation and/or access copies) based on selection. Thumbnail files are also generated during this micro-service.
Process submission documentation


Copy transfer submission documentation
Check for submission documentation
Move submission documentation into objects directory
Assign file UUIDs to submission documentation
Assign checksums and file sizes to submissionDocumentation
Sanitize file and directory names in submission documentation
Scan for viruses in submission documentation
Characterize and extract metadata on submission documentation

Processes any submission documentation included in the SIP and adds it to the /objects/ directory.
Process metadata directory


Set resume link after processing metadata directory
Move metadata to objects directory
Assign file UUIDs to metadata
Assign checksums and file sized to metadata
Sanitize file and directory names in metadata
GScan for viruses in metadata
Load finished with metadata processing link

Processes metadata.
Prepare DIP


Copy thumbnails to DIP directory
Copy METS to DIP directory
Set file permissions
Generate DIP

Creates a DIP containing access copies of the objects, thumbnails and a copy of the METS file.
Upload DIP


Upload DIP

Allows the user to choose to upload the DIP to either ICA-AtoM or CONTENTdm.
Upload DIP to ICA-AtoM


Upload DIP
Move to the uploadedDIPs directory

The user uploads the DIP to a selected description in ICA-AtoM.
Upload DIP to CONTENTdm


Restructure DIP for CONTENTdm upload
Select upload type (Project Client or direct upload)
Select target CONTENTdm server
Get list of collections on server
Select destination collection
Upload DIP to contentDM
Move to the uploadedDIPs directory

The user uploads the DIP to a selected description in CONTENTdm.
Prepare AIP


Remove files without linking information (failed normalization artifacts, etc.)
Verify checksums generated on ingest
Copy transfers metadata and logs
Remove empty manual normalization directories
Generate METS.xml document
Index AIP contents
Prepare AIP
Move to compressionAIPDecisions directory
Select compression algorithm
Select compression level
Compress AIP
Create AIP pointer file
Set bag file permissions
Removed bagged files

Creates an AIP in Bagit format. Creates the AIP pointer file. Indexes the AIP, then losslessly compresses it.
Store AIP


Move to the store AIP approval directory
Store AIP
Retrieve AIP Storage Locations
Store AIP location
Move to processing directory
Verify AIP
Store the AIP
Index AIP
Remove processing directory

Moves the AIP to /sharedDirectoryStructure/www/AIPsStore/ or another specified directory. Once the AIP has been stored, a copy of it is extracted from storage to a local temp directory, where it is subjected to standard BagIt checks: verifyvalid, checkpayloadoxum, verifycomplete, verifypayloadmanifests, verifytagmanifests.