Difference between revisions of "Archivematica 0.9 Micro-services"

From Archivematica
Jump to navigation Jump to search
Line 267: Line 267:
 
!style="width:70%"|'''Description'''
 
!style="width:70%"|'''Description'''
 
|-
 
|-
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Approve SIP creation'''<div class="mw-collapsible-content">
 
<br/>Approve SIP Creation
 
</div>
 
</div>
 
|
 
 
|-  
 
|-  
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Verify SIP compliance'''<div class="mw-collapsible-content">
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Verify SIP compliance'''<div class="mw-collapsible-content">
 
<br/>Set file permissions
 
<br/>Set file permissions
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Verify SIP compliance'''<div class="mw-collapsible-content">
 
 
<br/>Move to processing directory
 
<br/>Move to processing directory
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Verify SIP compliance'''<div class="mw-collapsible-content">
 
 
<br/>Verify SIP compliance
 
<br/>Verify SIP compliance
 
</div>
 
</div>
 
</div>
 
</div>
|
+
| Verifies that the SIP conforms to the folder structure required for processing in Archivematica. The structure is as follows: /logs/, /metadata/, /metadata/submissionDocumentation/, /objects/.
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Verify transfer compliance'''<div class="mw-collapsible-content">
 
<br/>Verify mets_structmap.xml compliance
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Failed compliance'''<div class="mw-collapsible-content">
 
<br/>Failed compliance. See output in dashboard. SIP moved back to SIPsUnderConstruction
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Rename SIP directory with SIP UUID'''<div class="mw-collapsible-content">
 
<br/>Rename SIP directory with SIP UUID
 
</div>
 
</div>
 
|
 
 
|-  
 
|-  
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Include default SIP processingMCP.xml'''<div class="mw-collapsible-content">
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Include default SIP processingMCP.xml'''<div class="mw-collapsible-content">
Line 314: Line 280:
 
</div>
 
</div>
 
</div>
 
</div>
|
+
| Includes a file named processingMCP.xml in the root of the transfer. This is a configurable xml file to pre-configure processing decisions. It can configure workflow options such as creating transfer backups, quarantining the transfer and selecting a SIP creation option.
 
|-  
 
|-  
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Remove cache files'''<div class="mw-collapsible-content">
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Remove cache files'''<div class="mw-collapsible-content">
Line 320: Line 286:
 
</div>
 
</div>
 
</div>
 
</div>
|
+
| Removes any Thumbs.db files
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Clean up names'''<div class="mw-collapsible-content">
 
<br/>Set file permissions
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Clean up names'''<div class="mw-collapsible-content">
 
<br/>Sanitize SIP name
 
</div>
 
</div>
 
|
 
 
|-  
 
|-  
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Check for Service directory
 
<br/>Check for Service directory
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
 
<br/>Check for Access directory
 
<br/>Check for Access directory
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Grant normalization options for pre-existing DIP
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Grant normalization options for no pre-existing DIP
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Move to workFlowDecisions-createDip directory
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
 
<br/>Find options to normalize as
 
<br/>Find options to normalize as
</div>
+
<br/>Create DIP directory
</div>
+
<br/>Normalize
|
+
<br/>Normalize access
|-
+
<br/>Normalize preservation
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Normalize  
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Process submission documentation'''<div class="mw-collapsible-content">
 
 
<br/>Remove files without linking information (failed normalization artifacts etc.)
 
<br/>Remove files without linking information (failed normalization artifacts etc.)
 +
<br/>Create thumbnails directory
 +
<br/>Normalize thumbnails
 +
<br/>Verify checksums generated on ingest
 +
<br/>Approve normalization
 
</div>
 
</div>
 
</div>
 
</div>
|
+
| Determins which normalization options are available for the SIP and presents them to the user as choices. Normalizes based on selection.
 +
* Creates access copies to be added to the DIP.
 +
* Creates preservation copies to be included in the AIP along with the original files.
 +
* Creates thumbnail files.
 +
If desired, the user can verify the quality of normalized files in /sharedDirectoryStructure/watchedDirectories/approveNormalization/.
 
|-  
 
|-  
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Process submission documentation'''<div class="mw-collapsible-content">
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Process submission documentation'''<div class="mw-collapsible-content">
<br/>Normalize submission documentation to thumbnail format
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Verify checksums generated on ingest
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
 
<br/>Copy transfers metadata and logs
 
<br/>Copy transfers metadata and logs
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Generate METS.xml document
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Normalize service files for access
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Normalize access
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Create DIP directory
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Normalize for preservation and access
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Normalize preservation
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
 
<br/>Remove files without linking information (failed normalization artifacts etc.)
 
<br/>Remove files without linking information (failed normalization artifacts etc.)
</div>
+
<br/>Normalize submission documentation to thumbnail format
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Move to approve normalization directory
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Approve normalization
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Move to processing directory
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Set file permissions
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Create thumbnails directory
 
</div>
 
</div>
 
|
 
|-
 
|  <div class="toccolours mw-collapsible mw-collapsed">'''Normalize'''<div class="mw-collapsible-content">
 
<br/>Normalize thumbnails
 
 
</div>
 
</div>
 
</div>
 
</div>

Revision as of 08:04, 24 August 2012

Main Page > Documentation > Technical Architecture > Micro-services > Archivematica 0.9 Micro-services

This page describes the key micro-services that are undertaken during transfer and ingest.

Transfer

Micro-service Description
Approve transfer Once the transfer has all its digital objects and has been formatted for processing, the user selects "Transfer complete" from the Actions drop-down menu.
Verify transfer compliance Verifies that the transfer conforms to the folder structure required for processing in Archivematica. The structure is as follows: /logs/, /metadata/, /metadata/submissionDocumentation/, /objects/.
Rename with transfer UUID Adds a unique universal identifier to the transfer folder name.
Include default Transfer processingMCP.xml Adds defaultTransferProcessing.xml file from /sharedDirectoryStructure/sharedMicroServiceTasksConfigs/ to the transfer directory. This xml file can configure workflow options such as creating transfer backups, quarantining the transfer and selecting a SIP creation option.
Workflow decision - create transfer backup The user can choose to create a complete backup of the transfer in case transfer or ingest are interrupted or fail. The transfer backup is placed in /sharedDirectoryStructure/transferBackups/ and will automatically be deleted once the AIP has been moved into storage.
Assign file UUIDs to objects Assigns a unique universal identifier to each file in the /objects/ directory.
Assign checksums and file sizes to objects Assigns a sha-256 checksum to each file in the /objects/ directory and calculates file sizes.
Verify metadata directory checksums Checks any checksum files that were placed in the /metadata/ folder of the SIP prior to ingest. Note that the filenames need to be named based on their algorithm: checksum.sha1, checksum.sha256, checksum.md5.
Generate METS.xml document Generates a basic mets file with a fileSec and structMap to record the presence of all objects in the /objects/ directory and their location in any subdirectories. Designed to capture the original order of the transfer in the event the user chooses subsequently to delete, rename or move files or break the transfer into multiple SIPs. The mets file is automatically added to any SIP generated from the transfer.
Extract packages Extracts objects from any zipped files or other packages.
Scan for viruses Uses ClamAV, parses the output and creates a PREMIS event. If a virus is found, the SIP is automatically placed in /sharedDirectoryStructure/failed/.
Sanitize object's file and directory names Some file systems do not support unicode or other special characters in filenames. This micro-service removes prohibited characters and replaces them with dashes. Original filenames are preserved in the PREMIS metadata.
Sanitize transfer name Same as above except does it for the transfer folder name.
Characterize and extract metadata Identifies and validates formats and extracts object metadata using the File Information Tool Set (FITS). Adds output to the PREMIS metadata.
Create SIP(s) The user chooses among a number of SIP creation options. See the user manual for details.


Ingest

Micro-service Description
Verify SIP compliance Verifies that the SIP conforms to the folder structure required for processing in Archivematica. The structure is as follows: /logs/, /metadata/, /metadata/submissionDocumentation/, /objects/.
Rename SIP directory with SIP UUID Adds a unique universal identifier to the SIP folder name.
Include default SIP processingMCP.xml Adds defaultSIPProcessing.xml file from /sharedDirectoryStructure/sharedMicroServiceTasksConfigs/ to the transfer directory. This xml file canbe used to configure SIP workflow options.
Remove thumbs.db files Removes any Thumbs.db files. May be expanded to others in future releases.
Sanitize object's file and directory names If user created new folder and/or file names during SIP creation, any prhobited characters are removed from these names and replaced with dashes.
Sanitize SIP name Same as above except that it's done for the SIP folder name.
Check for Service directory For digitization output workflows, checks to see if the SIP contains any service copies of master files. See the user manual for details.
Check for Access directory For digitization output workflows, checks to see if the SIP contains any access copies of master files. See the user manual for details.
Normalize The user can choose to create preservation and/or access copies of the ingested files based on rules in the transcoder database. These rules can be seen under the Preservation planning tab in the Archivematica dashboard.
Normalize access Creates access copies to be added to the DIP.
Normalize preservation Creates preservation copies to be included in the AIP along with the original files.
Approve normalization If desired, the user can verify the quality of normalized files in /sharedDirectoryStructure/watchedDirectories/approveNormalization/.
Check for submission documentation Checks for files in /metadata/submissionDocumentation/.
Move submission documentation to objects directory Moves the /submissionDocumentation/ directory from /metadata/ to /objects/ for ingest processing.
Assign file UUIDs to submission documentation Assigns a unique universal identifier to each file in the /objects/submissionDocumentation/ directory.
Assign checksums and filesizes to submission documentation Assigns a sha-256 checksum to each file in the /objects/submissionDocumentation directory and calculates file sizes.
Extract packages in submission documentation Extracts objects from any zipped files or other packages in the /objects/ directory.
Characterize and extract metadata on submission documentation Identifies and validates formats and extracts object metadata for files in the /objects/submissionDocumentation directory using the File Information Tool Set (FITS). Adds output to the PREMIS metadata.
Normalize submission documentation to preservation format Creates preservation copies of all files in the /objects/submissionDocumentation directory to be included in the AIP along with the original files.
Verify checksums generated on ingest Verifies checksums that were generated during transfer processing to ensure that the files have not been corrupted during transfer or ingest.
Remove empty directories Removes any empty directories from the SIP.
Generate METS.xml document Generates a METS file with PREMIS metadata. For more information on the METS file, see METS.
Copy transfers metadata and logs Copies all submission documentation included in the original transfer to the SIP. Copies all logs generated during transfer processing to the SIP.
Copy METS to DIP directory Creates a copy of the METS file in the DIP directory.
Generate DIP Moves the DIP to the DIP upload directory.
Upload DIP The user uploads the DIP to a selected description in the access system. See the user manual for details.
Prepare AIP Packages the SIP into an AIP using BagIt
Compress AIP Losslessly compresses the AIP for storage using p7zip.
Store AIP Moves the AIP to a specified directory. In the demo version of Archivematica the directory is /sharedDirectoryStructure/www/AIPsStore/. In other environments it can be a remote network mounted directory. The directory structure of the AIP store contains UUID quad directories and an index.html file listing the AIPs in storage. The index.html file is displayed in the dashboard in the Archival storage tab.

Once the AIP has been stored, a copy of the AIP is extracted from storage to a local temp directory, and is validated with the various BagIt checks: verifyvalid, checkpayloadoxum, verifycomplete, verifypayloadmanifests, verifytagmanifests.


Transfer

Micro-service Description
Approve Transfer

Set file permissions

Once the transfer has all its digital objects and has been formatted for processing, the user selects "Transfer complete" from the Actions drop-down menu.
Verify transfer compliance


Set file permissions
Move to processing directory
Set transfer type
Remove hidden files and directories
Attempt restructure for compliance
Verify transfer compliance

Verifies that the transfer conforms to the folder structure required for processing in Archivematica. The structure is as follows: /logs/, /metadata/, /metadata/submissionDocumentation/, /objects/. Standard transfers are restructured for compliance if pessary.
Rename with transfer UUID


Rename with transfer UUID

Directly associate the transfer with it's metadata by appending the transfer UUID to the transfer directory name.
Include default Transfer processingMCP.xml


Include default Transfer processingMCP.xml

Includes a file named processingMCP.xml in the root of the transfer. This is a configurable xml file to pre-configure processing decisions. It can configure workflow options such as creating transfer backups, quarantining the transfer and selecting a SIP creation option.
Assign file UUIDs and checksums


Assign file UUIDs to objects
Assign checksums and file sizes to objects

Assigns a unique universal identifier and sha-256 checksum to each file in the /objects/ directory.
Verify transfer checksums


Verify metadata directory checksums

Checks any checksum files that were placed in the /metadata/ folder of the SIP prior to ingest. Note that the filenames need to be named based on their algorithm: checksum.sha1, checksum.sha256, checksum.md5.
Generate METS.xml document


Generate METS.xml document

Generates a basic mets file with a fileSec and structMap to record the presence of all objects in the /objects/ directory and their location in any subdirectories. Designed to capture the original order of the transfer in the event the user chooses subsequently to delete, rename or move files or break the transfer into multiple SIPs. The mets file is automatically added to any SIP generated from the transfer.
Extract packages


Extract packages
Extract attachments

Extracts objects from any zipped files or other packages. Extracts attachments from maildir transfers.
Scan for viruses


Scan for viruses

Uses ClamAV, parses the output and creates a PREMIS event. If a virus is found, the SIP is automatically placed in /sharedDirectoryStructure/failed/.
Clean up names


Sanitize object's file and directory names
Sanitize Transfer name

Some file systems do not support unicode or other special characters in filenames. This micro-service removes prohibited characters and replaces them with dashes. Original filenames are preserved in the PREMIS metadata.
Characterize and extract metadata


Characterize and extract metadata
Identify Files ByExtension
Load labels from metadata/file_labels.csv

Identifies and validates formats and extracts object metadata using the File Information Tool Set (FITS). Adds output to the PREMIS metadata. Also identities file extensions, these identifications are used for normalization.
Complete transfer


Index transfer contents
Set file permissions
Move to completedTransfers directory

Index transfer contents, before marking the transfer as complete.
Create SIP from Transfer


Check transfer directory for objects
Create SIP(s)
Create SIP from transfer objects

The user chooses among a number of SIP creation options. See the user manual for details.

Ingest

Micro-service Description
Verify SIP compliance


Set file permissions
Move to processing directory
Verify SIP compliance

Verifies that the SIP conforms to the folder structure required for processing in Archivematica. The structure is as follows: /logs/, /metadata/, /metadata/submissionDocumentation/, /objects/.
Include default SIP processingMCP.xml


Include default SIP processingMCP.xml

Includes a file named processingMCP.xml in the root of the transfer. This is a configurable xml file to pre-configure processing decisions. It can configure workflow options such as creating transfer backups, quarantining the transfer and selecting a SIP creation option.
Remove cache files


Remove cache files

Removes any Thumbs.db files
Normalize


Check for Service directory
Check for Access directory
Find options to normalize as
Create DIP directory
Normalize
Normalize access
Normalize preservation
Remove files without linking information (failed normalization artifacts etc.)
Create thumbnails directory
Normalize thumbnails
Verify checksums generated on ingest
Approve normalization

Determins which normalization options are available for the SIP and presents them to the user as choices. Normalizes based on selection.
  • Creates access copies to be added to the DIP.
  • Creates preservation copies to be included in the AIP along with the original files.
  • Creates thumbnail files.

If desired, the user can verify the quality of normalized files in /sharedDirectoryStructure/watchedDirectories/approveNormalization/.

Process submission documentation


Copy transfers metadata and logs
Remove files without linking information (failed normalization artifacts etc.)
Normalize submission documentation to thumbnail format

Process submission documentation


Move to processing directory

Process submission documentation


Set file permissions

Process submission documentation


Copy transfer submission documentation

Process submission documentation


Check for submission documentation

Process submission documentation


Move submission documentation into objects directory

Process submission documentation


Assign file UUIDs to submission documentation

Process submission documentation


Assign checksums and file sizes to submissionDocumentation

Process submission documentation


Extract packages in submission documentation

Process submission documentation


Sanitize file and directory names in submission documentation

Process submission documentation


Scan for viruses in submission documentation

Process submission documentation


Characterize and extract metadata on submission documentation

Process submission documentation


Identify Files ByExtension

Process submission documentation


Normalize submission documentation to preservation format

Prepare AIP


Remove files without linking information (failed normalization artifacts etc.)

Prepare AIP


Verify checksums generated on ingest

Prepare AIP


Copy transfers metadata and logs

Prepare AIP


Generate METS.xml document

Prepare DIP


Copy thumbnails to DIP directory

Prepare DIP


Copy METS to DIP directory

Prepare DIP


Set file permissions

Prepare DIP


Generate DIP

Prepare AIP


Index AIP contents

Prepare AIP


Prepare AIP

Prepare AIP


Select compression algorithm

Prepare AIP


Select compression level

Prepare AIP


Compress AIP

Prepare AIP


Set bag file permissions

Prepare AIP


Removed bagged files

Store AIP


Move to the store AIP approval directory

Store AIP


Store AIP

Store AIP


Store AIP location

Store AIP


Move to processing directory

Store AIP


Store the AIP

Store AIP


Remove the processing directory

Upload DIP


Select target CONTENTdm server

Upload DIP


Get list of collections on server

Upload DIP


Select destination collection

Upload DIP


Select upload type (Project Client or direct upload)

Upload DIP


Restructure DIP for CONTENTdm upload

Upload DIP


Upload DIP to contentDM

Upload DIP


Upload DIP

Upload DIP


Move to the uploadedDIPs directory

Reject DIP


Move to the rejected directory

Reject SIP


Move to the rejected directory

Failed SIP


Move to the failed directory

Normalize


Find preservation links to run.

Normalize


Find access links to run.

Normalize


Find thumbnail links to run.

Process submission documentation


Find thumbnail links to run.

Process submission documentation


Find preservation links to run. |