Difference between revisions of "Archivematica 0.8 Micro-services"
Jump to navigation
Jump to search
(17 intermediate revisions by 2 users not shown) | |||
Line 1: | Line 1: | ||
[[Main Page]] > [[Documentation]] > [[Technical Architecture]] > [[Micro-services]] > Archivematica 0.8 Micro-services | [[Main Page]] > [[Documentation]] > [[Technical Architecture]] > [[Micro-services]] > Archivematica 0.8 Micro-services | ||
− | + | <blockquote style="background-color:orange;"> | |
+ | '''Note: The documentation on this page is out of date. Please see the [https://www.archivematica.org/docs official documentation page] for the latest.''' | ||
+ | </blockquote> | ||
− | =Transfer | + | This page describes the key micro-services that are undertaken during transfer and ingest. |
+ | |||
+ | =Transfer= | ||
+ | |||
+ | {| border="1" cellpadding="10" cellspacing="0" width=90% | ||
+ | |- | ||
+ | |- style="background-color:#cccccc;" | ||
+ | !style="width:30%"|'''Micro-service''' | ||
+ | !style="width:70%"|'''Description''' | ||
+ | |- | ||
+ | |Approve transfer | ||
+ | |Once the transfer has all its digital objects and has been formatted for processing, the user selects "Transfer complete" from the Actions drop-down menu. | ||
+ | |- | ||
+ | |Verify transfer compliance | ||
+ | |Verifies that the transfer conforms to the folder structure required for processing in Archivematica. The structure is as follows: ''/logs/'', ''/metadata/'', ''/metadata/submissionDocumentation/'', ''/objects/''. | ||
+ | |- | ||
+ | |Rename with transfer UUID | ||
+ | |Adds a unique universal identifier to the transfer folder name. | ||
+ | |- | ||
+ | |Include default Transfer processingMCP.xml | ||
+ | |Adds ''defaultTransferProcessing.xml'' file from ''/sharedDirectoryStructure/sharedMicroServiceTasksConfigs/'' to the transfer directory. This xml file can configure workflow options such as creating transfer backups, quarantining the transfer and selecting a SIP creation option. | ||
+ | |- | ||
+ | |Workflow decision - create transfer backup | ||
+ | |The user can choose to create a complete backup of the transfer in case transfer or ingest are interrupted or fail. The transfer backup is placed in ''/sharedDirectoryStructure/transferBackups/'' and will automatically be deleted once the AIP has been moved into storage. | ||
+ | |- | ||
+ | |Assign file UUIDs to objects | ||
+ | |Assigns a unique universal identifier to each file in the ''/objects/'' directory. | ||
+ | |- | ||
+ | |Assign checksums and file sizes to objects | ||
+ | |Assigns a sha-256 checksum to each file in the ''/objects/'' directory and calculates file sizes. | ||
+ | |- | ||
+ | |Verify metadata directory checksums | ||
+ | |Checks any checksum files that were placed in the /metadata/ folder of the SIP prior to ingest. Note that the filenames need to be named based on their algorithm: ''checksum.sha1'', ''checksum.sha256'', ''checksum.md5''. | ||
+ | |- | ||
+ | |Generate METS.xml document | ||
+ | |Generates a basic mets file with a fileSec and structMap to record the presence of all objects in the ''/objects/'' directory and their location in any subdirectories. Designed to capture the original order of the transfer in the event the user chooses subsequently to delete, rename or move files or break the transfer into multiple SIPs. The mets file is automatically added to any SIP generated from the transfer. | ||
+ | |- | ||
+ | |Extract packages | ||
+ | |Extracts objects from any zipped files or other packages. | ||
+ | |- | ||
+ | |Scan for viruses | ||
+ | |Uses [http://www.clamav.net/lang/en/ ClamAV], parses the output and creates a PREMIS event. If a virus is found, the SIP is automatically placed in ''/sharedDirectoryStructure/failed/''. | ||
+ | |- | ||
+ | |Sanitize object's file and directory names | ||
+ | |Some file systems do not support unicode or other special characters in filenames. This micro-service removes prohibited characters and replaces them with dashes. Original filenames are preserved in the PREMIS metadata. | ||
+ | |- | ||
+ | |Sanitize transfer name | ||
+ | |Same as above except does it for the transfer folder name. | ||
+ | |- | ||
+ | |Characterize and extract metadata | ||
+ | |Identifies and validates formats and extracts object metadata using the [http://code.google.com/p/fits/ File Information Tool Set (FITS)]. Adds output to the PREMIS metadata. | ||
+ | |- | ||
+ | |Create SIP(s) | ||
+ | |The user chooses among a number of SIP creation options. See the [[UM ingest|user manual]] for details. | ||
+ | |- | ||
+ | |} | ||
+ | </br> | ||
+ | |||
+ | =Ingest= | ||
{| border="1" cellpadding="10" cellspacing="0" width=90% | {| border="1" cellpadding="10" cellspacing="0" width=90% | ||
Line 11: | Line 71: | ||
!style="width:70%"|'''Description''' | !style="width:70%"|'''Description''' | ||
|- | |- | ||
− | | | + | |Verify SIP compliance |
− | | | + | |Verifies that the SIP conforms to the folder structure required for processing in Archivematica. The structure is as follows: ''/logs/'', ''/metadata/'', ''/metadata/submissionDocumentation/'', ''/objects/''. |
+ | |- | ||
+ | |Rename SIP directory with SIP UUID | ||
+ | |Adds a unique universal identifier to the SIP folder name. | ||
+ | |- | ||
+ | |Include default SIP processingMCP.xml | ||
+ | |Adds ''defaultSIPProcessing.xml'' file from ''/sharedDirectoryStructure/sharedMicroServiceTasksConfigs/'' to the transfer directory. This xml file canbe used to configure SIP workflow options. | ||
+ | |- | ||
+ | |Remove thumbs.db files | ||
+ | |Removes any [http://en.wikipedia.org/wiki/Windows_thumbnail_cache Thumbs.db] files. May be expanded to others in future releases. | ||
+ | |- | ||
+ | |Sanitize object's file and directory names | ||
+ | |If user created new folder and/or file names during SIP creation, any prhobited characters are removed from these names and replaced with dashes. | ||
+ | |- | ||
+ | |Sanitize SIP name | ||
+ | |Same as above except that it's done for the SIP folder name. | ||
+ | |- | ||
+ | |Check for Service directory | ||
+ | |For digitization output workflows, checks to see if the SIP contains any service copies of master files. See the [[UM digitization output|user manual]] for details. | ||
+ | |- | ||
+ | |Check for Access directory | ||
+ | |For digitization output workflows, checks to see if the SIP contains any access copies of master files. See the [[UM digitization output|user manual]] for details. | ||
+ | |- | ||
+ | |Normalize | ||
+ | |The user can choose to create preservation and/or access copies of the ingested files based on rules in the transcoder database. These rules can be seen under the '''Preservation planning''' tab in the Archivematica dashboard. | ||
+ | |- | ||
+ | |Normalize access | ||
+ | |Creates access copies to be added to the DIP. | ||
+ | |- | ||
+ | |Normalize preservation | ||
+ | |Creates preservation copies to be included in the AIP along with the original files. | ||
+ | |- | ||
+ | |Approve normalization | ||
+ | |If desired, the user can verify the quality of normalized files in ''/sharedDirectoryStructure/watchedDirectories/approveNormalization/''. | ||
+ | |- | ||
+ | |Check for submission documentation | ||
+ | |Checks for files in ''/metadata/submissionDocumentation/''. | ||
+ | |- | ||
+ | |Move submission documentation to objects directory | ||
+ | |Moves the ''/submissionDocumentation/'' directory from ''/metadata/'' to ''/objects/'' for ingest processing. | ||
+ | |- | ||
+ | |Assign file UUIDs to submission documentation | ||
+ | |Assigns a unique universal identifier to each file in the ''/objects/submissionDocumentation/'' directory. | ||
+ | |- | ||
+ | |Assign checksums and filesizes to submission documentation | ||
+ | |Assigns a sha-256 checksum to each file in the ''/objects/submissionDocumentation'' directory and calculates file sizes. | ||
+ | |- | ||
+ | |Extract packages in submission documentation | ||
+ | |Extracts objects from any zipped files or other packages in the ''/objects/'' directory. | ||
+ | |- | ||
+ | |Characterize and extract metadata on submission documentation | ||
+ | |Identifies and validates formats and extracts object metadata for files in the ''/objects/submissionDocumentation'' directory using the [http://code.google.com/p/fits/ File Information Tool Set (FITS)]. Adds output to the PREMIS metadata. | ||
+ | |- | ||
+ | |Normalize submission documentation to preservation format | ||
+ | |Creates preservation copies of all files in the ''/objects/submissionDocumentation'' directory to be included in the AIP along with the original files. | ||
+ | |- | ||
+ | |Verify checksums generated on ingest | ||
+ | |Verifies checksums that were generated during transfer processing to ensure that the files have not been corrupted during transfer or ingest. | ||
+ | |- | ||
+ | |Remove empty directories | ||
+ | |Removes any empty directories from the SIP. | ||
+ | |- | ||
+ | |Generate METS.xml document | ||
+ | |Generates a METS file with PREMIS metadata. For more information on the METS file, see [[METS]]. | ||
+ | |- | ||
+ | |Copy transfers metadata and logs | ||
+ | |Copies all submission documentation included in the original transfer to the SIP. Copies all logs generated during transfer processing to the SIP. | ||
+ | |- | ||
+ | |Copy METS to DIP directory | ||
+ | |Creates a copy of the METS file in the DIP directory. | ||
|- | |- | ||
− | | | + | |Generate DIP |
− | | | + | |Moves the DIP to the DIP upload directory. |
|- | |- | ||
− | | | + | |Upload DIP |
− | | | + | |The user uploads the DIP to a selected description in the access system. See the [[UM access|user manual]] for details. |
|- | |- | ||
− | | | + | |Prepare AIP |
− | | | + | |Packages the SIP into an AIP using [http://sourceforge.net/projects/loc-xferutils/ BagIt] |
|- | |- | ||
− | | | + | |Compress AIP |
− | | | + | |Losslessly compresses the AIP for storage using [http://p7zip.sourceforge.net/ p7zip]. |
|- | |- | ||
− | | | + | |Store AIP |
− | | | + | |Moves the AIP to a specified directory. In the demo version of Archivematica the directory is ''/sharedDirectoryStructure/www/AIPsStore/''. In other environments it can be a remote network mounted directory. The directory structure of the AIP store contains UUID quad directories and an ''index.html'' file listing the AIPs in storage. The ''index.html'' file is displayed in the dashboard in the '''Archival storage''' tab. </br> |
+ | Once the AIP has been stored, a copy of the AIP is extracted from storage to a local temp directory, and is validated with the various BagIt checks: verifyvalid, checkpayloadoxum, verifycomplete, verifypayloadmanifests, verifytagmanifests. | ||
|- | |- | ||
|} | |} |
Latest revision as of 14:14, 11 December 2019
Main Page > Documentation > Technical Architecture > Micro-services > Archivematica 0.8 Micro-services
Note: The documentation on this page is out of date. Please see the official documentation page for the latest.
This page describes the key micro-services that are undertaken during transfer and ingest.
Transfer[edit]
Micro-service | Description |
---|---|
Approve transfer | Once the transfer has all its digital objects and has been formatted for processing, the user selects "Transfer complete" from the Actions drop-down menu. |
Verify transfer compliance | Verifies that the transfer conforms to the folder structure required for processing in Archivematica. The structure is as follows: /logs/, /metadata/, /metadata/submissionDocumentation/, /objects/. |
Rename with transfer UUID | Adds a unique universal identifier to the transfer folder name. |
Include default Transfer processingMCP.xml | Adds defaultTransferProcessing.xml file from /sharedDirectoryStructure/sharedMicroServiceTasksConfigs/ to the transfer directory. This xml file can configure workflow options such as creating transfer backups, quarantining the transfer and selecting a SIP creation option. |
Workflow decision - create transfer backup | The user can choose to create a complete backup of the transfer in case transfer or ingest are interrupted or fail. The transfer backup is placed in /sharedDirectoryStructure/transferBackups/ and will automatically be deleted once the AIP has been moved into storage. |
Assign file UUIDs to objects | Assigns a unique universal identifier to each file in the /objects/ directory. |
Assign checksums and file sizes to objects | Assigns a sha-256 checksum to each file in the /objects/ directory and calculates file sizes. |
Verify metadata directory checksums | Checks any checksum files that were placed in the /metadata/ folder of the SIP prior to ingest. Note that the filenames need to be named based on their algorithm: checksum.sha1, checksum.sha256, checksum.md5. |
Generate METS.xml document | Generates a basic mets file with a fileSec and structMap to record the presence of all objects in the /objects/ directory and their location in any subdirectories. Designed to capture the original order of the transfer in the event the user chooses subsequently to delete, rename or move files or break the transfer into multiple SIPs. The mets file is automatically added to any SIP generated from the transfer. |
Extract packages | Extracts objects from any zipped files or other packages. |
Scan for viruses | Uses ClamAV, parses the output and creates a PREMIS event. If a virus is found, the SIP is automatically placed in /sharedDirectoryStructure/failed/. |
Sanitize object's file and directory names | Some file systems do not support unicode or other special characters in filenames. This micro-service removes prohibited characters and replaces them with dashes. Original filenames are preserved in the PREMIS metadata. |
Sanitize transfer name | Same as above except does it for the transfer folder name. |
Characterize and extract metadata | Identifies and validates formats and extracts object metadata using the File Information Tool Set (FITS). Adds output to the PREMIS metadata. |
Create SIP(s) | The user chooses among a number of SIP creation options. See the user manual for details. |
Ingest[edit]
Micro-service | Description |
---|---|
Verify SIP compliance | Verifies that the SIP conforms to the folder structure required for processing in Archivematica. The structure is as follows: /logs/, /metadata/, /metadata/submissionDocumentation/, /objects/. |
Rename SIP directory with SIP UUID | Adds a unique universal identifier to the SIP folder name. |
Include default SIP processingMCP.xml | Adds defaultSIPProcessing.xml file from /sharedDirectoryStructure/sharedMicroServiceTasksConfigs/ to the transfer directory. This xml file canbe used to configure SIP workflow options. |
Remove thumbs.db files | Removes any Thumbs.db files. May be expanded to others in future releases. |
Sanitize object's file and directory names | If user created new folder and/or file names during SIP creation, any prhobited characters are removed from these names and replaced with dashes. |
Sanitize SIP name | Same as above except that it's done for the SIP folder name. |
Check for Service directory | For digitization output workflows, checks to see if the SIP contains any service copies of master files. See the user manual for details. |
Check for Access directory | For digitization output workflows, checks to see if the SIP contains any access copies of master files. See the user manual for details. |
Normalize | The user can choose to create preservation and/or access copies of the ingested files based on rules in the transcoder database. These rules can be seen under the Preservation planning tab in the Archivematica dashboard. |
Normalize access | Creates access copies to be added to the DIP. |
Normalize preservation | Creates preservation copies to be included in the AIP along with the original files. |
Approve normalization | If desired, the user can verify the quality of normalized files in /sharedDirectoryStructure/watchedDirectories/approveNormalization/. |
Check for submission documentation | Checks for files in /metadata/submissionDocumentation/. |
Move submission documentation to objects directory | Moves the /submissionDocumentation/ directory from /metadata/ to /objects/ for ingest processing. |
Assign file UUIDs to submission documentation | Assigns a unique universal identifier to each file in the /objects/submissionDocumentation/ directory. |
Assign checksums and filesizes to submission documentation | Assigns a sha-256 checksum to each file in the /objects/submissionDocumentation directory and calculates file sizes. |
Extract packages in submission documentation | Extracts objects from any zipped files or other packages in the /objects/ directory. |
Characterize and extract metadata on submission documentation | Identifies and validates formats and extracts object metadata for files in the /objects/submissionDocumentation directory using the File Information Tool Set (FITS). Adds output to the PREMIS metadata. |
Normalize submission documentation to preservation format | Creates preservation copies of all files in the /objects/submissionDocumentation directory to be included in the AIP along with the original files. |
Verify checksums generated on ingest | Verifies checksums that were generated during transfer processing to ensure that the files have not been corrupted during transfer or ingest. |
Remove empty directories | Removes any empty directories from the SIP. |
Generate METS.xml document | Generates a METS file with PREMIS metadata. For more information on the METS file, see METS. |
Copy transfers metadata and logs | Copies all submission documentation included in the original transfer to the SIP. Copies all logs generated during transfer processing to the SIP. |
Copy METS to DIP directory | Creates a copy of the METS file in the DIP directory. |
Generate DIP | Moves the DIP to the DIP upload directory. |
Upload DIP | The user uploads the DIP to a selected description in the access system. See the user manual for details. |
Prepare AIP | Packages the SIP into an AIP using BagIt |
Compress AIP | Losslessly compresses the AIP for storage using p7zip. |
Store AIP | Moves the AIP to a specified directory. In the demo version of Archivematica the directory is /sharedDirectoryStructure/www/AIPsStore/. In other environments it can be a remote network mounted directory. The directory structure of the AIP store contains UUID quad directories and an index.html file listing the AIPs in storage. The index.html file is displayed in the dashboard in the Archival storage tab. Once the AIP has been stored, a copy of the AIP is extracted from storage to a local temp directory, and is validated with the various BagIt checks: verifyvalid, checkpayloadoxum, verifycomplete, verifypayloadmanifests, verifytagmanifests. |