Document (#40374)

Author
Mayo, D.
Bowers, K.
Title
¬The devil's shoehorn : a case study of EAD to ArchivesSpace migration at a large university
Source
Code4Lib journal. Issue 35(2017), [http://journal.code4lib.org]
Year
2017
Abstract
A band of archivists and IT professionals at Harvard took on a project to convert nearly two million descriptions of archival collection components from marked-up text into the ArchivesSpace archival metadata management system. Starting in the mid-1990s, Harvard was an alpha implementer of EAD, an SGML (later XML) text markup language for electronic inventories, indexes, and finding aids that archivists use to wend their way through the sometimes quirky filing systems that bureaucracies establish for their records or the utter chaos in which some individuals keep their personal archives. These pathfinder documents, designed to cope with messy reality, can themselves be difficult to classify. Portions of them are rigorously structured, while other parts are narrative. Early documents predate the establishment of the standard; many feature idiosyncratic encoding that had been through several machine conversions, while others were freshly encoded and fairly consistent. In this paper, we will cover the practical and technical challenges involved in preparing a large (900MiB) corpus of XML for ingest into an open-source archival information system (ArchivesSpace). This case study will give an overview of the project, discuss problem discovery and problem solving, and address the technical challenges, analysis, solutions, and decisions and provide information on the tools produced and lessons learned. The authors of this piece are Kate Bowers, Collections Services Archivist for Metadata, Systems, and Standards at the Harvard University Archive, and Dave Mayo, a Digital Library Software Engineer for Harvard's Library and Technology Services. Kate was heavily involved in both metadata analysis and later problem solving, while Dave was the sole full-time developer assigned to the migration project.
Content
Vgl.: http://journal.code4lib.org/articles/12239.
Theme
Formalerschließung
Auszeichnungssprachen
Object
EAD
Area
Archive

Similar documents (content)

  1. Carini, P.; Shepherd, K.: ¬The MARC standard and encoded archival description (2004) 0.20
    0.19598502 = sum of:
      0.19598502 = product of:
        0.81660426 = sum of:
          0.056697775 = weight(abstract_txt:case in 2830) [ClassicSimilarity], result of:
            0.056697775 = score(doc=2830,freq=1.0), product of:
              0.107831866 = queryWeight, product of:
                1.1112098 = boost
                4.807296 = idf(docFreq=981, maxDocs=44218)
                0.020185998 = queryNorm
              0.52579796 = fieldWeight in 2830, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.807296 = idf(docFreq=981, maxDocs=44218)
                0.109375 = fieldNorm(doc=2830)
          0.06994651 = weight(abstract_txt:challenges in 2830) [ClassicSimilarity], result of:
            0.06994651 = score(doc=2830,freq=1.0), product of:
              0.12403582 = queryWeight, product of:
                1.1917799 = boost
                5.155857 = idf(docFreq=692, maxDocs=44218)
                0.020185998 = queryNorm
              0.56392187 = fieldWeight in 2830, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.155857 = idf(docFreq=692, maxDocs=44218)
                0.109375 = fieldNorm(doc=2830)
          0.06425183 = weight(abstract_txt:project in 2830) [ClassicSimilarity], result of:
            0.06425183 = score(doc=2830,freq=1.0), product of:
              0.13417055 = queryWeight, product of:
                1.5180871 = boost
                4.378348 = idf(docFreq=1507, maxDocs=44218)
                0.020185998 = queryNorm
              0.4788818 = fieldWeight in 2830, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.378348 = idf(docFreq=1507, maxDocs=44218)
                0.109375 = fieldNorm(doc=2830)
          0.1987317 = weight(abstract_txt:archivists in 2830) [ClassicSimilarity], result of:
            0.1987317 = score(doc=2830,freq=1.0), product of:
              0.24881765 = queryWeight, product of:
                1.6879636 = boost
                7.3024383 = idf(docFreq=80, maxDocs=44218)
                0.020185998 = queryNorm
              0.7987042 = fieldWeight in 2830, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.3024383 = idf(docFreq=80, maxDocs=44218)
                0.109375 = fieldNorm(doc=2830)
          0.0890322 = weight(abstract_txt:metadata in 2830) [ClassicSimilarity], result of:
            0.0890322 = score(doc=2830,freq=1.0), product of:
              0.16676246 = queryWeight, product of:
                1.6924554 = boost
                4.881247 = idf(docFreq=911, maxDocs=44218)
                0.020185998 = queryNorm
              0.5338864 = fieldWeight in 2830, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.881247 = idf(docFreq=911, maxDocs=44218)
                0.109375 = fieldNorm(doc=2830)
          0.3379442 = weight(abstract_txt:archival in 2830) [ClassicSimilarity], result of:
            0.3379442 = score(doc=2830,freq=3.0), product of:
              0.28135616 = queryWeight, product of:
                2.1983473 = boost
                6.340301 = idf(docFreq=211, maxDocs=44218)
                0.020185998 = queryNorm
              1.2011261 = fieldWeight in 2830, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.340301 = idf(docFreq=211, maxDocs=44218)
                0.109375 = fieldNorm(doc=2830)
        0.24 = coord(6/25)
    
  2. Carpenter, K.E.: End of the war between print and electronics (1996) 0.13
    0.13419785 = sum of:
      0.13419785 = product of:
        0.8387366 = sum of:
          0.02414402 = weight(abstract_txt:their in 599) [ClassicSimilarity], result of:
            0.02414402 = score(doc=599,freq=1.0), product of:
              0.069867186 = queryWeight, product of:
                1.0954807 = boost
                3.1594994 = idf(docFreq=5101, maxDocs=44218)
                0.020185998 = queryNorm
              0.34557024 = fieldWeight in 599, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.1594994 = idf(docFreq=5101, maxDocs=44218)
                0.109375 = fieldNorm(doc=599)
          0.21875499 = weight(abstract_txt:inventories in 599) [ClassicSimilarity], result of:
            0.21875499 = score(doc=599,freq=1.0), product of:
              0.21053861 = queryWeight, product of:
                1.0979267 = boost
                9.499662 = idf(docFreq=8, maxDocs=44218)
                0.020185998 = queryNorm
              1.0390255 = fieldWeight in 599, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.499662 = idf(docFreq=8, maxDocs=44218)
                0.109375 = fieldNorm(doc=599)
          0.19511217 = weight(abstract_txt:archival in 599) [ClassicSimilarity], result of:
            0.19511217 = score(doc=599,freq=1.0), product of:
              0.28135616 = queryWeight, product of:
                2.1983473 = boost
                6.340301 = idf(docFreq=211, maxDocs=44218)
                0.020185998 = queryNorm
              0.6934704 = fieldWeight in 599, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.340301 = idf(docFreq=211, maxDocs=44218)
                0.109375 = fieldNorm(doc=599)
          0.40072542 = weight(abstract_txt:harvard in 599) [ClassicSimilarity], result of:
            0.40072542 = score(doc=599,freq=1.0), product of:
              0.45460212 = queryWeight, product of:
                2.7943695 = boost
                8.059301 = idf(docFreq=37, maxDocs=44218)
                0.020185998 = queryNorm
              0.88148606 = fieldWeight in 599, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.059301 = idf(docFreq=37, maxDocs=44218)
                0.109375 = fieldNorm(doc=599)
        0.16 = coord(4/25)
    
  3. Heastrom, M.: Descriptive practices for electronic records : deciding what is essential and imaging what is possible (1993) 0.12
    0.120115176 = sum of:
      0.120115176 = product of:
        0.60057586 = sum of:
          0.05995415 = weight(abstract_txt:challenges in 8332) [ClassicSimilarity], result of:
            0.05995415 = score(doc=8332,freq=1.0), product of:
              0.12403582 = queryWeight, product of:
                1.1917799 = boost
                5.155857 = idf(docFreq=692, maxDocs=44218)
                0.020185998 = queryNorm
              0.4833616 = fieldWeight in 8332, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.155857 = idf(docFreq=692, maxDocs=44218)
                0.09375 = fieldNorm(doc=8332)
          0.05745528 = weight(abstract_txt:while in 8332) [ClassicSimilarity], result of:
            0.05745528 = score(doc=8332,freq=1.0), product of:
              0.13801236 = queryWeight, product of:
                1.5396681 = boost
                4.44059 = idf(docFreq=1416, maxDocs=44218)
                0.020185998 = queryNorm
              0.4163053 = fieldWeight in 8332, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.44059 = idf(docFreq=1416, maxDocs=44218)
                0.09375 = fieldNorm(doc=8332)
          0.17034145 = weight(abstract_txt:archivists in 8332) [ClassicSimilarity], result of:
            0.17034145 = score(doc=8332,freq=1.0), product of:
              0.24881765 = queryWeight, product of:
                1.6879636 = boost
                7.3024383 = idf(docFreq=80, maxDocs=44218)
                0.020185998 = queryNorm
              0.6846036 = fieldWeight in 8332, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.3024383 = idf(docFreq=80, maxDocs=44218)
                0.09375 = fieldNorm(doc=8332)
          0.076313324 = weight(abstract_txt:metadata in 8332) [ClassicSimilarity], result of:
            0.076313324 = score(doc=8332,freq=1.0), product of:
              0.16676246 = queryWeight, product of:
                1.6924554 = boost
                4.881247 = idf(docFreq=911, maxDocs=44218)
                0.020185998 = queryNorm
              0.45761693 = fieldWeight in 8332, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.881247 = idf(docFreq=911, maxDocs=44218)
                0.09375 = fieldNorm(doc=8332)
          0.23651166 = weight(abstract_txt:archival in 8332) [ClassicSimilarity], result of:
            0.23651166 = score(doc=8332,freq=2.0), product of:
              0.28135616 = queryWeight, product of:
                2.1983473 = boost
                6.340301 = idf(docFreq=211, maxDocs=44218)
                0.020185998 = queryNorm
              0.84061307 = fieldWeight in 8332, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.340301 = idf(docFreq=211, maxDocs=44218)
                0.09375 = fieldNorm(doc=8332)
        0.2 = coord(5/25)
    
  4. Gracy, K.F.: Enriching and enhancing moving images with Linked Data : an exploration in the alignment of metadata models (2018) 0.11
    0.10842542 = sum of:
      0.10842542 = product of:
        0.4517726 = sum of:
          0.010347438 = weight(abstract_txt:their in 4200) [ClassicSimilarity], result of:
            0.010347438 = score(doc=4200,freq=1.0), product of:
              0.069867186 = queryWeight, product of:
                1.0954807 = boost
                3.1594994 = idf(docFreq=5101, maxDocs=44218)
                0.020185998 = queryNorm
              0.14810154 = fieldWeight in 4200, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.1594994 = idf(docFreq=5101, maxDocs=44218)
                0.046875 = fieldNorm(doc=4200)
          0.029977076 = weight(abstract_txt:challenges in 4200) [ClassicSimilarity], result of:
            0.029977076 = score(doc=4200,freq=1.0), product of:
              0.12403582 = queryWeight, product of:
                1.1917799 = boost
                5.155857 = idf(docFreq=692, maxDocs=44218)
                0.020185998 = queryNorm
              0.2416808 = fieldWeight in 4200, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.155857 = idf(docFreq=692, maxDocs=44218)
                0.046875 = fieldNorm(doc=4200)
          0.02872764 = weight(abstract_txt:while in 4200) [ClassicSimilarity], result of:
            0.02872764 = score(doc=4200,freq=1.0), product of:
              0.13801236 = queryWeight, product of:
                1.5396681 = boost
                4.44059 = idf(docFreq=1416, maxDocs=44218)
                0.020185998 = queryNorm
              0.20815265 = fieldWeight in 4200, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.44059 = idf(docFreq=1416, maxDocs=44218)
                0.046875 = fieldNorm(doc=4200)
          0.08517072 = weight(abstract_txt:archivists in 4200) [ClassicSimilarity], result of:
            0.08517072 = score(doc=4200,freq=1.0), product of:
              0.24881765 = queryWeight, product of:
                1.6879636 = boost
                7.3024383 = idf(docFreq=80, maxDocs=44218)
                0.020185998 = queryNorm
              0.3423018 = fieldWeight in 4200, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.3024383 = idf(docFreq=80, maxDocs=44218)
                0.046875 = fieldNorm(doc=4200)
          0.076313324 = weight(abstract_txt:metadata in 4200) [ClassicSimilarity], result of:
            0.076313324 = score(doc=4200,freq=4.0), product of:
              0.16676246 = queryWeight, product of:
                1.6924554 = boost
                4.881247 = idf(docFreq=911, maxDocs=44218)
                0.020185998 = queryNorm
              0.45761693 = fieldWeight in 4200, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                4.881247 = idf(docFreq=911, maxDocs=44218)
                0.046875 = fieldNorm(doc=4200)
          0.22123641 = weight(abstract_txt:archival in 4200) [ClassicSimilarity], result of:
            0.22123641 = score(doc=4200,freq=7.0), product of:
              0.28135616 = queryWeight, product of:
                2.1983473 = boost
                6.340301 = idf(docFreq=211, maxDocs=44218)
                0.020185998 = queryNorm
              0.7863215 = fieldWeight in 4200, product of:
                2.6457512 = tf(freq=7.0), with freq of:
                  7.0 = termFreq=7.0
                6.340301 = idf(docFreq=211, maxDocs=44218)
                0.046875 = fieldNorm(doc=4200)
        0.24 = coord(6/25)
    
  5. Trace, C.B.; Francisco-Revilla, L.: ¬The value and complexity of collection arrangement for evidentiary work (2015) 0.10
    0.10139934 = sum of:
      0.10139934 = product of:
        0.5069967 = sum of:
          0.019511316 = weight(abstract_txt:their in 2164) [ClassicSimilarity], result of:
            0.019511316 = score(doc=2164,freq=2.0), product of:
              0.069867186 = queryWeight, product of:
                1.0954807 = boost
                3.1594994 = idf(docFreq=5101, maxDocs=44218)
                0.020185998 = queryNorm
              0.27926293 = fieldWeight in 2164, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.1594994 = idf(docFreq=5101, maxDocs=44218)
                0.0625 = fieldNorm(doc=2164)
          0.040865388 = weight(abstract_txt:involved in 2164) [ClassicSimilarity], result of:
            0.040865388 = score(doc=2164,freq=1.0), product of:
              0.12588255 = queryWeight, product of:
                1.2006191 = boost
                5.194097 = idf(docFreq=666, maxDocs=44218)
                0.020185998 = queryNorm
              0.32463107 = fieldWeight in 2164, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.194097 = idf(docFreq=666, maxDocs=44218)
                0.0625 = fieldNorm(doc=2164)
          0.036715332 = weight(abstract_txt:project in 2164) [ClassicSimilarity], result of:
            0.036715332 = score(doc=2164,freq=1.0), product of:
              0.13417055 = queryWeight, product of:
                1.5180871 = boost
                4.378348 = idf(docFreq=1507, maxDocs=44218)
                0.020185998 = queryNorm
              0.27364674 = fieldWeight in 2164, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.378348 = idf(docFreq=1507, maxDocs=44218)
                0.0625 = fieldNorm(doc=2164)
          0.16059946 = weight(abstract_txt:archivists in 2164) [ClassicSimilarity], result of:
            0.16059946 = score(doc=2164,freq=2.0), product of:
              0.24881765 = queryWeight, product of:
                1.6879636 = boost
                7.3024383 = idf(docFreq=80, maxDocs=44218)
                0.020185998 = queryNorm
              0.6454504 = fieldWeight in 2164, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.3024383 = idf(docFreq=80, maxDocs=44218)
                0.0625 = fieldNorm(doc=2164)
          0.24930519 = weight(abstract_txt:archival in 2164) [ClassicSimilarity], result of:
            0.24930519 = score(doc=2164,freq=5.0), product of:
              0.28135616 = queryWeight, product of:
                2.1983473 = boost
                6.340301 = idf(docFreq=211, maxDocs=44218)
                0.020185998 = queryNorm
              0.886084 = fieldWeight in 2164, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                6.340301 = idf(docFreq=211, maxDocs=44218)
                0.0625 = fieldNorm(doc=2164)
        0.2 = coord(5/25)