"Digitale Langzeitarchivierung: Ein Thema fur die Digitalen Geisteswissenschaften?". We argue that the so-called Digital Humanities fail to meet conventional criteria to be an accredited field of study on a par with Literature, Chemistry, Computer Science, and Civil Engineering, or even a specialized professorial emphasis such as Ancient History or Nuclear Physics. The argument uses long-term digital preservation as an example to argue that Digital Humanities proponents' case for their research agenda does not merit financial support, emphasizing practical aspects over subjective theory.
Progress towards practical long-term preservation seems to be stalled. Preservationists cannot afford specially developed technology, but must exploit what is created for the marketplace. Economic and technical facts suggest that most preservation work should be shifted from repository institutions to information producers and consumers. Prior publications describe solutions for all known conceptual challenges of preserving a single digital object, but do not deal with software development or scaling to large collections. Much of the document handling software needed is available. It has, however, not yet been selected, adapted, integrated, or deployed for digital preservation. The daily tools of both information producers and information consumers can be extended to embed preservation packaging without much burdening these users. We describe a practical strategy for detailed design and implementation. Document handling is intrinsically complicated because of human sensitivity to communication nuances. Our engineering section therefore starts by discussing how project managers can master the many pertinent details. To make discoveries one must not be too anxious about errors. One must be willing to state a theory clearly and crisply and say as the physicists do: I have worked on this for a long time, and this is what I have come up with; now tell me what, if anything, is wrong with it. And before one ever gets that far, one has usually found many of one's own attempts faulty and discarded them. What is most needful is by no means certainty but rather, to quote Nietzsche's happy phrase, "the courage for an attack on one's convictions.” Kaufman
Focusing on end users' needs rather than those of archiving institutions.
How can an author store digital information so that it will be reliably intelligible, even years later when he or she is no longer available to answer questions? Methods that might work are not good enough; what is preserved today should be reliably intelligible whenever someone wants it. Prior proposals fail because they generally confound saved data with irrelevant details of today's information technology---details that are difficult to define, extract, and save completely and accurately.We use a virtual machine to represent and eventually to render any data whatsoever. We focus on a case of intermediate difficulty---an executable procedure---and identify a variant for every other data type.This solution might be more elaborate than needed to render some text, image, audio, or video data. Simple data can be preserved as representations using well-known standards. We sketch practical methods for files ranging from simple structures to those containing computer programs, treating simple cases here and deferring complex cases for future work. Enough of the complete solution is known to enable practical aggressive preservation programs today.
In ancient times, wax seals impressed with signet rings were affixed to documents as evidence of their authenticity. A digital counterpart is a message authentication code fixed firmly to each important document. If a digital object is sealed together with its own audit trail, each user can examine this evidence to decide whether to trust the content - no matter how distant this user is in time, space, and social affiliation from the document's source.We propose an architecture and design that accomplish this: encapsulation of digital object content with metadata describing its origins, cryptographic sealing, webs of trust for public keys rooted in a forest of respected institutions, and a certain way of managing information identifiers. These means will satisfy emerging needs in civilian and military record management, including medical patient records, regulatory records for aircraft and pharmaceuticals, business records for financial audit, legislative and legal briefs, and scholarly works.This is true for any kind of digital object, independent of its purposes and of most data type and representation details, and provides every kind of user - information authors and editors, librarians and collection managers, and information consumers - with autonomy for implied tasks. Our prototype will conform to applicable standards, will be interoperable over most computing bases, and will be compatible with existing digital library software.The proposed architecture integrates software that is mostly available and widely accepted.
E-business, information serving, and ubiquitous computing will create heavy request traffic from strangers or even incognitos. Such requests must be managed automatically. Two ways of doing this are well known: giving every incognito consumer the same treatment, and rendering service in return for money. However, different behavior will be often wanted, e.g., for a university library with different access policies for undergraduates, graduate students, faculty, alumni, citizens of the same state, and everyone else. For a data or process server contacted by client machines on behalf of users not previously known, we show how to provide reliable automatic access administration conforming to service agreements. Implementations scale well from very small collections of consumers and producers to immense client/server networks. Servers can deliver information, effect state changes, and control external equipment. Consumer privacy is easily addressed by the same protocol. We support consumer privacy, but allow servers to deny their resources to incognitos. A protocol variant even protects against statistical attacks by consortia of service organizations. One e-commerce application would put the consumer's tokens on a smart card whose readers are in vending kiosks. In e-business we can simplify supply chain administration. Our method can also be used in sensitive networks without introducing new security loopholes.
The Safeguarding ... series in D-Lib Magazine is intended to suggest technology to help manage digital intellectual property. That technology can contribute only in a complex of administrative, legal, contractual, and social practices is broadly accepted; we now pause to examine how efforts to fit in the technological component are progressing and what next needs attention. Among concerns for responsive and responsible management of intellectual property, technical aspects are surely secondary to prominent issues of public policy, law, and ethics. The latter are beginning to be addressed both in legislative processes and also by academic investigators. For the technical community, we assert that we can design offerings with sufficient flexibility that we need not wait for policy decisions which might affect software to administer rules chosen or to hinder unacceptable behavior. The current article projects technical directions without designing solutions. It emphasizes managing the data -- how it is stored, protected, and communicated.
Efforts to place vast information resources at the fingertips of each individual in large user populations must be balanced by commensurate attention to information protection. For distributed systems with less-structured tasks, more-diversified information, and a heterogeneous user set, the computing system must administer enterprise-chosen access control policies. One kind of resource is a digital library that emulates massive collections of paper and other physical media for clerical, engineering, and cultural applications. This article considers the security requirements for such libraries and proposes an access control method that mimics organizational practice by combining a subject tree with ad hoc role granting that controls privileges for many operations independently, that treats (all but one) privileged roles (e.g., auditor, security officer) like every other individual authorization, and that binds access control information to objects indirectly for scaling, flexibility, and reflexive protection. We sketch a realization and show that it will perform well, generalizes many deployed proposed access control policies, and permits individual data centers to implement other models economically and without disruption.
The Vatican Library is an extraordinary repository of rare books and manuscripts. Among its 150,000 manuscripts are early copies of works by Aristotle, Dante, Euclid, Homer, and Virgil. Yet today access to the Library is limited. Because of the time and cost required to travel to Rome, only some 2000 scholars can afford to visit the Library each year. Through the Vatican Library Project, we are exploring the practicality of providing digital library services that extend access to portions of the Library's collections to scholars worldwide, as an early example of providing digital library services that extend and complement traditional library services. A core goal of the project is to provide access via the Internet to some of the Library's most valuable manuscripts, printed books, and other sources to a scholarly community around the world. A multinational, multidisciplinary team is addressing the technical challenges raised by that goal, including • Development of a multiserver system suitable for providing information to scholars worldwide. • Capture of images of the materials with faithful color and sufficient detail to support scholarly study. Protection of the on-line materials, especially images, from misappropriation. • Development of tools to enable scholars to locate desired materials. • Development of tools to enable scholars to scrutinize images of manuscripts. • In this paper, we provide an overview of the project, a description of the system being developed to satisfy its needs, and a discussion of how the technical challenges are being addressed.