Tuesday, 26 May 2009

Geocities being rescued by Archive Team

Yahoo's decision to pull the plug on geocities sites shows how ephemeral web content can be. There should be a means for individual users to obtain their geocities stuff first, so long as they get the message to do something in time. In a more comprehensive approach, the Archive Team are busy downloading all the geocities websites they can find, though what happens after that is unclear. The Archive Team have a page with more info. and there's a nice article in The Register, complete with geocities styling.

Thursday, 21 May 2009

Whitepaper from MITH/HRC/Emory born-digital literary mss. project

I mentioned a white paper coming out of the 'Approaches to managing & collecting born-digital literary materials for scholarly use' (phew, long title!) grant in this post a while back. The paper is available here and definitely worth a read.

Monday, 18 May 2009

Digital Repository Workshop at Oxford

I attended a University of Oxford Digital Repositories Steering Group Workshop a couple of weeks ago (I've been on honeymoon so I'm only just now getting around to posting it!). The internal workshop was subtitled "Tools and Infrastructure" and provided an opportunity for the various repository projects around the University to present what they were doing and then some discussion on how repositories might go forward. All good and interesting stuff!

Details on the day are available and I've also put my talk on slideshare.

Tuesday, 12 May 2009

Standing on the shoulders of Giants?

Just attended the Repositories and Preservation Programme meeting in Aston Birmingham I would really recommend the talk The Institutional Perspective - How can institutions most effectively exploit work in the Repositories and Preservation field? given by Jeff Haywood- University of Edinburgh.

I would like to think this may kickstart a process to find methods by which current projects could more easily use and build on the outputs of previous projects and create a framework to more easily exchange code and ideas.

Jeff's talk was given even more currency as in the afternoon session Rachel Heery give a presentation (its on slideshare) in Repositories Roadmap Session launching her just published Digital Repositories Roadmap Review: towards a vision for research and learning in 2013

The question however remains: by then will Standing on the shoulders of Giants still be a distant concept?

Monday, 11 May 2009

What is this thing anyway?

The first step in doing anything useful with a digital accession is to answer just that question. The next is generally "now that I know, what 'stuff' do I need to recover the data and how might I do that?". With some items, it's easy enough. With others it can be rather more challenging. Alex Eveleigh just pointed me at Mediapedia - a database being developed at The National Library of Australia to help folk identify media. Best of all, it sounds like it's been designed to help people find things by easily determined characteristics (e.g. physical measurements) rather than relying on the user to know, more or less, what they are looking for. Super idea.

Thursday, 7 May 2009

The best introduction to digital preservation ever?

Recently I've been involved in organising a series of 'digital preservation roadshows' for archivists in the UK. The series is aimed at archivists who have had little chance to look at digital preservation issues; the idea is to provide an introduction and some practical, inexpensive, tips and tools that will help people make a start. We've had a lot of positive feedback from the first roadshow (Gloucester, April 29th), though there were calls for a little more easing into the topic and its terminology that we'll need to take on board for the next event (York, June 26th). I wonder whether our delegates might have appreciated this video, produced by the Digital Preservation Europe project. If it's not the *best* introduction to digital preservation I've seen, it's most definitely the funniest. It made me laugh anyway :-) .

Monday, 27 April 2009

Wahcade Emulator Front-End

Now, you might think I've gone mad putting a link here to Wahcade. Either that or you'll think I've too much time on my hands and spend all my time digging out old games to play on this arcade machine manager. While making an arcade machine, case and all, sounds like a lot of fun (and one day when life is less busy I might just give it a go), I'm really making a note here to flag it as an interesting example of what we might do in reading rooms. I guess it is a bit like those media centre PCs (Mythbuntu for example) or the BL's "Turning the Pages" only this is a front-end for emulators.

Imagine that a collection is a "rom" (the rom being the image of a chip containing, in Wahcade's case, a game but for us could be a disk image from a donor's PC). The user picks from a list of roms and then the reading room "arcade" starts an emulator and away you go. Before you know it the dumb terminal is a replica Mac System 6 desktop complete with donor file system, etc.

Be neat wouldn't it?

Friday, 17 April 2009

Terrabyte Terror

I knew I'd come to the right place for work one morning when I was talking with Renhart and Susan and mentioned that I was very excited at having discovered a Maplin just minutes walk from the office. Instead making a hasty retreat from the conversation, both of them gave me knowing smiles and agreed that Maplin was wonderful!

From my early days watching my Dad teach electronics I've loved the smell of soldering, the look of components and the idea that you can make your own set of LEDs flash just for the fun of it. Thumbing through the Maplin catalogue with a cup of tea was once one of my favourite past times. But these days, more and more, I get a sense of dread as I check out the special offers.

Why?

Let me give you an example: 1TB External Drive, £99.
Here is another: 1TB Internal Drive, £89.

You read that right - 1 TerraByte of storage for under £100! Doesn't that make you quake? Probably not, but I can't help but wonder how long it will be before we have to accession a 1TB drive. What do we do with it? Do we even know what that amount of detritus accumlated over, well how long? a lifetime? a couple of evenings with iTunes? We don't know how long it'll take the average person to fill up a 1TB drive. Do we have the capacity to store 1TB of data and even if we do, how sustainable is that?

You could argue that since storage like this is so cheap, we can rest assured that our own storage costs will be less, so we always keep up with the growth of consumer storage. It is a fair point, but how many preservation-grade storage devices can manage 10p a GB? None I imagine, and for good reason. There is a whole lot more to a preservation system than a disk and a plastic case - it takes more than 1TB to keep 1TB safe for a start! (Mind you, I couldn't help but smile at Maplin's promise of "Peace of mind with 5 year limited warranty").

If we cannot keep up with the storage then, what do we do? A brute force method would be to compress the data, but then bit rot becomes a much more worrying issue (and it is pretty worrying already). We could look for duplicates - how many MP3 collections will include the same songs for instance and should we keep them all (if any)? What if it is the same song, with a different encoding/bitrate/whatever? What about copies of OSs - all those i386 directories? (Though arguably an external drive will not contain an OS, so we wont save space there).

We probably don't need or want to keep all of those 1000GBs, but how will we identify what to preserve? Susan and Renhart came up with some answers to this with their brilliant Paradigm project - which I'll paraphrase as "encourage the creators to curate their own data" - and I'm hopeful that will happen, but what if it doesn't? Will we see "personal data curation" and "managing information overload" added to the National Curriculum anytime soon? I hope so!

All of which finally gives me reason to stop worrying about cheap terrabytes! Data is going to keep growing and someone is going to have to help manage all that stuff. I guess that is where we fit in.

Monday, 6 April 2009

Validating normalised dates in XML

I had some fun (hmm..., maybe that's not the right word) a year or two ago with regular expressions, trying to come up with something that could validate the kinds of 'normalised' dates that archivists use. You know the ones. The fuzzy dates, the approximates, the uncertains, the 'it was in this decade, but I can't be more precise than that' date, the 'I can tell you the start-date, but not the end-date' (and vice-versa) date. To add to this, we now have the very precise dates associated with born-digital materials - down to the second complete with timezone. In the event, my problem was dispatched by the folks working on PREMIS, who created a union type that brings together some regular expressions to provide a fix (not perfect, but that's regular expressions for you). Just recently the Library of Congress have mounted some pages in the Standards section of their website, where they have put together a nice statement of the problem, as well as pubishing the union type and an XML document with some test dates. See their Extended Date Time Format page.

Friday, 3 April 2009

Draft data dictionary and schema for document significant properties

A data dictionary and related schema has been drafted for those documents that are largely text, but where creators can specify formatting, such as fonts, colours, text size and page layout; where they can embed images and other items; and where there might take advantage of application features, such as the ability to create annotations or page thumbnails. Specifically targetted formats are: OpenDocument Text, PDF, Staroffice, MS Works, MS Word and Wordperfect. Significant properties relating to appearance, behaviour, content and structure are recorded, and it's anticipated that this metadata could be plugged into PREMIS 2.0's objectCharacteristicsExtension.

The designers, from the California Digital Library and Harvard's University Library, are seeking comments from the digital preservation community. Semantic units are: PageCount, WordCount, CharacterCount, ParagraphCount, Line Count, TableCount, GraphicsCount, Language, Fonts, FontName, IsEmbedded, Features. You can see the current schema in full at http://www.fcla.edu/dls/md/docmd.xsd

This looks like a useful addition to preservation metadata, provided tool support for extracting the information and populating metadata records follows. I think the list of values for 'Features' - isTagged, hasLayers, hasTransparancy, hasOutline, hasThumbnails, hasAttachments, hasForms, hasAnnotations - may need extending (hasFootnotes, hasEndnotes?), and it would be good to see some definitions and examples of the existing values.

I wonder if we need a different data dictionary and schema for slideshows? This one might be adequate with some additions to cover things like animations, timings, etc. Seeing this data dictionary also reminds me that we need to look at where the Planets folk are up to on their significant properties work (XCDL/XCEL).

Thursday, 2 April 2009

Digital preservation for individuals and small organisations

Hoppla is a prototype archiving toolkit, with in-built digital preservation capacity; it's not yet available to test, but it sounds very promising. It's being designed specifically for home and small office users by developers at the Institute of Software Technology and Interactive Systems at Vienna University of Technology. It supports preservation at the bit-stream level and includes functions for managing format obsolescence through migration pathways (don't know which objects/pathways). The system also records metadata about preservation actions and object characteristics. If this is reliable, unobtrusive and low maintenance, perhaps we could roll it out to some of the creators the Library works with. It's possible that the acquisition modules could be useful to accessioning archivists too.

Tuesday, 17 March 2009

Shared marginalia for any webpage

Just stumbled on a tool called reframe it. It's available as a Firefox extension, and allows you to add marginalia to any website, whether it has a comments feature of not. Looks like you can share your comments with specific groups, or more widely. Reviews suggest it may need a little fine-tuning, but it could be a useful tool for researchers.

Tuesday, 24 February 2009

Shoot those files!

Just wanted to record a pointer to Manfred Thaller's 'shoot' tool, which I saw demonstrated at a Planets preservation planning event sometime last year (see the posting). It's a handy way to get a feel for which formats suffer most loss of functionality when damaged.

Odds and ends from day one of the digital lives conference

The digital lives conference provided a space to digest some of the findings of the AHRC-funded digital lives project, and also to bring together other perspectives on the topic of personal digital archives. At the proposal stage, the conference was scheduled to last just a day; in the event one day came to be three, which demonstrates how much there is to say on the subject.

Day one was titled 'Digital Lifelines: Practicalities, Professionalities and Potentialities'. This day was intended mostly for institutions that might archive digital lives for research purposes. Cathy Marshall of Microsoft Research gave the opening talk, which explored some personal digital archiving myths on the basis of her experiences interviewing real-life users about their management of personal digital information.

Next came a series of four short talks on 'aspects of digital curation'.

  • Cal Lee, of UNC Chapel Hill, emphasised the need for combining professional skills in order to undertake digital curation successfully. Archives and libraries need to have the right combination of skills to be trusted to do this work.
  • Naomi Nelson of MARBL, Emory University, told a tale of two donors. The first donor being the entity that gives/sells an archive to a library and the second being the academic researcher. Libraries need to have a dialogue with donors of the first type about what a digital archive might contain; this goes beyond the 'files' that they readily conceive as components of the archive, and includes several kinds of 'hidden' data that may be unknown to them. The second donor, 'the researcher', becomes a donor by virtue of the information that the research library can collect about their use of an archive. Naomi raised interesting questions about how we might be able to collect this kind of data and make it available to other researchers, perhaps at a time of the original researcher's choosing.
  • Michael Olson of Stanford University Libraries spoke of their digital collections and programmes of work. Some mention of work on the fundamentals - the digital library architecture (equivalent to our developing Digital Asset Management System - DAMS - which will provide us with resilient storage, object management and tools and services that can be shared with other library applications). Their digital collections include a software collection of some 5000 titles, containing games and other software. I think that sparked some interest from many in the audience!
  • Ludmilla Pollock, Cold Spring Harbour Laboratory, told us about an extensive oral history programme giving rise to much digital data requiring preservation. The collection contains videos of the scientists talking about their memories and has a dedicated interface.
After, we heard from a panel of dealers in archival materials: Gabriel Heaton of Sotheby's, Julian Rota of Bertram Rota and Joan Winterkorn of Bernard Quaritch. I was curious to hear if the dealers had needed to appraise archives conatining obsolete digital media. Digital material is still only a tiny proportion of collections being appraised by dealers, and it seems that what little digital material they do encounter may not be appraised as such (disk labels are viewed rather than their contents). While paper archives are plentiful, perhaps there's not much incentive to develop what's needed to cater for the digital (many archivists may well feel this way too!). What's certain is that the dealer has to be quite sure that any investment in facilitating the appraisal of digital materials pays dividends come sale time.

Inevitably, questions of value were a feature of the session. The dealers suggest that archives and libraries are not willing to pay for born-digital archives yet; perhaps this stems from concerns about uniqueness and authenticity, and the lack of facilities to preserve, curate and provide access. It's not like there's actually much on the market at the moment, so perhaps it's a matter of supply as much as demand? Comparisons with 'traditional' materials were also made using Larkin's magic/meaningful values:

"All literary manuscripts have two kinds of value: what might be called the magical value and the meaningful value. The magical value is the older and more universal: this is the paper [the writer] wrote on, these are the words as he wrote them, emerging for the first time in this particular magical combination. We may feel inclined to be patronising about this Shelley-plain, Thomas-coloured factor, but it is a potent element in all collecting, and I doubt if any librarian can be a successful manuscript collector unless he responds to it to some extent. The meaningful value is of much more recent origin, and is the degree to which a manuscript helps to enlarge our knowledge and understanding of a writer’s life and work. A manuscript can show the cancellations, the substitutions, the shifting towards the ultimate form and the final meaning. A notebook, simply by being a fixed sequence of pages, can supply evidence of chronology. Unpublished work, unfinished work, even notes towards unwritten work all contribute to our knowledge of a writer’s intentions; his letters and diaries add to what we know of his life and the circumstances in which he wrote.”

Philip Larkin 'A Neglected Responsibility: Contemporary Literary Manuscripts', Encounter, July 1979, pp. 33-41.
The 'meaningful' aspects of digital archives are apparent enough, but what of the 'magical'? Most, if not all, contributors to the discussion saw 'artifactual' value in digital media that had an obvious personal connection, whether Barack Obama's Blackberry or J.K. Rowling's laptop. What wasn't discussed so much was the potential magical value of seeing a digital manuscript being rendered in its original environment. I find that quite magical, myself. I think more people will come to see it this way in time.

Delegates were then able to visit to digital scriptorium and audiovisual studio at the British Library.

After lunch, we resumed with a view of the 'Digital Economy and Philosophy' from Annamaria Carusi of the Oxford e-Research Centre. Some interesting thoughts about trust and technology, referring back to Plato's Phaedrus and the misgivings that an oral culture had about writing. New technologies can be disruptive and it takes time for them to be generally accepted and trusted.

Next, four talks under the theme of digital preservation.

  • First an overview of the history of personal films from Luke McKernan, a curator at the British Library. This included changes in use and physical format, up to the current rise of online video populating YouTube, and its even more prolific Chinese equivalents. Luke also talked about 'lifecasting', pointing to JenniCam (now a thing of the past, apparently), and also to folk who go so far as to install movement sensors and videos throughout their homes. Yikes!
  • We also heard from the British Library's digital preservation team, about their work on risk assessment for the Library's digital collections (if memory serves, about 3% of the CDs they sampled in a recent survey had problems). Their current focus is getting material off vulnerable media and into the Library's preservation system; this is also a key aim in our first phase of futureArch. Also mention of the Planets and LIFE projects. Between project and permanent posts, the BL have some 14 people working on digital preservation. If you count those working on webarchiving, audiovisual colections, digitisation, born-digital manuscripts, digital legal deposit, etc., areas, who also have a knowledge of this area, it's probably rather more.
  • William Prentice offered an enjoyable presentation on audio archiving, which had some similar features to Luke's talk on film. It always strikes me that audiovisual archiving is very similar to digital archiving in many respects, especially when there's a need to do digital archaeology that involves older hardware and software that itself requires management.
  • Juan-José Boté of the University of Barcelona spoke to us about a number of projects he had been working on. These were very definitely hybrid archives and interesting for that reason.

Next, I chaired a panel of 'Practical Experiences'. Being naturally oriented toward the practical, there was lots for me here.

  • John Blythe, University of North Carolina, spoke about the Southern Historical Collection at the Wilson Library, including the processes they are using for digital collections. Interestingly, they have use of a digital accessioning tool created by their neighbours at Duke University.
  • Erika Farr, Emory University, talked about the digital element of Salman Rushdie's papers. Interesting to note that there was overlap of data between PCs, where the creator has migrated material from one device to another; this is something we've found in digital materials we've processed too. I also found Rushdie's filenaming and foldering conventions curious. When working with personal archives, you come to know the ways people have of doing things. This applies equally to the digital domain - you come to learn the creator's style of working with the technology.
  • Gabby Redwine of the Harry Ransom Center, University of Texas at Austin gave a good talk about the HRC's experiences so far. HRC have made some of their collections accessible in the reading room and in exhibition spaces, and are doing some creative things to learn what they can from the process. Like us, they are opting for the locked down laptop approach as an interim means of researcher access to born-digital material.
  • William Snow of Stanford University Libraries spoke to us about SALT, or the Self Archiving Legacy Toolkit. This does some very cool things using semantic technologies, though we would need to look at technologies that can be implemented locally (much of SALT functionality is currently achieved using third-party web services). Stanford are looking to harness creators' knowledge of their own lives, relationships, and stuff, to add value to their personal archives using SALT. I think we might use it slightly differently, with curators (perhaps mediating creator use, or just processing?) and researchers being the most likely users. I really like the richness in the faceted browser (they are currently using flamenco) - some possibilities for interfaces here. Their use of Freebase for authority control was also interesting; at the Bod, we use The National Register of Archives (NRA) for this and would be reluctant to change all our legacy finding aids and place our trust in such a new service! If the NRA could add some freebase-like functionality, that would be nice. Some other clever stuff too, like term extraction and relationship graphs.

The day concluded with a little discussion, mainly about where digital forensics and legal discovery tools fit into digital archiving. My feeling is that they are useful for capture and exploration. Less so for the work needed around long-term preservation and access.

Thursday, 12 February 2009

MinivMac

Seems that one of the things we wrestle with when preserving old stuff for use in the future is the question of what I guess is called "transforming content" - the process by which a thing made usable to a reader by (literally) transforming it into a different format (Word 5.5 for DOS (download direct from Microsoft) to Word 2007 for example) - and "preserving environments" - which is where you make DOS and Word 5.5 for DOS and the document available to the reader and let them go back in time.

There are pros and cons to both and the best thing will be to do both. Some readers will want, for example, to experience the pain of using Word for DOS, others will care only for the content of the document and want to read it with their new personal computer (we have to spell it out now since the great PC/Mac debate - folks, a Mac IS a PC!).

Why am I saying all this? Mostly because I sit opposite a wall of shelves that will one day form a museum of old kit and those old machines have kept the subject on my mind for a bit. I have also been experimenting with virtual machines (for reasons beyond emulation) and emulators. Finally, I'm saying all this because Susan tells me this blog is the place to keep and share things that might be useful and so I wanted to log that Apple make their old software available including the OSes and that MinivMac and this Mac-On-A-Stick project looks like they may one day be useful to us. (And if you're a Mac user, check out System 7 via Mac-On-A-Stick - it really isn't much different! :-))

Friday, 6 February 2009

Behold yesterday's snow family

The abundance of snow got me snapping yesterday. Everybody else too from what I saw. Afterwards, many headed home and started downloading, editing and uploading (well, maybe they skipped the editing part). The web is laden with snow-related images, moving and still, from affected parts of the UK. Perhaps we'll be acquiring some of them in personal archives in a few years time. Just for fun, here's one more. I'm particularly proud of the snow dog.

Wednesday, 4 February 2009

Academic Earth

Academic Earth presents 'thousands of video lectures from the world's top scholars'. So far, contributors are from top U.S. universities: Berkeley, Harvard, MIT, Princeton, Stanford and Yale. There is scope for expansion and the Academic Earth team are inviting new partners to contribute.

This is a great idea, but my main reason for linking to Academic Earth is that I rather like the interface. It feels very clean and it's easy to navigate.

Tuesday, 3 February 2009

KEEP project (FP7)

The latest round of FP7 projects in cultural heritage, digital libraries and preservation have started. The Bibliothèque National de France's KEEP project may be interesting - "KEEP addresses the problems of transferring digital objects stored on outdated computer media onto current devices through portable emulators for accurate rendering of both static and dynamic digital objects".

Tuesday, 20 January 2009

Imaging alien and ancient floppies on a contemporary Windows PC

This Omniflop software was designed as a tool for archiving data residing on older floppy formats. It understands a number of older floppy disk formats and can read/write disk images that can be passed to an appropriate emulator. It also claims to be able to work out disk formats it has not previously encountered. Could be worth a closer look.

Library of Congress 'Bagit' transfer tools on Sourceforge

The Library of Congress has released some open source tools for transferring archives at (see Library of Congress Transfer Tools at Sourceforge). The tools support the Bagit specification, and can check for missing, extra and duplicate files as part of the process. Transfers can be made using rsync, http and ftp. The verify script also checks hash values of files against those listed in the Bagit manifest.

Wednesday, 7 January 2009

Metrics toolkit for assessing usability of archival websites

There was some discussion yesterday on the EAD list about evaluating the effectiveness of online finding aids. Wendy Duff drew attention to a project she is involved in, called 'Archival Metrics', which has put a toolkit together for undertaking user evaluation of online finding aids. There are also other archival metrics toolkits available for download. Am wondering whether this will be useful to us as interface developments become a larger part of our work.

Tuesday, 6 January 2009

A copyright agenda for C21?

Tim Padfield signaled yesterday (nra posting) that the Intellectual Property Office are undertaking a new review of copyright legislation in the UK. An issues paper, designed to start a debate, was published on the IPO's website in December. This new review is intended to build on the work of the 2006 Gowers Review; the paper says that work arising from the Gowers Review continues to be taken forward (format shifting is specifically mentioned, I'm pleased to say), but that there has been change enough in the 'technological and commercial landscape of the creative industries' to merit new consideration.

The issues paper is succinct and there are a few interesting ideas in it, such as:
  • to what degree should creators be able to control the way in which their work is re-used?
  • how can exceptions to copyright remain effective while contractual and technical measures override them?
  • does a personal blog deserve the same copyright protection as the work of a best-selling author?
The paper solicits responses to the following four questions:

1. Does the current system provide the right balance between commercial certainty and the rights of creators and creative artist? Are creative artists sufficiently rewarded/protected through their existing rights?

2. Is our current system too complex, in particular in relation to the licensing of rights, rights clearance and copyright exceptions? Does the legal enforcement framework work in the digital age?

3. Does the current copyright system provide the right incentives to sustain investment and support creativity? Is this true for both creative artists and commercial rights holders? Is this true for physical and online exploitation? Are those who gain value from content paying for it (on fair and reasonable terms)?

4. What action, if any, is needed to address issues related to authentication? In considering the rights of creative artists and other rights holders is there a case for differentiation? If so, how might we avoid introducing a further complication in an already complicated world?

Monday, 5 January 2009

Chris Rusbridge's '12 Files of Christmas'

Just before the Christmas break, Chris Rusbridge of the DCC offered to recover obsolete files. Today is the 'closing date' for submissions. I hope Chris gets some interesting submissions that we can hear about in due course. See Chris' blog entry for more info.

Hypertext and self-destructing poems


A few of us attended the Flair Symposium at the Harry Ransom Center in Austin in November. It was an opportunity to hear about literary archives from the perspective of the writers that create them and the archivists that curate them.

While at the HRC, I also found some time to take in their (now just closed) exhibition - The Mystique of the Archive. The exhibition described the contents of a literary archive, exploring its journey from creator to archival repository to scholar. Two items in particular stuck with me. First, the plot chart for Norman Mailer's Harlot's Ghost, and second the display of Michael Joyce's Afternoon. The first has all the magical qualities of a traditional manuscript - it's a window on the composition process. The second is one of few exhibits of born-digital archive material that I have seen. Afternoon is a work of hypertext fiction created using software called Storyspace. You can buy both the current versions of the work and the software used to create it from Eastgate. I have never read a work of hypertext fiction, but I understand that the format presents the reader with a variety of paths through the work. Matthew Kirschenbaum of MITH is working with the archive of Deena Larsen, a hypertext author who encourages the extension of her works; this introduces all sorts of interesting questions about the (potential) scope of the archive.

Matt was also the first user of the Joyce archive at the HRC and he writes about this in his book Mechanisms. He writes too of the importance of the physical medium on which data is inscribed and the use of data forensic techniques to recover and examine data. There's much in Mechanisms that chimes with issues and approaches we have encountered, though the terminology is very different to that used by archivists. At the FLAIR Symposium, Matt spoke about the Agrippa Files (also the subject of a chapter in Mechanisms) and there's a website dedicated to this. The work that's been done on Agrippa is an interesting example of data recovery and emulation. The poem was a digital work designed to scroll on the screen, accompanied by the odd sound effect. It was also crafted to self-destruct by encrypting itself after a single reading. The poem has been recovered, using dd, from a 3.5" disk used with an early 90s Mac. On the website, you can now view a video of the poem running in the mini vMac emulator booted with a System 7 boot disk.

The NEH has funded a collaboration between MITH, HRC and Emory on digital literary manuscripts. The grant supports a series of site visits, which will enable the partners to share experiences and to develop a fuller proposal for the curation of contemporary literary archives. We were lucky enough to meet with the good folk of this collaboration at Texas. As well as hearing about MITH's work with the Larsen archive, we learned of their work on the Jonathan Larson archive and a project which investigates the preservation of virtual worlds (see Kari Kraus speak on this at the Metaverse conference). We also heard from Erika Farr and Naomi Nelson about Emory's experiences with the digital elements of the Rushdie archive, and from Gabriela Redwine about the Ransom Center's work on the digital materials in its collections (including the Joyce and Mailer archives). The stories from all three institutions stem from hands-on engagement with born-digital archives, which makes their telling all the more valuable; others will be able to read about this work when a paper is published toward the end of the NEH grant.

Thursday, 6 November 2008

Digital Preservation Policies Study published

JISC appointed Charles Beagrie to develop this Digital Preservation Policies Study back in March with a pretty small timeframe to deliver the goods. It's well worth a look if you're putting your own digital preservation policy together, especially if you're operating in the HE sector. The study provides a handy template of policy clauses that can be adapted for local needs. That should help people get started.

If I'm honest, it's the emphasis on context and mappings that I like best about the work (as an archivist I'm bound to like that ;-), right?). By examining policy documents in the areas of research, teaching and learning, information, library and records management, the study has identified how digital preservation supports the work of universities. This alignment of digital preservation policy to the business of the University is critical to answering questions about why digital preservation matters. Anyone needing to make the case for digital preservation should take a look at the detailed mappings to these wider university policies in the appendices.

Friday, 24 October 2008

Slides from Digital Archives meeting

These are some slides from a talk given at a meeting on Digital Archives hosted by the Andrew W. Mellon Foundation back in September. They give an overview of what the futureArch project is about.
20080903arsenalsofnemesis 04
View SlideShare presentation or Upload your own. (tags: digital archives)

Thursday, 23 October 2008

iPres presentations now online

This year iPres was hosted by the British Library. Much food for thought, so much so that choosing between parallel sessions was something of a challenge. Good news then that the presentations and full papers are now available online.

Wednesday, 1 October 2008

Fun with tag clouds

Not the traditional form of indexing an archive, I know, but it seems to me that automagically extracted metadata formed into tag clouds would be a marvelous way of navigating through some digital archives.

We could present clouds at different levels of granularity - at the collection level, in series and lower levels all the way down to the item. We could even present clouds across multiple aggregations, be they of series, collections or items. This could be fun.

For some digital archives, I think tag clouds are probably a 'must'. Poorly structured and overly large email archives are a good candidate.

One of the downsides of the 'hybrid archive' is that we can't necessarily generate tag clouds that draw on all the contents of the archive. All 'physical' material and non-textual digital formats are excluded unless these things are already tagged by creators. They can, of course, be tagged later by cataloguers and/or users. I guess that we need to recognise that imbalance in our user interface, to help our users get to grips with the nature of research in a hybrid archive.

I know that automatic metadata extraction may have shortcomings, but I'd really like to see a fusing of standardised subject headings with tag clouds. We can have the best of both worlds, surely?

There have been lots of examples of tag clouds about recently, including TagCrowd and Wordle.

This is a Tag Crowd entry for this blog...



created at TagCrowd.com


Saturday, 27 September 2008

XML Schema for archiving email accounts


I attended several great sessions at the Society of American Archivists conference last month. There is a wiki for the conference, but very few of the presentations have been posted so far...

One session I particularly enjoyed addressed the archiving of email - 'Capturing the E-Tiger: New Tools for Email Preservation'. Archiving email is challenging for many reasons, which were very well put by the session speakers.

Both the EMCAP and CERP projects were introduced in the session.

EMCAP is a collaboration between state archives in North Carolina, Kentucky, and Pennsylvania to develop means to archive email. In the past, the archives have typically received email on CDs from a variety of systems, including MS Exchange, Novell Groupwise and Lotus Notes. One of the interesting outcomes of this work is software (an extension of the hmail software - see sourceforge) that enables ongoing capture of email, selected for archiving by users, from user systems. Email identified for archiving is normalised in an XML format and can be transformed to html for access. The software supports open email standards (POP3, SMTP, and IMAP4) as well as MySQL and MS SQL Server. The effort has been underway for five years and the software continues to be tested and refined.

CERP is a collaboration between the Smithsonian Institution Archives and Rockefeller Center Archives. This context has more in common with archiving email in the Bodleian context, where an email account is more likely to be accessioned from its owner in bulk than cumulatively. Ricc Ferrante gave an overview of the issues encountered, which were similar to our experiences on the Paradigm project and in working with creators more generally.

CERP has worked with EMCAP to publish an XML schema for preserving email accounts. Email is first normalised to mbox format and then converted to this XML standard using a prototype parser built in squeak smalltalk, which also has a web interface (seaside/comanche). The result of the transformation is a single XML file that represents an entire email account as per its original arrangement. Attachements can be embedded in the XML file, or externally referenced if on the larger side (over 25kb). If I remember rightly, the largest email account that has been processed so far is c. 1.5GB; we have one at the Library that's significantly larger and I'd like to see how the parser handles this. It will be interesting to compare the schema/parser with The National Archives of Australia's Xena. The developers are keen to receive comments on the schema, which is available here.

Monday, 11 August 2008

Keeping the user experience in the browser

A few days back, news that Barrack Obama is using Scribd prompted me to take another look at this document sharing site. I'm interested in user interfaces at the moment because developing interfaces for curators and researchers wanting to use hybrid archives will be an important part of futureArch's work.

iPaper
I'm quite taken with iPaper's features. I should explain that iPaper is the flash technology developed by Scribd to display documents published to the sites by its users. Documents load quickly within the browser and the functionality is similar to that of Adobe Acrobat. You can search text, there's a thumbnail view, you can zoom, launch the document to display at full screen, turn the pages, etc. It also has some 'social' features (which may or may not be useful in this context), and it claims to be a more secure document format than PDF.

iPaper can display files encoded in a number of formats, a feature that may well prove useful for developing a browser-based interface to your typical born-digital archive. Imagine wanting to view a handful of items in a collection, each of which is in a different format. Rather than launching a different application to render each format (even if these are available as browser plug-ins), the user could access all the items using a single lightweight viewer that can be embedded in the web page itself. This normalisation for presentation simplifies access for repositories and users, and for those users interested primarily in content, a single viewer would provide a convenient and predictable experience that requires no software installations.

So far iPaper supports these formats:

* Adobe PDF (.pdf)
* Adobe PostScript (.ps)
* Microsoft Word (.doc/ .docx)
* Microsoft PowerPoint (.ppt/.pps/.pptx)
* Microsoft Excel (.xls/.xlsx)
* OpenOffice Text Document (.odt, .sxw)
* OpenOffice Presentation Document (.odp, .sxi)
* OpenOffice Spreadsheet (.ods, .sxc)
* All OpenDocument formats
* Plain text (.txt)
* Rich text format (.rtf)

For users interested in qualities beyond content, technologies such as iPaper may be less useful. On example is users who require an experience that reflects the context of creation; use of the iPaper format and viewer requires the transformation of the original item into iPaper format and its rendering in an environment quite different to that of its creation. Another example is users wishing to use specialist analytical tools, which might be domain- or format-specific.

Scribd
There are some aspects of Scribd itself that appeal to me too. A collection overview pane sits next to a list of child items in the collection, each of which have a thumbnail, page-count, format indicator and brief abstract. I'm sure we could do something similar in a view for archival collections, though an archive repository would more likely point to series descriptions from the collection level, and to items at lower levels of the archive's hierarchy. Scribd's item-level view works well too: the document is displayed in the page (using the iPaper viewer) and a little metadata is available on the right - some tags, rights information (creative commons), relevant categories, etc., and since this is web 2.0, users are able to add their comments.

Other possibilities?
I've also been investigating another means of enabling browser-based delivery for the kinds of file formats found in a born-digital archive. More and more of us want to create and view data at the network level and we want to be able to do it with a variety of devices. This can only mean that more options are going to be available, but, as ever, the creative direction is unlikely to correlate exactly with our requirements.

One possibility is javascript lightboxes, so long as the usual accessibility issues are addressed. Many of these tools are designed for image galleries, but there are some with other functionality too. Highslide is one I've spent some time looking at, and now version 4 is just out (five days old) it might be time to take another look. Perhaps the subject of a future post.

Thursday, 31 July 2008

DPC's preservation planning workshop

Earlier in the week I attended a DPC workshop on preservation planning, which was largely constructed of material coming out of the European project called Planets, which is now half-way through its four-year programme. There were also interesting contributions from Natalie Walters of the Wellcome Library and Matthew Woollard of the UK Data Archive.

A preservation system for the Wellcome Library?
Much of what Natalie had to say about the curation of born-digital archives chimed with our experiences here. Unlike us though, Wellcome are in the process of evaluating 'off the shelf' systems to manage digital preservation. They put out a tender earlier this year and received five responses that seem, in the main, to demonstrate a misunderstanding of archival requirements and the immaturity of the digital curation/preservation marketplace. One criticism was that the responses offered systems for 'access' or 'institutional repositories' (of the kind associated with open access HE content - academic papers and e-theses). This is something we also felt when we evaluated the Fedora and DSpace repositories on the Paradigm project (admittedly, this evaluation becomes a bit more obsolete day by day). Balancing access and preservation requirements has long been an issue for archivists, since we often have to preserve material that is embargoed for a period of time. I still believe that systems providing preservation services and systems providing researcher access are doing different things, but we do of course need some form of access to embargoed material for management and processing purposes. I also find the adoption of new meanings for words, like 'repository' and 'archive', tricky to negotiate at times. These issues aside, one of the systems offered seems to have held Wellcome's interest and I'll be keen to find out which one when this information can be revealed.

Preservation policy at UKDA
Matthew spoke about the evolution of preservation policy at the UKDA, which had no preservation policy until 2003 despite celebrating its 40th anniversary last year. The first two editions of the policy were more or less exclusively concerned with the technical aspects of preserving digital material, specifying such things as acceptable storage conditions and the frequency with which tape should be re-tensioned. The latest (third) edition embraces wider requirements including organisational/business need, user requirements (designated community and others), standards, legislation, technology and security. The new policy increases emphasis on data integrity and archival standards, it defines archival packages more closely to provide for their verification, and it pays attention to the curation of metadata describing the resources to be preserved.

If I understood correctly, the UKDA preserves datasets in their original form (SIP), migrates them to a neutral format (AIP1) and creates usable versions from the neutral format (AIP2). All these versions are preserved and dissemination versions of the dataset are created from AIP2. The degree of processing applied to a dataset is determined by applying a matrix which assigns a value on the basis of likely use and value. These processes feel similar to those evolving here, though we need to do more work to formalise them.

Matthew also showed us a nice little diagram from 1976, which was created to document UKDA workflow from initial acquisition of a dataset to its presentation to the final user. The fundamentals of professional archival, or OAIS-like, practice are evident. The UKDA's analysis of its own conformance with the OAIS model undertaken under the JISC 04/04 Programme is worth a look for those who haven't seen it.

Towards the end of the talk Matthew reminded us that having written a policy, one must implement it. It's not normally possible to implement every new thing in a policy at once, but the policy is valueless without mechanisms in place to audit it. Steps must be taken to progress those aspects of the policy that are new and to audit compliance more generally. The policy must also be available to relevant audiences who can evaluate the degree to which the archive complies with its own policy for themselves. I found this a very useful overview of the key issues involved in developing a preservation policy and the resulting policy itself is very clear and concise.

Planets tools for preservation planning
It's great to see the promise of Planets starting to be realised, especially since we plan to build on the project's work in relation to characterising material, planning and executing preservation strategies. Andreas Rauber kicked things off with an overview of the Planets project, which helped to demonstrated how the various components fit together.What is uncertain at the moment is how the software and services being developed by Planets will be sustained beyond the project's life. Neither is it clear what licensing model/s will be adopted for different components in the project, since there are the needs of commercial partners to consider as well as those of national archives, libraries and universities.

Plato
Christoph Becker gave us an overview of Plato, a tool which allows the user to develop preservation strategies for specific kind of objects. In Plato, users can design experiments to determine the best available preservation strategy for a particular type of material. This involves a formal definition of constraints and objectives, which includes an assessment of the relative importance of each of these factors. Factors might include:

* object migration time - max. 1 second
* object migration cost - max £0.05 per object
* preserve footnotes - 5
* preserve images- 5
* preserve headings - 4
* open format required - 5
* preserve font - 3
* and so on...

These are expressed in an 'objective tree', which can be created directly in Plato or in the Freemind mind mapping tool and uploaded to Plato. Objective tress can be very simple, but the process of creating a good and detailed objective tree is quite demanding (we had a go at doing this ourselves in the afternoon). In future we should be able to build on previous objective trees as these are developed and that will ease the process. For the moment the templates provided are minimal because the Plato team don't want to preempt user requirements!

The user must also supply a sample of material which can be used to assess the effectiveness of different strategies. This should be the bare minimum of objects required to represent the range of factors expressed in the objective tree. The user then selects different strategies to apply to the sample material, sets the experiment in motion, and compares the results against the objective tree. The process of evaluating results is manual at present, but there are plans to begin automating aspects of this too. Once the evaluation is complete, Plato can produce a report of the experiment which should demonstrate why one preservation strategy was chosen over another in respect of a particular class of material.

Plato is available for offline use, which will be necessary for us when processing embargoed material, but it is also offered as an online service where users can perform experiments in one place and benefit from working with the results of experiments performed by others.

Characterisation
The Planets work on characterisation was introduced by Manfred Thaller. This work develops two formal characterisation languages - the extensible characterisation extraction language (XCEL) and the extensible characterisation description language (XCDL). The work should make it possible to perform more automatically determine whether a preservation action, such as migration, has preserved an object's essential characteristics (or significant properties). It is expected that the Microsoft family of formats, PDF formats and common image formats will treated before the end of the project.

One of the interesting aspects of the characterisation work is developing an understanding of what is preserved or not in a particular process and how a file format impacts on this. Thaller demonstrated this (using a little tool for *shooting* files) by deliberately causing a small amount of damage to a png file and a tif file. A small amount of damage to the png file had severe consequences for its rendering, while the tif file could be damaged much more extensively and still retain some of its informational value. Thaller also used the example of migrating a MS Word 2003 document to the Open Document Text format. The migration to ODT seemed to lose a footnote in the document. Thaller then showed the same MS Word 2003 document migrated to PDF, where the footnote appears to be retained. In actual fact the footnote isn't lost in the migration to ODT, it's just not rendered. On the other hand, the footnote is structurally lost in the PDF file, but visually present. Thaller is proposing a solution which allows structure and appearance to be preserved.

Testbed
The final element of planets on show was the testbed developed at HATII, demonstrated by Matthew Barr. The testbed looks very useful and, like Plato, will be available for use online and offline. There did seem to be some overlap in aims and functionality with Plato, but there are differences too. It's essential objectives seem similar - users should be able to perform experiments with select data and tools, evaluate those experiments and draw conclusions to inform their preservation strategy; the testbed will also tools and services to be benchmarked. It struck me is that the process of conducting an experiment was simpler than with Plato, since a granular expression of objectives is not necessary. It's more quick and dirty, which may suit some scenarios better, but will the result be as good? Aspects I found particularly interesting were the development of a corpora and the ability to add new services (tools are deployed and accessed using web services) for testing.

Monday, 21 July 2008

Seeking a Software Engineer

We are looking for a Software Engineer to work on the futureArch project. You can read the job advertisement at the vacancies section of the University of Oxford's website; there is also a link to the further particulars from here. Closing date is 29 August 2008.

Wednesday, 16 July 2008

Annotating sound and video

At the JISC innovation forum earlier this week, I was fortunate enough to run into an improptu demo of Synote by Mike Wald of ECS, who had hijacked the British Library's sound archive project stand. Well, perhaps 'hijack' is a little strong - the BL demo was pretty much done and Peter Findlay was happy to tune in to what Mike was showing.

Synote is a rather nifty tool which lets users add annotations to specific points in a digital sound or video recording. These annotations might be notes, tags, or images; they act like bookmarks - they can be returned to easily as and when the need arises. Synote uses a transcript of the audio, which can be generated by speech recognition software if the audio is clean enough, or compiled by hand if not. The transcript plays alongside the content, and the users' annotations are highlighted in it; clicking on a word in the transcript allows the user to skip ahead to the bookmark and, of course, the transcript is searchable. It's been designed as a teaching and learning tool, but I think it has a lot of possibilities as a means of interacting with audio and video content present in archival collections. The project has a sourceforge page, so hopefully we'll be able to have a go ourselves in due course.

Wednesday, 2 July 2008

Greetings!

Hello world, as they say. Welcome to this new blog, which is a place for those of us working with born-digital archives at the Bodleian Library to share our thoughts, frustrations and successes. We'll also be making a note of interesting or useful things we stumble upon.

We've been working on issues relating to the long-term preservation of digital archives for a few years now. If you take a look at our Paradigm and Cairo projects, that should give you an idea of the kinds of issues we're dealing with.

This blog is being born as we launch an important phase of development at the Library. We're about to begin the futureArch project, which will see us move the curation of born-digital archives and manuscripts from a series of small projects to a sustainable activity integrated with other aspects of the Library's operations. When futureArch concludes, in just over three years time, we aim to have embedded the curation of born-digital archives and manuscripts into the way we do things.

That's probably more than enough for a first post...