Showing posts with label autogenerated metadata. Show all posts
Showing posts with label autogenerated metadata. Show all posts

Thursday, 11 March 2010

scat @ Gloucestershire archives

Yesterday, I enjoyed a very scenic drive through the Cotswolds to attend a workshop at Gloucestershire Archives (GA). GA have been working on digital curation for a few years now, using their historical digital archives - a discrete and reasonably unproblematic set of data - as a testbed for developing approaches that might help them with the more modern digital records created by the local authority in due course.

Yesterday's event was the culmination of a six month project 'Digital curation: from ingest to trusted storage' and was designed to provide attendees with lots of hands-on time, using the 'SCAT' tool developed through the project. The project was funded by the Society of Archivist's research fund, but builds on previous work funded by CyMAL to develop GAip - a software tool which packages digital objects ready for ingest to preservation storage. There is not much in the way of a web presence for GAip, or its successor project just yet, but the slides from Viv's presentation at the Society of Archivist's digital preservation road show give a flavour of GAip at least.

SCAT
The main output of the recent project is a tool called SCAT (Scat is Curation And Trust). The tool is written in perl and currently runs only in Linux environments; with a few modifications it should also run on a Windows platform. SCAT provides an interface to a number of open source digital curation tools that exist out in the wild; by loading a file or directory into SCAT, it is possible to apply these curation tools to them. Among the tools represented are:
  • Bagit - the Library of Congress tool mentioned elsewhere on this blog
  • GAip - GA's own packaging tool, which creates a Bagit-conforming package. This is used by GA to package its digital archives.
  • DROID - The National Archives' tool for file format identification
  • Jhove - the tool developed by JSTOR and Harvard for object identification, validation and metadata extraction
  • NLNZ metadata extraction tool - extracts basic metadata from some popular formats
  • FITS - identifies and validates files, and extracts technical metadata. It is a wrapper for a number of third-party tools (Jhove, EXIFtool, NLNZ, DROID, Ffident and the file utility). The intersting things about FITS is that is also attempts to normalise and consolidate the metadata output from these tools. Pete is using FITS at the moment to generate certain file-level metadata for dissemination purposes.
  • Antiword - a reader for Word files
  • Imagemagick - a tool that can do many wonderful things with image files
  • xmllint - used for validating XML files against their schemas, etc.
  • Unix's file utility.
  • SWORD deposit to repository (GA have been experimenting with an eprints instance in this project)
  • tools for fixity checking employing the MD5 and SHA1 algorithms.
  • document conversion tool (destination formats odt and pdf - possibly openoffice.org?)
This list isn't comprehensive, but it's enough to show you that there is a strong open source philosophy underpinning the project.

SCAT is still very much alpha code, and Viv Cothey (its developer) intends to do a bit of tidying up before putting it out on the web. It's really been designed to provide a hands-on learning space for archivists, and is not conceived as a ready-to-use system for digital curation. As a learning environment, I think it is very effective, providing a workbench which can call up a whole host of tools that the archivist can experiment with.

GAip packages
It's worth saying a little something about the package produced by GAip. Using the tool, the archivist can decide to package:
  • a single item
  • a collection of materials as single items
  • a collection of materials as a bundle, retaining their directory structure.
In a GAip package you will find a:
  • copy of the original source data
  • a sidecar metadata record for each data item (GAip uses XMP for this, and the metadata includes dublin core, and other, metadata expressed in rdf)
  • an inventory of the files contained in the package including their hash values 
The package is a compressed tar file (.tar.gz, although gaip uses a .gaip file extension). Each package is identified by a unique timestamp.

This is not a million miles away from our approach with the BEAM  ingest tool.

Building on SCAT
Could SCAT be developed into something more? There are a number of areas that would need to addressed. Some of the ones raised in discussions yesterday include:
  • the need to support multiple users (both GAip and SCAT are conceived as single-user software at present)
  • a better approach to unique, and persistent, identifiers
  • a method by which data objects can accrue additional metadata in their XMP sidecars beyond that supplied when the GAip package is created
  • workflow
If any of this is of interest, then I'm sure Viv Cothey would be pleased to hear from potential collaborators.

Thursday, 9 July 2009

Automatic Metadata Workshop (Long post - sorry!)

Sometimes I wonder if automatic metadata generation is viewed a little like the Industrial Revolution; which is to say that it is replacing skills and individuals with large scale industry.

I do not think it is really like that at all, being much more about enabling people to manage the ever increasing waves of information. It isn't saying to a weaver, "we can do what you can only faster, better and cheaper"; it is saying "here is something to help you make fabrics from this intangible intractable ether".

What got on this philosophical tract? The answer, as ever, is a train journey - in this case the ride home from Leicester, having attended a JISC-funded workshop on Automatic Metadata Generation. Subtitled "Use Cases" the workshop presented a series of reports outlining potential scenarios in which automatic metadata generation could be used to support the activities of researchers and, on occasion, curators/managers.

The reports have been collated by Charles Duncan and Peter Douglas at Intrallect Ltd. and the final report is due at the end of July.

The day started well as I approached the rather lovely Beaumont Hall at the University of Leicester and noted with a smile the acronym on a sign - "AMG".

Now, I'm from Essex so it is in my genes to know that AMG is the "performance" wing of Mercedes and looking just now at the AMG site, it says:

"Experience the World of Hand Crafted Performance"

a slogan any library or archive could (and should) use!

(Stick with me as I tie my philosophising with my serendipitous discovery of the AMG slogan)

I couldn't help but think a AMG-enabled (our sort, not the car sort) Library or Archive is like hand crafting finding aids, taking advantage of new technology, for better performance. I also thought that most AMG drivers don't care about the science behind getting a faster car, but just that it is faster - think about it...

Where was I? Oh yes.

The workshop!

It was a very interesting day. The format was for Charles, Peter or, occasionally the scenario author, to present the scenarios followed by an opportunity for discussion. This seemed to work well, but it was unfortunate that more of the authors of the scenarios themselves were unable to attend and give poor Charles and Peter a break from the presenting!

The scenarios themselves were around eight metadata themes:
  1. Subject-based
  2. Geographic
  3. Person-related
  4. Usage-related
  5. File Formats
  6. Factual
  7. Bibliographic
  8. Multilingual/Translated
I'll not cover all the scenarios here, but you are encouraged to visit the project Wiki where you can find more information and look out for the final report, but here are some things I got from the day:
  1. AMG to enhance discovery through automatic classification, recommendations on the basis of "similar users" activity ("also bought" function), etc. Note that this is not "by enhancing text-based searching".
  2. AMG could encourage more people to self-deposit (to Institutional Repositories) by automatically filling in the metadata fields in submission forms (now probably isn't the time to discuss the burden of metadata not being the only reason people don't self-deposit! :-)).
  3. AMG to help produce machine-to-machine data and facilitate queries. The big example of this was generating coordinates for place names to enable people with just place names to do geospacial searches, but there are uses here for generating Semantic Web-like links between items.
  4. AMG for preservation - the one I guess folks still reading are most familiar with. Identifying file formats, using PRONOM, DROID & JHOVE, etc. to identify risks, etc.
  5. AMG at creation. Metadata inserted into the digital object by the thing used to create it - iTunes grabbing data from Gracenote and poplating ID3 tags in its own sweet way, a digital camera recording shutter speed and appeture size, time of day and even location and embedding that data into the photo.
  6. The de facto method of AMG was to use Web services - with a skew towards REST-based services - which probably brings us back to cars - REST being nearer the sleek interior of a car than SOAP which exposes its innards to its users.
  7. Just in time AMG (JIT AMG - now there's a project acronym). When something like a translation service is expensive why pay to have all your metadata translated to a different language when you may be able to just do the titles and give your users the option to request (and get the result instantly) a translation if they think it useful.
  8. You might extend JIT AMG and wonder if it is worth pushing the AMG into the search engine? Text search engines already do that - the full-text being the bulk of the metadata - so what if a search engine were also enabled to "read" a music manuscript (a PDF or a Sibelius file for example) and you search for a sequence of notes. Would there be any need to put that sequence of notes into a metadata record if the object itself can function as the record (if you'll forgive the pun!)?
So what does all that mean for us?

Well, it is pretty clear that futureArch must rely on automatic metadata creation at all stages in the archival life cycle and a tool-chain to process items is a feature on diagrams Renhart has shown me since I arrived. It just would not be possible to manage a digital accession without some form of AMG - anyone fancy hand-crafting records for 11,000 random computer files? (Which are, of course, not random at all - representing as they do an individuals own private "order").

I worry slightly about the Web service stuff. For a tool to be useful to futureArch we need a copy here on our servers. First and foremost this ensures the privacy of our data and secondly we have the option then of preserving the service.

(Not to mention that a Web service probably wouldn't want us bombarding it with classification requests!)

(Fortunately the likes of DROID have already gone down the "engine" and "datafile" route favoured by anti-virus companies and let us hope that pattern remains!).

I quite like the idea of resource as metadata object, but I suspect it remains mostly unworkable. It was by accident rather than design that text-based documents, by virtue of their format, contain a body of available metadata. Still, I imagine image search engines are already extracting EXiF data and how many record companies check MP3s ID3 tags to trace their origins...? ;-)

At the end of the workshop we talked a bit about how AMG can scare people too - the Industrial Revolution where I started. To sell AMG technologists talk of how it "reduces cataloging effort", but in an economic climate looking for "reductions in cost" it is easy for management to assume the former implies the later, not realising that while the effort per item may go down, there are much more items!

Whether or not this is true remains to be seen, but early indications suggest AMG isn't any cheaper - just as any new technology isn't. It is just a different tool, designed to cope with a different information world; an essential part of managing digital information.

Yep, it is out of necessity that we will become the Automatic Metadata generation... :-)

Wednesday, 1 October 2008

Fun with tag clouds

Not the traditional form of indexing an archive, I know, but it seems to me that automagically extracted metadata formed into tag clouds would be a marvelous way of navigating through some digital archives.

We could present clouds at different levels of granularity - at the collection level, in series and lower levels all the way down to the item. We could even present clouds across multiple aggregations, be they of series, collections or items. This could be fun.

For some digital archives, I think tag clouds are probably a 'must'. Poorly structured and overly large email archives are a good candidate.

One of the downsides of the 'hybrid archive' is that we can't necessarily generate tag clouds that draw on all the contents of the archive. All 'physical' material and non-textual digital formats are excluded unless these things are already tagged by creators. They can, of course, be tagged later by cataloguers and/or users. I guess that we need to recognise that imbalance in our user interface, to help our users get to grips with the nature of research in a hybrid archive.

I know that automatic metadata extraction may have shortcomings, but I'd really like to see a fusing of standardised subject headings with tag clouds. We can have the best of both worlds, surely?

There have been lots of examples of tag clouds about recently, including TagCrowd and Wordle.

This is a Tag Crowd entry for this blog...



created at TagCrowd.com