Saturday, November 22, 2008

Pre-installed VM

I think building my own content management systems was a highly worthwhile series of exercises. I learned something about the technological components of each system I built. I think this knowledge provided some insight for how individual installations worked. By building a series of content management systems, I have begun to get an idea of features and aspects common to all of the systems. With this familiarity with some general outlines for how content management systems are installed, much of the mystery of the processes and technologies associated with installing a content management system has been dispelled. The repetitive aspects of building my own system were helpful. Repeating technical procedures (creating a virtual machine, downloading software, modifying files, setting permissions) developed for me a familiarity and comfort level with those procedures. I would not have such a level of familiarity or comfort with those technical procedures without practicing them repeatedly. Building a series of content management systems increased the opportunity to make errors and encounter difficulties in the build processes. The challenges these errors and difficulties presented were valuable learning experiences, particularly helping to develop analytical and problem solving skills and strategies when working with technology. Although these challenges were frequently frustrating, I think they were valuable learning experiences.

That said, it is also true that downloading pre-installed virtual machines would provide more time to concentrate on experimenting with the installed system. I think this is especially true for students taking IRLS 675 while concurrently taking another course (like, for me, this semester, IRLS 671). Generally this semester I was primarily focused upon the timely completion of all my coursework. Maybe this focus was partly because it had been some years since I had taken two concurrent college-level courses. While focused on completing my coursework in a timely manner, I found it challenging to "shift gears" to allocate time to experiment. An general issue here is that of a student learning activity (experimentation) which seems inherently difficult to incorporate into a formalized course structure. Incorporating experimentation would be worthwhile, because there would be much to learn with these systems. My idea has been to go back to these systems and get more deeply involved with exploring them at some time subsequent to the DigIn coursework. I think IRLS 675 did a good job of introducing the basics of various content management systems, which is what I think was the course objective. Pre-installed VM would provide more time to concentrate on the collections, but what would this achieve? It is hard conceive of a realistic learning objective beyond that of the level of introduction to the systems. For instance, it seems like an entire semester or more could be spent on Drupal. So I don't know what would be the demonstrable benefit of giving the students a marginal amount of time more to work with the installed systems, especially at the cost of reducing their learning experience with applied technologies. Moreover, without a good familiarity with technological routines associated with installing the content management systems, I think I would have a significant increase in difficulty with experimenting with these systems after the DigIn Program concludes.

I think it is relevant to consider what are the objectives for student learning in IRLS 675. What skill sets does DigIn want its graduates to have? I suppose this is a difficult question. Apparently there is a lack of clear consensus on what the competencies for digital collection managers should be. (Maybe an increased clarification of the competencies for digital collection managers can be achieved though an audit of skills/experience requirements specified in a representative sample of relevant position vacancy announcements.) The DigIn page says the compelling issue is that an "explosion of digital information and the growth of online digital resources has led to a shortage of individuals with an understanding of the disciplines of libraries, document management and archives who also have the technical knowledge and skills needed to create, manage and support digital information collections." If the DigIn Program is satisfied with this characterization of the compelling issue, then certainly making students proficient with the technological procedures of building content management systems is relevant to the "technical knowledge" described in this characterization.

Saturday, October 25, 2008

Service Providers

Service providers who specialize in a topic provide useful functionality limiting the service to a specialized topic. For example, the Avano marine and aquatic sciences service provider (http://www.ifremer.fr/avano) clearly conveys what are the topic strengths of the service (marine and aquatic sciences.) Avano has a very useful Browse Archives feature, which displays the collections Avano draws from, links to the collections, the number of items harvested from each collection, and a description of the collection. Avano data providers include the Alfred Wegener Institute for Polar and Marine Research, the Aquatic Commons, and the Open Marine Archive.

The Scientific Commons (http://www.scientificcommons.org/) is a service provider which harvests a wide variety of scientific topics. The Scientific Commons has harvested 23 million records from 9 million authors in 950 archives. Searching a topic produces very large numbers of results (somewhat Google-like.) The search results can be filtered through only a limited number of options (Year, Language, and Sort Mode.) I think large-scale service providers such as the Scientific Commons call to question whether such services providers scale well for topic searching. The data providers to Scientific Commons are unidentified.

PLEIADI: Portal for Italian Literature in Open and Institutional Archives (http://www.openarchives.it/pleiadi/index.php?sel_lang=english) incorporates a variety of functionality in a multi-featured portal. Features include personal accounts, recent news, a wiki, RSS, and related web resources. Searching PLEIADI, it appears that many of its data providers are Italian university repositories.

Generally, I think it is most important for a service provider to harvest relevant data providers. It seems to me that service providers specializing in a topic can be more useful (especially searching on topics) than large-scale comprehensive service providers. In addition to harvesting relevant data providers, I think it is important that service providers have powerful search interfaces that provide users with sufficient filtering options and other search options.

Monday, October 20, 2008

EPrints

I thought it might be a good idea to try using the Library of Congress Subject Headings to describe my collection (a collection of evaluation reports in PDF) in EPrints. The first two levels of Library of Congress Subject Headings were insufficiently granular to distinguish individual collection items. In this sense, I think the problem with using LCSH with my EPrints collection is similar to the problem with LCSH in highly-specialized libraries. If I use EPrints again for this collection, I may create an alternate subjects taxonomy. A higher degree of granularity (distribution across more subjects) would make the subject browse feature more useful. Because my collection items are closely related in terms of subject, maintaining consistent subject terms has not been a problem. I used "uncontrolled keywords" to enhance the level of description. The uncontrolled keywords are not, however, by default browsable.

Monday, October 6, 2008

Drupal

I think Drupal should be a suitable content management system for my collection, but I probably need to work with Drupal more get the collection to where it needs to be. My collection consists of reports in PDF format which are best described with usual bibliographic categories (title, author, publisher, etc.) Primarily, I think the collection should be browsable by title, author, and keyword. I think Drupal should be able to do this, but I did not achieve this functionality within the time allotted.

Sunday, September 21, 2008

IRLS 675 Tech Assignments Pace

The pace of the IRLS 675 tech assignments are relatively slower than the pace of assignments from IRLS 671. Taking two courses requires much more work than taking one course. I would like to have more time to experiment with the technologies introduced in IRLS 675. A dear friend of mine who is very skilled in information technology was of the opinion that extended amounts of time (several weeks?) could be spent on many (any?) of the technologies surveyed in IRLS 675. I'm probably still adjusting to taking two classes, and will become more efficient with delivering the assignments as my method improves. I expect some of those taking only IRLS 675 (and not IRLS 671) might find the pace of IRLS 675 too slow. If so, maybe they could further delve into experimenting with the technologies. For me, I'd say I've got quite enough to do, and to keep the IRLS 675 tech assignments paced as they have been.

Saturday, September 13, 2008

Review of "LibData to LibCMS: One Library's Evolutionary Pathway to a Content Management System"

Review of "LibData to LibCMS: One Library's Evolutionary Pathway to a Content Management System," by Paul F. Bramscher and John T. Butler, in Library Hi Tech, Volume 24, Number 1, 2006.

In this article, Bramscher and Butler chronicle and analyze the evolution of the University of Minnesota Libraries (UML) website during the period 2003-2005. In 2003, the UML website was no longer comprised of static HTML pages, but based upon a data repository of 40 relational database tables. The software application used to manage this repository was LibData. The UML developed a succession of software applications ("authoring mechanisms") with which content could be created from the databases and published on the UML website. The first of these authoring mechanisms (Research QuickStart) allowed authors to utilize relational database functionalities, but provided limited options for content presentation. The second authoring mechanism (CourseLib and PageScribe) added course-support functionality, and style options through cascading style sheets. The third authoring mechanism (LibCMS) allowed for a greater variety of content to be published, including HTML, server-side scripts, XML, and RSS.

The authors refer to evidence in the literature that the failure rate of content management systems (CMS) is high. (Of course, the failure rate for information technology projects generally is high. See the Chaos Report of the Standish Group at http://www.projectsmart.co.uk/docs/chaos-report.pdf.) The authors caution against having unrealistic expectations for a CMS as a cure-all for a library's challenges or "slam-dunk transformative technology." Similar caution was echoed in John Blyberg's Drupal presentation at the Summer 2007 American Libraries Association Convention. The authors state that a CMS should be expected to facilitate the customization of web pages, the management of the uniformity of web pages, and the publishing of content (especially by staff without web developer expertise.)

The authors ingeniously formulate a triangular relationship diagram from the term "content management system." Using this formulation, they diagram tensions and dynamics inherent to the development and management of a CMS. Such tensions include those between content and systems (database normalization and certain cases of query efficiency,) and those between management and systems (the ecology of other systems, including operating systems, server upgrade cycles, and organizational security and authentication requirements.) Other challenges seem to fall outside the "content, management, system" triangle. Understanding the realities of the local social context (for instance, who performs what work) is important for the successful implementation of a CMS.

The authors detail the factors considered in determining whether to use an open-source CMS or to buy or outsource one. The factors considered will be familiar to those acquainted with the issue, and I will not detail them here. Generally, the arguments for open-source related to flexibility and control, and the arguments for buying or outsourcing related to convenience.

The UML CMS project began with the identification of requirements for the system, resulting in a list of less than ten core requirements. Core requirements included a versioning mechanism, a draft area, a simplified publishing routine, server-side and client side programming language support, breadcrumb navigation and sitemap, a self-indexing mechanism, secure sockets layer (SSL) protocol support, and ease in building application programming interfaces (APIs.) Weakness analysis was performed to identify deficiencies or non-desirable features of system models. The CMS the UML chose was LibCMS.

The UML made several design choices in the development of LibCMS. One design choice was to leave the root node of the system unused, which allowed for the system to be scalable to other libraries in the UML system. Another design choice reducing the occurrence of broken URLs was to disassociate URLs from file system structure. LibCMS incorporated server-side markup routines using PHP and Apache server to display links to navigation, associated pages, database resources, and other CMS. LibCMS was designed to convert all page source code to Base64, an encoding scheme which allowed the databases to process a great variety of characters, including foreign character sets, quotes, and double quotes. The description of the self-archiving feature of LibCMS was interesting. LibCMS opens an HTML socket to retrieve web pages it sends as they would appear to the clients. LibCMS then strips the pages of HTML tags and client-side scripts, compresses trailing and leading spaces, and inserts the results into a database field.

I think this article effectively detailed many issues involved with planning and developing a CMS in a large research library. The complexity and scope of issues involved might give me pause when inclined to criticize a library website. Many of the technical concepts presented were challenging but particularly interesting.

Monday, September 8, 2008

Assignment 5


Pete Williams et al. survey challenges technological innovation presents in archiving personal collections at the British Library in "Digital Lives: Report of Interviews with the Creators of Personal Digital Collections." The article can be found at http://www.ariadne.ac.uk/issue55/williams-et-al/. Challenges identified by Williams include the issue of how do digital collection managers identify and collect digital information content in a variety of formats and diverse locations. Digital content in personal collections includes documents, articles, digital images, audio recordings, web pages, blogs, and email. The content of blogs and web-based email present particular challenges, because such content is physically located on hosting servers, rather than on the content creators' personal computers.


Initially I had thought to base this assignment on a collection of digitized photographs. As such, the collection would have been an extension of the content I used in the IRLS 672 term project. However, I think the issues raised by Williams are intriguing. I would like to base this assignment using the Drupal content management system on a diverse collection of personal documents I have created in a range of activities over the years. Working with such a collection would be interesting to me because I am not readily aware of how such a collection will be described and organized, and because the variety of object formats might present interesting problems. I am particularly interested to see to what degree I can link disparate objects of the collection.


The digital objects include MS Word documents, PDF documents, image files, emails, and blogs. I think it will be especially interesting to see if I can integrate the emails into the collection, because the emails have a stack of files associated with them. The documents include a variety of subject matter, including professional activities, travel plans, maps, business transactions, and hobbies.


I think the question of who might access this collection might have less relevance than with other collections. Since this is a personal collection, I suppose only I would access it. If this is problematic, then I can use a collection of digitized photographs instead. But if the issue of who might access the collection is a deficiency, I think this deficiency is mitigated by interesting problems of incorporating digital objects of diverse format and subject matter into an integrated collection.


The diversity of subject matter will present interesting problems in the development of a taxonomy. I hope this assignment will present an opportunity to experiment with other methods of classification, such as tagging. I am tending to agree with Clay Shirky's argument in "Ontology is Overrated," that effective classification can organically grow from a population of diverse user tags.