Saturday, November 22, 2008

Pre-installed VM

I think building my own content management systems was a highly worthwhile series of exercises. I learned something about the technological components of each system I built. I think this knowledge provided some insight for how individual installations worked. By building a series of content management systems, I have begun to get an idea of features and aspects common to all of the systems. With this familiarity with some general outlines for how content management systems are installed, much of the mystery of the processes and technologies associated with installing a content management system has been dispelled. The repetitive aspects of building my own system were helpful. Repeating technical procedures (creating a virtual machine, downloading software, modifying files, setting permissions) developed for me a familiarity and comfort level with those procedures. I would not have such a level of familiarity or comfort with those technical procedures without practicing them repeatedly. Building a series of content management systems increased the opportunity to make errors and encounter difficulties in the build processes. The challenges these errors and difficulties presented were valuable learning experiences, particularly helping to develop analytical and problem solving skills and strategies when working with technology. Although these challenges were frequently frustrating, I think they were valuable learning experiences.

That said, it is also true that downloading pre-installed virtual machines would provide more time to concentrate on experimenting with the installed system. I think this is especially true for students taking IRLS 675 while concurrently taking another course (like, for me, this semester, IRLS 671). Generally this semester I was primarily focused upon the timely completion of all my coursework. Maybe this focus was partly because it had been some years since I had taken two concurrent college-level courses. While focused on completing my coursework in a timely manner, I found it challenging to "shift gears" to allocate time to experiment. An general issue here is that of a student learning activity (experimentation) which seems inherently difficult to incorporate into a formalized course structure. Incorporating experimentation would be worthwhile, because there would be much to learn with these systems. My idea has been to go back to these systems and get more deeply involved with exploring them at some time subsequent to the DigIn coursework. I think IRLS 675 did a good job of introducing the basics of various content management systems, which is what I think was the course objective. Pre-installed VM would provide more time to concentrate on the collections, but what would this achieve? It is hard conceive of a realistic learning objective beyond that of the level of introduction to the systems. For instance, it seems like an entire semester or more could be spent on Drupal. So I don't know what would be the demonstrable benefit of giving the students a marginal amount of time more to work with the installed systems, especially at the cost of reducing their learning experience with applied technologies. Moreover, without a good familiarity with technological routines associated with installing the content management systems, I think I would have a significant increase in difficulty with experimenting with these systems after the DigIn Program concludes.

I think it is relevant to consider what are the objectives for student learning in IRLS 675. What skill sets does DigIn want its graduates to have? I suppose this is a difficult question. Apparently there is a lack of clear consensus on what the competencies for digital collection managers should be. (Maybe an increased clarification of the competencies for digital collection managers can be achieved though an audit of skills/experience requirements specified in a representative sample of relevant position vacancy announcements.) The DigIn page says the compelling issue is that an "explosion of digital information and the growth of online digital resources has led to a shortage of individuals with an understanding of the disciplines of libraries, document management and archives who also have the technical knowledge and skills needed to create, manage and support digital information collections." If the DigIn Program is satisfied with this characterization of the compelling issue, then certainly making students proficient with the technological procedures of building content management systems is relevant to the "technical knowledge" described in this characterization.

Saturday, October 25, 2008

Service Providers

Service providers who specialize in a topic provide useful functionality limiting the service to a specialized topic. For example, the Avano marine and aquatic sciences service provider (http://www.ifremer.fr/avano) clearly conveys what are the topic strengths of the service (marine and aquatic sciences.) Avano has a very useful Browse Archives feature, which displays the collections Avano draws from, links to the collections, the number of items harvested from each collection, and a description of the collection. Avano data providers include the Alfred Wegener Institute for Polar and Marine Research, the Aquatic Commons, and the Open Marine Archive.

The Scientific Commons (http://www.scientificcommons.org/) is a service provider which harvests a wide variety of scientific topics. The Scientific Commons has harvested 23 million records from 9 million authors in 950 archives. Searching a topic produces very large numbers of results (somewhat Google-like.) The search results can be filtered through only a limited number of options (Year, Language, and Sort Mode.) I think large-scale service providers such as the Scientific Commons call to question whether such services providers scale well for topic searching. The data providers to Scientific Commons are unidentified.

PLEIADI: Portal for Italian Literature in Open and Institutional Archives (http://www.openarchives.it/pleiadi/index.php?sel_lang=english) incorporates a variety of functionality in a multi-featured portal. Features include personal accounts, recent news, a wiki, RSS, and related web resources. Searching PLEIADI, it appears that many of its data providers are Italian university repositories.

Generally, I think it is most important for a service provider to harvest relevant data providers. It seems to me that service providers specializing in a topic can be more useful (especially searching on topics) than large-scale comprehensive service providers. In addition to harvesting relevant data providers, I think it is important that service providers have powerful search interfaces that provide users with sufficient filtering options and other search options.

Monday, October 20, 2008

EPrints

I thought it might be a good idea to try using the Library of Congress Subject Headings to describe my collection (a collection of evaluation reports in PDF) in EPrints. The first two levels of Library of Congress Subject Headings were insufficiently granular to distinguish individual collection items. In this sense, I think the problem with using LCSH with my EPrints collection is similar to the problem with LCSH in highly-specialized libraries. If I use EPrints again for this collection, I may create an alternate subjects taxonomy. A higher degree of granularity (distribution across more subjects) would make the subject browse feature more useful. Because my collection items are closely related in terms of subject, maintaining consistent subject terms has not been a problem. I used "uncontrolled keywords" to enhance the level of description. The uncontrolled keywords are not, however, by default browsable.

Monday, October 6, 2008

Drupal

I think Drupal should be a suitable content management system for my collection, but I probably need to work with Drupal more get the collection to where it needs to be. My collection consists of reports in PDF format which are best described with usual bibliographic categories (title, author, publisher, etc.) Primarily, I think the collection should be browsable by title, author, and keyword. I think Drupal should be able to do this, but I did not achieve this functionality within the time allotted.

Sunday, September 21, 2008

IRLS 675 Tech Assignments Pace

The pace of the IRLS 675 tech assignments are relatively slower than the pace of assignments from IRLS 671. Taking two courses requires much more work than taking one course. I would like to have more time to experiment with the technologies introduced in IRLS 675. A dear friend of mine who is very skilled in information technology was of the opinion that extended amounts of time (several weeks?) could be spent on many (any?) of the technologies surveyed in IRLS 675. I'm probably still adjusting to taking two classes, and will become more efficient with delivering the assignments as my method improves. I expect some of those taking only IRLS 675 (and not IRLS 671) might find the pace of IRLS 675 too slow. If so, maybe they could further delve into experimenting with the technologies. For me, I'd say I've got quite enough to do, and to keep the IRLS 675 tech assignments paced as they have been.

Saturday, September 13, 2008

Review of "LibData to LibCMS: One Library's Evolutionary Pathway to a Content Management System"

Review of "LibData to LibCMS: One Library's Evolutionary Pathway to a Content Management System," by Paul F. Bramscher and John T. Butler, in Library Hi Tech, Volume 24, Number 1, 2006.

In this article, Bramscher and Butler chronicle and analyze the evolution of the University of Minnesota Libraries (UML) website during the period 2003-2005. In 2003, the UML website was no longer comprised of static HTML pages, but based upon a data repository of 40 relational database tables. The software application used to manage this repository was LibData. The UML developed a succession of software applications ("authoring mechanisms") with which content could be created from the databases and published on the UML website. The first of these authoring mechanisms (Research QuickStart) allowed authors to utilize relational database functionalities, but provided limited options for content presentation. The second authoring mechanism (CourseLib and PageScribe) added course-support functionality, and style options through cascading style sheets. The third authoring mechanism (LibCMS) allowed for a greater variety of content to be published, including HTML, server-side scripts, XML, and RSS.

The authors refer to evidence in the literature that the failure rate of content management systems (CMS) is high. (Of course, the failure rate for information technology projects generally is high. See the Chaos Report of the Standish Group at http://www.projectsmart.co.uk/docs/chaos-report.pdf.) The authors caution against having unrealistic expectations for a CMS as a cure-all for a library's challenges or "slam-dunk transformative technology." Similar caution was echoed in John Blyberg's Drupal presentation at the Summer 2007 American Libraries Association Convention. The authors state that a CMS should be expected to facilitate the customization of web pages, the management of the uniformity of web pages, and the publishing of content (especially by staff without web developer expertise.)

The authors ingeniously formulate a triangular relationship diagram from the term "content management system." Using this formulation, they diagram tensions and dynamics inherent to the development and management of a CMS. Such tensions include those between content and systems (database normalization and certain cases of query efficiency,) and those between management and systems (the ecology of other systems, including operating systems, server upgrade cycles, and organizational security and authentication requirements.) Other challenges seem to fall outside the "content, management, system" triangle. Understanding the realities of the local social context (for instance, who performs what work) is important for the successful implementation of a CMS.

The authors detail the factors considered in determining whether to use an open-source CMS or to buy or outsource one. The factors considered will be familiar to those acquainted with the issue, and I will not detail them here. Generally, the arguments for open-source related to flexibility and control, and the arguments for buying or outsourcing related to convenience.

The UML CMS project began with the identification of requirements for the system, resulting in a list of less than ten core requirements. Core requirements included a versioning mechanism, a draft area, a simplified publishing routine, server-side and client side programming language support, breadcrumb navigation and sitemap, a self-indexing mechanism, secure sockets layer (SSL) protocol support, and ease in building application programming interfaces (APIs.) Weakness analysis was performed to identify deficiencies or non-desirable features of system models. The CMS the UML chose was LibCMS.

The UML made several design choices in the development of LibCMS. One design choice was to leave the root node of the system unused, which allowed for the system to be scalable to other libraries in the UML system. Another design choice reducing the occurrence of broken URLs was to disassociate URLs from file system structure. LibCMS incorporated server-side markup routines using PHP and Apache server to display links to navigation, associated pages, database resources, and other CMS. LibCMS was designed to convert all page source code to Base64, an encoding scheme which allowed the databases to process a great variety of characters, including foreign character sets, quotes, and double quotes. The description of the self-archiving feature of LibCMS was interesting. LibCMS opens an HTML socket to retrieve web pages it sends as they would appear to the clients. LibCMS then strips the pages of HTML tags and client-side scripts, compresses trailing and leading spaces, and inserts the results into a database field.

I think this article effectively detailed many issues involved with planning and developing a CMS in a large research library. The complexity and scope of issues involved might give me pause when inclined to criticize a library website. Many of the technical concepts presented were challenging but particularly interesting.

Monday, September 8, 2008

Assignment 5


Pete Williams et al. survey challenges technological innovation presents in archiving personal collections at the British Library in "Digital Lives: Report of Interviews with the Creators of Personal Digital Collections." The article can be found at http://www.ariadne.ac.uk/issue55/williams-et-al/. Challenges identified by Williams include the issue of how do digital collection managers identify and collect digital information content in a variety of formats and diverse locations. Digital content in personal collections includes documents, articles, digital images, audio recordings, web pages, blogs, and email. The content of blogs and web-based email present particular challenges, because such content is physically located on hosting servers, rather than on the content creators' personal computers.


Initially I had thought to base this assignment on a collection of digitized photographs. As such, the collection would have been an extension of the content I used in the IRLS 672 term project. However, I think the issues raised by Williams are intriguing. I would like to base this assignment using the Drupal content management system on a diverse collection of personal documents I have created in a range of activities over the years. Working with such a collection would be interesting to me because I am not readily aware of how such a collection will be described and organized, and because the variety of object formats might present interesting problems. I am particularly interested to see to what degree I can link disparate objects of the collection.


The digital objects include MS Word documents, PDF documents, image files, emails, and blogs. I think it will be especially interesting to see if I can integrate the emails into the collection, because the emails have a stack of files associated with them. The documents include a variety of subject matter, including professional activities, travel plans, maps, business transactions, and hobbies.


I think the question of who might access this collection might have less relevance than with other collections. Since this is a personal collection, I suppose only I would access it. If this is problematic, then I can use a collection of digitized photographs instead. But if the issue of who might access the collection is a deficiency, I think this deficiency is mitigated by interesting problems of incorporating digital objects of diverse format and subject matter into an integrated collection.


The diversity of subject matter will present interesting problems in the development of a taxonomy. I hope this assignment will present an opportunity to experiment with other methods of classification, such as tagging. I am tending to agree with Clay Shirky's argument in "Ontology is Overrated," that effective classification can organically grow from a population of diverse user tags.

Wednesday, August 6, 2008

Where I was when the course started and how far I have come

Aside from factual knowledge and technical skills I have learned in IRLS 672, I would say that my understanding of how the component aspects of a typical digital technology implementation relate and fit together have greatly improved. When I started the IRLS 672 course, I had prior limited exposure to many of the concepts that would comprise the course content. I have watched libraries undergo a remarkable transformation in the last 15 years. I was aware of the importance of learning about aspects of this transformation, but it was difficult to determine where to begin. I have begun to develop a good foundation on which to build skills. The focus of my professional development has improved. As a result of my studies, I have gained heightened awareness of the exciting potential for work with digital collections.

When IRLS 672 started, I had read about the Linux operating system, and installed the Debian Linux operating system on a spare computer. I had limited familiarity with live-disk operating systems, such as Knoppix. My use of Linux was limited to graphical user interface (GUI) desktop applications. I had acquaintance with some concepts of networking from workplace exposure and reading. My knowledge of servers was limited to a conception of the role of a server in the server/client model. I had written static HTML pages. I had no knowledge of or experience with technology planning. I had a lot of experience using commercial, GUI-based database products. I had little exposure to project management concepts. I knew that PHP sometimes stood for the "P" in LAMP.

So I think I have come a long way since the beginning of the course. I have a good basic knowledge of using the command line interface (CLI) to navigate, configure, and administer Linux. I have a knowledge of basic networking. I have learned a lot about how HTML relates to XML and PHP. I have worked with web servers and database servers. I have learned something about technology planning and project management. I have learned a lot about how all these things fit together. Generally, I would say that I have developed a good foundation for working with digital technology applications. I can build on this foundation by continuing to learn about the topics introduced in this course.

Readings on Project Management

Personally I found the Project Management Body of Knowledge (PMBOK) Guide to be an interesting and thorough analysis of critical components of project management. The PMBOK Guide typically creates a matrix of nine knowledge areas (scope, time, cost, quality, human resource, communications, risk, procurement, and integration) and five process groups (initiating, planning, executing, controlling, and closing.) It is interesting how general principles such as these can apply to a wide range of projects. I think a good point was made in the readings regarding the challenge of choosing the correct level of detail to implement the PMBOK. I suspect it is this challenging choice which leads to under-implementation of PMBOK and similar principles in many library projects.

Scope management is important to avoid project escalation. I thought particularly interesting the example of project escalation documented in the Mark Keil article "Pulling the Plug: Software Project Management and the Problem of Project Escalation." Keil effectively documented how social and psychological factors can contribute to project escalation. I think such social and psychological factors frequently have significant influence on project outcomes. The Bas de Bar videos "Software Project Management in 15 Minutes" provided practical advice for managing such social and psychological factors. Bas de Bar asserts that project stakeholders are key to the success of a technology project. These stakeholders are primarily driven by tacit expectations of a project. It is the task of project management to divine and address these expectations through the definition of requirements. A key tool to define the scope of a project is the development of a work breakdown structure. I think the work breakdown structure would be a useful tool in many library projects.

It seems to me that much of project management deals with the allocation of limited resources. This resource allocation of limited resources is effectively illustrated by the Tradeoff Triangle diagram, as described in the Microsoft Solutions White Paper MSF Process Model V 3.1. The three aspects of the Tradeoff Triangle are resources, schedule, and features. These three aspects are assigned exclusively the respective qualities of fixed, chosen, and adjustable. So, for example, given fixed resources and a chosen schedule, a technology project product's features should be realistically projected to be adjustable. I think that understanding such realities can aid in the management of stakeholder expectations through the definition of realistic requirements.

Sunday, August 3, 2008

A MySQL concept I found difficult to understand


These two MySQL queries return the same value:



select photographer_lname, image_title

from photographer

right join image

on photographer.photographer_id=image.photographer_id;




select photographer_lname, image_title

from image

left join photographer

on photographer.photographer_id=image.photographer_id;



In the first query, the right join returned all the rows from the second table (image), even if there was no value from the first table (photographer.)

In the second query, the left join returned all the rows from the first table (image), even if there was no value from the second table (photographer.)

Left and right refer to the first and second tables. Left- first table, right- second table.


Below, this is the general example for all rows returned, from the table in CAPS:

select column1, column2

from TABLE1

left join table2


select column1, column2

from table1

right join TABLE2

Sunday, July 27, 2008

Databases

Relational databases seem to be a quite complex subject. The theory behind relational databases can be complex. Practical implementations of the theory are more simple. As a practical matter, many-to-many entity relationships are to be avoided. Interestingly, many-to-many entity relationships are not problematic theoretically. But relational database management systems (RDMS) are unable to implement such relationships. Many-to-many entity relationships are resolved with intersection entities. Database modeling involves conceptualizing the database structure, by identifying entities (relevant things), attributes (descriptive aspects of each thing), and relationships (associations between things.) Relationships generally have up to three aspects, a name, optionality, and degree. For me, I am still having difficulty understanding optionality, which is whether a relationship is optional for mandatory. For instance, two examples in this week's readings were: 1. an author can write 0 or more books, and 2. a plant may be given one or more waterings (optional.) I do not understand 1. how an author can be an author having written 0 books, and 2. how it is optional that a plant be given a watering. (Maybe the plant is outside?)

With respect to normalization, I followed the normalization procedure detailed in the Three Normal Forms tutorial. It was slow-going at first, but after a while I think I got the hang of it. However, starting with a database model of four entities, the normalization process proliferated their number to ten. I think I normalized the database correctly, but, for the purposes of the class assignment, I am using a modified schema with four entities. So I think there are different levels at which a database can be implemented, depending upon its purpose.

Sunday, July 20, 2008

Technology Planning

It seems often easy to forget that technology is an end to itself. Great are the demands of keeping a library technologically up-to-date. As libraries become increasingly technological, technophobia will reduce.

Technology plans can help as coordinating documents. Not infrequently such documents are created, filed, and forgotten. Staff buy-in correlates with whether or not the technology plan is a "living document." (Stephens, 2004)

Indeed, libraries need to know their communities. In practice, libraries now spend an increasing amount of time studying their "competitors" (Amazon, Google, etc.) I think this is because the latter is now changing faster than the former. (Sager, 1999)

Government library environmental scanning can be a complex task. Formal explicit identification of user needs can be difficult. (Sager, 1999)

The Schools and Libraries Division of the Universal Service Fund (E-Rate) provides discounted computing and networking resources for schools and libraries. E-Rate was authorized under the Telecommunications Act of 1996. Technology planning can improve the efficiencies of expenditures, competitive bidding, auditing processes, and procurements. (Bertot, 2002)

The 1995 Standish Group Study underscores the challenges to technology project, and puts those of E-Rate into perspective. "Poor project planning" was identified as the primary cause of project failure. "Poor project planning" specifically entailed that "risks were not addressed" or the "project plan was weak." The extent which technology plans address technology project risks is unclear. Such risk include "slippage from the schedule, "changes in the scope of technology, functionality, or business case," cost overruns, and changes in key individuals including managers or sponsors. (Whittaker, 2007)

Technology plans should acknowledge realities referenced in the 2003 OCLC Environmental Scan. These realities include that users are usually satisfied with information search results from open-web search engines, libraries can expect increasing competition for funding, and (most surprisingly) libraries are still spending a relatively small part of their budgets (3%) on digital content. (OCLC, 2003)

Sunday, July 13, 2008

XML and the demo server

I went about learning XML using the W3Schools XML tutorial at http://www.w3schools.com/xml/default.asp, which covered the basics of XML, including the use of XML, the XML tree structure, XML syntax, XML elements, XML validation, viewing XML, and more.

I also went through the Document Type Definition (DTD) tutorial, which describes the current de facto schema standard for XML documents. The tutorial included DTD elements, attributes, entities, validation, and examples.

I also went through the XML Schema tutorial. XML Schema is an XML-based alternative do DTDs.

I also viewed Mark Long's CBT tutorials XML Basics and XML Documents. I think they we useful reinforcement to the W3Schools tutorials.

Overall I think the lessons on XML were pretty straightforward. As I use XML, I will become more proficient.

---

This week I accessed the demo server under SSH protocol with the client utility Putty.

I transferred the server to headless mode, first testing for keyboard or mouse boot hang-ups by unplugging the keyboard and mouse and restarting the server. Although it returned an error message, it did reboot. So I took off the monitor too.

I found the server's web site in var/www, and examined the default website configuration location from the documentroot directive in /etc/apache2/sites-available/default.

I created a group webdev for developing the website, added my username to webdev, changed group ownership of /www from root to webdev, changed /www permissions to 775, changed /www subdirectories to 775, and changed /www files to 664. Then I used the file transfer utility WinSCP to transfer files from my UA website to /var/www on my demo server. I received file sharing permissions errors while transferring the files, but the files did transfer.

I also create a user web space enable through the Apache module Userdir. I created a directory public_html and an HTML file index.html. I enabled the Userdir module with the command a2enmod userdir.

Besides the demo server, there is an issue I want to look into about the VM server. I want to know why I cannot ping from the VM server to other machines.

Thursday, July 3, 2008

How I Went About Learning HTML, and a Brief Report on the Installation of My Demo System

I went about learning HTML from the W3Schools website at http://www.w3schools.com/html/default.asp. I found it very helpful as a self-paced tutorial, because it provided examples of each HTML element, and contained a live interactive feature which displayed the effects of changing the code. The layout of the W3Schools site is very efficient, showing all relevant course content listed in a column on the left side of the page. This layout enabled me to easily find the HTML code I might be looking for. The W3Schools site does a good job of incorporating necessary advertising into the website in a way that is non-distracting. Mercifully, the website lacks flashing advertisements. Each HTML element receives its own page, consistent in format with the other pages. HTML elements appearing on pages other than their own are liked to their own page for efficient reference. I reviewed the HTML Basics and Advanced Sections, and the Cascading Style Sheets Section. There is a lot of good stuff there, worth going back to.

The installation of the demo server progressed without incident. I made a selection for manually identifying the keyboard to Linux, which was different from the class installation procedure. After answering a few questions to Linux about my keyboard, I and had Linux automatically identify the keyboard. I installed LAMP and SSH, edited etc/apt/sources.list, updated the system with aptitude, and installed WebAdmin. I configured static IP addresses on the virtual machine server and the "actual" (?) machine server, and successfully pinged them. Setting up name resolution in my hosts file, I searched "system32" in Windows XP Search, located the file \windows\system32\drivers\etc\hosts, and made the appropriate edits. I noted numerous addresses of malicious websites associated to the local host address by the anti-malware tool StopZilla, rendering those malicious websites ineffective. Good lookin' out, StopZilla.

Sunday, June 29, 2008

Reflections on the Variety of Presentation of the Unit 6 Material

Two aspects of the variety of formats of presentation of the Unit 6 material which struck me as interesting had to do with the podcasts.

The first aspect had to do with my expectations about the accessibility of content with respect to a delivery format. Prior to viewing the Harvard Computer Science Podcast, my expectation of the podcast format was that of exclusively audio content delivery. Given this expectation, I was skeptical that the information content could be effectively conveyed to me in audio-only format. I was relieved to discover my notions of podcasting to be dated, and thought the actual audio-video presentation of the Harvard Computer Science Podcasts to be highly effective. However, I began to wonder why this should be the case. Because overall, the video component of the podcast did not substantively contribute additional quantitative information absent from the audio. Much of the video content was an image of a professor delivering a lecture. The same amount quantifiable content would have been delivered if the format were solely audio, but I'm sure I would not have judged it to be an effective presentation. My guess is the distinction speaks to a certain vague subjective factor in effective content delivery. I think much of this has to do with content user expectations. I expect to see a lecture if I am listening to a lecture. Watching the video of the lecture added something to the audio information which I think was mainly of a subjective quality. But the addition, at least for me, was very important. But on the other hand, maybe I'm making too much of this. I do listen to radio broadcasts of baseball games.

The second aspect had to do with the RSS component of podcast format, and the effect of this technology on my understanding of the models of networking architecture. Specifically, what is the designation of an architectural component which, to me, is neither precisely a client nor a server? The content of the Harvard Computer Science Podcasts is hosted by the content creators, presumably by servers of Harvard. But users request this content not from Harvard servers, but from iTunes. So my question is, what is the name for functional role of iTunes in this instance? Maybe I am making too much of this, but I think what it adds to the simple client-server model is interesting. This is the second time I have had the impression that RSS creates interesting effects with networking architecture. At the Computers in Libraries Conference in 2007, I listened to a presentation by a librarian at the Rand Corporation. He was using RSS in an innovative way, providing a greater degree of freedom for Rand staff to edit web pages on the Rand LAN. According to him, his utilization of RSS allowed Rand staff to edit web pages without messing up the web page architecture.

I would say the Unit 6 material is relatively "deep," and the variety of presentations of the same repeated content was helpful. For instance, I found the Warriors of the Net animation to be very helpful for me to conceptualize TCP/IP. The same material, for me, was effectively reinforced in detail in Nemeth's Linux Administration Handbook. I found the Wikipedia articles on local area networks to be very helpful gaining a conceptual understanding of the applied assignment.

Friday, June 20, 2008

Installation Routine Experiences

Installation Routine Experiences

I completed the assigned installation routines without significant issues. I built a server GUI by installing the X window system and the IceWM window manager, and verified installation of the server GUI via http://localhost/.

I installed Ubuntu Desktop on a virtual machine (VM) free of any major issues. I think I am going to move the .iso files of my Ubuntu operating systems from my machine's C Drive to its D Drive, because the C Drive is getting full. While on the subject of drives, I need to read about installing VMware virtual machines on a portable flash drive. If the result would mean being able to carry around a portable Linux OS, then such a thing would be useful.

I utilized wget and dpkg to install the Webmin admin GUI. Notable are the large variety of available modules (http://www.webmin.com/standard.html). I ran ifconfig to find the virtual server's IP address on the second line of eth0. I might be making a mistake, but according to my notes the server's IP address changed from a previous day. Hmm. I connected to the Webmin GUI on the IceWM window manager on the X window system through Port 10000 from my host machine browser via https://localhost:10000.

Generally all these routines completed without major issues. My big issue for this week was the installation of vim, which kept me busy Monday and Tuesday. The issue took many hours to resolve, and is detailed in the Unit 5 Activities section.

Saturday, June 14, 2008

The New York Times: Charging by the Byte to Curb Internet Traffic

The New York Times
Printer Friendly Format Sponsored By

June 15, 2008
Charging by the Byte to Curb Internet Traffic
By BRIAN STELTER

Some people use the Internet simply to check e-mail and look up phone numbers. Others are online all day, downloading big video and music files.

For years, both kinds of Web surfers have paid the same price for access. But now three of the country’s largest Internet service providers are threatening to clamp down on their most active subscribers by placing monthly limits on their online activity.

One of them, Time Warner Cable, began a trial of “Internet metering” in one Texas city early this month, asking customers to select a monthly plan and pay surcharges when they exceed their bandwidth limit. The idea is that people who use the network more heavily should pay more, the way they do for water, electricity, or, in many cases, cellphone minutes.

That same week, Comcast said that it would expand on a strategy it uses to manage Internet traffic: slowing down the connections of the heaviest users, so-called bandwidth hogs, at peak times.

AT&T also said Thursday that limits on heavy use were inevitable and that it was considering pricing based on data volume. “Based on current trends, total bandwidth in the AT&T network will increase by four times over the next three years,” the company said in a statement.

All three companies say that placing caps on broadband use will ensure fair access for all users.

Internet metering is a throwback to the days of dial-up service, but at a time when video and interactive games are becoming popular, the experiments could have huge implications for the future of the Web.

Millions of people are moving online to watch movies and television shows, play multiplayer video games and talk over videoconference with family and friends. And media companies are trying to get people to spend more time online: the Disneys and NBCs of the world keep adding television shows and movies to their Web sites, giving consumers convenient entertainment that soaks up a lot of bandwidth.

Moreover, companies with physical storefronts, like Blockbuster, are moving toward digital delivery of entertainment. And new distributors of online content — think YouTube — are relying on an open data spigot to make their business plans work.

Critics of the bandwidth limits say that metering and capping network use could hold back the inevitable convergence of television, computers and the Internet.

The Internet “is how we deliver our shows,” said Jim Louderback, chief executive of Revision3, a three-year-old media company that runs what it calls a television network on the Web. “If all of a sudden our viewers are worried about some sort of a broadband cap, they may think twice about downloading or watching our shows.”

Even if the caps are far above the average users’ consumption, their mere existence could cause users to reduce their time online. Just ask people who carefully monitor their monthly allotments of cellphone minutes and text messages.

“As soon as you put serious uncertainty as to cost on the table, people’s feeling of freedom to predict cost dries up and so does innovation and trying new applications,” Vint Cerf, the chief Internet evangelist for Google who is often called the “father of the Internet,” said in an e-mail message.

But the companies imposing the caps say that their actions are only fair. People who use more network capacity should pay more, Time Warner argues. And Comcast says that people who use too much — like those who engage in file-sharing — should be forced to slow down.

Time Warner also frames the issue in financial terms: the broadband infrastructure needs to be improved, it says, and maybe metering could pay for the upgrades. So far its trial is limited to new subscribers in Beaumont, Tex., a city of roughly 110,000.

In that trial, new customers can buy plans with a 5-gigabyte cap, a 20-gigabyte cap or a 40-gigabyte cap. Prices for those plans range from $30 to $50. Above the cap, customers pay $1 a gigabyte. Plans with higher caps come with faster service.

“Average customers are way below the caps,” said Kevin Leddy, executive vice president for advanced technology at Time Warner Cable. “These caps give them years’ worth of growth before they’d ever pay any surcharges.”

Casual Internet users who merely send e-mail messages, check movie times and read the news are not likely to exceed the caps. But people who watch television shows on Hulu.com, rent movies on iTunes or play the multiplayer game Halo on Xbox may start to exceed the limits — and millions of people are already doing those things.

Streaming an hour of video on Hulu, which shows programs like “Saturday Night Live,” “Family Guy” and “The Daily Show With Jon Stewart,” consumes about 200 megabytes, or one-fifth of a gigabyte. A higher-quality hour of the same content bought through Apple’s iTunes store can use about 500 megabytes, or half a gigabyte.

A high-definition episode of “Survivor” on CBS.com can use up to a gigabyte, and a DVD-quality movie through Netflix’s new online service can eat up about five gigabytes. One Netflix download alone, in fact, could bring a user to the limit on the cheapest plan in Time Warner’s trial in Beaumont.

Even services like Skype and Vonage that use the Internet to transmit phone calls could help put users over the monthly limits.

Time Warner would not reveal how many gigabytes an average customer uses, saying only that 95 percent of customers use under 40 gigabytes each in a month.

That means that 5 percent of customers use more than 50 percent of the network’s overall capacity, the company said, and many of those people are assumed to be sharing copyrighted video and music files illegally.

The Time Warner plan has the potential to bring Internet use full circle, back to the days when pay-as-you-go pricing held back the Web’s popularity. In the early days of dial-up access, America Online and other providers offered tiered pricing, in part because audio and video were barely viable online. Consumers feared going over their allotted time and bristled at the idea that access to cyberspace was billed by the hour.

In 1996, when AOL started offering unlimited access plans, Internet use took off and the online world started moving to the center of people’s daily lives. Today most Internet packages provide a seemingly unlimited amount of capacity, at least from the consumer’s perspective.

But like water and electricity, even digital resources are finite. Last year Comcast disclosed that it was temporarily turning off the connections of customers who used file-sharing services like BitTorrent, arguing that they were slowing things down for everyone else. The people who got cut off complained and asked how much broadband use was too much; the company did not have a ready answer.

Thus, like Time Warner, Comcast is considering a form of Internet metering that would apply to all online activity.

The goal, says Mitch Bowling, a senior vice president at Comcast, is “ensuring that a small number of users don’t impact the experience for everyone else.”

Last year Comcast was sued when it was disclosed that the company had singled out BitTorrent users.

In February, Comcast departed from that approach and started collaborating with the company that runs BitTorrent. Now it has shifted to what it calls a “platform agnostic” approach to managing its network, meaning that it slows down the connection of any customer who uses too much bandwidth at congested times.

Mr. Bowling said that “typical Internet usage” would not be affected. But on the Internet, “typical” use is constantly being redefined.

“The definitions of low and high usage today are meaningless, because the Internet’s going to grow, and nothing’s going to stop that,” said Eric Klinker, the chief technology officer of BitTorrent.

As the technology company Cisco put it in a recent report, “today’s ‘bandwidth hog’ is tomorrow’s average user.”

One result of these experiments is a tug-of-war between the Internet providers and media companies, which are monitoring the Time Warner experiment with trepidation.

“We hate it,” said a senior executive at a major media company, who requested anonymity because his company, like all broadcasters, must play nice with the same cable operators that are imposing the limits. Now that some television shows are viewed millions of times online, the executive said, any impediment would hurt the advertising model for online video streaming.

Mr. Leddy of Time Warner said that the media companies’ fears were overblown. If the company were to try to stop Web video, “we would not succeed,” he said. “We know how much capacity they’re going to need in the future, and we know what it’s going to cost. And today’s business model doesn’t pay for it very well.”

My Experiences with the Text Editors and the Hands-on Assignments

The appearance of the new VMware Workstation 8.04 "feature," the Recovery Menu, suggested for me an investigation into the differences between shutdown and halt. (M)an shutdown says shutdown arranges for the system to be brought down in a safe way. Users on the system are notified. (S)hutdown options include shutdown -P, which requests that the system be powered off after it has been brought down. (M)an halt says when halt is called without options, it simply invokes shutdown. Apparently some halt options cause the system to stop more abruptly than with shutdown. This is underscored by the following post on a Unix (Solaris) forum, found by googling "difference between shutdown and halt":


The shutdown command runs various shutdown scripts and then finally invokes the halt command. Not running those scripts could be a problem depending on what they do. As one example, if you are running an Oracle database, you could lose transactions or even mangle the database....


------------------------------------


Wanting to more frequently use Vim, I downloaded a MS DOS version at http://www.vim.org/download.php. The download was a .zip file. I used a Windows utility to download and unzip the file. (I did not figure out how to download and unzip in DOS. Sorry, maybe next time.) File downloaded and unzipped, I went to the DOS command line and found it, in directory of downloaded files. I did not want the file in a download directory, but in the program files directory. So I tried to moved it to my program files directory with the DOS command

move c:\Downloads\vim71w32\*.* c:\Program_Files

But I got this message:

The process cannot access the file because it is being used by another process.

0 file(s) moved.

What was the other process that would be using the Vim file, preventing me from moving it? I decided at that point to go the easy route and figure out the issue another day. I opened the download folder window, opened the c: window, and dragged and dropped the Vim folder icon from the former into the latter. The DOS version of Vim resembles that of Linux. The help feature does not work. Maybe I will fix it. Anyway, I now have a version of Vim in Windows, which I have the idea of using for typical word processing tasks, such as producing this blog. We'll see how it goes.

Saturday, June 7, 2008

Ubuntu Live CD Experiences

Not long ago I burned a Debian image file on CD. If I recall correctly, at that time I downloaded the image file in binary mode. So initially I tried downloading the Ubuntu file in binary mode. This did not seem to go well. The file was downloading very slowly, and it occurred to me the generally user-friendly instructions did not prompt me switch to binary mode. So I cancelled that process, followed the instructions downloading the file in text mode and all went well. At first the Ubuntu Desktop CD would not boot. So I went into the BIOS and changed the boot order, and that did the trick.

Initially I went through the tutorials in the order they were presented. The CBT videos were a good introduction to the Linux file systems, especially because I could see demonstrations of the commands. The sections of Learning the Shell which list the contents of the Linux directories was useful as a reference to which I frequently returned. I frequently referred to the 15 commands in An Introduction to the Linux Command Line Interface. But I would include the halt command, since I needed it properly shut down the virtual server. Conveniently, being a virtual server, it was running within Windows, so I was able exit the Linux window to look up the halt command.

Generally the topics covered made sense to me. I expect it will be a process of increasing familiarity with them. Typical directories in a Linux file system and basic commands were introduced. I expect I will be referring back to the Learning the Shell tutorial pages about the various directories until I gain a greater familiarity. There are other details I need to review in the not-to-distant future, such as octal numbers for assigning permissions, details of find and grep This should be just a matter of consulting the man pages. Learning the Vi editor looks like it is going to be somewhat of a project. Vi is on the virtual server, so I can practice using Vi while still running Windows. Vi Help is not on the Ubuntu CDs. But it is on the Web at http://vimdoc.sourceforge.net/htmldoc/help.html#help.txt.

Sunday, June 1, 2008

Notes on the Linux File System, Disk Drives and Device Nodes

Below are my notes from Sections 2 and 3 of the Introduction to Linux tutorial. The sections are titled "A Look Around the File System" and "Disk Drives and Device Nodes". Also included are some notes from Section 1, the introductory chapter.

I find these notes useful as identification of concepts covered. I'm posting them in case someone besides myself should find them useful.


www.distrowatch.com
www.ibiblio.org/pub/Linux/distributions
www.linuxhq.com
www.AnchorPointBooks.com/etc/linux.html

contents of linux file systems
cygwin- cygnus? - X terminal emulator
ls- list current directory
bin- executable files
linked files end in @
boot- operating system files
dev- device files
etc- system configuration files
home- login home directories
lib- shared libraries
lost+found- used if the system crashes, hold fragments of files
mnt- mount
opt- optional programs
proc- holds running operating system files
root- home directory of the super user
sbin- commands and utilities used by the super user
swap- swap programs in and out of memory as they are running
temp- temp files
usr- catchall, all kinds of stuff
var- datafiles which change when files are running

etc- system configuration files
ls | more - runs list files in current directory and pipes results to more
more, return- scrolls one line at a time
more, space bar - scroll one page at a time
cat- writes a file to the screen

passwd
user name:password:user ID number:group ID number:full name:home directory:name of the program run upon login
-password is in the shadow file, used to be encyrpted in passwd
-usually do not edit passwd by hand

group and shadow files
in /etc
su -superuser
shadow
user name:encrypted password:date(since 1970):restrictions on changing password:number of days to change password:reminder time:number of grace days:date account became disabled:empty field
group
group name:group password:group ID number:can be used to list users in the group

file permissions
every file has an owner, is a member of a group, a type, a set of permissions (read, write, execute),
type: file, directory, symbolic link, device nodes, others too
t rwx rwx rwx - type of file, owner, group, others
ls -l permissions of files in a directory
number of links to the file

change ownership and permissions
touch -if file exists is bring the modification date stamp up to date on a file, or if the file does not exist it creates an empty file
umask -default permission settings, in the form of three octal digits

chmod -change permission settings for a file
chmod 777 myfile give read, write, excecute perm to everyone
or do it without using octals:
chmod O+r myfile give read permission to others
chmod gu+rw myfile give read, write permissions to group and owner (u is owner, (user))
chmod g-w myfile remove write permission from group
chmod a+rW read, write permissions to all

change group
done as superuser
chgrp rpc myfile -change myfile to the group rpc
change owner
chown root myfile -change owner of myfile to root

partition the disk, create filesystems, mount the filesystems

Linux file system
each Linux partition:
boot block (creates a temporary device node so the system can start running)
super block (filesystem information)
inode list (list of disk locations with file locations and sizes)
directory is a file with a list of entries with filenames and index into the inode table
data blocks

device nodes in /dev
device node is also called a device special file
drives are addressed through their device nodes
IDE are hda, hdb, hdc, etc.
hda with 3 partitions - hda1, hda2, hda3
SCSI are sda, etc.
flobby is fd0
ls -l /dev/hda -examine the device node
b - block device like a disk drive, c - a characher device like a terminal (tty device)

partitioning/fdisk
fdisk /dev/fd0 -to partition a floppy
m - menu, follow the menu instructions
ID number -a hex value indicating the file system for the partition (fdisk only makes partitions, not filesystems. The operating systems make the filesystems.)
w -saves the changes

extended and swap partitions
more than 4 partitions, need an extended partition
boot partition -first parition
extended part - second partion
then use logical partitions within the extended partition
a -flags the bootable partition

creating filesystems
operating systems can use each other's file systems but each must create it own filesystem
mkfs /dev/fdo -create a linux filesystem on a floppy disk
creates a new and empty superblock and inode table
linux filesystem is ext (ext ention of the Minux filesystem)
ext 3 is the latest version (with journaling, which helps in system restoration after crashes)
mkfs -t ext3 /dev/hdb1 if ext2 is the default, this overrides it, selecting ext3 for the first partition of the second IDE drive
mkfs -c -t ext3 /dev/hdb1 check for bad sectors
mkfs -t msdos /dev/fdo create a dos floppy
after creating the filesystems, attach them to Linux (mount the filesystem)

mounting filesystems
attaching the filesystem to Linux
create a directory and use the directory as a mount point
for a floppy disk:
mkfs /dev/fdo
mkdir floppy
ls -ld floppy see information for the directory "floppy"
unmounted, the directory is empty
the mount command attaches the floppy disk to the directory
mount /dev/fd0 floppy
now there is a file, lost+found, in the directory "floppy"
umount unattaches the directory "floppy" from the floppy disk, now the directory "floppy" is empty again
/mnt directory with mount points
you can set relationships between the directories and the device nodes in a filesystem table

/ect contains tables of filesystems to be mounted
fstab filesystems table
the filesystems table
device node, directory onto which the disk is to be mounted, filesystem type
iso9660 cd roms
/dev/pts psuedo terminal, a new window
can mount on machine's filesystem onto another, so another machine can be address as if it were local
last two columns of single digits
first digit backup utility (dump)
second digit order of check for integrity (fsck)
eject unmounts and ejects cdrom
mount -a mounts everything in the the fs table

manuals
man more options for the more command
q exits the man page

www.tldp.org Linux documentation project- good place for help