Tuesday, November 15, 2005

Learning cool new things

One of the things I love most about my job as a librarian is the enormous variety of content I get to work with. By partnering with content specialists for most if not all of our digital library projects, I get introduced to research areas I previously didn't know much about: the rise and fall of the "company town," victorian literature, "the commons," etc. One of the more recent topics is The Chymistry of Isaac Newton, a project in which I'm only tangentially involved. But this is really cool stuff. You can learn more about it too in a Nova episode titled Newton's Dark Secrets, premiering tonight in most areas.

Saturday, November 05, 2005

True folksonomic thesauri?

I saw on the Simile mailing list recently a thread discussing possible uses of relationships between user-supplied tags in an information system. This idea is intriguing to me. I've long believed we don't use the relationships recorded in our library-land controlled vocabularies in our systems for end-users to anywhere near their potential. A digital library collection I've been involved with demonstrates ways in which we might use these relationships. The methodology used is documented in this paper.

Yet I'd never thought about relationships for folksonomic vocabularies before. I think it's a fantastic idea, however. The same strategies for improving end-user discovery based on term relationships can be used no matter where these relationships come from. Relationships determined by methods such as this could be used in the same way human-generated relationships in a formal thesaurus could be used. I wonder if these relationships might be even more important in a folksonomic environment, as a method by which the vocabulary control us library folk hold so dear could be achieved.

Wednesday, October 26, 2005

Hierarchical catalog records

The October issue of D-Lib Magazine has an article describing an experimental FRBR implementation in the Perseus Digital Library, by David Mimno, Gregory Crane, and Alison Jones, entitled Hierarchical Catalog Records. I'm thrilled to see reports of experiments like this be shared in outlets as widely read as D-Lib. I'm also extremely happy to see this particular experiment happen outside of the MARC environment. I've been becoming more and more convinced as some experiments we're conducting with MARC records for sound recordings progress that FRBRization really is a revolution in resource description. Statistics abound estimating the small percentage of works that exist in more than one expression, and expressions that exist in more than one manifestation. While I don't doubt these numbers at all, I think they way in which they're presented minimizes the enormity of the task ahead of us to reach a true FRBR environment.

I believe the same is true of efforts to use MARC for FRBRized records. The MARC format could be adopted for this purpose. But is it in our best interests to do so? Using MARC makes the task seem less scary, that it won't be that difficult. But it is a difficult task, and we're fooling ourselves if we pretend otherwise. I wonder if we aren't better off addressing the issue head-on, admitting to a change with a new base record format. The change would be one of mind-set, rather than functionality.

I've mentioned I believe the FRBRization task is difficult. I don't believe difficult means impossible in this case, however. We don't yet have a good sense of the cost associated with such a conversion, so any claim to its value will be tempered by that uncertainty. But I am convinced of that value, and I believe studies like that of the Perseus Digital Library are vital in demonstrating it. No cost can be justified without first understanding the associated benefit. We have a great deal more work to do to reach that understanding.

Sunday, October 23, 2005

You know you go to too many conferences when...

...you look down halfway through the morning session and realize you're wearing the name tag from the wrong conference. Sigh.

Saturday, October 22, 2005

Separating data entry from data structure

I believe we've fallen extremely short in at least one area of potential for improving our cataloging and metadata creation systems -- user interfaces. We're still stuck in a mindset developed in the early days of the MARC format, whereby data is entered in the exact form in which it needs to be stored. When Web-based OPACs and cataloging modules emerged, cursory attempts to "improve" the interface appeared, but the changes were almost exclusively surface changes (labeling, etc.), and not implemented with community involvement.

But of course current technology provides many possibilities for a design layer in between the data entry interface and the data storage format. Metadata creation by humans is expensive. We need to do everything we can to design data entry interfaces that speed this process along, that help the cataloger to create high-quality data quickly. Visual cues, tab completion, and keyboard shortcuts are just a few simple tricks that could help. More fundamental approaches like automatic inclusion of boilerplate text and integration of controlled vocabularies could provide enormous strides forward.

Yet with all of this potential, I frequently (WAAAAAY too frequently) have conversations with librarians where it becomes clear they're focused exclusively on the data output format. It never even occurred to them that a system could do something with entered data that doesn't require cataloger involvement. (Man, I knew we librarians were control freaks, but this really takes the cake.) Of course, librarians aren't on the whole system designers. That's OK. But all librarians still need to be able to think creatively about possibilities. I'm convinced that the way forward here is to take the initiative to develop systems that demonstrate this potential, that show everyone what is possible with today's technology. Everyone has vision, yet that vision always has limits. By demonstrating explicitly a few steps forward from where we are, vision can then expand that much further.

Sunday, October 09, 2005

Museums and user-contributed metadata

It's funny how often, once one starts thinking about a subject, one finds examples of it absolutely everywhere. I've been thinking about user-contributed metadata for a while now in the context of a digital music library project, where we can provide innovative types of searching, if only we could find a way to make the creation of the robust metadata that drives it cost-effective. I wrote about this topic recently, inspired by OCLC's Wiki WorldCat pilot service.

So imagine my pleasure when, catching up on my reading this weekend, I came across "Social Terminology Enhancement through Vernacular Engagement" by David Bearman and Jennifer Trant in September's D-Lib Magazine. (Yes, I do know it's no longer September. Thanks for asking.) I'm thrilled to hear about this initiative, especially how well-developed it seems to be. I haven't yet followed the citations in the article to read any of the project documentation, but it certainly looks extensive. In the digital library (and museum!) world, I firmly believe ongoing documentation such as this associated with a project can be of as much or even more value than formally-published reports.

Two features strike me about the "Steve" system described here, that make it clear to me there are many ways to implement systems collecting metadata from users. It also makes me realize these decisions need to be made at the very beginning of a project, as they drive all other implementation decisions. The first is an assumption that the user interacting with the system is charged with the task of description rather than simply reacting to something they see and perceive either as an error or an omission. The user is interacting with the system for the purpose of contributing metadata; finding resources relevant to an information need is not the point. I suppose different users end up contributing with this model than with one that allows users to comment casually on resources they find in the course of doing other work. Different users might affect the "authoritativeness" of the metadata being contributed, but I wonder to what degree.

The second feature I find notable is that the system is designed to be folksonomic; there is no attempt at vocabulary control. Us library folk tend to start from the assumption that controlled vocabulary is better than uncontrolled and move on from there. At first glance, some of the reports from this project seem to resist that assumption, and start from the beginning looking for a real comparison. I'm anxious to read on.

Thursday, October 06, 2005

User-contributed metadata

OCLC recently announced the Wiki WorldCat pilot service. What a fantastic idea! Too bad I'm having trouble trying it out. I looked at a few books in Open WorldCat (via Google), including this one that I read recently and the book shown on the Wiki WorldCat page (The Da Vinci Code), and I didn't see the reviews tab or the links to add a table of contents or a note shown on the project page. Hmm. I wonder what I'm missing.

But, anyhoo... incorporating user-contributed metadata into library systems is something I've been thinking about for a while. Librarians tend to be pretty wedded to the notion of authority, that as curators of knowledge we're the best qualified folks out there to perform the documentation of bibliographic information. Assuming for a moment that this is true for some data elements, there are still several classes of data that could easily benefit from end-user involvement.

The first is detailed information from specialized domains. I work on a number of projects related to music. Information such as exactly which members of a jazz combo play on any given piece on a CD or the date of composition of a relatively obscure work is the sort of thing our catalogs could be providing to serve as research systems instead of just finding systems. But this sort of metadata is expensive to create; it requires research and domain expertise on the part of the cataloger. Many of our users, however, do have this specialized knowledge and love to share it.

Other information that might be appropriate for supplying by end-users could be tables of contents, instrumentation of a musical work, language of a text, and others of this type of "objective" information. Before you say, "But what about standard terminology, spelling, capitalization?!?" in a panicked voice, consider basic interface capabilities in 21st-century systems such as picking values from provided lists rather than typing them in.

But should we restrict ourselves to these more obvious of elements? I've been hoping for some time to be able to test various degrees of vetting of user-contributed metadata to a digital library system. I have in mind a completely open Wiki-type system, one that simply sends a suggestion to a cataloger, and a number of options in between. I suspect the quality of the user-contributed metadata will be overall much higher than critics assume. Yet even if it isn't, what sort of trade-off between quality and quantity are we willing to make? Traditional cataloging operations don't have extensive quality control operations, perhaps because QC is expensive work. And catalogers make mistakes, every day, just like the rest of us. Assuming a system where users can correct errors, how quickly will errors (made by a cataloger or by another end-user) be found and corrected? Will the "correct" data win out in the end? Surely these issues are worth a serious look.

Tuesday, September 27, 2005

The more things change...

I've just finished reading the short volume The MARC music format : from inception to publication / by Donald Seibert. MLA technical report, no. 13. Philadelphia : Music Library Association, 1982. The book is an account of the decision-making process involved in designing and implementing the MARC music format. I was both heartened and discouraged to read arguments in support of implementing MARC that mirror closely arguments I and others make today for moving beyond MARC.

The rationale behind the MARC music format reads full of hope, for improved access for users and higher quality data. Yet many of the improvements mentioned have not come to fruition. I'm heartened to see the vision represented here for the type of access we can and should be providing. Yet I'm discouraged to see more evidence that we haven't achieved this level of access in the time since the MARC format was implemented. I believe this serves to remind us that many factors other than database structure contribute to the success of a library system.

I also learned a valuable lesson reading this text that ideas and potential alone are not enough to convince everyone that any given change is a good idea. A large percentage of librarians out there have heard these very arguments before and seen them not pan out. I do believe, however, that this time can be different. (Yes, I know how that sounds...) Computer systems are much more flexible than they were when the MARC music format was first implemented, and can be designed to alleviate more of the human effort than before. We've learned a great deal from automation and implementation of the MARC format that we can build on in the next generation library catalog. We have a long road ahead of us, but I think it's time to address these issues head-on once again. I'd like to believe we can leverage the experience of those like Donald Siebert involved in the first round of MARC implementation, together with experts in recent developments, to make progress towards our larger goal.

Sunday, September 18, 2005

The next big thing in searching?

At a conference last week, I heard Stephen Robertson of Microsoft Research Cambridge speak about the primacy of text in information retrieval, whether for text, images, or any other type of medium. He made a statement in the talk that the first generation of information retrieval systems operated on Boolean principles, and the second generation (our current systems) provide relevance-ranked lists. This may be a truism in the IR world, but it's something I hadn't thought about in these terms before. Our library sytems certainly are primitive in terms of searching, and they operate on the Boolean model. But I hadn't thought of relevance ranking as the "next step" - probably because the control freak in me is suspicious of a definition of "relevance" not my own. But I think it's fine to look at the progression of IR systems in this way.

So what's the third generation? Where are we going next? I think the next step is grouping in search results. Grouping is where I see the power of Google-like search systems merging with library priorities like vocabulary control. Imagine systems that allow the user to explore (and refine) a result set by a specific meaning of a search term that has multiple meanings, by format, or by any number of other features meaningful to that user for that query at that time. I picture highly adaptive systems far more interactive than those we see today. Options for search refinement alone, I don't believe, go far enough, as they require the user to deduce patterns in the result set. I believe systems should explicitly tell users about some of those patterns and use them to present the result set in a more meaningful way. Search engines like Clusty are starting to incorporate some of these ideas. It remains to be seen if they catch on.

FRBR assumes this sort of grouping can be provided, using the different levels of group 1 entities. Discussions of FRBR displays frequently talk about presenting Expressions with a language for textual items, with a director for film, or with a performer for music, allowing users to select the Expression most useful to them before viewing Manifestations. What's missing is how the system knows what bits of information would be relevant for distinguishing between Expressions, since these bits of information will be different for different types of materials, and sometimes even with similar types of materials. We have a ways to go before the type of system I'm imagining reaches maturity.

Wednesday, September 07, 2005

Dangers of assumptions

Over the holiday weekend, I read the paper by Thomas Mann, Will Google’s Keyword Searching Eliminate the Need for LC Cataloging and Classification? Mann presumes to know exactly what is possible (not just currently implemented) in a search engine - the paper is stuffed full of absolutes: "cannot," "only," and "will not." The paper seems to focus on Google as simply taking in words in a query, looking them all up in a word-by-word index of all documents, and performing some sort of relevance ranking on documents that contain the search terms. It not only assumes that Google takes this simplistic approach, it rejects that any further capabilities are even possible in a search engine.

I believe this is a thoroughly (and perhaps, in this case, deliberately) naive assessment of the situation. Just because library catalogs offer only simple fielded searching and straightforward keyword indexes doesn't mean all retrieval systems do the same. Mann ignores the possibility of a layer between the user's query and the word-by-word index. He states, "having only keyword access to content is that it cannot solve the problems of synonyms, variant phrases, and different languages being used for the same subjects." This statement confuses "keyword access" (just looking something up in a full-text index) with a system that uses a keyword index among other things for searching. Google could (and right now, does, with the ~ operator [thanks Pat, for the heads up on this!], and who of us library folk is to say they won't do this by default in Google Print) do synonym expansion on search terms before sending the query to the full-text index. Point is, it's not impossible to do this in a search system. The same idea goes for finding items in other languages - translation before the search is actually executed could be done. Ordering, grouping (yes, grouping!), and presentation of search results in this environment would require some advanced processing, but that's doable too.

Of course, there is a difference between what's possible and what's actually implemented in Google today. Mann's language confuses the two, by stating (incorrectly) what's possible using as evidence what's implemented. What's implemented today is the functionality in the Web search engine, but we shouldn't assume the same functionality will drive Google Print. This article uses rhetoric to stir the librarians up for their cause. But it does us a disservice by making false assumptions and obscuring the facts. There are arguments to be made for why libraries are still essential and relevant today. But rabble-rousing with partial truths isn't the way to make them.

Monday, August 29, 2005

Google Print and Fair Use

Having thought some about the "copying" aspect of Google Print, it would now be prudent to think about exceptions to the exclusive right of copyright holders to reproduce a work. Google's stance seems to be that their activities fall under the scope of the Fair Use exception to copyright. Fair Use is by far a straightforward concept, and comparatively very few cases have served to clarify the issue. Here's the text of section 107 of the copyright act, which describes the fair use exception:


Notwithstanding the provisions of sections 106 and 106A, the fair use of a copyrighted work, including such use by reproduction in copies or phonorecords or by any other means specified by that section, for purposes such as criticism, comment, news reporting, teaching (including multiple copies for classroom use), scholarship, or research, is not an infringement of copyright. In determining whether the use made of a work in any particular case is a fair use the factors to be considered shall include —

(1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;

(2) the nature of the copyrighted work;

(3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and

(4) the effect of the use upon the potential market for or value of the copyrighted work.

The fact that a work is unpublished shall not itself bar a finding of fair use if such finding is made upon consideration of all the above factors.


Note that whether the copyright owner objects or not is not a factor to be considered when determining fair use. That copyright owner could file a lawsuit, but the fair use claim is evaluated on these four factors only.

So how does Google Print stack up against the four factors?

(1) Purpose and character. Commercial vs. educational is singled out here, and certainly Google's use is commercial. But that's not the only purpose or character allowed to be considered. A lawyer for Google could claim that their service, meeting people's information needs and directing them to a copyright holder when a work meets that information need, is a Good Thing. They could then go on to argue that making money of this is secondary, but lots of folks wouldn't believe that.

(2) Nature of the copyrighted work. This is hard to pin down due to the scope of what's being digitized. Books that have been out of print for 45 years and aren't widely available in the used book marked would evaluate differently according to this criteria than Harry Potter. (Yes, research libraries collect fiction too.)

(3) Amount of the work. Again, tricky. Google is digitizing (copying) the entire work, and, presumably, using the entire work to create their index. The counter-argument seems to be they're only showing a small part to users of their service, but I don't believe that applies here. The exclusive right is the copying part, not what you show to other people.

(4) Effect on the market. Here is where only showing snippets to end-users comes in to play. Certainly the effect on the market is potentially severe if one could download, print, read a whole book from Google instead of purchasing it. The recording industry feels that way about file sharing, but there are many who disagree, claiming file sharing actually stimulates purchasing. (Sorry no citations right now, but there are gobs of studies out there on both sides of this issue.) I imagine Google would claim that by showing snippets they're telling users about resources they didn't know about before, and are thus adding to the market. This will be an interesting argument to follow.

My conclusion is that the fair use claim is far from a slam dunk in either direction. Personally, I'd love to see this litigated (and found in favor of Google!) to start what I consider to be much-needed reform in copyright law.

IANAL. Any misinterpretations or flawed analyses are entirely mine, and the result of me trying to pretend I know something about this stuff.

Sunday, August 28, 2005

Musings on the state of coyright

The recent brou-ha-ha (wow, I think that’s the first time I’ve ever written that word down!) over Google Print has me thinking about copyright law. I am not a lawyer. I have no legal training or education. I have picked up a bit about copyright law while working in the area of digital libraries for the past five years, however. I think what I think I know is accurate, but hey, I'm wrong a reasonable amount of the time.

The publishers who have objected to the Google Print project say that the project violates copyright law by scanning the books in question (copying, which is the first exclusive right granted to copyright holders by section 106 of U.S. copyright law) to index them. So how is this different than Google’s Web index? Well, in creating the Web index Google caches Web pages too. Caching may not actually be the right word there – Google probably more actively, intentionally, or permanently creates a copy than Random J. User’s Web browser does. One could argue there’s some sort of difference between the caching done by Google of Web pages and scanning page images of printed books, but it seems to me this difference is a matter of degree rather than of real substance. So if the digitization for Google Print is a copyright violation, does that mean all Web search engines are copyright violations?

Let’s take this exercise one step further. Indexes have been around for a very long time: the Readers’ Guide to Periodical Literature, Academic Search Premier, the MLA International Bibliography, and on ad infinitum. I admit to being ignorant as to whether these more traditional indexes tend to operate with the blessing of the copyright holders (although many of them are actually produced by publishers to cover their content), but surely not all of them do, and the library world isn’t exactly abuzz with these copyright holders crying foul. One difference is that the processing that happens to create these more traditional indexes (although this may no longer be true today!) is entirely an intellectual exercise. Any “copying” of the work done to create the index is purely in a person’s head. Is this difference one of degree or of substance?

To go yet another step further, library catalogs use a copyrighted item to create a new representation – is there an argument there that catalog records are derivative works? Obviously we’re in danger of descending into the ridiculous here, but the need for some sort of balance is clear. The concept of balance between the rights of the creator of a work and the benefit to the public good from its use is inherent in copyright law. Too bad the specifics of maintaining this balance are in language that languishes far behind current technologies.

I think it will take a copyright challenge to a large for-profit like Google (rather than to even the most resource-rich library) to overhaul copyright law, to bring it up to the times. Google seems to me to have the desire and the resources to present a reasonable defense, and persist through a legal battle rather than settling the short-term problem through an agreement with publishers. But, as I’ve said, I’m wrong a reasonable amount of the time.

Thursday, August 11, 2005

A billion and one, a billion and two...

The OCLC folks are all abuzz with the addition today of the billionth holding to Worldcat, as reported all over. This is obviously an enormous milestone for OCLC and for libraries in general. Kudos are in order for all of us, I think!

The union catalog has transformed the way libraries provide access to their material. A billion holdings in one database seems to me to be proof positive of that. But OCLC Research staff and many others, researchers and practicioners, aren't content with the functionality our current union catalogs offer. The enormous wealth of data represented by those one billion holdings has the potential to be used in innumerable ways. I believe OCLC's FRBR activities are excellent examples of the sorts of creative things we can do with this data to better serve our users. We've made huge strides in access to materials, yet we have many miles to go.

UPDATE: I've discovered today the misfortune of having a book on The Monkees be Worldcat's one billionth holding. We're going to have a country of librarians walking around for two weeks now with that damn theme song stuck in our heads!

Wednesday, August 10, 2005

Keeping up with technology

Podcasting, Web services, RDF, Flickr, Ebooks. Buzzwords, right? All of these are extremely useful technologies or applications, but I don't actually use any of them. Each has its place, each is good at solving certain types of problems. None, of course, is a magic wand that makes everything in life easier.

I follow a number of library- and technology-related blogs. Many of them hype a certain technology that is meaningful to the blogger for their particular needs. I learn a huge amount from these bloggers, the information they provide, and the fervor with which they provide it. But rarely do I go out and try any of the technologies being described just to see what they are. A few peak my curiosity and I go check them out, but for the majority I just mentally file the information away for when I have a problem the technology in question solves. There's just too much going on in this environment right now to really delve in and learn everything new that comes along. Each of us picks up on the emerging technologies most relevant to us in our personal or professional lives. Other technologies are only relevant to us at a later time, but hearing about them before we need them reminds us of the vast range of possibility out there. Sharing our experiences helps others both to adopt them right away when appropriate, but also to adopt them later as the need grows.

Tuesday, August 02, 2005

To each their own "metadata"

I was introduced to someone today as the "Metadata Librarian," and received a reaction I seem to get a lot: "Oh, metadata, huh? Someday I'll understand that." On my optimistic days, I want to respond "Would you like to go get a cup of coffee and chat?" On my cynical days, "You've got an opportunity here to learn something new! Take it!"

Everyone has their talents and areas of difficulty. We're all really good at some things and equally bad at others. Me, I'm completely spatially inept. It once took me 3 hours to put together a futon frame (with instructions). I'm fine with that, because I know my talents lie elsewhere, although I do often think it would be nice to be handy. Despite my lack of innate talent in some areas, I've never thought I simply can't learn any of it. Little by little I'll learn to fix things around the house. I'll never be able to paint with any level of inspiration, but with a whole lot of practice I might be able to use color effectively or produce a still life that is recognizable. One might think metadata is uninteresting. That's cool. I find a lot of stuff out there uninteresting. But don't think it's unlearnable.

Part of the problem here is that "metadata" isn't a monolithic concept. Depending on one's perspective, it can mean virtually anything. To lots of people, all they need is descriptive metadata, and maybe even some version of qualified Dublin Core their content management solution provides them. GIS specialists delve deeply into an area of metadata many know very little about. For many, text encoding is the metadata world, of extremely rich depth and subtlety. I had an interesting conversation recently with a colleague about the definition of "structural" metadata. By some definition, TEI markup is structural metadata, indicating the stucture of the text by surrounding that text with tags. Does that same logic apply to music encoding? Music markup languages specify the musical features themselves, rather than "surrounding" them with metadata. But certainly there's some similarity to text markup. The boundary between structural metadata and markup isn't the same to everyone. Similarly, there are times when I use the word metadata to refer to something that might more accurately be "data," and when I use it to refer to something that might be "meta-metadata."

All of these views are valid. I'm constantly reminding myself of this. Often when my first reaction is that someone doesn't get it, it's really their view not quite meshing with mine. It's important that we have some common terminology and meanings, but I believe there's room for perspective as well. I can get better at my job if I listen more closely to these perspectives.

Thursday, July 28, 2005

Music subject headings

Wow, the blog has been lonely lately, hasn't it? It's almost as if there are too many things floating around in my head that I can't get any of them fully formed enough to even write about here.

One of those many floating thoughts has been subject headings for music. Many traditional schemes, like LCSH, make a distinction between headings used for works about music, and headings used for music itself. For example, "Symphonies" is used for music scores and recordings of symphonies. But "Symphony" is assigned to texts about symphonies.

Obviously at first glance the distinction between the two forms is subtle. Even if a user realized the potential for this distinction being made (!), it would be difficult for that user to determine which form to use in which case. In my library catalog, a subject browse on "symphonies" lists first an entry for 5407 matches, then second, "see related headings for: symphonies." Clicking on the latter yields a screen saying "Search topics related to the subject SYMPHONIES," but no way to actually do that. This is probably because the authority record for symphonies has no 550s specifying any related headings. Geez. Both because the system shows this anyways and because there are no related headings. [Yet another NOTE: the mechanism for specifying a heading is broader or narrower than another heading in the MARC authority format is ridiculously complicated. No wonder the relationships between LCSH headings are so poor.] This same screen is also where one would view the scope note for the heading "symphonies":

Here are entered symphonies for orchestra. Symphonies for other mediums of performance are entered under this heading followed by the medium, e.g. Symphonies (Band); Symphonies (Chamber orchestra). Works about the symphony are entered under Symphony.

OK. So to find out if "symphonies" is what I'm looking for, I need to click "see related headings for: symphonies"? Riiiiight. Sure, my catalog could handle this better. Not many do.

This distinction isn't always so obvious to specialists, either. I've been reading up on the topic for a project and I'm struck by how rarely it's made explicit. A huge majority of writings simply assume they're talking about one, the other, or both, but never say so. Many others indicate they're discussing one or the other but provide examples of both. I myself recently forgot the distinction at a critical juncture. :-)

I'm wondering if this distinction between headings for works about music and works of music is still needed in modern systems. [NOTE: I don't consider any of the MARC catalogs I'm familiar with to be "modern systems"!] We certainly now have mechanisms to make this distinction in ways other than a subject string. Most of me says this is an outdated mechanism. But in a huge library catalog covering both types of materials, the distinction does need to be made in some way. I'm still pondering over exactly which way that should be.

Monday, July 11, 2005

Structure standards and content standards

It's funny how related things seem to come in spurts in our lives. Or maybe it's just that once we notice something once, it's easier to notice again. In my standard metadata spiel, I, like many others, distinguish between structure standards that tell you what "fields" (for lack of a better term) to record, and content standards that tell you how to structure values in those fields. The latter can be either rules for structuring content or actual lists of permissible entries. It's an extremely useful distinction. Yet I've been noticing recently that it's frequently misunderstood, or that the distinction is implicit in a conversation rather than explicit.

One place this trend caught my eye recently was in a blog post by Christopher Harris on using LII's RSS feed to generate MARC records, and subsequent comments and posts by several people, including Karen Schneider of LII. Most of the ensuing discussion was about keeping the two data sources in sync, which of course is important to plan for. But I noted a conspicuous absence of content standards in the discussion. MARC records, of course, do not have to adhere to AACR2 practices. In fact, there are millions of non-AACR2 records (mostly created pre-AACR2 and never upgraded for practical reasons) in our catalogs. But today if one is creating a MARC record, it would be prudent to either use AACR2 or have a compelling argument against it. Yet neither of those options appeared in this discussion. Reading between the lines, I suspect the transformation should be reasonably straightforward, but one shouldn't have to read between the lines to know.

I suppose what I'm really saying here is that when talking about these sorts of activities, we need to completely define the problem to be solved before a solution can be determined. And that includes dealing with content standards in addition to structure standards. Explicitly. Knowing which standards (or lack of them) are in use in the source data and which are expected in the target schema. Planning for moving between them. This is an extremely interesting topic, and I personally would love to see more discussion about it.

Oh, and, for the record, I'm with Karen that one would want to be careful about putting lots of records for things like LII content into our MARC catalogs. My vision (imperfectly focused, unfortunately!) is that because the format (and the content standard that is normally used with it) doesn't describe this type of material well, and the systems in which we store and deliver our MARC records don't provide the sort of retrieval we might desire for these materials, our users would be better served by a layer on top of the catalog that also provides retrieval on other information sources better suited to describing these materials. This higher-level system would provide some basic searching but most importantly lead a user down into specific information sources that best meet his needs. We have lots of technologies and bits of applications that might be used for this purpose. I wonder what will emerge.

Wednesday, July 06, 2005

So what's up with RDF?

Kevin Clarke posted on his blog last week some thoughts on the recent ALA conference and a session on metadata interoperability. A discussion has ensued from this about RDF, with commentary by Leigh Dodds and a follow-up post by Kevin. I've learned a great deal from this exchange. I've always felt that I was missing something with RDF, that I needed a discussion on a much more practical level than those I'd been exposed to in order to understand what it could do for me better than the tools I already use. I've heard smart people I like and respect make comments like those by Kevin, Bill Moen, Dorothea Salo, and Roy Tennant quoted in these blog postings, and felt a bit of comfort that I wasn't the only one who felt left out. But it's not enough to have company in the "huh?" camp - I want to understand. I want to be able to make a reasoned argument against RDF, or embrace it for tasks it does better (in my world) than other things. Yet I've never felt like I can do either of those things. For now, I'll follow discussions such as this one in order to slowly absorb all the angles.

And all of this banter reminds me I need to learn RelaxNG and finally figure out what the deal is with topic maps. Anybody have a few extra hours in their day they're willing to send my way? :-)

Tuesday, July 05, 2005

Addition of dates to existing name headings

The Library of Congress Cataloging Policy and Support Office recently announced a review and request for comments on a potential change of policy regarding addition of dates to existing personal name headings. Currently, dates are only added in certain situations, and once a heading is established, dates are never added to it after the fact. Personal name headings are frequently created while an individual is alive, leading to headings such as (from the CPSO proposal):

100 1# $a Bernstein, Leonard, $d 1918-

This heading then was not changed when Bernstein died in 1990. The CPSO proposal notes that libraries, including LC, receive frequent comments and complaints from users regarding the "out of date" nature of headings of this sort.

In discussion of this policy on the AUTOCAT listserv, the question arose as to whether name authority files served to simply generate unique headings for an person, or if they served a wider biographical function. Certainly historically the former is true. But many, including the CPSO, are recognizing that increasingly we may be well served by delving into the latter. We have an opportunity here to become more useful and relevant to the wider information community. To take that opportunity might seem to be a no-brainer.

However, the current cataloging infrastructure makes the implementation of this change challenging, to say the least. As authority data is replicated in local catalogs and the shared environment, and most integrated library systems store actual heading strings in bibliographic records rather than pointers to authority records, changing a heading would then require notifying all libraries that a change has been made, propagating that change from one library to the rest, then continuting to propagate that change in every local system to all affected bibliographic records. Clearly this mechanism is anachronistic in today's networked world, where relational databases are so entrenched as to be considered almost quaint. I fully understand the practical implications of the CPSO implementing this policy. Yet I believe that it is the right thing to do. We as librarians simply must have a vision for what we're trying to accomplish, and work tirelessly towards that goal. While we must keep the practical considerations in mind, we can't let them dictate all of our other decisions. Let's set the policy to do the right thing, and insist on systems that support our goals.

Tuesday, June 28, 2005

Back from ALA

ALA Annual in Chicago was the usual flurry of old friends, Powerpoint presentations, and exposure to topics new to me. The blogger's get-together put together by the It's all good folks and generously hosted by OCLC was certainly one of the highlights. I've been a bit tentative in promoting my blog to date, so it was nice to mingle with other bloggers and talk shop (and beyond!). Another major highlight was Kevin Clarke's presentation on XOBIS. Nice to finally meet you, Kevin!

I spent most of my time at ALA attending presentations I "had" to attend--those related to my daily work. I was able to spend a small amount of time expanding my horizons, but I wish I could have done more. And this schedule is without being involved in any ALA committees that meet during the conference. There is simply too much going on to take advantage of it all.

On another note: on the trip home I started reading, but didn't finish, Martha Yee's recent paper outlining a "MARC 21 Shopping List." I should hold any substantial comment until I finish the article, but so far I'm impressed. The approach of looking very precisely at the criticism of MARC and current cataloging practice to determine what exactly is being criticized, I believe, is long overdue. I do find myself thinking of counter-arguments to some of the conclusions, however. But intelligent discourse is absolutely what we should be striving for!