Thursday, August 02, 2007
Can LIS education learn something from CS?
But seeing some discussion recently about whether or not some of the more technical positions in libraries should require an MLS, I'm wondering if there's something we can learn from the CS community. Technical jobs commonly require a BS in Computer Science (note: NOT a graduate degree, whether it be research-based or professional), or demonstrated expertise in the task at hand, say, programming. That expertise can be demonstrated through that degree, through various certification programs, or by showing code one has written. While I suspect some would argue that the MLS is equivalent to those certification programs, I'm not so sure. A certification program for, say, Windows server administration, would be based on many practical tasks, and we don't see many of those in our library schools.
While I do agree we should be teaching the theory of things and then apply the practice on top of it, I see our library schools failing our students by ONLY teaching the former and providing no opportunity for the latter. Even single undergraduate programming classes manage to teach both. Can't we learn something from that?
Monday, July 30, 2007
Cutting through the rhetoric about subject headings
It seems to me that a great many of our disagreements in the library realm have at their root people talking past each other, each side meaning something different by a given term or two, but not cognizant of that fact. I see LCSH as a prime example of this phenomenon. A great deal of debate occurs over whether precoordinated subject strings or postcoordinated subject strings are more useful. But I see a fundamental difference in the way various participants in these discussions define “postcoordinated.”
One definition is that postcoordinated headings have no subdivisions at all; in LCSH-speak, have no --. The other definition is that postcoordinated headings are “faceted” (to introduce another term that complicates the issue); that each heading reflects only one characteristic of the work, such as “topic,” “place,” or “date.” The difference here is the difference between “subdivisions” and “facets.” These two concepts are not identical. A common criticism of postcoordinated headings is that they would not represent the essential distinction between concepts like “History--Philosophy” and “Philosophy--History.” While (ignoring the syntax; whether or not the double dashes are used is a style issue) this would be true according to the first definition, it’s not necessarily true according to the second—a “topical” facet may very well represent a complex concept. I’ve never seen a discussion on this issue in which this distinction is made clear to both sides. It’s unclear to me whether the “traditional” definition of postcoordinate allows the faceted interpretation or requires the subdivision interpretation, but I think what’s needed here is clarification of current definitions rather than historical ones.
I don’t have all the answers in this debate, nor does anyone else at this point. My inclination is toward the postcoordinate side, although I do very much want to keep an open mind on the issue. I’d like to see a well-reasoned argument for a postcoordinate system presented according to the facet definition (something I’ve long been wanting to write but find this is one of the many issues that have trouble finding their way from my brain to a shareable form). I personally read arguments for precoordinate indexing and think to myself, “We can do all of this with postcoordinate headings if we had systems that operated reasonably.” (Big IF there, considering our current state of affairs!) We need to have more room to experiment with these options to see if my interpretation is a good one. The Endeca use of precoordinated strings shows powerful promise; we need more large-scale implementations of systems working off of postcoordinated data to allow us to compare both user functionality and cataloging time (a much-forgotten but essential factor) of the two approaches. I want data, darn it! We can only go so far with the philosophical argument; to get beyond our current roadblock we need to see what will happen if we follow the various paths available to us.
Sunday, July 08, 2007
Everything I know about librarianship I learned from Star Wars
Forgive the hyperbole—of course it’s not everything. But hear me out.
In celebration of meeting a major deadline and milestone in my career, I took some time for myself this weekend and watched the original Star Wars trilogy. (Yes, the good one.) This is something of a ritual for me, albeit one I’ve only performed only one other time in recent memory. It stretches back to junior high days when my brother and I, when we had a day off of school, would frequently watch all three movies right in a row off of a somewhat wobbly VHS tape made from early HBO airings. (If you’ve ever done this, you know just how very boring Jedi gets in the middle, but I digress.) Nowadays, three movies in three days is about all I have patience for (and Jedi still got boring), but it was comforting nonetheless.
While watching, I found myself saying lines out loud before they were said on screen, an annoying habit of mine. A few of these lines struck me as interesting, however. While Star Wars isn’t exactly the pinnacle of Western philosophy, my brain made some funny connections between the storylines and dialogue I know so well and librarianship. Here are a few examples:
“Use the force, Luke.” (Disembodied Obi-Wan voice to Luke, in the first movie.) We need to trust ourselves as skilled professionals. We know what we’re doing. Most of us in the library profession are in it because we love the work and believe we really can make a difference. This heavy personal investment in our work gives us the luxury to rely on our instincts in many cases, pushing forward with initiatives that are simply the right thing to do. Now we need to back up that instinct with reasonable plans, budget justifications, and all that administrative stuff, but I really do believe the best ideas come out of pure inspiration and vision, facilitated by the connections between us.
“R-2, you know better than to trust a strange computer.” (C3PO, Empire.) Now, the computer turned out to be right in this case, but we as librarians consider it part of our job to promote the effective evaluation of information. Many of the discussions today around this issue take an adversarial tone, as if the goal is to spot the misinformation and quash it. But we simply can’t just look at it as ferreting out the bad. We must not be judgmental. Instead, this evaluation can and should be just a routine part of our information flow. We simply need to evaluate everything. The source is only one factor among many that should be considered.
“What I told you is true, from a certain point of view.” (Blue-energy Obi-Wan dude, Jedi.) The role of perspective in truth or falsehood could provide more commentary on the evaluation of information theme, but I’ll take it in a slightly different direction – the role of metadata records in libraries. I find myself talking about this topic, inspired by Carl Lagoze, a great deal (and I believe on this blog before). A metadata record is necessarily a surrogate for a resource, and thus inherently takes a certain perspective on that resource in what it includes, what it leaves out, and the vocabularies it uses. We need to dispel ourselves of the myth that our records can or should be all things to all people, and instead focus on defining the views our metadata records need to support.
And of course the overall theme of the movies that a relatively small, smart, dedicated movement can effect sorely needed, large-scale change gives me good feelings for the future of libraries as well. So there you go. Little did Mr. Lucas know he was providing the library profession with a model to help guide our work. :-)
Oh, and I learned that my dog is strangely fascinated by Ewoks. Go figure.
Saturday, June 16, 2007
Re-imagining browsing
In a recent issue of Educause Review, Robert Kieft describes work at the Tri-Colleges outside
I like this idea, as when I do shelf browsing, I do open up books, read a random page or two, and just generally poke around to see if the book looks interesting. I also believe this is an area in which our current catalogs don’t remotely match the physical experience.
But I also think there’s more to browsing than having something catch your eye then explore it further. When shelf browsing, we don’t look at everything – we only look at select things. In a bookstore, a cover might catch one’s eye, but this is less likely to happen on shelves with plain library bindings. A title or an author might be the hook that causes one to pick up one book and not another. To some extent, this type of browsing activity is random.
But what if it were less random? Can we re-imagine (or at least extend) our notion of browsing to make it more targeted? I like to think of browsing both as the ability to “look inside” a resource to learn more about it before committing to it, and as a “more like this” feature that introduces me to resources that I didn’t previously know existed. Shelf browsing of course does this but is also obviously limited by physical constraints, so that only one aspect of a work can be brought out by a classification scheme that locates it on the shelf. This isn’t news to anyone—see for example, the much-blogged-about Everything is Miscellaneous by David Weinberger for a discussion of what he calls first-, second-, and third-order methods of organization.
What I’m interested in is the ability of our catalogs to bring out more flexibility in browsing. If I’m looking at a resource, I want to be able to note one feature of it, and instantly get other resources that share that feature. I’ll then want to be able to add or subtract features to exert control over the size and characteristics of my results set. Say I’ve just read The Godfather. I might want more mafia fiction after that. But that’s a long list that I might consider too broad. Perhaps I’d limit myself to novels about the Sicilian mafia, or set in the 1940s or ‘50s. Then after I select a work or two, I may want to move on to real-life mafia stories, but only from the mobster’s perspective (not law enforcement’s). I’d likely find Henry Hill’s story in Nicholas Pileggi’s Wiseguy, which could lead me to watching the movie Goodfellas, and start wondering what happened to Hill. From there I might start exploring the history of the witness protection program, then FBI training methods, then perhaps all the way to Thomas Harris’ Silence of the Lambs and the film that was made based on the novel! And now imagine we can do all of that in a few clicks, without having to type anything in, or to know any of those books or films exist. And imagine if this could happen across multiple databases of content (including those from outside the library sector, such as IMDB), without me ever having to know that.
Wednesday, June 06, 2007
Partnership announced between CIC libraries and Google
The CIC Libraries are now participating in Google Book Search, announced today.
Sunday, May 06, 2007
DC and RDA - the beginning of a beautiful friendship?
An interesting announcement was made this past week that the DC and RDA communities will be working together to do the following:
- development of an RDA Element Vocabulary
- development of an RDA DC Application Profile based on FRBR and FRAD
- disclosure of RDA Value Vocabularies using RDF/RDFS/SKOS
The news isn’t exactly taking the library community by storm, but the commentary I have seen has all been of the “this is a good thing, I’ll follow this with interest” theme. But something bothers me about this plan, and I’m having trouble deciding exactly what it is I find, well, wrong in some way.
There’s nothing in the announcement that indicates the development of RDA proper will be affected by this work; in fact, the indication in the announcement that funding will be sought for the activities outlined implies the work is a long way off, likely entirely too late to have any real effect on RDA. This seems to be to be entirely backwards – trying to harmonize DC principles with RDA after the fact. Didn’t the DC community learn its lesson about the pitfalls of this approach when developing the Abstract Model, only realizing long after developing a metadata element set that it would benefit from an underlying model?
This general approach failed miserably with the DC Libraries Application Profile. There, the application profile developers wanted to use some elements from MODS, but weren’t able to because MODS doesn’t conform to the DCMI Abstract Model. So basically what the DC community said here was that application profiles are great, they form the fundamental basis of DC extensibility, but, oh yeah, you can’t actually use elements from any other standards unless they conform to the Abstract Model, even though are no approved encodings for even DC itself more than two years after the Abstract Model was released. OK then. Way to foster collaboration between metadata communities.
But maybe the DC community will change paths and realize flexibility and collaboration get more users than intellectual rigor. It’s sad in some respects, but true. It wouldn’t be the first time there’s been a major shift in the direction of the DCMI. I don’t know that this is a good thing, but I’ll certainly follow it with interest.
Saturday, April 14, 2007
ALA Draft Digitization Principles
One big-picture issue I don't think is clear, however, is the label "digitization." The principles are for the most part not about the digitization (conversion from analog to digital format) process, nor do I think they should be. They're more about the properties of "digital libraries" as a whole, which have content that was once analog, content that is born digital, and perhaps even metadata about objects that aren't digital at all. These principles seem to describe systems and organizations more than just objects.
The "expand the scope even further" commentary is also particularly apt. Coming from ALA, the focus on "libraries" could, as one comment on the blog mentions, to exclude other producers and maintainers of digital content, even others in the cultural heritage sector such as archives and museums. The direction I'd like to see these principles expand is related to (buzzword warning) interoperability. (Don't fall asleep--although that term that is often empty in its usage it really does describe some essential concepts.) My reading of the principles seems to focus inward, on developing maintaining digital collections within a single institution or close consortium. But we have an opportunity now to move away from the traditional (another buzzword warning!) "silo" approach to libraries, and create systems that operate in a much more open fashion, promoting re-use and exchange of content and metadata in new and unexpected ways. The digital libraries we maintain shouldn't just be accessible through our well-designed interfaces intended for a human to interact with - we need to supplement that access with additional methods. These methods are constantly expanding, and it will be difficult for us to keep up, but we can't ignore them.
Thursday, April 12, 2007
Weighing in on speaker compensation
choose as invitees to engage in coffee conversations, eclectic dinner meetups, and learn from each of our communities through attending others' presentations? Certainly, not every invited speaker will be able to (or in some cases, want to) stay much longer than his or her talk, but we need models that encourage them to do so rather than making it difficult for them.
Saturday, March 24, 2007
Staying out of it, for now
Many of the discussions are ebbing to the point where it doesn’t make sense for me to weigh in, but even if they weren’t, my inclination is to sit back and watch rather than participate. Perhaps I’m a bit disillusioned thinking that my words will have very little impact. I’m glad the discussion is ongoing, and I still believe that most people out there are reasonable, and can behave themselves while engaging in professional discourse. But many of the current discussions have taken on an “us vs. them” bent, and when that happens I tend to stay away, avoiding situations where emotions and stereotypes start getting in the way of dialogue.
I am sad for feeling this way, but we all must make decisions about how best to spend our precious time. I believe I have a great deal to offer these discussions, particularly in expanding their scope beyond just “catalogs” to the types of digital library systems I’m involved with building and with types of materials not well-served by the MARC/AACR2/LCSH/etc. suite that’s the main focus of these discussions. I’ll probably post some thoughts to this blog (please do comment – I’m not afraid of discussion overall!) but for now I feel the time and intellectual investment of communicating in some of these other forums is best left to others. I’m hoping my aversion to participation is temporary, and that as I get caught up both with others thoughts and my own, I’m able to jump back into the community. See you soon.
Tuesday, January 30, 2007
Modular Vocabularies
Spilker, John D. “Toward an International Music Thesaurus” Fontes Artis Musicae 52/1, January - March 2005: 29-44.
The mention in the article of the debate about whether to include a facet for “content subjects” which would include “extra-musical associations” (i.e., topics, things the music is “about”) gave me pause, however. Looking more closely, I also see that they propose a “philosophies and religions” facet, ask the question if “distant” terms that appear in source vocabularies as compound terms relating them to music (e.g., “astrology and music”) should be retained in any way, and copy terms from an existing instrument vocabulary into their own instrument facet. This duplication of effort bothers me a great deal. I see this phenomenon fairly often—projects that try to do everything end up doing nothing very well. LCSH does this, trying to shove everything under the heading “subject” and not making very good distinctions between what type of subjects those terms are. Even in faceted vocabularies, I see communities (like the one described in this article) try to include everything that might be needed to index material for that community. The emerging Ethnographic Thesaurus, facing an enormous task in developing a vocabulary for a very large and diverse field, shows signs of this as well, but I know the editors are considering these issues as they move forward with development. I think this would work much better if these communities focused instead on only those facets that they have particular expertise in, and “borrowed” the rest from other communities.
There’s often an assumption in the library world that a record needs to use a single “subject” vocabulary. But as we move forward, surely that’s a constraint (even if it’s only perceived) we can break out of. There’s no reason vocabulary for different facets (notice how I’m assuming a faceted, post-coordinate structure here) has to come from the same vocabulary. Let’s leave each specific vocabulary to its experts, and not try to have musicians developing terminology for religion and astronomers developing terminology for book bindings.
There are many details to work out, for example, the user implications of one vocabulary using singular forms by default and another using plural (there are standards for such things, but let’s be realistic about how many vocabularies that are otherwise useful we’d throw out of consideration because of the tense of its headings), but there are technological means for doing this. I’d hate to see an inordinate focus on the (potentially many) small challenges derail the larger, necessary, move in a more flexible direction.
Monday, December 18, 2006
I *love* this
The best is the "genre" browse (but take out those --Young adult fiction subdivisions and move them to an audience facet). It's not a short list, but it's not too long either. It would be interesting to arrange these hierarchically and see if navigating that list made any sense to users. And "settings"! How cool to be able to locate fiction that takes place in the Pyrenees. This is what library catalogs should do for our users.
I'm also intrigued by the "character" browse. This is something I've never thought of before. My general rule for browsing facets is to only include facets that have a (relatively) small number of categories, each with a (relatively) large number of members. At first, I didn't think characters met this requirement. Then I clicked on Captain Ahab, and I realized just how many works of fiction there are about him! Great works inspire derivatives, and exploring those is a fun way to guide new reading, in my opinion. It would be interesting to have access to a browse list of all characters in some situations, and only those with a large number of works (note works here, not publications) in other situations. Exploring which situations warrant which presentation would be another interesting line of inquiry.
The next improvement I want to see is allowing users to combine these facets (and others) dynamically so I can find Psychological fiction set in the Pyrenees, then narrow it to works after 1960, then remove the Pyrenees requirement, then add in Captain Ahab to the requirements that are left.... ad nauseum. Our catalogs need to support discovery of new works, not just those we already know the author and title. Systems like this are light years (sci fi fan here!) ahead of LCSH-style "browsing". I want more!
(Note to OCLC - the link to "Known problems" is broken. I'm interested to find out what challenges you've faced when building this beta system. I have a very strange idea of fun.)
I *love* this
The best is the "genre" browse (but take out those --Young adult fiction subdivisions and move them to an audience facet). It's not a short list, but it's not too long either. It would be interesting to arrange these hierarchically and see if navigating that list made any sense to users. And "settings"! How cool to be able to locate fiction that takes place in the Pyrenees. This is what library catalogs should do for our users.
I'm also intrigued by the "character" browse. This is something I've never thought of before. My general rule for browsing facets is to only include facets that have a (relatively) small number of categories, each with a (relatively) large number of members. At first, I didn't think characters met this requirement. Then I clicked on Captain Ahab, and I realized just how many works of fiction there are about him! Great works inspire derivatives, and exploring those is a fun way to guide new reading, in my opinion. It would be interesting to have access to a browse list of all characters in some situations, and only those with a large number of works (note works here, not publications) in other situations. Exploring which situations warrant which presentation would be another interesting line of inquiry.
The next improvement I want to see is allowing users to combine these facets (and others) dynamically so I can find Psychological fiction set in the Pyrenees, then narrow it to works after 1960, then remove the Pyrenees requirement, then add in Captain Ahab to the requirements that are left.... ad nauseum. Our catalogs need to support discovery of new works, not just those we already know the author and title. Systems like this are light years (sci fi fan here!) ahead of LCSH-style "browsing". I want more!
(Note to OCLC - the link to "Known problems" is broken. I'm interested to find out what challenges you've faced when building this beta system. I have a very strange idea of fun.)
Friday, December 08, 2006
True confessions
In light of this and other related events, I've been thinking a bit about what I do get done and why. I believe I've been spoiled by having jobs for a number of years now where I find the work interesting. It's a whole lot easier to get work done when it's engaging and I care about the outcome. I find the tasks I find interesting are the ones I end up working on for the most part, leaving the ones I find un-interesting until right before a deadline.
So what does this mean for libraries? I think it means that we need to make sure to allow our staff to step up and get involved in projects as deeply as interests them. There are many of us out there who get motivated by understanding and buying into the big picture. Don't "protect" your staff from those high-level discussions - allow them to participate as much as they see fit. Sure, there are lots of folks in library-land that are just interested in the paycheck. We need to meet their needs too. But reward those who think beyond the next five minutes - they're going to be running the place soon enough.
Wednesday, November 15, 2006
Children's Book Week
Reading all the touching stories of favorite childhood books across the biblioblogosphere in honor of Children's Book Week has guilted me into posting my own contribution. I still smile when I think of The Little Old Man Who Could Not Read, Irma Simonton Black (Author), Seymour Fleishman (Illustrator). It's a story of a man (who cannot read) who goes to the grocery store and selects items based on the box size and color, trying to match them to products he knows he has at home. Of course, he ends up with an amusing assortment of unintended purchases. The story is touching and the illustrations really make the point. Like many books from my childhood, I think it's out of print (and I see it was first published in 1968, before I was born), but it looks like Amazon can hook you up with a copy, as could many local libraries.
Tuesday, November 07, 2006
More structured metadata
I want more automation. Throwing more money at a manual cataloging process is not a reasonable solution. First of all, it would take waaaaaaayyyyy more money than we can even dream of getting, and second, much metadata creation is not a good use of human effort. Let’s automate everything we can, saving our skilled people for the tasks current automation means are furthest from performing adequately. Let’s get more objective types of metadata, such as pagination, from resources themselves or from their creators (including publishers). Let’s build systems that make data entry and authority control easy. Yes, there will be some mistakes. There will be mistakes if the whole thing is done by humans too. Are catching the few mistakes that will happen from these automated processes more important than devoting our human effort to that extra few resources? More automation means more data total, and the sorts of discovery services I have in mind need lots of that data.
I want more consistency. Users can’t find what’s not there. While we can’t prescribe all records for all resources everywhere have to have a large number of features (I’m against metadata police!), the more of those features that are there mean more discovery options for those users. Imagine a system that provides access to fiction based on geographic setting. Cool, huh? I read one book recently set in Cape Breton Island and can’t wait to get my hands on more. We can’t do that very well today because that data is in very few of our records, and when it is there, isn’t always in the same place. The more consistent we are with our metadata, the better able we’ll be to build those next-generation systems.
I want more structure. I’m a big fan of faceted browsing. The ability to move seamlessly through a system, adding and removing features such as language, date, geography, topic, instrumentation (hey, I’m a musician…), and the like based on what I’m currently seeing in a result set is something I believe our users will be demanding more and more. But we can’t do this if that information isn’t explicitly coded. Instrumentation (e.g., “means of performance”) as part of a generic “subject” string isn’t going to cut it. Geographic subdivisions (even in their own subfield) that are structured to be human- rather than machine-readable also aren’t going to cut it. Nor are textual language notes, [ca. 1846?], or most GMDs. Many of these things can be parsed, and turned into more highly structured data with some degree of success. But why aren’t we doing it that way in the first place? More structure = better discovery capabilities.
What this all means is I’m glad there are lots of extremely bright people with all sorts of perspectives and skills thinking about improved discovery for library materials, but that doesn’t necessarily mean throwing out metadata-based searching. The sorts of systems I envision require more, more highly structured, more predictable, and higher-quality metadata. I want more, not less.
I’ll stand on one last (smallish) soapbox before wrapping this up. In many communities (including both search engines and libraries), discussions about retrieval possibilities often center around textual resources. However, not everything that people are interested in is textual. That’s of course not a surprise, but I’m shocked at how often discovery models are presented that rely on this assumption. I’m all for using the contents of a textual resource to enhance discovery in interesting ways, but we need systems that can provide good retrieval for other sorts of materials too. Let’s not leave our music, our art, our data sets, our maps hanging out to dry while we plow forward with text alone.
Sunday, October 29, 2006
Thinking bigger than fixing typos
There are many ways our cataloging systems could better promote quality records and make it more difficult to commit simple errors. I’ll mention just two here: spell checking and heading control. We hear frequent complaints about the lack of spell checking in our patron search interfaces, but few talk about this feature of being useful to catalogers. And I’m not talking about a button that looks over a record before saving it—I’m talking about real interactive visual feedback that helps a cataloger fix a typo right when it happens. Think Word with its little red squiggly lines—they show up instantly so all you have to do it hit backspace a few times while you’re thinking about this particular field and not miss a beat. If it’s not really an error, the feedback is easy to ignore. Word also has a feature whereby it can automatically correct a misspelling as you type based on a preset (and customizable) list of common typos. Features like this require a bit more attention to make sure the change isn’t an undesired one, but for most people in most cases it saves a great deal more time than it takes, and the feature can be tuned to an individual’s preferences. Checking the entire record after the fact requires a higher cognitive load—turning back to a title page, remembering what you were thinking when you formulated the value for that field, checking an authority file a second time, etc., and is less helpful than real-time feedback.
Heading control is the second area in which our systems could make it easy to do the right thing. Easier movement between a bibliographic record and an authority file, widgets that fill in headings based on a single click or keystroke, and automatic checks that ensure a controlled value matches an authority reference before leaving the field can all help the cataloger avoid simple typographical errors in the first place and make the sort of treasure hunt common typo lists provide less necessary.
Consider also the enormous duplication of effort we’re expending by hundreds of individuals at hundreds of institutions all looking up the same typos in our catalogs and all editing our own copies of the same records. This local editing makes an already tough version control problem worse by increasing the differences between hundreds of copies of a record for the same thing. We have way more cataloging work to do than we can possibly afford, and duplication of effort like this is an embarrassingly poor use of our limited resources. The single most effective step we can take to improve record quality is to stop this insanity we call “cooperative cataloging” today and adopt a streamlined model whereby all benefit instantaneously and automatically from one person fixing a simple typo.
Tuesday, October 17, 2006
Grant proposals
The trick is that in order to write that convincing proposal, you have to do a significant amount of the project, even before you write the proposal and before you get any money. Most of the important decisions, such as what metadata standards you will use, must be made before you write the proposal, both to convince a funding agency you know what you are doing and to develop reasonable cost figures. To make these decisions, an in-depth understanding of the materials, your users, the sorts of discovery and delivery functionalities you will provide, and the systems you will use are all necessary. Coming to those understandings is no small task, and is one of the most important parts of project planning. Don’t think of grant money as “free”—think of it as a way to do something you were going to do anyways, just a bit faster and sooner.
Saturday, September 30, 2006
Librarians in the Media
The article contains the usual rhetoric about caution in evaluating the “authority” of information retrieved by Web search engines, the need for advanced search strategies to achieve better search results, and the bashing of keyword searching. Here, as in so many other places, the subtext is that “our” (meaning libraries’) information is “better” – that if only you, the lowly ignorant user, would simply deem to listen to us, we can enlighten you, teach you the rituals of “quality” searching and location of deserving resources rather than that drivel out there on the Web, that could be written by (gasp!) any yahoo out there.
Of course we know it’s not that simple. But the oversimplification is what’s out there. We’re not doing ourselves any favors by portraying ourselves (or allowing ourselves to be portrayed) as holier-than-thou, constantly telling people they’re not looking for things the right way or using the right things from what they do find, even though they thought they were getting along just fine. We simply can’t draw a line in the sand and say, “the things you find through libraries are good and the things you don’t are suspect.” There are really terrible articles in academic journals, and equally terrible books, many published by reputable firms. There are, on the other hand, countless very good resources out there on the Web, discoverable through search engines. And the line between the two is becoming ever more blurry as scholarly publishing moves towards open access, libraries are putting their collections online, government resources are increasingly becoming Web-accessible, and search engines gain further access to the deep Web.
The first strategy I feel we should be taking is to move discussion away from focusing on the resource and its authority to the information need. Evaluating an individual resource is of course important, but it’s not the first step. Let’s instead talk first about all the resources and search strategies that can meet a given need, rather than always focusing on resources and search strategies that can’t meet that need. There are many, many ways a user can successfully locate the name of the actor in the movie he saw last night, identify a source to purchase a household item at a reasonable price, find a good novel to read on a given theme, or learn more about how the War of 1812 started. Let’s not assume every information need is best met by a peer-reviewed resource, and make those peer-reviewed resources and the mediation services for them we can offer more accessible when these resources and our services are appropriate to meet those information needs. Let’s be a part of the information landscape for our patrons, rather than telling them we sit above it.
Saturday, September 02, 2006
On "authority"
Britannica’s objections to the Nature article arise from a different interpretation of the words “accuracy” and “error.” The refutations by Britannica fall into two general categories. The first is the disputation of certain factual statements, mostly when such facts were established by research. Here, these facts aren’t truly objective, rather, they’re a product of what a human is willing to believe based on the evidence. Different humans will draw different conclusions based on the same evidence. And then there’s the other human element: mistakes. We make them, both those of us who work for Britannica and those who work for Nature. The “error” rates Nature reported for both sources are astonishingly high. Certainly not all of these are true mistakes, maybe not even very many of them, but they exist, in every resource humans create, despite any level of editorial oversight.
Second, and more prevalent, are differing opinions among reasonable people, even experts in a given domain, about what is appropriate at what isn’t to include in text written for a given audience. Anything but the most detailed, comprehensive coverage of a subject requires some degree of oversimplification (and maybe even those as well). By some definition, all such oversimplifications are “wrong” – it’s a matter of perspective and interpretation whether or not they’re useful to make in any given set of circumstances. Truth is circumstantial, much as we hate to admit it.
I’d say the same principles apply to library catalog records. First, think about factual statements. At first glance, something like a publication date would seem to be an objective bit of data that’s either wrong or right. But it’s not that simple. There are multitudes of rules in library cataloging governing how to determine a publication date and how to format it. Interpretation of those rules is necessary, therefore often two different reasonable decisions based on them as to what the publication date is are possible. In cases where a true mistake has been made, our copy cataloging workflows require huge amounts of effort to distribute corrections among all libraries that have used the record with that mistake. Only sometimes is a library correcting a mistake able to reflect this correction in a shared version of a record, and no reasonable system exists to populate that correction to libraries that have already made their own copy of that record. The very idea of hundreds of copies of these records, each slightly different, floating around out there is ridiculous in today’s information environment. We’re currently stuck in this mode for historical reasons, and a major cooperative cataloging infrastructure upgrade is in order.
More subjective decisions are not frequently recognized as such when librarians talk about cataloging. We talk as if one would only follow the rules, the perfect catalog record would be produced, and that if two people were to just follow the same rules, they would produce identical records. But of course that’s not true. There will always be individual variation, no matter how well-written, well-organized, or complete the instructions. Librarians complain about “poor” records when subject headings don’t match their ideas of what a work is about. But catalogers don’t (and of course can’t) read every book, watch every video, or listen to every musical composition they describe. Why have we set up a system whereby we spend a great deal of duplicate effort overriding one subjective decision with another, based on only the most cursory understanding of the resources we’re describing, and keeping multiple but different copies of these records in hundreds of locations? How, exactly, does this promote “quality” in cataloging?
An underlying assumption here is that there is one single perfect cataloging record that is the best description of an item. But of course this isn’t true either. All metadata is an interpretation. The choices we make about vocabularies, level of description, and areas of focus all preference certain uses over others. I’m fond of citing Carl Lagoze’s statement that "it is helpful to think of metadata as multiple views that can be projected from a single information object." Few would argue with this statement taken alone, yet our descriptive practices don’t reflect it. It’s high time we stopped pretending that the rules are all we need, changed our cooperative cataloging models to do it truly cooperatively, and use content experts rather than syntax experts to describe our valuable resources.
Tuesday, August 08, 2006
What about dirty OCR?
I often hear discussions as part of the digital project planning process about how best to approach full-text searching of documents. A common theme of these discussions is whether or not “dirty” (uncorrected, raw) OCR is acceptable or not. The “con” position tends to argue that OCR is only so effective (say, 95%) and that the errors made can and will adversely affect searching. The “pro” position is that some access is better than none, and OCR is a relatively cheap option for providing that “some” access.
The con position has some convincing arguments. Providing some sort of full text search sends a very strong implication that the search works – and if the error rate in the full text is more than negligible, it could be said that implied promise has been broken. Error rates themselves are misleading. A colleague of mine likes to use the following (very effective, in my opinion) example, noting that error rates refer to characters, but we search with words:
Quick brown fix jumps ever the lazy dog.
In this case, there are two errors (fix and ever), out of 40 characters (including spaces), for an accuracy rate of 95%. However, only 75% (6 of 8) words are correct in that example.
So uncorrected OCR has some problems. But the costs of human editing of OCR-ed texts are high – too high to be a valuable alternative in many situations. Double- and triple-keying (two or three humans manually typing in a text while looking at scanned images) tends to be cheaper than OCR with human editing, but these cost savings are typically achieved by outsourcing the work to third-world countries, promoting ethical concerns for many. And both of the human-intervention options themselves represent a non-zero error rate. No solution can reasonably yield completely error-free results.
I’ll argue that the appropriate choice lies, as always, in the details of the situation. How accurate can you expect the OCR to be for the materials in question? 90% vs. 95% vs. 99% makes a big difference. What sorts of funds are available for the project? Are there existing staff available for reassignment, or is there a pool of money available for paying for outsourcing? TEST all the available options with the actual materials needing conversion. Find out what accuracy rate can be achieved via OCR with all available software. Ask editing and double-keying vendors for samples of their work based on samples from the collection. Do a systematic analysis of the results. Don’t guess as to which way is better. Make a decision based on actual evidence, and make sure you get ample quantities of that evidence. Results from one page, or even ten pages, are not sufficient to make a reasoned decision. Use a larger sample, based on the size of the entire collection, to provide an appropriate testbed for making an informed choice between the available options. Too often we assume a small sample represents actual performance and accept quick support of our existing preferences as evidence of their superiority. To make good decisions about the balance of cost and accuracy, we must use all available information, including accurate performance measures from OCR and its alternatives.