Wednesday, December 1, 2010

Radio Ballet

← This book led me to these:



Monday, November 29, 2010

Cross-references: what good are they?

Authority work has been one of my long-time pet projects. Authority records, properly linked to headings in bibliographic records, make it much, much easier to globally update headings as changes occur. Authority records can help keep headings in bibliographic records consistent, and consistent headings allow users to search those headings in the catalog and get what they're looking for. Even if users don't know a thing about name and subject headings and just use keyword searches, hyperlinked consistent headings allow users to click on the headings and retrieve everything else that has that same heading (depending on system settings - and, actually, I'm not quite sure what our setting are like). It's very important that the headings are consistent, because, if they aren't, clicking on the link isn't necessarily going to bring everything up. The OPAC doesn't know that the hyperlink "Tiger" and the previous authorized form "Tigers" should be considered the same thing.

There's one thing about authorities that bothers me, though. When I was in library school, one of the touted benefits of using authority records was their cross-references. If a user doesn't know that the authorized form used by their library happens to be "Airships" and not "Blimps," the cross-references are supposed to help them find the records they're looking for anyway.  The problem is that this assumes that users are doing browse searches. Anecdotal evidence (and quite possibly actual studies, which I haven't tried looking up) says that this isn't true. Instead, users, including a lot of librarians, are probably using keyword searches. True, they may be subject keyword or author keyword searches, but they're still keyword searches and, as far as I know, there is no ILS out there that searches cross-references in authority records in addition to text within bibliographic records. I had heard that SirsiDynix Symphony does somewhat, but, from what I can tell, "somewhat" means that, if the keyword search retrieves nothing, users are redirected to a browse search for that word. That can work well enough in some cases. If users don't automatically assume that the redirection is a completely failed search and actually click on the cross-reference hyperlink. And only if the keyword search retrieves absolutely nothing.

It would be nice if the cross-references of any authority record to which headings in a bibliographic record are linked were searched in subject/author/genre keyword searches (maybe even general keyword searches). If an ILS exists that can do this, I'd love to hear about it. And I'd like to know why more don't.

Friday, November 12, 2010

Current big global editing project

I'm halfway through my current large global editing project that is cleaning up the name headings (flipping subfield q and d so that they're in the correct order), deleting obsolete subfield w's in access points, and fixing obsolete indicators in several fields (100, 700, 110, 710, 260) in our oldest records. I looked at the numbers, and I think it'll take 10 more days of work to finish the whole project up. Not bad.

After this project is done, I think I'll go back to concentrating more on straightening up our authority records and name and subject headings - a never-ending job.

While I was doing some subfield q and d flipping, it occurred to me that the technique I was using could be used to fix other problems we have. Since the technique took a bit of work and a lot of testing for me to figure out in the first place, and since every step must be done in a particular order, I decided to save myself future pain by posting instructions, complete with screenshots, in our staff wiki. That'll keep me from having the reinvent the wheel when I finally get around to doing those other fixes.

Monday, November 8, 2010

Fun with GIMP

I figured it was time for a non-cataloging related post.

On the left is an image I recently edited nearly to death in GIMP.

And here's the image as it was before I GIMPified it.

The edited image is made up of 5 layers (actually 6, but only 5 of them make up the visible parts of the image - I kept the original image as a background layer that I could copy in order to create additional layers).

Originally, I tried using the "cartoon" filter to create the black lines I wanted, but I didn't entirely like the results and the filter, however nice, didn't give me enough control. Since I'm still limited almost entirely to using a touchpad, I don't have much fine editing ability, either.

I created an effect similar to the cartoon filter by copying the original image and applying the photocopy filter. Then I selected according to color and selected all the true black areas of the layer with the photocopy filter applied. I inverted the selection, cut everything that was selected, and then made the selected area transparent. I repeated those steps with another layer with the photocopy filter applied, only this time I selected gray areas. I repeated the steps again for another gray. For all those layers, I made the remaining ares of color (the lines leftover from the photocopy filter) completely black, either with levels or with the colorify tool.

Then I decided to mess with color. I may not wear them, but I love bright colors, so I created a new layer and used the Color Balance tool until I got something I liked. However, I only really liked it on my shirt, so I deleted and made transparent every part of that layer but the shirt.

I still wanted to punch up the rest of the colors in the picture, though, so I created another layer and used, I think, the Hue-Saturation tool until I got something I liked. I thought I'd end up doing the walls separately from my face, but I ended up liking that particular color effect on both areas. However, my face had gotten a bit patchy-looking, and I wanted to smooth that out. I tried out a few tools but ended up liking the Oilify filter the best.

I didn't entirely like the hard lines (the result of the stuff I did with the photocopy layers) along my jawline, some areas near my mouth, and on my neck, so I used the eraser tool to get rid of those. I can do that much, even with a touchpad.

And that's basically how I did that image. It's nice to know that I can still use GIMP a little, even with a touchpad - there are just a few limits to what I can do. Drawing in GIMP, no, but editing a photograph? That I can do.

Also: yes, my NaNoWriMo novel is not going well. As has happened every time I've taken part, my writing speed has tanked. I'm hoping I can get it back up again - there are still several weeks left in the month.

Tuesday, October 26, 2010

More deduping, plus an explanation of why it is necessary

I did 6 or 7 more deduplication tests using MARCEdit, with no true success, but a little minor success. On the plus side, I can produce an overzealous list of duplicate records that includes true duplicates and a few that only look like duplicates (for example, same title, but one is a newer edition than the other). That at least gives us a list to work from, I suppose, although matching on ISBN would give a more accurate and probably more complete list.

In case you're wondering (I know this and my last post are somewhat technical), duplicate records are records that are basically for the exact same title - it was published by the same publisher, published on the same year, etc. When we get e-book record files from vendors, we sometimes get records for the same title from multiple vendors. Some of the vendors have records with OCLC numbers in them, some don't, and sometimes they might have OCLC numbers in them but not the same ones that another vendor used (yes, OCLC has duplicate records, lots and lots of them). When we load them, we end up with multiple records for basically the same thing. Ideally, we'd like to have an e-book that is available from multiple vendors accessible on one record.

That's where record deduplication comes in. Right now, we could do our deduping by searching each and every e-book title in the catalog and clearing up duplicates as we come across them. This is not a good idea - we have tens of thousands of e-books, and the number will only grow. The tests I've been doing are part of an attempt to automate deduplication, or at least come up with a list of potential duplicates so that we could avoid having to search every single title in our e-book collection.

I think I'm going to start reading articles on record deduplication. I probably should have done this earlier - if I find something right away that could help us, I'm going to kick myself.

Trying to dedupe records...and failing

While provider neutral e-book records are a nice idea, it's a little hard to do in practice when you're dealing with vendor e-book record packages. Today will be Round 3 of me trying to figure out how to dedupe our records without having to go through each title one by one.

In theory, deduplication could be done at the record loading stage, using, for instance, ISBNs as a second match point. In practice, this probably wouldn't go well, unless we decided to have print and electronic formats on one record - by matching on ISBN, we would end up matching our e-book records with our print records. There are probably other issues with this method that I haven't even thought of. I could do some testing, but I haven't really focused much on this method of record deduplication yet.

Instead, I've mostly been looking at methods of deduplication using MARCEdit. The obvious method, using MARCEdit's deduplication tool and trying to dedupe on ISBNs, has so far failed. I'm either using the tool wrong, or it's not working the way it should. The first day I started experimenting, I remember having some success by matching on main title information. I think I might try that again today. Unfortunately, that would result in multiple editions of one title being considered dupes. If it also lists actual dupes, it would still be better than nothing. Instead of having to search hundreds of titles, maybe we'd only have to search a few dozen. Or so I hope...

Friday, October 8, 2010

Vacation, catalog maintenance

Wow, it's been almost a month and a half since my last post.  My vacation had a little to do with that, but the rest was just...pre-vacation near burn-out, maybe?

My vacation went great. It took me a while to get comfortable with my niece, since I've never really been around babies before, but now I find I feel sad that I won't get to see her very often. At the very least, everyone in her family but her mom and dad is going to miss out on her first birthday - so sad!

Being back at work feels a little weird, but that'll wear off. With SCUUG only a week away, I've been reminding myself how to use MARCEdit for catalog maintenance by working on a project I started looking into right before my vacation. An unknown number of name headings in our catalog are messed up, with subfield d coming before subfield q, rather than after. I had been ignoring this problem, but now it's starting to interfere pretty significantly with my batch authority searching and loading process.

An example of the problem:
Babcock, C. J. $d 1894- $q (Clarence Joseph),

Should be:
Bacock, C. J. $q (Clarence Joseph), $d 1894-

In the past, I occasionally fixed these by hand as I came across them. However, this is annoying, and also bad for my wrist. Global editing is a good thing, and this looked like something that should be fixable globally. I just wasn't sure how.

It turns out it's possible with MARCEdit, and I figured out how to do it all on my own. Woohoo! I'm planning on running the fix for all the oldest records in our catalog (nearly 200,000 I think) over the course of a few weeks. That should take care of most, if not all, of the problem, and then I can get back to batch searching and loading authority records.

While playing with all of that, I also learned the first few steps for a new tool in MARCEdit that allows you to extract certain records from a larger file, edit the smaller file of records, and (in theory) re-insert the edited records back into the larger file. This will be great for all kinds of projects, once I figure out how to get the reinsertion part to work.