Advertise with Googlier.com The Open Library Blog https://blog.openlibrary.org A web page for every book Sun, 23 Aug 2026 22:54:30 +0000 en-US hourly 1 https://wordpress.org/?v=6.8.9 https://blog.openlibrary.org/files/2016/02/OL-logo.jpg The Open Library Blog https://blog.openlibrary.org 32 32 Building a Library Patrons Choose to Revisit https://blog.openlibrary.org/2026/08/23/building-a-library-patrons-choose-to-revisit/ https://blog.openlibrary.org/2026/08/23/building-a-library-patrons-choose-to-revisit/#respond Sun, 23 Aug 2026 17:44:14 +0000 https://blog.openlibrary.org/?p=2707 forward by Mek, Open Library’s program lead:
This year, through Google Summer of Code, the Open Library Team had the privilege of collaborating with Tanishq Sangwan to help readers avoid dead ends and discover more of the books they’re looking for. This collaboration was particularly fruitful because of Tanishq’s strong product sense, his commitment to metrics, and his ability to move quickly from idea to prototype. Tanishq has an impressive ability to cover a lot of ground while consistently bringing thoughtful, ambitious, new ideas to the table. He’s also a clear and proactive communicator, making him an especially well rounded and effective collaborator. I’m excited for the world to see what Tanishq accomplished this summer, and even more excited for countless readers to experience the benefits of his work firsthand.

This year, Open Library published a call for Proposals to tackle one of its most persistent problems: patrons registering, borrowing (or trying to), and disappearing. My proposal, “Personalizing the Patron Experience,” set out to fix that drop off by giving new patrons a taste-driven onboarding flow and a personalized dashboard to land on.

Going in, we thought the gap was mostly about personalization – that patrons weren’t returning because there was nothing tailored to bring them back. Once we looked deeper into the real funnel, instead of assumptions, the picture shifted: a large share of patrons were dropping off way earlier, before personalization could even matter. 

Why now?

In 2023, more than 500,000 books were removed from the lending library. That single event created a lasting mismatch between the value patrons expect from Open Library and the value they are able to access when they show up. A library catalog matters so long as it is useful every time someone visits it – and for a growing share of visits, it wasn’t.

My proposal, built with feedback from my mentors, focused on the post-registration onboarding experience: the sequence of moments right after a patron signs up, when they either find a reason to stay or quietly leave.

Finding a Deeper Problem: A Drop-off Funnel Analysis

Our initial hypothesis was that focusing on the post-registration onboarding experience was the best point of intervention to help patrons connect with more value. Before racing forward and writing code based on this assumption, we evaluated our hypothesis against analytics about the current journey of Open Library patrons.

Our findings from this research helped us shape a revised plan that would ultimately lay groundwork to make our onboarding work as impactful as possible:

1. Patrons fail to find a readable book more often than they succeed.

Analytics show that for every successful “Borrow” click on Open Library, patrons clicked “Locate” twice, attempting to access a book that wasn’t available to borrow. That means patrons who came looking for a specific book were unable to read it roughly two-thirds of the time. This is a problem we needed to address first because no amount of post-registration onboarding will compensate for an initial experience that fails to connect most patrons with a book that is useful to them.

2. Registration itself works against onboarding. 

Before patrons could begin an onboarding flow, they first had to make it through a complicated, high-friction, confusing registration process that brought them to archive.org for activation. Just as bad, the registration process did not remember or preserve the reason a patron registered an account to begin with. In the 1/3 of cases where patrons found a book of interest, they would have to survive a confusing registration process and remember why the registered before accessing that value.

3. Day zero is a dead end. 

After registering, a patron lands on a virtually empty “My Books” page. It’s a ghost-town: There’s nothing to do, nothing to return to, and, unsurprisingly, very few patrons find reason to re-engage after that first visit. This is the crux of the retention problem: onboarding doesn’t fail at signup, it fails in the days immediately after.

Starting with the ability to measure outcomes

Before building any patron-facing feature, we needed a way to actually know whether what we shipped was having a positive impact – not just a hunch that it probably helped. So the first thing built wasn’t a UI feature at all: it was a micro A/B testing framework (#12789) to give the team a way to run experiments and get data-backed answer without guessing.

Helping more patrons succeed

In a library where two-thirds of book clicks are leading to dead ends, the immediate question shouldn’t be how we can optimize for the minority of cases that succeed, it’s: how can we flip the calculus so the majority of patrons succeed. In service of helping patrons succeed in finding books of value to them, we brainstormed four solutions.

  1. Suitable alternatives banner (#12743): For years, book pages on Open Library have had “related books” carousels, buried below the fold. We designed an intervention where, when a patron lands on a book page with no read options, they’re shown a banner with a link to explore similar books that are available to read now.
  2. Offering patrons a “Search Inside” option alongside “Preview” (#12881): Roughly a million book pages on Open Library are preview-only, meaning they cannot be borrowed but patrons can explore a sample of the book. Each of these previews include a “search inside” feature that is useful for researchers, but is tucked away within the Bookreader UI. We prototyped a solution that shows an interactive “Search Inside” button next to the existing “Preview” button. Our hypothesis is, if more patrons knew about the search-inside functionality, then the sum of previews and search-inside actions would be greater than just having the preview button.
  3. Fix the confusing “Locate” button: Several years ago, the Open Library team used to show a grey “Unavailable” button when a book was not readable. Staff hypothesized that it may be more useful to replace this with a “Locate” button that brings patrons to Worldcat where they could see if the book of interest was available in a nearby library. After making this change, many patrons opened support tickets expression confusion at being brought to another site that asked them to register. Furthermore, many international patrons reported that Worldcat had no availability options that were relevant to them. As a result, we decided to replace the “Locate” button with a “Check options” button that brings them to the Open Library book page, and then to show an explicit “Worldcat” button in case the patron wants to check nearby libraries. The benefit of this change is that more patrons will see the “Suitable alternatives” banner and have more access options to choose from, such as buying the book should they choose. We currently have a prototype of this change and are in the process of collecting feedback from the community.
  4. Consolidating “Buy” options (#13113, #12914): The sidebar was scattered, with multiple CTAs for books. These were unified into a single, clean “Buy” dropdown so the interested patrons can easily access all purchase options, while uninterested patrons aren’t overwhelmed by this information.

Fixing registration so every success counts

For many years, after a patron signed-up, they would receive an activation email that redirected them to archive.org and then required them to return to Open Library and then login. This process was confusing because it spanned two domains. During our collaboration, Mek switched this flow so that activation emails brought you directly to Open Library and automatically logged you in. But what then?

Prior to my work, the registration process would forget the action the patron was trying to accomplish, making success less likely. To address this problem, I designed a banner (#9409) that remembers a patron’s registration intent and allows them to seamlessly continue after they’ve registered. Since implementing this feature, more than 10% of registered patrons choose to use this to continue where they left off.

Giving patrons reasons to return

As a result of the above interventions, we expect our metrics to show a greater percentage of patrons are succeeding to connect with book they love — whether that’s taking advantage of search inside, finding suitable alternative books when their first choice is unavailable, or having a registration process that works with them instead of against them.

The final step (and our initial goal for GSoC) is to give patrons reasons to return. This means every time a patron visits their “My Books” page, they should encounter value. We outlined two efforts that dovetail with our original onboarding plans:

  1. The “Continue Reading” carousel: a consolidated view of books the patron may wish to continue reading.
  2. The “Build Your Library” section: a collection of ways for patrons to add books to their reading log.

We were able to complete the “Continue Reading” initiative during the course of GSoC. We produced specifications and mockups for the “Build Your Library” phase, which we plan to continue working on as follow-ups.

The “Continue Reading” carousel (#13256)

Over the past several years, the Internet Archive has moved from a default two-week loan period towards more flexible and equitable models that allow patrons to borrow book for the duration of session, so long as the book is being actively used. On one hand, this means fewer patrons are leaving books sitting unused on their digital desks. On the other hand, it means going to your “My Books” page is less likely to show active loans to continue reading. To address this problem, we’ve been building a “Continue Reading” carousel with two phases:

  1. Phase I: Merge Active Loans & Loan History (#13272).
    • Because many loans expire after about an hour and then vanish into a separate history table, this phase makes it so that both your active loans and your loan history get merged into a single coherent view, making it easier for you to continue reading where you left off.
  2. Phase II: Enhance “Continue Reading” by adding open-access books (#13273, #13274).
    • One gap of our current system is that open-access books don’t get recorded under loan history and so never show up under your My Books page. Similar to our “recent searches” history, we’re first using local storage and then exploring a dedicated database table to help patrons keep track of the open access books they’ve began reading. We intend for these settings to be manageable via the loan / reading history page.

Next steps: The “Build Your Library” section (#13255)

After GSoC, we plan to continue development on the second phase of improvements for the “My Books” page, to give patrons clear calls-to-action and new experiences for adding books to their library. The goal it to give patrons the tools to, “Make Open Library Your Library“. Rather than designing a blocking onboarding flow that is only done once or even skipped, we decided to try adding a section to the My Books page that increases discovery of our import options (Goodreads, barcode scanner, search) so that patrons can add books any time they visit their library. As more patrons use these import options and grow their reading logs, their activity contributes towards book discovery for everyone. The library grows not just through its books, but through the collective reading experience of its readers.

The idea of the new “Build Your Library” section is to meet patrons where they already are, based on which of three groups they fall into, and give each a seamless action to bring their reading into Open Library at once rather than book by book:

  • New readers: a flow to explore genres and bulk-add books to shelves and lists directly from that exploration, instead of searching for each book individually, opening its page, and adding it one at a time.
  • Readers with a library elsewhere: the import tool already exists but is underutilized. This means improving that experience and expanding it to cover other popular platforms, so a patron can export from wherever they already track their reading and import straight into Open Library.
  • Readers with a physical collection: the barcode scanner already exists but is underutilized too. This means improving it for bulk scanning, so a patron can log a shelf of physical books in one sitting instead of one at a time.

How patrons discover what to add next, how shelves get organized and recommended, and how the page evolves as a patron’s reading history grows are all part of this epic too. This is the epic I’d point a future contributor toward first.

This unlocks two things down the line. Once patrons have a real library instead of an empty shelf, there’s finally enough signal to build genuine recommendations – curated to taste, not generic – since bulk-adding from these three flows is what generates that signal in the first place. It also revives the existing activity feed, which is underused today for the same reason: it only has something to show once patrons are active. As more patrons bulk-add their libraries, the feed becomes more valuable to the whole library.

Impact

Already, several of our interventions have improved the experience of thousands of patrons across the Open Library.

  • Preserve Intent banner: ~500 clicks/day against total shown, a ~12% click-through rate – patrons redirected through registration are following the banner back to what they were originally trying to do.
  • Preview / Search Inside visibility: ~13,800 previews and ~1,100 Search Inside opens per day since the buttons became more prominent.
  • Unavailable-book alternatives banner: ~2,500 clicks/day, patrons routing to a readable alternative instead of a dead end.

Having fun along the way

One of the parts I enjoyed about working on Open Library was the autonomy to work on features that I thought were useful. One example was the “Stopped Reading” bookshelf (#12400) that many patrons had been requesting for years. There was another case where I took the initiative to improve the design the highly visited “My Loans” page (#12912) to make it easier for patrons to navigate their books.

What I’ve learned

The instinct going in was to treat this as a personalization problem – better empty states, taste capture, recommendations. Open Library already has plenty of features like that, loved by a smaller set of patrons – it’s easy to keep adding to that pile. The funnel numbers pushed me to stop designing from “what’s a good addition” and start from the patron’s actual seat: what’s stopping them from getting anywhere at all? That shift mattered more than any single feature – it’s why the summer became about removing friction patrons were already hitting, not adding something new to discover.

Looking back, a lot of the highest-impact work shared the same idea: don’t ask patrons to do something new, save what they’re already doing. Preserving intent through registration (#9409) means a patron who clicked “borrow” doesn’t have to re-find that book after signing up – we just remember what they were already trying to do. The Continue Reading epic (#13256) applies the same idea to reading itself, picking up on what a patron’s already reading instead of asking them to curate a shelf. Neither asked patrons to learn something new – they just stopped losing track of what was already in motion.

I also came away with a much better appreciation for measuring before shipping. We didn’t get to actually run experiments through the A/B testing architecture yet, but having it in place – ready to give real answers instead of hunches – is something I’m genuinely excited to put to use.

Finally, I’m proud that we we’ve been able to move quickly and prototype many features that span the patron experience, however, with a large open source project, it’s been a learning that we can only move as quickly as the review process and the feedback we’re able to get from the community. There are several efforts are in review or mid-flight, worth tracking post-GSoC:

  • #13113 – Consolidating buy options (open)
  • #12914 – Locate → Check WorldCat options (open)
  • #13281, #13282, #13283 – the three Continue Reading phases (all open)

None of this ends with GSoC. Both epics – Continue Reading and Build Your Library – are far from done, and I plan to keep working on them past the program, seeing the remaining phases through rather than leaving them as a hand-off.

Acknowledgments

Thank you to my mentor, Mek Karpeles, and the Open Library staff and community – they’ve been extremely helpful and supportive throughout, even before GSoC. The drop-off funnel analysis that shaped nearly every decision in this project came out of that early framing, and the steady feedback across a summer of shipping into a live, high-traffic library catalog made all of it possible.

A special thanks to Lokesh Dhakar and Ray Berger, who encouraged and supported me in getting started with Open Library back in February – which eventually led to spending the whole summer building with the team.

]]>
https://blog.openlibrary.org/2026/08/23/building-a-library-patrons-choose-to-revisit/feed/ 0
Helping Patrons Discover Books https://blog.openlibrary.org/2026/08/23/helping-patrons-discover-books/ https://blog.openlibrary.org/2026/08/23/helping-patrons-discover-books/#respond Sun, 23 Aug 2026 13:52:56 +0000 https://blog.openlibrary.org/?p=2697 A forward by Mek, Open Library’s program lead:
This year, the Open Library team and I were fortunate to collaborate with Chisom, as part of Google Summer of Code, to make millions of books more discoverable to readers. Chisom entered this year’s Google Summer of Code program a motivated and capable software developer and continued to impress us with her focus, proactivity, and problem solving. It was a joy working with Chisom and rewarding to witness her make consistent forward progress, rise to growth opportunities, and — as a result of hard work and initiative — achieve an excellent outcome that I believe will benefit millions of Open Library patrons. I encourage you explore how she strategically approached this challenge and created a general purpose tool and patterns that will allow others to continue her work into the future.

My name is Chisom Nnamani, and this summer I had the opportunity to join the Internet Archive’s Open Library team as a Google Summer of Code (GSoC) contributor. This was my first experience contributing to a large open-source project, and I could not have asked for a better place to start. As someone who cares deeply about making books accessible to everyone, I was drawn to Open Library’s mission of providing a free webpage where anyone, anywhere can discover and access published works. You can view my initial proposal here.

My GSoC project focused on a problem that sounds simple at first: helping readers find books by genre.

Millions of books available, the challenge is discovery

Over the years, Open Library has accumulated millions of free-form book labels that have never gone through a standardization process. As a result, a book about science fiction might be labeled “science fiction,” “science-fiction,” “sci-fi,” “scifi,” or dozens of other variations. Without a way to recognize and merge these synonyms, books become scattered across hundreds of different labels like a needle in a hundred haystacks.

Search termResults returnedBooks missed vs. best
“Science Fiction”17,900 hits— (baseline)
“science-fiction”16,497 hits~1,421 books
“sci-Fi”2,721 hits~15,179 books

The result shows how a reader can miss hundreds or even thousands of books simply by using a different term to describe the same genre. The books are there. The problem is that the catalogue does not always connect them.

Messy labeling means messy recommendations

The problem goes deeper than the search box. For years, Open Library has used a subjects field to describe what a book is about. These subjects are stored as plain text strings in a flat list, with no consistent rules about how they should be written or organized.

As a result, different kinds of information can end up sitting side by side. A book can have its genre, characters, places, themes, and other descriptions all represented as separate subject strings. There is no structure telling the catalogue how these descriptions relate to one another.

The screenshots below show what this looks like on real book pages. The Hobbit has “Fantasy,” “Fantasy fiction,” and “Juvenile fantasy fiction” as separate tags on the same page:

Figure 1: The Hobbit’s subjects include multiple variants of “fantasy” with no major classification.

The screenshot below shows another example: And Then Were None has over 30 subject tags with no structure distinguishing genre, subgenre, language, audience, character, and other types of information. You can also see multiple variations of “mystery,” including “Mystery fiction,” “Mystery & Detective,” and “Fiction, mystery & detective, general.”

Figure 2: And Then Were None — 30+ subjects with no type distinction between genre, subgenre, language, audience, character.

Together, these examples revealed the larger problem I wanted to address. Open Library had a huge amount of useful information about its books, but it lacked a consistent structure for connecting related genres and descriptions.

That became the starting point for my GSoC project: building a more structured way for Open Library to describe books by genre and subgenre, and eventually using that structure to make browsing and discovery better for readers.

The recipe for organizing 860K books

When I began GSoC, the Open Library team had already identified several high-impact label categories (tag types) that could benefit from this kind of cleanup, including genres and subgenres, audiences, content warnings, and formats. What we didn’t yet have were mappings from our existing messy labels to these new, cleaner categories, or a common software framework that contributors could use to define these mappings and perform the cleanup.

I began by working with genres and subgenres. The first step was to define the categories we wanted to recognize. We then needed to connect the many ways these concepts already appeared in Open Library’s catalogue to a consistent set of canonical labels – the standardized labels we want those variations to map to. For example, different subject descriptions might refer to the same genre using slightly different wording or formatting.

From there, I worked on expanding and refining the mappings so that more of the catalogue’s existing subject descriptions could be connected to the appropriate genres and subgenres.

But the goal was not to build something that only worked for genres. As the project evolved, we built a common core that could support different types of labels through the same process: define a vocabulary, create mappings, analyze existing data, and eventually migrate the cleaned information back into Open Library.

This separation between the tag type and the shared tooling became an important part of the project. Genres and subgenres were the first categories I worked on, but the same framework can be used for other categories as contributors begin cleaning and structuring them.

Here’s what that mapping looks like for a few genres:

Existing subject strings on Open LibraryCanonical genre
“Fantasy”, “fantasy fiction”, “Juvenile fantasy fiction”Fantasy
“Mystery fiction,” “Mystery & Detective,” “Fiction, mystery & detective, general”Mystery
“Science fiction,” “science-fiction,” “sci-fi,” “Science Fiction Literature”Science Fiction

Several different subject strings can now point to the same canonical genre.

Once these mappings were in place, the next step was to give each canonical label a structured representation in Open Library. In the Tags project, a Tag is an Open Library data object representing a defined label, such as a genre or subgenre. This gives the canonical concept its own consistent identity instead of treating every variation of a subject string as a separate concept.

I then created the canonical Tags for the genres and subgenres we had defined.

Fantasy, now represented as a genre Tag in Open Library, rather a plaintext subject string.

Steampunk, a subgenre represented as a Tag.

With the genre and subgenre labels defined and represented as Tags in Open Library, the next challenge was connecting them to the millions of existing works in the catalogue.

Open Library already had millions of works with years of existing metadata. I could not simply assign these new Tags manually to every work. The next challenge was figuring out how to connect the information that was already there to this new structure, and then safely apply those connections across the catalogue.

So I built migration tooling that could analyze existing subjects, identify matches using the genre and subgenre mappings, and connect those matches to the appropriate Tag keys on each work.

Before thinking about millions of records, I first needed to understand what the migration would actually find. I ran the matching process against Open Library’s April data dump, a monthly snapshot of the catalogue’s data, and found 869,461 works with genre matches and 51,526 works with subgenre matches.

Those numbers changed the way I thought about the project. This was no longer just about creating a better vocabulary. It was about applying that vocabulary across millions of works while making sure the information already there was not accidentally changed or lost.

I worked on the migration scripts, the shared utilities behind them, and the changes to Open Library’s work schema needed to store the new genre information. I also validated the migration on a smaller pilot before moving toward the production run.

Safely running a large-scale migration

One of my biggest lessons from this project was that writing the code is only part of the job. When you are changing a large, live system, you have to think about what happens when the code actually runs.

  • What happens if something fails halfway through?
  • How do you know the migration did what you expected?
  • How do you avoid changing records that should not be changed?
  • How do you test an operation that will eventually touch hundreds of thousands of works?

I used dry runs and small pilots before larger operations. I added ways to track progress and designed the migration so that it could be run in controlled batches. Along the way, I also encountered some of the less glamorous parts of working with production systems, from authentication and request limits to unexpected differences in the data itself.

One of my favourite lessons from the project is that production engineering requires trust.

Before you can make a change at scale, you have to earn the right to trust your own tools.

Translating better data into better discovery

In our GSoC project, fixing book labels was always a means to an end: improving how readers discover books. Many patrons come to the Open Library looking for a specific book. But not every reader arrives knowing exactly what they want to read next. With consistent genre data in place, we could begin to ask a different question:

What if readers could browse and discover books by genre instead of having to already have a book in mind?

Search and filtering can help when you already know what you are looking for. But discovery is different. Sometimes you just want to browse.

That question became the idea behind Genre Explorer.

Taking inspiration from Drini Cami’s Library Explorer – a system that uses Dewy Decimal classification numbers to digitally emulate the organized bookshelves of a physical library – we imagined a more visual way for readers to explore the Open Library’s book catalogue, using the same genre and subgenre structure I was building for the tagging project. Instead of presenting genres as another long list of links, we imagined something closer to the experience of walking into a bookstore.

Genres could act as bookcases.
Subgenres could become shelves.

A reader could choose a genre, step inside it, explore its subgenres, and discover books along the way. The idea was to make genre browsing feel less like searching through metadata and more like browsing a library. I developed the initial concept and built a clickable prototype to explore how this experience could work.

From there, Mek and I continued developing the idea together. We reviewed the experience, explored how it could fit with Open Library’s existing components, and refined the concept into an interactive version now available on the testing site.The interactive version follows the same idea: genres act as bookcases, and entering a genre reveals its subgenres as shelves.

What I find most exciting about Genre Explorer is that it grew out of the original tagging problem, but takes the idea one step further. The canonical Tags give Open Library a consistent way to describe books. That structure can support better search and filtering, while Genre Explorer explores what it could look like when the same information is used to help readers browse.

It was an unexpected direction for my GSoC project. I came in focused on the data and infrastructure behind genre information. Along the way, I started thinking beyond how books are described to how that work could become something a reader actually experiences.

Takeaways

When I started this project, I expected to learn more about software engineering. I did, but not always in the ways I expected.

One of my biggest lessons was learning to slow down and understand a system before trying to change it. I learned to look at messy data and find the patterns hidden inside it, to test my assumptions against real examples, and to treat small experiments as part of the engineering process rather than as steps before the “real” work begins.

I also learned to think beyond the implementation. Throughout the project, I kept coming back to a simple question: Does this actually make the experience better for the person using it? That question shaped how I thought about the tagging system, and eventually led to the idea of Genre Explorer. It reminded me that good engineering is not only about building something that works. It’s also about understanding if and why something should exist in the first place.

Working with Open Library also gave me my first real experience contributing to a large open-source project. I had to learn how to navigate an unfamiliar codebase, communicate ideas clearly, ask questions when I was unsure, respond to feedback, and make decisions when there was no obvious answer. I was not doing this work in isolation. My mentor, Mek, pushed me to think beyond the code and focus on the larger problem we were trying to solve. Open Library contributors and maintainers, including Jim, Drini, Liz, and Katrina helped me understand different parts of the systems I was working with. Every review, discussion, and debugging session became part of the learning process.

Looking back, I think that may be one of the most valuable things I am taking away from GSoC: learning how to become useful in a system that existed long before I arrived.

Next steps

By the end of GSoC, we were able to add genre tags to more than 50,000 works. We also built common infrastructure to standardize the tag migration process and enable others to contribute to the greater cleanup process.

The next stage is to extend this process to add subgenre tags to works and to index these genre and subgenre tags in Open Library’s search engine, so readers can find and explore books by genre.

Genres and subgenres are only the beginning. Open Library has other high-impact label categories that will benefit from the same approach, including audiences, moods, content warnings, and content formats. Because we built this project to have a shared core for defining vocabularies, creating mappings, analyzing existing data, and migrating cleaned information, future contributors can use the same tooling to work on these categories rather than building a new system from scratch.

This shared core is the legacy I hope will last beyond this GSoC project: not just cleaner genre and subgenre data, but a reusable foundation that makes it easier for Open Library and its contributors to continue turning messy catalogue labels into structured information that can improve how readers discover books.

]]>
https://blog.openlibrary.org/2026/08/23/helping-patrons-discover-books/feed/ 0
Google Summer of Code Contributors Improve Open Library’s Patron Experience https://blog.openlibrary.org/2026/07/15/google-summer-of-code-contributors-improve-open-librarys-patron-experience/ Wed, 15 Jul 2026 19:53:28 +0000 https://blog.openlibrary.org/?p=2684 This year, as part of Google Summer of Code (GSoC), the Internet Archive is collaborating with two outstanding contributors to make it easier for patrons to find relevant books on Open Library.

Tanishq Sangwan, a 19-year-old from Gurugram, India, and Chisom Nnamani of Lagos, Nigeria, are two of 1,141 software developers from around the world who have been selected through GSoC to hone their engineering capabilities with open-source organizations.

“The Internet Archive’s focus for 2026 is: tools for participation. Participation must be earned by building an experience patrons want to return to.” said Mek, program lead for Open Library. “The work that both Chisom and Tanishq are doing is central to creating a more reliable Open Library experience, where readers can repeatedly discover, access, and enjoy books.”

Sangwan, who just completed his second year of college studying artificial intelligence, is focusing on the journey of patrons who join Open Library and use it once, but do not return. His objective is to understand where patrons encounter barriers and identify opportunities to create a more useful, lasting experience.

Sangwan, who just completed his second year of college studying artificial intelligence, is focusing on the journey of patrons who join Open Library and use it once, but do not return. His objective is to understand where patrons encounter barriers and identify opportunities to create a more useful, lasting experience.

Sangwan comes into the project with two years of experience from ZNotes, an educational organization that provides free notes and videos from students around the world.

“I built this passion and got this amazing feeling when my work was making an impact on people and they were receiving some value,” he said. “Now, Google Summer of Code is a wonderful opportunity to connect with open-source organizations and Open Library where its work directly impacts people’s lives.”

“This year’s collaboration is important because lots of patrons discover Open Library, but too often don’t always find the books they want,” Mek said.

The work begins when a patron lands on a book that is unavailable for reading. Soon, instead of reaching a dead end, patrons will be presented with nearby books on the same shelf that are available now.

When patrons find a relevant, available book, a simpler registration process will help them get started with fewer steps and return to the book they found. Furthermore, Sangwan is helping patrons connect with new book recommendations on an ongoing basis by introducing an activity feed to the account page.

“I like to hear about the patron psychology, how they’re interacting with the platform, and what’s going in their mind from the first moment to the very last,” he said. “We’ll be researching and conducting interviews with lost patrons so we can connect this bridge between patrons and the millions of books in our catalog.”

Chisom Nnamani
Chisom Nnamani

Nnamani, who already has certifications in Data Analytics and Data Engineering, just completed her sophomore year pursuing a second degree in Computer Science. At Open Library, Nnamani is leading a major cleanup effort to add structured tag data to books so they can be searched by genre and subgenre. Her work is paving the way for the addition of a wide variety of new searchable tags, including: moods, fiction and non-fiction, content warnings, and literary formats, such as memoirs, biographies, and more. By cleaning up messy data and enabling better genre and subject browsing, Chisom is helping remove barriers preventing patrons from discovering books they love.

“Today, many of our subject pages feel computer generated and can’t compare to the beautiful, curated experiences you find at small book stores,” says Mek. “The work Chisom is leading to map the messy subject tags we have to clear genres and subgenres will help us offer patrons a more useful and satisfying browsing experience.”

“I really care about books being accessible to people,” Nnamani said. “In Nigeria, we have limited access to physical libraries, so Open Library is something that matters. It gives everyone the opportunity to come and read any kind of book and gain insights.”

“This project is inspiring to me,” Nnamani said. “I like to work on projects where I can connect the data and infrastructure in ways that contribute to the organization’s goals. I enjoy helping to solve complex problems at the intersection of systems and data.” 

Open Library Fellows work remotely, but meet regularly online with Internet Archive staff and mentors. At the end of the summer, each contributor will publish a blog post explaining their technical journey and experience gained.

Since 2005, Google’s Summer of Code has supported more than 23,000 students from 123 countries with stipends to receive mentorship and contribute 48 million lines of code to over 1,000 open source organizations worldwide.

]]>
Security Incident Disclosure https://blog.openlibrary.org/2026/04/28/security-incident-disclosure-2024-04-28/ Tue, 28 Apr 2026 22:42:18 +0000 https://blog.openlibrary.org/?p=2676 Early, around 7:30am Pacific, on Tuesday, April 28th, high database load was detected on OpenLibrary.org. Investigation revealed a set of at least 38,703 residential IP addresses performing a coordinated sqlinjection attack on a vulnerable openlibrary.org endpoint, resulting in exfiltration of emails and encrypted passwords of 175,080 legacy accounts, registered before March, 2011. This table has not been used for authentication since 2016, however we advise affected accounts to change their passwords on any relevant platforms.

The attack was identified and mitigated within a four hour window. Impact was limited due to the obscurity of the attack which could only process a single account query per malicious request. This is an old, no-longer-in-use table that was formerly used for Open Library sign-in prior to switching to use Archive.org login credentials in 2016.

Details

Prior to 2016, Open Library maintained its own login system distinct from Archive.org, which used a legacy account database table. In 2016, both for improved security and patron convenience, the Open Library website switched to a unified model where authentication is performed using archive.org credentials and not legacy Open Library credentials. Since this date no new Open Library patron account passwords have been stored within Open Library’s legacy account database.

Today’s incident only affects a subset of legacy accounts whose credentials are no longer in active use. Furthermore, no plaintext passwords were compromised – all passwords in this table were both salted and encrypted.

Remediation & Impact

Upon discovery, the identified exploited path was blocked at the nginx level and a security fix was then patch deployed to our servers. All accounts in the no-longer-in-use legacy `account` table have had their encrypted password fields cleared. We are releasing a tool to check whether your email was affected by the breach.

Check If Your Account is Affected

If your email is on this list, out of an abundance of concern, we recommend changing your password for any service that matches the password used when registering your OpenLibrary account.

Open Library’s Security Policy

The Open Library team routinely monitors security alerts, performs sqlinjection audits, and responds seriously to security reports we receive. We believe strongly in a full transparency policy when incidents occur, both so our patrons have the best information to make decisions, are able to understand our responses, and so our developer community can help report and address issues.

Followup

Followup details and actions will be updated via our post-mortem. Please feel warmly invited to direct questions and concerns to info@archive.org and report security vulnerabilities to security@archive.org and mek@archive.org.

Apologies & Gratitude

Thank you for your patience and understanding and our sincere apologies for the poor behavior of these malicious actors and the impact this has on our community. As AI tooling makes it easier for malicious actors to attack websites like ours, our team will continue to proactively take steps to put our patrons’ privacy and security first.

The Open Library Team

Mek, Drini, Jim, Lisa

]]>
Improving Search by 10% https://blog.openlibrary.org/2026/04/16/improving-search-by-10-percent/ Thu, 16 Apr 2026 04:18:15 +0000 https://blog.openlibrary.org/?p=2655 Today’s challenge is to find “The Secret of Secrets“, by Dan Brown using Open Library search. It’s not impossible, but it is not easy… And it’s not just because the Я is backwards on the cover.

If you search for “the secret of secrets“, you won’t find the right match on the first two pages of results.

If you search for “the secret of secrets dan brown“, the result is still 7th in the list.

In this example, our current search algorithm is biasing too heavily on returning book results that have lots of editions to vouch for them, as well as other boosting factors (like star ratings) that don’t always produce desired result.

What search algorithm would perform better? And how do know whether one approach is better than the other?

These are key questions Drini Cami — core maintainer of Open Library search — has been investigating this month.

Is Search Improving?

In order to know whether we’re making changes that improve the quality of our search results, we can’t just change the algorithm, type in a search, and see if the result is better in that one case. We need to apply some consistent framework across a collection of challenging queries and measure how the system — on a whole — performed before versus after.

In Open Library’s case, Drini maintains a Search Evaluation Spreadsheet that measures 100 common searches—everything from “Harry Potter” and “Little Prince” to “The Secret Garden” and “Narnia”. These searches come from our server logs (i.e. popular searches from patrons). It also contains challenging cases that we’ve seen underperform in the past.

For each search query we’ve collected, we define what we expect the “correct” search result to be and then check how often the correct result appears in the top 3 search results (across the search algorithms we’re considering).

Multiplicative Instead of Additive

For the technical crowd, this change (PR #12357) adjusts Work Search’s Solr eDisMax tuning to use multiplicative boosting (via boost) instead of additive boosting (via bf) to reduce over-weighting of popularity signals relative to textual relevance:

  • Replaces bf-based additive boosts with an eDismax boost function expression.
  • Expands/adjusts qf and phrase-boost parameters (pfpf2) to change match weighting and proximity scoring.

Overall, the change has improved the relevance of our test results by roughly 10%:

Early Anecdotes:

In our testing playground

  • “The secret of secrets” now appears as the 3rd result.
  • “My Life”, by Bill Clinton went from 101th to 19th
  • “laws field guide” went from 14th to 1st

Expect these improvements to be live on the main site early next week. Happy reading!

Future Opportunities: Exact Match versus Browser

Since releasing this blog post, we received a great question internally by Sawood Alam, from the Wayback Machine team who asks:

Have you measured how would this change affects the discovery use-case where a patron is not after one specific document in mind, but wants to find out what options are out there (as opposed to the lookup use-case where they already know what they are looking for)?

And this is indeed something the Open Library team has been considering. Sawood is pointing out that there are [at least] two modes of searching:

  1. by Exact match
  2. Browsing

One may browse in a variety of different ways, but for simplicity I’d like to refer to this as “searching by proxy”. That is to say, instead of searching for an exact book by title, a patron may endeavor to discover a suitable book by any number or combination of proximal qualities like author, topic, format, genre. An example is a search for nonfiction books about UFOs that are advertised as textbooks and published before 1950.

For browsing queries of this flavor, it’s difficult to know (in advance) what book(s) should appear in the top-three position in search results as the answer will often be subjective to the searcher.

As a result, an additional approach will need to be instrumented and added to our existing process that (a) accurately identifies when a search term is for a proximal quality rather than an exact (e.g. title) match and, (b) introduces secondary evaluation metric, such as:

  • Success Rate: How often any result is clicked — perhaps called Query Success Rate (QSR)
  • Relevance: When a result is clicked… how often is this click for a record in a top 3 position — something like Mean Reciprocal Rank (MRR)

Further research is required to understand how these metrics may be combined in a recipe that results in the best experience for patrons. For instance, [when] is it more important to improve relevance versus the general distribution of clicks? Maybe “better” means increasing the ratio of searches-to-clicks by 20% rather than increasing the number of clicks in the top-3 position by 25% (if search-to-clicks were to drop by 5%).

Suffice to say, as policy changes make it more challenging for some readers to find the exact book(s) they are looking for, it becomes increasingly meaningful to be able to suggest suitable alternatives — and be able to measure how effective we are at making relevant recommendations. Measuring browse cases is something we expect to work towards in the coming months.

]]>
Canada Reads Awards: My Journey Championing Canadian Content on Open Library  https://blog.openlibrary.org/2026/04/15/canada-reads-awards/ Wed, 15 Apr 2026 16:05:16 +0000 https://blog.openlibrary.org/?p=2647 By Catherine Gosztonyi

Hi! I’m Catherine, a curious and avid librarygoer, currently pursuing a career transition into the Library and Information Sciences field.

I joined the Open Library Librarian’s Team as a Librarian-In-Training during the summer of 2025 to learn about the world of online libraries, contribute to an open source project, and be part of a library community. My initial objective was to learn how to work with book metadata in MARC format but what I ended up learning, and what I’m still learning, has far surpassed my expectations. Before I took on this role, I didn’t know how to code in HTML or Markup (or any code, really). I had never contributed to an open-source project or submitted a GitHub ticket – and now I’ve done all three! I’ve also created a curated collection, tagged books by subject to further categorize them within the Open Library catalog, and learned how to create book carousels using the subject tags to visually display a collection of books on a page.

As I write about and reflect on my Open Library journey, I’m so proud of what I’ve learned and accomplished, and of the wonderful time I’m having engaging with the Open Library community.

Getting Started & Deciding How to Contribute

I discovered Open Library through the Internet Archive while researching volunteer opportunities within libraries. Open Library‘s mission resonated with me, and I felt called to join the community to help make knowledge free and accessible. I submitted my application and not long after, I was invited to the community Slack.

When I joined the Open Library Slack channel as a Librarian-In-Training, I was ready and excited to get started but realized I had no idea where or how to begin. I gained my bearings by reading through the instructional documentation and the Librarians In Training (LIT) Guide, letting the information guide me as I explored and familiarized myself with the website’s layout, behaviours, content, and pages. My interest piqued when I read the How to Create Curated Collections guide and saw how well the Star Wars and Star Trek collections had been curated. Inspired, I decided to contribute a curated collection of my own, though I was unsure what my collection would showcase. It was while going through the list of all the curated collections that I noticed a distinct lack of Canadian literary awards representation. As a Canadian, this was something I needed to remedy! I decided to highlight Canadian content on Open Library by creating a collection for the CBC’s Canada Reads Awards. This is a literary award I am very familiar with, having read many of the championed books, and one that I continuously use to inspire my own reading list.

About the CBC’s Canada Reads Awards

Canada Reads is a radio program broadcast yearly on CBC Radio One in which five celebrity judges champion and debate a book in the hope of it winning and being crowned as Canadian’s must-read book of the year. The radio show takes place over five days in five hour-long segments with one book voted off each day. Canada Reads was launched in 2002 and is still going strong. It also has a French equivalent on Radio Canada called Le combat national des livres.

My familiarity with the books and subject matter helped me understand the scope of the project and visualize the ways in which I wanted to present the collection. My vision for this curated collection was to have displayed, on one page and by each award year, the book winners and contenders along with the people who championed them. It was quite a learning process to set up this collection, and I’m so happy with the way it turned out.

Establishing the Collection

Before I even created a curated collection page for the awards, I compiled a list of all the Canada Reads books in my notes, separated by year from 2002 – 2025, noting the winners and contenders and who championed the books. Thankfully, this was an easy step as all the books are very well documented on the CBC’s website. I brought this information into Open Library by creating personal lists in my Open Library account for each award year and manually searching for and adding the books to their respective lists. With over 100 books to add, this was a tedious, time-consuming, and unsustainable process. I knew there had to be a more efficient way to complete this task, but I didn’t know how. I chipped away at it for a while and learned as I went, also taking the time to update the book metadata to ensure the accuracy of records, editions and general book information. However, I was getting overwhelmed by the process, my questions were adding up, and I couldn’t seem to find the answers I was looking for in the documentation. It was time to turn to the community for help. I was so relieved to discover the weekly Open Library Community Zoom calls on Tuesdays, where I would get to ask my questions to actual human beings – and have conversations! 

Open Library Community Calls

Joining that first community call was instrumental in my Librarian-In-Training journey and for the development of my Canada Reads curated collection. I was immediately welcomed, supported, and made to feel like a part of the community. I had the opportunity to present my collection for the first time, share my vision, and explain my pain points. I was met with enthusiasm and incredible insights that eased my overwhelm and helped me move forward. I hadn’t realized how, in vocalizing my project to the community, it would help me feel so connected, supported, and invested in seeing my project through. 

Every community call I have attended has provided me with the tidbits of information needed to help me develop the skills required to execute my vision. These community calls have been an invaluable resource and remain a delightful part of my week.

Among the pieces of information I received during these calls was news of the bulk editing tool, which enabled me to automatically add, from a typed list, all the Canada Reads books by year to my personal lists, saving me a lot of time and manual labour. I also learned the intricacies of editing book metadata according to industry standards. And I learned how to set up a curated collection page along with the coding required to create book carousels, which visually enhance the look of my collection pages. Each step has helped me establish my collection and display the information the way I envisioned. I now have a good understanding of the technical knowledge required to set up a collection page, and though there is still plenty to learn, I feel a lot more confident in my technical skills than I did when I started.

Leveling Up the Collection with Subject Tags

Having gained these new skills and technical knowledge, I was encouraged during a community call to take my Canada Reads collection to the next level by adding unique subject tags to each book championed during Canada Reads. This not only categorized the books into a Canada Reads subject page and increased their searchability across the Open Library catalog, it also served as the basis for the coding required to create book carousels.

Subject tags are part of a book’s metadata and are added according to the subjects found in the book as well as any subject the book might be associated with. For Canada Reads, I used three unique subject tags to categorize each championed book into a specific Canada Reads subject, as described in the table below: 

Canada Reads Subject Tags
Subject TagDescription
Collection:Canada ReadsGeneral subject tag added to all the books championed during Canada Reads over the years
Collection:Canada Reads WinnerSpecific subject tag added only to the Canada Reads book winners 
Collection:Canada Reads YearCollection:Canada Reads 2025Specific subject tag added to all the Canada Reads books according to the year in which they were championed

For each subject that is tagged, Open Library automatically creates a subject page that groups all the tagged books and displays them on the page. The subject page displays a book carousel of all the tagged books. In the section below the carousel, the metadata of those books is displayed by category, showcasing the publishing history, related subjects, places, people, and times. It is a great place to see the book metadata in one place and the variety of topics within a collection. With the right librarian privileges, these subject pages can be edited and spruced up.

I was granted editing permission to maintain the Collection:Canada Reads subject page and decided to set it up in the same way as I had done for my curated collection. I did so for visual consistency as well as ease of editing. By using the same format on both pages, I can copy any updates I make on one page and paste them onto the other. There are technical differences between the Collection:Canada Reads subject page and the Canada Reads Awards curated collection page, but the only visual difference is the logo, which I specifically changed to help me distinguish between the two pages when I’m working on them.

I wanted to include book carousels on my pages because they add a dynamic and interactive element that, to me, feels like a more realistic library experience than scrolling through a wall of text on a webpage. Although the coding for the book carousels confused me at first, I was guided through my confusion during a community call. In these calls, others walked me through the coding process and explained the technical setup of the subject tags. I was able to understand how the unique Collection:Canada Reads subject tags I had been adding would form the basis of the code required to create the book carousels. I used the code below, only changing the ‘collection:‘ and ‘title=’ elements to match the collection I wanted to have displayed on the page.

{{QueryCarousel(‘subject:(“collection:Canada Reads 2025″)’, search=False, has_fulltext_only=False, title=”Canada Reads 2025”, limit=20)}}
Screenshot of part of Canada Reads 2025 collection on Open Library

Celebrating My Canada Reads Collection 

From August 2025 to November 2025, I was focused on building my Canada Reads collection pages, attending the Community calls on a weekly basis, and learning everything I could to make them come to life. I took on this project knowing it would be an excellent challenge, and one I am continuing to meet with openness and curiosity. 

This collection has taken on a life of its own. Although I never expected to share it with anyone, I am tremendously glad I joined that first community call to share it with the Open Library team. Since then, I have established two Canada Reads collection pages, updated the metadata on many books, written this blog post, and even got the opportunity to present my collection at the 2025 Open Library Community Celebration in November 2025!

I couldn’t have predicted what I’ve accomplished with this collection, nor the recognition I’ve received. I’m incredibly proud of the work I’ve done and so, so grateful for the Open Library team’s support, without which I wouldn’t have gotten this far. 

Next Steps for the Collection

This collection is a work in progress and as of this blog’s publication date, I have completed the award years from 2016 – 2025. Over the coming months, my goal is to add the remaining books from 2002 – 2015 as well as clean up the metadata for each book. I will also add the 2026 books once Canada Reads has concluded in April 2026.

This is a manual process that can be fairly time consuming as editing book metadata varies greatly depending on the popularity of the book and the number of editions it has. Although I would love for my collection to be complete, I am making sure not to rush the process and to take the necessary time to correctly input the information and create a collection that is accurate – and one that I am proud of.

My hope for this collection is for people to enjoy it and hopefully, be inspired to read some Canadian content. 

If you would like to contribute to Open Library as a librarian, you can fill out this form and join the Slack channel and the weekly community call. Other contributors will provide mentorship and help you get started. 

]]>
Featuring Nazar Kotsur: Tracking Ukrainian Books Missing from the Digital Commons https://blog.openlibrary.org/2026/03/24/featuring-nazar-kotsur-tracking-ukrainian-books-missing-from-the-digital-commons/ Tue, 24 Mar 2026 16:27:35 +0000 https://blog.openlibrary.org/?p=2642 Open Library volunteers regularly work behind the scenes to build collections and improve access to books from around the world. One of these is Nazar Kotsur, who has contributed as a volunteer librarian since 2022.

A student pursuing a bachelor’s degree in Japanese language and literature, Nazar first learned about Open Library through a language-learning group that shared a list of online resources. Drawn to its open-source mission, he became involved as a volunteer librarian. 

Among many projects, Nazar is curating a still-growing list of 64 books in Ukrainian that don’t have a readable copy in Open Library, as well as a collection-in-progress of Ukrainian literature for students

The list began as a personal resource Nazar uses to track books that appear to be missing from the internet—books or specific editions for which he has not been able to locate a PDF or other digital file.

These were “some interesting books I found, I like or I want to read and I was compiling the list because I couldn’t find the specific editions,” Nazar says. 

The list contains works spanning topics from the Ukrainian fight for independence to the country’s history and culture, as well as fiction and literature. 

Many of these books have not yet been preserved digitally. While some of the books are modern and might be present in a library, others are from the 1920s or 1930s and could be difficult to find even in physical form. A few of the books on the list are public domain works, which have files that Nazar hopes to later add to Ukrainian Wikisource. 

The work has become all the more urgent as recently, the biggest Ukrainian online library, Chtyvo, closed. It contained 87,000+ books, some of them not available anywhere else, including many public domain works. Nazar, who is currently collaborating with others to preserve records from that site as well, says this was a big loss for the Ukrainian humanities and for readers. 

Nazar is no stranger to open source projects. Today, Nazar is also a part-time Wikisource development consultant for Wikimedia Ukraine.

Wikisource is a project of the Wikimedia Foundation that aims to build a freely accessible online library of source texts, including translations of those works in many languages.

At Wikisource, Nazar organizes proofreading contests and finds new contributors by reaching out to students and professors to help speed up preservation efforts.

“My contributions in Open Library right now mostly consist of fixing the books that were proofread on Wikisource and adding IDs so that the Read button [on Open Library] becomes available,” Nazar says.

This is important because when patrons click the Read button on Open Library, they get redirected to the online book reader with scanned PDF or DJVU files of the edition. This enables them to see the pages from editions as images with the original printed text, color, previous owners’ notes in the margins and often for old books, various defects. This is excellent for preservation, but can make reading harder, especially if for those with less-than-perfect vision or who are reading on a small screen.

“Open Library is a great project that has done a lot to preserve various books in digital form and make them available for reading to people around the world,” Nazar says. “Wikisource is quite similar to Open Library in that its goal is to preserve books and a lot of the files we work on actually come from Open Library, but the way Wikisource goes around this task is different.” 

Wikisource editors transcribe the pages of scanned books into digital form, preserving the original structure and formatting. That allows for a better reading experience, in which readers can configure the font and size of the text. The ability to resize text to fit the screen size is more convenient on small screens or when a book has burned-out letters, water damage, or other defects. Most important, these texts can be put into text-to-speech software so that visually impaired people can access them too.

Once a Wikisource ID is added to the edition in Open Library, it will redirect users to the text in Wikisource, where it may be more convenient to read.

Nazar continues to add Wikisource IDs to the books in the lists and many others. 

If you would like to contribute to projects like this one, you can indicate your interest in volunteering as a librarian in this form and connect with Nazar in the librarians Slack channel.

]]>
Featuring Ben Deitch, Engineering Fellow https://blog.openlibrary.org/2026/03/03/featuring-ben-deitch-engineering-fellow/ Tue, 03 Mar 2026 00:10:27 +0000 https://blog.openlibrary.org/?p=2630 By Elizabeth Mays & Mek

Open Library is powered by a global community of volunteers, a small team of staff, and several extraordinary, handpicked volunteer fellows who are picked to work alongside staff to tackle ambitious, high impact projects. This week, we’re featuring the work of Engineering Fellow Ben Deitch, who has made a dramatic impact on the Open Library initiative since 2024. 

Ben’s numerous engineering contributions have strengthened Open Library’s experience for hundreds of thousands of patrons. With the mentorship of senior staff engineer Drini Cami, Ben wrote code that enables patrons to:

  • Find exact book editions from their Reading Logs
  • Search which books were read within any given year
  • Discover interesting books based on a sophisticated reddit-style trending algorithm
  • Mark books as “Want to Read” from author pages

Prior to Ben’s work, reading logs would show works instead of editions. Ben added the ability to view editions on a reading log. This enables users to track the precise editions of books they have read. The log will also show the right cover for each edition. 

Ben also implemented a basic “fuzzy search” for the Solr Search engine, making the overall search system much more tolerant of spelling errors and bringing it closer to modern standards for search engines so that patrons don’t hit dead ends. 

In another project, Ben coded in a new, image-based preview for user created book lists, which appears on users’ My Books pages. This feature also enables patrons to see the first few books in a list at a glance. 

Ben can also be credited for enabling the “already read” books to be viewed by the year in which they were read

Historically, search results and book pages have featured a “Want to Read” button that patrons can click to keep track of books of interest. Ben extended book results so patrons can also click “Want to Read” from the author’s page.

In the past, when carousels were rendered for the homepage, facets were included – like language – that were accidentally being dropped when additional results were loaded. Ben fixed this issue so books results were relevant even when loading multiple pages.

Finally, Ben fixed an librarian issue to alter the display message after the merging author entries so merging two author records no longer show as successful when there was actually an error

Ben’s work can be found across the Open Library – from the search page, to the home page, the author’s page, and the my books page. We are grateful to Ben and couldn’t be more proud of his contributions to our Open Library.

You can send Ben kudos here on LinkedIn!

]]>
A Community-Curated Nancy Drew Collection https://blog.openlibrary.org/2026/01/30/a-community-curated-nancy-drew-collection/ Fri, 30 Jan 2026 18:46:39 +0000 https://blog.openlibrary.org/?p=2614 A team of volunteer Open Librarians have worked together to organize the many Nancy Drew book series into a beautiful collection on Open Library

If you’re excited about this collection, you can direct your thanks to Open Library volunteer Emily, who proposed the project. A few months ago, Emily put out a call in Open Library’s librarian Slack channel to see if other librarians might be interested in teaming up. Today, the collection is live and ready for the benefit of the public.

A collaborative approach was second nature for Emily, a librarian and educator who recently completed a Master of Information program. 

“Almost all of our projects were group work to help prepare us to work together and collaborate in libraries,” she said.

To organize the project, Emily built a detailed Google Document with information, ideas, questions about methodology and choices the team would need to make as a group. Participants added thoughts and notes asynchronously before the call.

An initial Zoom call then brought the team of volunteers together in real time. The call was held in a time zone that worked for the international contributors, who came from Tokyo, Pakistan and the western U.S.

“I think it was really important to do a video call to start things off, just to really just humanize everyone,” Emily said. “Like you see everyone a little bit, you hear their voices, you know that you’re working together. You know that you’re a team, and that helps everyone stay motivated.”

Maahin, located in Pakistan, worked on two series of the collection, Nancy Drew and the Clue Crew and Nancy Drew Notebooks. She had long wanted to become a librarian. “Being in a place where there’s no scope of reading and related professions, Open Library is the best chance to contribute in book-related tasks, and it motivated me finding you can contribute to it remotely,” she said.

On the kickoff call, contributors aligned on preliminary decisions, discussed how to divide the work and shared their reasons for contributing.

How best to build the collection required some sleuthing. Contributors explored various methods to build collections and tag large numbers of works. They considered using Python scripts to automate finding books and adding metadata, but determined the approach was impractical given the extensive metadata cleaning and large-scale review this project required, combined with limitations in their current technical expertise. In addition, they experimented with alternative versions of the current carousel code. However, they found that these new versions would result in a lag when users loaded the page. Contributors wanted to make sure the collection would be accessible to anyone, regardless of their Internet speed.

Because this was such a large collection with so many different series, Emily checked behind the scenes to learn how similar collections in Open Library had been built. 

With that info in view, a decision was made to manually tag books’ subject fields with a collectionid: tag for each series. 

Nichole, who focused on the Nancy Drew: Girl Detective and Nancy Drew on Campus series of the collection, joined the project out of a desire to learn.

“I was new to the Open Library and wanted to learn how to create collections and hone my metadata editing skills,” she said “I also noticed that we had a Hardy Boys series but not a Nancy Drew series, which felt like a gap.”

Working on metadata taught Nichole about source verification. Most of her previous metadata assignments involved checking single documents or websites, so she assumed the task of editing metadata for a series of books would be straightforward. But this project required evaluating and aggregating information from multiple sources. 

“It was surprisingly challenging to confirm basic facts (like how many editions of a Nancy Drew book exist and how they were published) and find reliable information.”

Another volunteer, Liz, consulted portions of the book, “Girl Sleuth: Nancy Drew and the Women Who Created Her” by Melanie Rehak. The biography identified another of the major challenges for the collection–many of the Nancy Drew books in Open Library had been attributed to the wrong ghostwriter, instead of the pen name Carolyn Keene. The pen name refers to numerous ghostwriters across different works and editions of the books (such as reissues). But most of the books in the library attributed the wrong ghostwriter as the author. 

Lead Staff Librarian Lisa Seaberg helped to correct conflated author metadata, a common issue when a pseudonym is shared by multiple authors and when multiple authors have the same name.

The collection as it stands represents many hours of metadata cleanup from each of the contributors. 

For future groups collaborating on collections, Emily suggests live-demo-ing how to edit the metadata before asking people to do it. “I think we all had to go on an individual journey of reading the documentation and figuring out how to do it,” Emily said. 

Contributors each had their own reasons for helping to bring the Nancy Drew collection to life. 

“For me, there’s definitely some nostalgia,” Nichole said. “I grew up during the era of Nancy Drew PC games and remember playing Nancy Drew: The Phantom of Venice. But I also appreciate Nancy Drew as a character. Growing up, I read a lot of detective fiction and became familiar with detectives like Sherlock Holmes, Hercule Poirot, Miss Marple, S. S. Van Dine and Hajime Kindaichi. Nancy Drew feels unique — not just in age and life experience but also in personality and technique.”

Emily grew up on a small dairy farm in rural Canada, without access to many TV channels. “It was just so nice to have these stories that I could access — this huge wealth of narratives about a woman who was really curious,” Emily said. “It was just a really good role model for me growing up.”

Maahin joined for the chance to do library work. “I am excited to work on all kinds of tasks including documentation and cataloguing and every other thing that is related to books and library,” she said.

Work on the collection and cleanup of metadata in the current collection is still ongoing, with continuing opportunities to contribute. Some of these include adding series tags or special featured collections, such as the books that inspired the Nancy Drew computer game series.

Emily also aspires to make the books appear in order in the series. (Currently the order is tied to the date of the most recent edition.) “If we could find a developer who can find a solution to help us make all these books appear in order in the series, that would be wonderful,” Emily said. 

The project took months from start to finish. The work to clean metadata and get the first eight series of 500-plus books into the collection was substantial, but rewarding.

“It was work to learn how to do it, but it is so satisfying to have built something and help share the things that helped you have a love of reading with other people,” Emily said. “And it’s been really wonderful to connect with similarly minded people as well.”

If you would like to contribute to this or future collections projects at Open Library, fill out this volunteer form.

Image of a few series in the new Nancy Drew Collection on Open Library. Shows carousel selections from Nancy Drew on Campus and Nancy Drew Girl Detective.


Lessons Learned:

  • How to Build Collections: For now, manually tagging books’ subject fields with a collectionid: tag for each series, and copying the code from past multi-series collections, is the most expedient way to build a collection.
  • Human Connection Matters: Meeting fellow librarians, combined with defined asynchronous processes, can help a collaborative project go smoothly. 
  • Live Training Could Save Time: Future projects would benefit from a short live demonstration of metadata editing at the outset. This could reduce the learning curve and help volunteers feel confident contributing sooner.
]]>
Celebrating Our Community in 2025 https://blog.openlibrary.org/2025/12/25/celebrating-our-community-in-2025/ Thu, 25 Dec 2025 08:56:56 +0000 https://blog.openlibrary.org/?p=2600 Highlights From the 2025 Open Library Community Celebration

This year, staff, fellows, and volunteers made a number of improvements to Open Library. Here are some highlights of contributors’ accomplishments in 2025, as presented in the annual Community Celebration.

  • Ray Berger, volunteer Developer Experience Lead celebrated his fifth year with Open Library, having reviewed and merged more than 100 pull requests in 2025. This year Ray launched https://docs.openlibrary.org, a new searchable portal for developer documentation. To improve website performance, Ray upgraded the Open Library Search APIs to use FastAPI, a more performant, modern web framework. He has also helped modernize the code base to make it easier for developers to contribute. 
  • GSoC Engineering Fellow Sandy Chu worked with staff members Drini Cami and Mek to enable on-the-fly translations in BookReader. Now, books in BookReader can be translated into more than 40 languages. This project also enabled read-aloud capabilities in BookReader, which helps to close the accessibility gap for international readers. 
  • Engineering volunteers David Ragipi and Krishna Gowsami redesigned book lists to add a “follow” button next to usernames. The feature increased the number of patrons following each other from 182 to more than 3K followers in 2025.
  • Under Mek’s mentorship, GSoc Engineering Fellow Roni Bhakta developed a prototype of Lenny, a self-hostable digital library for storing and lending EPUBs. Lenny gives libraries and individuals a lightweight way to host and securely lend the EPUBs they own.
  • Engineering Fellow Ben Deitch worked with Drini to develop a new trending algorithm built on Solr that uses hour-by-hour statistics to give patrons fresh, timely books that are getting high interest at any given moment on Open Library. This feature replaced trending views that changed infrequently and also gives insight into what’s trending in a given subject.
  • Engineering Fellow Stef Kischak worked with Drini to develop a script that scrapes Wikimedia APIs for Wikisource ebooks to import. Stef also made efforts to improve the import pipeline and identify orphaned editions to edit.
  • Librarian Fellow Jordan Frederick imported reading levels metadata that improved the K-12 reading collection. Jordan also fixed metadata, split wrongly merged records, and created tutorials for Open Library patrons. 
  • New librarian volunteer Catherine Gosztonyi created the still-growing Canada Reads Awards collection on Open Library. 
  • Internet Archive Staff Member Lisa Seaberg celebrated the 684 volunteer librarians in the Open Library community Slack channel. Volunteer librarians improve the catalog by adding metadata, author info, images, and collections. Lisa also recognized multiple superlibrarians who reviewed hundreds of thousands of merge requests, curated special collections, and mentored librarians in training.
  • Volunteer communications lead Elizabeth Mays, with team Nick Norman, Ella Cuskelly, and Jordan Frederick, doubled the number of blog posts published in 2025 and streamlined a process for writing and approving blog posts. The group also defined standard starter tasks for future volunteers toward projects that will enable more frequent social content. 
  • Staff member Drini Cami highlighted the work of the developer community on a unified read button, language-aware autocomplete and carousels, full-text list search, new librarian features, Wikipedia links on author pages, and Wikidata integration. Drini presented staff improvements such as a data-quality tool that lets librarians see which popular books are missing metadata, streamlined special access for patrons with qualifying print disabilities, grid view, and security improvements to prevent cyberattacks. 

Watch the replay of our 2025 Community Celebration or these slides to learn more about these upgrades. 

Previous Community Celebrations

This is Open Library’s sixth Community Celebration to recognize contributors, who come from more than 20 countries. Catch up on past years’ events at these links:

2024, 2023, 2022, 2021, 2020

Get Involved

If you’d like to get involved, indicate your interest in volunteering with Open Library in this interest form. We’ll be in touch to connect you to the community Slack and weekly call. 

]]>