Bias in Information Search

A cartoon of a megaphone in front of a number of document icons floating in front of it over a blue gradient background.

Image copyright Adobe Stock

Inevitable bias

Millions of Google searches are conducted every hour, and most of us take it for granted that the information we need will show up on the first page of results. However, all information search tools, whether they are search engines like Google, emerging AI tools like ChatGPT, or article databases from a library, are inherently biased in terms of the information they include, how the information is organized, and how the results are displayed.

Algorithms and AI

Algorithms are essentially the formulas that search engines follow to retrieve and sort results. Although often seen as neutral and objective, algorithms are created by and for humans, who, of course, have biases. It’s inevitable that prejudices and subjectivities make their way into algorithms. Biases also affect how algorithms develop over time as they receive new inputs, such as users’ searches and choices of results to view. In addition, with more and more content being generated by AI, our perspective can be influenced by the way AI tools can flatten the diversity of the real world into the common denominators they are trained to identify and generate, as demonstrated in a study of the stereotypes depicted by the image generator Midjourney (Turk, 2023).

The image is a mosaic of faces and buildings that are racialized, displaying many non-white stereotypes with exaggerated features. The mosaic is 5 by 4 square images. It was generated by the Midjourney generative graphics tool.

Images generated by Midjourney

Embedded bias

Dr. Safiya Noble’s groundbreaking research into racist and sexist algorithmic bias in the 2010s documented how Google’s results for search terms like “black girls” (as well as for other terms related to women of colour) returned sexualized images and websites, especially compared to the search for “white girls.” This work and others have shed light on specific, blatant examples of racism displayed through search engines’ algorithms, leading to changes over the past 5-10 years.

However, the underlying problem of bias remains when predictive algorithms perpetuate existing stereotyped representation or exclusion of marginalized and oppressed groups. For example, a more recent study (Vlasceanu & Amodio, 2022) conducted Google image searches for the seemingly neutral term “person” in countries across the world and found a correlation between male-dominated images appearing as the default corresponding with higher levels of gender-based inequality.

Algorithms of Oppression: Safiya Umoja Noble

Bias in academic search tools

Bias exists not only in search engines like Google and generative chatbots like ChatGPT, but also in search tools provided by libraries, such as the Sofia Discovery Tool and article databases from publishers like EBSCO, ProQuest, and Elsevier. It’s important to understand that some search tools are better than others for retrieving certain types of information—for example, you should definitely use an article database to find peer-reviewed articles rather than just Google or ChatGPT! However, no search tool is fully neutral and objective. This is due to decisions made about which information sources the tool searches, how the information is organized, and how the results are retrieved and ranked.

Library subject headings

A clear example of bias can be seen in the use of subject headings and classification systems used to organize content in library systems, which include outdated and hegemonic terminology. Subject headings can be very useful when searching because they bring together books on related topics using standardized terminology—like tags. However, many subject headings continue to reflect the white supremacist and patriarchal culture from which they originated. For example, the heading “Indians of North America” is used for books about Indigenous Peoples in North America, especially on older books, instead of or alongside accepted and respectful collective terms like “Indigenous,” and specific terms like First Nations, Métis, and Inuit, or the names of individual nations.

Also, on the library shelves, most books about First Nations, Métis and Inuit communities and thought are found in the E classification area, for “History of North America.” This represents an erasure of living peoples, philosophies, and knowledge.

This gif shows where to find some problematic library of congress headings such as the "five civilized tribes after 1907" starting from the Library of Congress website and looking through the corresponding pdf detailing the subheadings of the LoC classification system. The gif is captioned “Where to find some of the headings in the Library of Congress classification system discussed in this section.”

Where to find some of the headings in the Library of Congress classification system discussed in this section.

There are ongoing efforts to change the ways in which academic libraries organize information (for example, Westmount Public Library is working on updating its catalogue subjects associated with Indigenous people), but in some ways, the job is just as big as shifting entire academic institutions and Western-centric worldviews.

While there isn’t a way to simply remove or avoid biases, being aware of the prejudices and limitations underlying search tools is an important aspect of being effective and critical reader and writer.

Structure and some text adapted from University of Arizona under a CC-BY 4.0 international license, some text adapted from Concordia University Library presentation "Researching Indigenous Topics in the Library" by Michelle Lake, with permission and feedback.

Resources

References

Baer, Andrea. “Is Google Neutral: Unpacking Algorithmic Bias.” Workshop template. 2023.

Bush, Olivia. Google Statistics in Canada. Made In CA. 2024.

Noble, Safiya Umoja. Algorithms of Oppression: How Search Engines Reinforce Racism. 2018.

Turk, Victoria. How AI Reduces the World to Stereotypes. Rest of the World. 2023.

University of Arizona Libraries. The Complexity of Searching.

University of Victoria Libraries. Changes in the Library of Congress Subject Headings. 2022.

Vlasceanu, Madalina, & Amodio, David M. Propagation of Societal Gender Inequality by Internet Search Algorithms. Proceedings of the National Academy of Sciences of the United States of America, Volume 119, Issue 29. e2204529119. 2022.

Westmount Public Library. Modification to Catalogue Subjects Associated with Indigenous Peoples. 2021.