Web Fragility

Introduction

The internet is part of our daily lives, but it is surprisingly fragile. The hyperlinks that connect the internet together can easily break if the pages they link to are moved, deleted, or renamed. When links break, this is called link rot. Even if the link still works, the content of the site can change significantly if the website is taken over by new people. Simply updating a site to keep it current can mean removing information that other pages are linked to. This is called content drift. Estimates place the lifespan of a web page somewhere between a few months and a few years. While this has the potential to be inconvenient on a personal level, it also poses a serious problem for people who rely on stable links to build bodies of knowledge. If links referenced in articles become inaccessible, it is impossible to verify the information that was cited, which undermines that article and its conclusions – whether that article is a piece of local news, a supreme court ruling, or the results of a medical study. Archivists and historians argue that preserving the content of the internet is important because of its historical significance, and that if we fail to do so, we could lose much of today’s knowledge, culture, and art as web pages deteriorate or disappear.

Web Archiving

Web archiving is one response to the threats of link rot and content drift – a process that uses a computer program (usually a web crawler) which navigates the internet and takes snapshots of the websites it encounters in order to preserve those pages and their relationships to other pages on the internet. The snapshots are then stored in an archive in a format that can later be accessed and consulted, even if the original website no longer exists. These snapshots allow you to read and interact with them as though they were live web pages. The most well-known web archive is the Internet Archive’s Wayback Machine, which began in 1996 and currently contains over 70 petabytes of data. Many institutions, including Concordia (see Concordia's Archive-It page), have ongoing web archiving projects that preserve online journalism, government webpages, and the pages of organizations related to a variety of topics.

click image to enlarge

Right to be Forgotten

While there are clearly many reasons why web archiving is an urgent and important project, there are some people who object to online content becoming part of the historical record, especially when it comes to personal information. People who have embarrassing or incriminating information about them posted online can face serious consequences in their relationships and jobs. In 2014, the European Union’s Court of Justice ruled that individuals have the right to ask that outdated or inaccurate material be delisted from search engines like Google. Though that protection may seem like a good thing, some experts feel that this threatens the integrity of the public record and is a kind of selective editing of history, especially since web archives have been used as evidence in legal cases. They argue that it is important to be able to check facts about events and people’s actions, particularly when it comes to holding government and public officials accountable. Canada does not yet have clear legislation about whether search engines must de-list content at the request of individuals, but a recent Federal Court of Appeal ruling in 2023 took a step in that direction, deciding that search engines are subject to existing Canadian privacy laws.

Image source: Source: by Jacob Lund Photography on Noun Project. Creative Commons Attribution-Non-Commercial-No Derivatives 2.0 licence.

Internet Infrastructure

Though it is easy to forget that the internet has physical parts other than the device you are using, the physical infrastructure that the internet relies on is also vulnerable. The cables and satellites that transport data (also called the internet backbone) and the server farms that store data are some of the most essential parts that keep the internet going. These are vulnerable to digital security breaches, as well as power fluctuations and physical threats ranging from shark attacks to natural disasters. When major servers or services go down, it can cause significant chunks of the internet to become inaccessible for hours at a time. If servers are sufficiently damaged or if the internet were downed for long enough, lots of valuable information could be lost. An important part of digital preservation initiatives is keeping backup copies of web archives, so that if a server is damaged or inaccessible, another in a different location will still have a copy.

Video: The early internet is breaking - here’s how the World Wide Web from the 90s on will be saved by Quartz

Conclusion

While web archiving offers a partial solution to preserving the ever-changing internet, it isn’t perfect. Many web crawlers struggle to capture certain media types and content created on social media platforms. There are also copyright issues with web crawlers: anyone running a web archiving project should be getting permission from the creators of websites before taking a snapshot, since the creators hold the copyright for that content. In practice, though, it is sometimes very difficult to get in touch with website creators, especially for websites that are no longer being maintained and are at the greatest risk of disappearing. Some people are advocating for changes in research and publishing processes that would require an academic or journalist to submit archived copies of any resource that they link to, ensuring that their citations remain stable and informative. The internet is still a relatively new technology, and best practices around preserving the internet’s information are still being worked out. We will see new solutions emerge as the internet continues to change and grow.

Quiz

Activities

Let's reflect on web archiving.

Write down your thoughts on the following questions:

  • Have you ever been a member of an online community or maintained a personal website or blog?
  • Does that site and its content still exist?
  • Do you wish that that content had been captured either for your personal archive or by a public web archive like the Wayback Machine?
  • If it was stored in a public web archive, would you want to have a way to get it removed?

This activity should take about 15-20 minutes to complete.

This activity gives you the chance to see web archiving in action.

Use the open source web archiving site Conifer to capture a website, and then compare your capture with other versions of the site preserved in the Internet Archive’s
Wayback Machine.

This activity should take about 15 minutes to complete.

Are you an undergraduate student participating in the FutureBound program?

After reviewing the content and completing the activities for one of the Things, you can complete a
reflection form to have it count as one activity towards your Digital Capabilities & Mindsets certificate!

Resources

Right to be Forgotten

An episode of the Radiolab podcast about how journalists at one publication decide to unpublish some of the content they previously published.

See more

Harvard Law Review study about link rot

Article that discusses how to make legal scholarship more permanent.

See more

How to Use the Wayback Machine

This introduction to using the Internet Archive's Wayback Machine includes information about searching the archive and how web pages are captured.

See more

References

Zittrain, J., Albert, K., and Lessig, L. Perma: Scoping and Addressing the Problem of Link and Reference Rot in Legal Citations. Harvard Law Review 2014, 127(176).

Ott, D. E. Reference Hygiene and Death on the Internet – Decay, Rot, Half-Life, Deterioration, and Corruption. JSLS : Journal of the Society of Laparoscopic & Robotic Surgeons. 2022; 26(1): e2021.00082.

Nguyen, H. and Weber, M. S. Internet Archives as A Tool for Research: Decay in Large Scale Archival Records. 2015 IEEE International Congress on Big Data, New York City, NY, USA: IEEE, 724-727.

Zittrain, J., Bowers, J. and Stanton, C. The Paper of Record Meets an Ephemeral Web: An Examination of Linkrot and Content Drift within The New York Times. SSRN 2021, May 4.

Brügger, N. Web Archiving - Communication. Oxford Bibliographies 2018, August 28.

Quartz. The Early Internet Is Breaking - Here’s How the World Wide Web from the 90s on Will Be Saved [Youtube Video]. 2019 July 18.

Internet Archive. How to Use the Wayback Machine [YouTube Video]. 2021 January 13.

Zittrain, J. The Internet Is Rotting. The Atlantic 2021 June 30.

Anat, B.-D. and Amram, A. The Internet Archive and the Socio-Technical Construction of Historical Facts. Internet Histories 2022; 6(1-2): 179–201.

McDonald, S. The Right to Be Forgotten: The Potential Effects on Canadian. Dalhousie Journal of Interdisciplinary Management 2019, 15.

Webster, M. Right to Be Forgotten [Podcast]. RadioLab 2019 August 23.

Vavra, A. N. The Right to Be Forgotten: An Archival Perspective. The American Archivist 2018, 81(1): 100–111.

Grandhi, S. A., Plotnick, L. and Hiltz, S. R. An Internet-Less World?: Expected Impacts of a Complete Internet Outage with Implications for Preparedness and Design. Proceedings of the ACM on Human-Computer Interaction 2020, 4(no. GROUP): 1–24.

Greenstein, S. The Basic Economics of Internet Infrastructure. Journal of Economic Perspectives 2020, 34(2): 192–214.

Vox. How Does the Internet Work? [YouTube Video]. Glad You Asked S1 E2 2020 January 8.