Throughout the Internet’s decades-old history, many web pages have come and gone. Though billions use the Internet to globally send information, a lot of its data has been permanently erased. Take a look at the early years of the Internet, for example: the first web pages went up in 1991, yet, the Internet Archive — a non-profit that strives to preserve what goes on the Internet — only started operations in 1996.
Information will not last if people do not actively try to save it. The Internet allows people an easy way to store and share data, but that same data can cease to exist in an instant. A study by the Pew Research Center shows that even data from the last year can disappear quickly as eight percent of web pages from 2023 are now inaccessible.

Source: Pexels
The Internet’s Original Purpose
The Internet’s global accessibility was not originally meant for the average citizen’s web surfing. It was instead designed for American army personnel to communicate. The U.S. Defense Department’s Advanced Research Projects Agency (ARPA) connected American “mainframes at universities, government agencies, and defence contractors” in 1969 via its network ARPANET, which held almost 60 “nodes,” intersections of data transmission, by the mid-1970s.
ARPANET’s growth came with a hurdle: with increasing global usage in places like Norway and the United Kingdom, data needed to move quicker across borders. In 1974, researchers Bob Kahn and Vint Cerf created a new, faster, and universal way for computers to communicate: the transmission-control protocol (TCP). The TCP “allowed computers to speak the same language,” which allowed ARPANET to introduce the Internet protocol (IP).

Source: Pexels
Not Just Web Pages and Personal Files Stored on the Internet
The Internet offers options to preserve many important pieces of information. Many people store their files to look back on later by using services like Google Drive and Dropbox. However, some of these preservation efforts come with larger implications.
For example, Internet Archive Canada and Library and Archives Canada launched a project in July to preserve 100,000 “out-of-copyright” publications from the 1200s to 1920. Individuals have even used other digital technologies to preserve languages, like efforts to share resources of endangered languages using AI. This, however, is not exempt from the ever present threat of Internet systems going down. So, how exactly can the world tackle this enormous task of retaining something that never stops growing?

Source: Pexels
Problems with Accommodating all the Old and New Data
In 2024, 5.35 billion people are using the Internet; 97 million of which were new users in 2023 alone. The Internet also holds an estimated 1.09 billion websites, and a new one pops up every three seconds. In the 50 years since TCP/IP implementation, this never-ending growth is leading to an uncontrollable amount of data. In its handbook, the Digital Preservation Coalition outlines some major issues faced in saving Internet data.
- Threats to digital materials
Anything on the Internet needs to survive unscathed without deterioration. This proves difficult when technology like hard drives can wear down or — like floppy disks — become incompatible with current computers. Any data needs to be saved as soon as possible in its original form, original context, and original usage. - Organizational Issues
Organizations need to decide whether an in-house team or third party does the preservation, and resources need to be available in either case. This also requires other considerations such as legality, security, selection process, team structure, and team adaptability. Tasks of this importance also need a high level of collaboration to run smoothly. - Resourcing Issues
Everything mentioned thus far comes at a financial cost. Data preservation requires a large group of professionals from all sorts of backgrounds, who need to use the most up-to-date skill sets. The team(s) also need an adequate facility to store the data.
Source: Pexels


