By the late 1980s the Internet already connected universities, laboratories and government agencies. Using it, however, meant learning specific commands, knowing which machine held which file, and mastering separate programs for electronic mail, file transfer and remote login. Information stored on computers of different makes was, in practice, walled off from everyone who did not know where to look. Finding a document depended on knowing its location in advance, and following a reference from one text to another meant repeating the search by hand. The Web answered a problem of organization rather than of transmission: it offered a common way to name, publish and link documents that could run over networks already in place.
It helps to keep two terms apart. The Internet is the infrastructure, a set of interconnected networks that exchange data under shared protocols, a story told in the article on the first computer networks. The World Wide Web is one service running on that infrastructure, made of pages and other resources identified by addresses and connected by links. Email and file transfer are Internet uses that do not depend on the Web at all.
How it works
The Web rests on three complementary technical ideas, worked out by Tim Berners-Lee and colleagues at CERN, the European particle physics laboratory near Geneva.
- URL (Uniform Resource Locator, now treated more broadly as a URI): a standardized address that identifies a resource and says how to reach it.
- HTTP (Hypertext Transfer Protocol): the protocol by which a client program, the browser, asks a server for a resource, and the server answers with the content or an error message.
- HTML (Hypertext Markup Language): a markup language, meaning a set of tags placed inside text to mark headings, paragraphs, lists and, above all, links to other resources.
In practice, when someone types an address, the browser uses the domain name system to find which computer answers to that name, sends an HTTP request, receives an HTML document and displays it, then fetches the images, style files and other items the document mentions. The link is the central element. It turns a pile of isolated documents into a network of references that a reader can follow without knowing anything about the destination machine. Part of the design's strength was its simplicity and openness: anyone could write a server or a browser by following the same published conventions.
Antecedents of hypertext
The idea of texts joined by associations is older than networked computers. In 1945 the engineer Vannevar Bush described the Memex in an essay titled As We May Think, published in The Atlantic. It was a hypothetical microfilm-based desk on which a user would record trails of association between documents. It was never built as described. In the 1960s Ted Nelson coined the word hypertext and began the Xanadu project, which envisioned two-way links and tracking of authorship and payment for quoted passages; Xanadu never became a widely used service. In December 1968 Douglas Engelbart and his team at the Stanford Research Institute gave a public demonstration in San Francisco of an interactive system with text editing, windows, the mouse and linked documents, later nicknamed the "mother of all demos." In the 1980s programs such as Apple's HyperCard popularized linked cards on personal computers, but only for local files, with no global network behind them. Berners-Lee himself had written a small personal linking program in 1980 while working at CERN.
Historical context
In March 1989 Berners-Lee, working in CERN's computing division, circulated a document titled Information Management: A Proposal. It addressed how to keep track of information in an organization with high turnover among researchers and many incompatible systems. There was no immediate formal approval. The proposal was reworked with the Belgian systems engineer Robert Cailliau into a formal management proposal in November 1990, and by the end of 1990 Berners-Lee had, on a NeXT computer, the first browser, which could also edit pages, and the first server. A simpler line-mode browser, written by Nicola Pellow during a student work placement at CERN, could run on many kinds of machines. In August 1991 Berners-Lee announced the software on Internet newsgroups, and it began circulating beyond CERN.
In its first years the Web competed with other tools for navigating information. Gopher, created at the University of Minnesota in 1991, organized documents in hierarchical menus and was widely used in academia. In 1993 the university announced that commercial users of its Gopher server software would need a paid license. The move drew a hostile reaction and is often cited as one factor in Gopher's decline. On April 30, 1993, CERN issued a statement relinquishing its intellectual property rights in the basic Web client, server and code library, so that anyone could use, modify and redistribute them without royalty payments, a decision that lowered the risk of fragmentation and licensing costs.
The next step was the interface. In 1993 the National Center for Supercomputing Applications (NCSA) at the University of Illinois began releasing Mosaic, developed by a team that included Marc Andreessen and Eric Bina; the first public version appeared in January 1993 for Unix systems. Mosaic showed text and images in the same window and was soon available for several operating systems, which helped widen the audience beyond physicists and specialists. Andreessen and other team members later joined a new company whose browser, Netscape Navigator, appeared in late 1994. Microsoft entered the market with Internet Explorer in 1995, and the rivalry that followed, marked by incompatible extensions and antitrust litigation, became known as the "browser wars."
In October 1994 Berners-Lee founded the World Wide Web Consortium (W3C), initially hosted at the Massachusetts Institute of Technology, to coordinate the evolution of standards in an open way. Successive versions of HTML and HTTP were documented in public specifications, some of them published as RFCs (Requests for Comments) by the Internet community. Cascading style sheets (CSS) separated presentation from content, and scripting languages running in the browser allowed more interactive pages.
As pages multiplied, directories and search engines appeared. These services indexed content and ranked results, and they changed how people found information: instead of following links from a curated list, users began querying an index. Commercial enthusiasm fed a wave of investment in Internet companies whose share prices peaked around March 2000 and then fell over the following years, an episode called the dot-com bubble. Many companies failed, yet much of the infrastructure and many of the practices created in those years stayed in use.
In the following decade the phrase "Web 2.0" came to describe sites built around user-produced content: blogs, collaborative encyclopedias, social networks and video platforms. Techniques for updating part of a page without reloading it gave sites behavior closer to that of desktop programs. With the arrival of smartphones, discussed in the article on mobile phones, browsing moved largely to small screens, and many services were reorganized into apps and into pages that adapt to different screen sizes.
Other information networks
The Web was not the only path. In France, the Minitel, tested in Brittany from 1980 and introduced nationally in 1982 over the telephone network, offered directory lookups, catalogs, messaging and ticket booking on inexpensive terminals, and reached millions of users. France Télécom retired the service on June 30, 2012. Commercial online services and electronic bulletin boards also flourished before and alongside the Web, typically run by a single operator. What set the Web apart was the combination of open standards, decentralized publishing and the absence of a single operator controlling who could take part.
Impact and limitations
The Web lowered the cost of publishing and accessing information and reshaped commerce, education, journalism, public services and scientific research. It also created problems that remain under debate. Access is uneven: it depends on infrastructure, income, language and skills, and a considerable part of the world's population still lacks regular connections. Ease of publishing brought the challenge of misinformation and of judging the origin of content. The collection of browsing data by trackers raises privacy questions and has prompted specific legislation in several countries. Digital memory is fragile, too: addresses stop working and pages vanish, which is why libraries and archives, including the Library of Congress, run Web preservation programs. On more recent trends, such as the concentration of traffic on a few platforms or the addition of automated features to search, there is no consensus, and this account reflects only what was known as of its latest review.
Connections to other technologies
The Web builds on decades of earlier work on networks, covered in the first computer networks, and on machines able to display graphics, discussed in the personal computer. Its markup and scripting languages belong to the path traced in programming languages. In the other direction, the Web became a principal source of data for recent work in computational intelligence.
