Concepts / HTTP Protocol Fundamentals

HTTP Protocol Fundamentals

curl and wget are command-line tools for retrieving files from the web on Unix-like systems using HTTP or FTP protocols.

  • Programming

Why Command-Line Retrieval Matters

A graphical browser is not always available or convenient. You may be working on a remote server, automating a task, or retrieving many files at once. curl and wget let you fetch webpages and files directly from a Unix-like command line. Both tools can retrieve content using HTTP or FTP protocols.

The central difference to remember at the beginning is simple: curl prints retrieved content to the terminal unless you use the -O flag to save it, while wget automatically saves files using their original names.

Tracing the HTTP Exchange

HTTP GET requestHTTP responsedisplay or savecurl or wgetcommand-line toolWeb serverremote contentTerminal or diskretrieved data
How does a command-line tool send an HTTP request, receive the server response, and turn that response into terminal output or a local file?

When you run curl or wget, the command sends an HTTP GET request to the server named in the URL. The request identifies the content the tool wants. The server locates that content and sends a response back. The command-line tool then either displays the received content or saves it to disk, depending on the command and its options.

The response also provides a status. A 200 OK response means the file was successfully retrieved. A 404 Not Found or another error status indicates that retrieval failed, and the source material states that no file will be saved in that case.

Saving One File with curl

bash

The -O option tells curl to save the retrieved content in the current directory using the filename from the remote URL. In this source example, the remote file is cover.jpg, so curl saves it locally with that same filename. The file is a binary image, but the same retrieval approach also works for plain text files.

Following a curl Download

You need to retrieve the source example's cover.jpg file and keep the remote filename locally.

Choose curl: Use curl as the command-line retrieval tool.

Add -O: The -O option instructs curl to use the filename supplied by the remote URL instead of printing the response to the terminal.

Send the request: curl sends an HTTP GET request for cover.jpg to the remote server.

Save the response: If the server responds with 200 OK, the retrieved file is saved in the current directory as cover.jpg.

The expected local result is a file named cover.jpg in the current directory.

Retrieving Files with wget

bash

wget also sends a request for the remote file and saves the result in the current directory. Like curl -O, it uses the remote filename, so the source example saves cover.jpg as cover.jpg locally. wget can retrieve webpages as well as other remote files.

identify resourceserver processes requestreturn contentsave resultRemote URLrequested resourceHTTP GETrequestServer responsetext or binary datacurl or wgetretrieval toolLocal filenamecurrent directory
How does data move from a remote web server through curl or wget to a saved file on the local disk?

Choosing Between the Tools

Taskcurlwget
Retrieve one fileUse curl with -O to save the remote filenameUse wget; it automatically saves the remote filename
Display retrieved contentWithout -O, curl prints content to the terminalwget is described primarily as saving retrieved files
Download a webpage or remote fileCan retrieve web contentCan retrieve webpages and remote files
Download a website hierarchyThe source identifies curl as flexible for general-purpose data transferUse -r for recursive downloading; use -l to limit depth

Both tools perform the basic download task through the same HTTP request-response cycle. The practical choice depends on the retrieval task. curl is presented as more flexible for general-purpose data transfer, while wget is particularly useful when the goal is to save files automatically or traverse a website recursively.

Following Links Recursively

bash

The -r option enables recursive downloading in wget. After wget retrieves a webpage, it parses the HTML for links and resource references. It then retrieves those connected pages and resources, repeating the process recursively. This can create a local mirror of the website structure and is useful for archiving websites, backing up documentation, or preparing resources for offline viewing.

follow linkretrieve referencefollow link recursivelyStarting webpageinitial URLLinked webpagefollowing a linkNext linked pagecontinued traversalLinked resourcepage asset
What happens when wget follows links from one webpage and downloads connected pages and resources recursively?

Diagnosing Download Results

receive response200 OK404errorRun retrievalcommandcurl or wget200 OKfile retrievedServer statusresponse result404 Not Foundretrieval failedOther error statusretrieval failed
How can you distinguish a successful retrieval from an HTTP error when using a command-line download tool?
  • Running curl without -O when the goal is to save a file

    Without -O, curl prints the retrieved content to the terminal rather than saving it to disk.

    Fix: Add -O when you want curl to save the file using the remote filename.

  • Assuming an error response produced a valid local download

    The source states that a 404 or another error status means retrieval failed and no file will be saved.

    Fix: Check the reported HTTP result and expect a 200 OK response for a successfully retrieved file.

  • Using a single-file command when the task requires a website hierarchy

    A single retrieval does not perform the recursive link traversal described for wget.

    Fix: Use wget with -r for recursive downloading, and use -l when you need to limit traversal depth.

Verify both parts of the result: first, confirm that the server response indicates success with 200 OK; second, confirm that the expected filename was saved in the current directory. For curl, also check that you used -O when saving was intended.

Practice the Decision

MEDIUM

You need to retrieve one binary image and save it with the filename supplied by the remote URL. Then you need to download a connected website hierarchy while limiting how deeply links are followed. Which tool and option would you choose for each task?

Hints
  • For the single image, decide whether curl needs an option to save instead of printing.
  • For the website hierarchy, look for wget's recursive option and its depth-control option.

What do you think happens?

What happens if you run curl for a remote file but leave out -O?

  • curl saves the file using the remote filename
  • curl prints the retrieved content to the terminal
  • curl recursively downloads linked pages
Reveal answer

Answer: curl prints the retrieved content to the terminal.

The -O flag is what tells curl to save the file using the remote filename. Recursive website downloading is a wget capability enabled with -r.

Key Takeaways

  1. curl and wget retrieve web content from the command line through HTTP or FTP.
  2. Both tools send an HTTP GET request and process the server response.
  3. Use curl -O when curl should save a file with the remote filename; without -O, curl prints content to the terminal.
  4. wget automatically saves retrieved files with their original names and supports recursive downloading with -r.
  5. A 200 OK response indicates successful retrieval, while a 404 Not Found or another error status indicates failure.

Key Takeaways

  • curl and wget are command-line tools for retrieving web content without relying on a graphical browser.
  • The HTTP process is a request-response cycle: the tool sends an HTTP GET request, the server returns a response, and the tool displays or saves the data.
  • curl needs -O to save a remote file with its original filename, while wget saves retrieved files automatically.
  • wget -r follows links and resource references recursively, and -l can limit the traversal depth.
  • Use the HTTP status and the expected local filename together to verify whether retrieval succeeded.