Concepts / Web Scraping Basics

Web Scraping Basics

curl and wget are command-line tools for retrieving files from the web on Unix-like systems using HTTP or FTP protocols.

  • Programming

Why Command-Line Retrieval Matters

Web scraping basics begin with retrieving a resource from the web. On a Linux, Unix, or Macintosh system, you may be working on a remote server, automating a task, or retrieving many files without keeping a graphical browser open. curl and wget let you fetch webpages and files directly from the command line.

curl and wget are command-line tools for retrieving files from the web on Unix-like systems using HTTP or FTP protocols.

Both tools can retrieve plain-text files and binary files. The retrieval method is the same at a high level: the command requests a URL, the server responds, and the tool displays or saves the received data.

The Request-Response Path

When you run curl or wget, the command sends an HTTP GET request to the remote server. The request identifies the URL of the resource you want. The server receives the request, locates the requested file, and sends the file data back. The command-line tool then either saves the received data to disk or displays it in the terminal, depending on the command and its options.

HTTP GET requestreturns resourcesave or displaycurl or wgetcommand lineWeb serverremote resourceFile dataHTTP responseLocal filesystemsaved file
How does a command-line tool send a request to a web server and receive a file in response?

Saving One File with curl

curl is a general-purpose command-line data-transfer tool. To save a retrieved file, use the -O option. The uppercase O tells curl to use the remote filename. Without -O, curl prints the retrieved content to the terminal instead of saving it to disk.

bash

In this example, curl requests cover.jpg from www.py4e.com and saves it in the current directory with the same filename. The same mechanism works for plain-text and binary files: the received data is transferred from the remote URL, and -O directs curl to save it using the remote name.

specifiesretrievessaved with -Ocover.jpg URLremote resourceHTTP GETcurl requestFile datatext or binarycover.jpgcurrent directory
What data moves from the remote URL to the local filesystem when curl is used with the -O flag?

What do you think happens?

What happens when curl is run without the -O option?

  • It automatically saves the file with its remote filename
  • It prints the retrieved content to the terminal
  • It recursively downloads linked pages
  • It reports a 404 response
Reveal answer

Answer: It prints the retrieved content to the terminal.

curl requires the -O option to save the file. Without that option, curl displays the retrieved content instead of saving it to disk.

Retrieving Files with wget

wget also retrieves webpages and remote files from the command line. Unlike curl, wget automatically saves a downloaded file with its original name. Therefore, a basic wget command does not need an option equivalent to curl's -O for this task.

bash

This command downloads cover.jpg and saves it in the current directory using the remote filename. In the HTTP cycle, wget still sends a request and receives a response; its practical difference here is that saving with the original name happens automatically.

supportssupportssupportscurlgeneral data transferOne file-O saves remote namewgetfile retrievalOne filesaves original nameWebsite hierarchy-r follows links
What is the practical difference between curl and wget for retrieving one file, downloading recursively, or scripting a request?

Following Links Recursively

wget has a recursive downloading capability that is enabled with -r. When wget encounters a webpage, it parses the HTML to find links and resource references. It then downloads the linked pages and resources, repeating the process through the website's link structure. This can create a local mirror of the website structure.

bash

The command above illustrates the placement of -r for recursive downloading. The source material defines -r as the option that enables recursive behavior. It also explains that -l can limit the number of levels wget follows, so recursion can be controlled rather than allowed to continue through every reachable level.

follows linkfollows linkfinds referencefinds referenceStarting webpagewget -rLinked webpagediscovered linkLinked resourcepage assetAnother webpagediscovered linkReferenced filepage resource
How does wget follow links and expand from one webpage into a set of downloaded pages and assets?

Choosing the Retrieval Tool

Retrieval taskPractical choiceReason
Save one remote file with its remote filenamecurl -Ocurl requires -O to save the file using the remote filename
Retrieve one file with automatic savingwgetwget automatically saves files with their original names
Perform general-purpose data transfercurlcurl is described as more flexible for general-purpose data transfer
Download a website hierarchywget -rwget can follow links recursively
Limit recursive link depthwget with -lThe -l option controls how many levels wget follows

Use the task, not the tool name, as the starting point for your choice.

curl and wget overlap because both can retrieve web files through the same request-response pattern. The useful distinction is their emphasis: curl is more flexible for general-purpose data transfer, while wget provides convenient file saving and recursive website downloading.

Diagnosing Retrieval Failures

Debug a failed retrieval by tracing the stages of the request. First identify the URL being requested. Then inspect the response reported by the tool. A successful 200 OK response means the file was retrieved; a 404 Not Found or another error status means the request failed and no file will be saved. If the response succeeds, confirm that the expected remote filename was used: curl needs -O for this behavior, while wget saves with the original name automatically.

receives responsesuccesserrorconfirm filenameRun retrievalcommandcurl or wgetResponse status200 OK or error200 OKfile retrievedExpected filenamesaved result404 Not Foundtool reports failure
How can a learner trace a failed download and verify that the response and saved result are correct?
  • Running curl without -O when the goal is to save the file

    Without -O, curl prints the retrieved content to the terminal instead of saving it.

    Fix: Use curl -O followed by the URL when the file should be saved with its remote filename.

  • Expecting wget to behave like curl without automatic file saving

    wget automatically saves files with their original names.

    Fix: Use the basic wget URL form for a remote file when its original name is suitable.

  • Treating a reported HTTP error as a successful download

    A 404 or another error status means the retrieval failed, and no file will be saved.

    Fix: Read the response status reported by the tool and correct the failed request before relying on the file.

  • Using ordinary single-file retrieval when the task requires a website hierarchy

    Following links recursively is a wget capability enabled with -r.

    Fix: Use wget -r for recursive downloading and consider -l when the depth must be limited.

Practice the Decision

EASY

You need to retrieve one binary image and save it using the filename supplied by the remote URL. Which command-line tool and option would you choose, and what would happen if you omitted the option?

Hints
  • The task requires saving one file with its remote filename.
  • curl needs a specific uppercase option for saving instead of displaying the response.
  • wget automatically saves with the original name.
MEDIUM

A documentation site contains linked webpages and resources that must be retrieved as a local website structure. Which tool and option match this task? What option can limit how many levels of links are followed?

Hints
  • Look for the tool described as supporting recursive downloading.
  • The recursive option is -r.
  • The level-control option is -l.

Working Retrieval Checklist

  1. Choose the URL and decide whether the task is one-file retrieval, general data transfer, or recursive website downloading.
  2. Use curl with -O when curl should save the response using the remote filename.
  3. Use wget for automatic saving with the original filename or for recursive downloading with -r.
  4. Read the reported response status: 200 OK indicates successful retrieval, while 404 Not Found or another error indicates failure.
  5. After success, confirm that the expected file was saved under the expected remote filename.

Key Takeaways

  • curl and wget retrieve web resources from the command line through HTTP or FTP.
  • Both tools follow an HTTP request-response cycle and can retrieve text or binary files.
  • curl needs the -O option to save a file using its remote filename; otherwise it prints the content to the terminal.
  • wget automatically saves files with their original names and supports recursive downloading with -r.
  • A 200 OK response indicates successful retrieval, while a 404 Not Found or another error status indicates that the download failed.