Web Scraping Basics
curl and wget are command-line tools for retrieving files from the web on Unix-like systems using HTTP or FTP protocols.
Why Command-Line Retrieval Matters
Web scraping basics begin with retrieving a resource from the web. On a Linux, Unix, or Macintosh system, you may be working on a remote server, automating a task, or retrieving many files without keeping a graphical browser open. curl and wget let you fetch webpages and files directly from the command line.
curl and wget are command-line tools for retrieving files from the web on Unix-like systems using HTTP or FTP protocols.
Both tools can retrieve plain-text files and binary files. The retrieval method is the same at a high level: the command requests a URL, the server responds, and the tool displays or saves the received data.
The Request-Response Path
When you run curl or wget, the command sends an HTTP GET request to the remote server. The request identifies the URL of the resource you want. The server receives the request, locates the requested file, and sends the file data back. The command-line tool then either saves the received data to disk or displays it in the terminal, depending on the command and its options.
Saving One File with curl
curl is a general-purpose command-line data-transfer tool. To save a retrieved file, use the -O option. The uppercase O tells curl to use the remote filename. Without -O, curl prints the retrieved content to the terminal instead of saving it to disk.
In this example, curl requests cover.jpg from www.py4e.com and saves it in the current directory with the same filename. The same mechanism works for plain-text and binary files: the received data is transferred from the remote URL, and -O directs curl to save it using the remote name.
What do you think happens?
What happens when curl is run without the -O option?
Reveal answer
Answer: It prints the retrieved content to the terminal.
curl requires the -O option to save the file. Without that option, curl displays the retrieved content instead of saving it to disk.
Retrieving Files with wget
wget also retrieves webpages and remote files from the command line. Unlike curl, wget automatically saves a downloaded file with its original name. Therefore, a basic wget command does not need an option equivalent to curl's -O for this task.
This command downloads cover.jpg and saves it in the current directory using the remote filename. In the HTTP cycle, wget still sends a request and receives a response; its practical difference here is that saving with the original name happens automatically.
Following Links Recursively
wget has a recursive downloading capability that is enabled with -r. When wget encounters a webpage, it parses the HTML to find links and resource references. It then downloads the linked pages and resources, repeating the process through the website's link structure. This can create a local mirror of the website structure.
The command above illustrates the placement of -r for recursive downloading. The source material defines -r as the option that enables recursive behavior. It also explains that -l can limit the number of levels wget follows, so recursion can be controlled rather than allowed to continue through every reachable level.
Choosing the Retrieval Tool
| Retrieval task | Practical choice | Reason |
|---|---|---|
| Save one remote file with its remote filename | curl -O | curl requires -O to save the file using the remote filename |
| Retrieve one file with automatic saving | wget | wget automatically saves files with their original names |
| Perform general-purpose data transfer | curl | curl is described as more flexible for general-purpose data transfer |
| Download a website hierarchy | wget -r | wget can follow links recursively |
| Limit recursive link depth | wget with -l | The -l option controls how many levels wget follows |
Use the task, not the tool name, as the starting point for your choice.
curl and wget overlap because both can retrieve web files through the same request-response pattern. The useful distinction is their emphasis: curl is more flexible for general-purpose data transfer, while wget provides convenient file saving and recursive website downloading.
Diagnosing Retrieval Failures
Debug a failed retrieval by tracing the stages of the request. First identify the URL being requested. Then inspect the response reported by the tool. A successful 200 OK response means the file was retrieved; a 404 Not Found or another error status means the request failed and no file will be saved. If the response succeeds, confirm that the expected remote filename was used: curl needs -O for this behavior, while wget saves with the original name automatically.
Running curl without -O when the goal is to save the file
Without -O, curl prints the retrieved content to the terminal instead of saving it.
Fix:
Use curl -O followed by the URL when the file should be saved with its remote filename.Expecting wget to behave like curl without automatic file saving
wget automatically saves files with their original names.
Fix:
Use the basic wget URL form for a remote file when its original name is suitable.Treating a reported HTTP error as a successful download
A 404 or another error status means the retrieval failed, and no file will be saved.
Fix:
Read the response status reported by the tool and correct the failed request before relying on the file.Using ordinary single-file retrieval when the task requires a website hierarchy
Following links recursively is a wget capability enabled with -r.
Fix:
Use wget -r for recursive downloading and consider -l when the depth must be limited.
Practice the Decision
You need to retrieve one binary image and save it using the filename supplied by the remote URL. Which command-line tool and option would you choose, and what would happen if you omitted the option?
Hints
- The task requires saving one file with its remote filename.
- curl needs a specific uppercase option for saving instead of displaying the response.
- wget automatically saves with the original name.
A documentation site contains linked webpages and resources that must be retrieved as a local website structure. Which tool and option match this task? What option can limit how many levels of links are followed?
Hints
- Look for the tool described as supporting recursive downloading.
- The recursive option is -r.
- The level-control option is -l.
Working Retrieval Checklist
- Choose the URL and decide whether the task is one-file retrieval, general data transfer, or recursive website downloading.
- Use curl with -O when curl should save the response using the remote filename.
- Use wget for automatic saving with the original filename or for recursive downloading with -r.
- Read the reported response status: 200 OK indicates successful retrieval, while 404 Not Found or another error indicates failure.
- After success, confirm that the expected file was saved under the expected remote filename.
Key Takeaways
- curl and wget retrieve web resources from the command line through HTTP or FTP.
- Both tools follow an HTTP request-response cycle and can retrieve text or binary files.
- curl needs the -O option to save a file using its remote filename; otherwise it prints the content to the terminal.
- wget automatically saves files with their original names and supports recursive downloading with -r.
- A 200 OK response indicates successful retrieval, while a 404 Not Found or another error status indicates that the download failed.