Building Network Applications with Python
curl and wget are command-line tools for retrieving files from the web on Unix-like systems using HTTP or FTP protocols.
Retrieving Files Without a Browser
When you need a file from the web, a graphical browser is not always the most useful tool. You may be working on a remote server, automating a task, or retrieving many files at once. On Unix-like systems, curl and wget let you retrieve webpages and files directly from the command line. Both tools can use HTTP or FTP, and both can download plain text as well as binary data.
The important distinction is not whether curl or wget can retrieve a file. Both can. The practical distinction is how they save content and which retrieval tasks their features support most directly.
Following the Request
curl and wget follow the same fundamental HTTP request-response pattern. Your command identifies a URL, and the command-line tool sends an HTTP GET request for that resource. The server receives the request, locates the requested file, and sends a response back. The tool then either displays the received content or saves it to disk, depending on the command and its options.
A successful response has a 200 OK status. In that case, the requested file is successfully retrieved. A 404 Not Found response, or another error status, indicates that the retrieval failed; the tools report the failure and no file is saved according to the source example.
What do you think happens?
Suppose curl retrieves a URL without the -O option. Where does the received content go?
Reveal answer
Answer: It is printed to the terminal.
curl requires the -O flag to save the file. Without -O, curl prints the content to the terminal instead.
Saving One Remote File
curl is a command-line tool that retrieves content from the web. The -O option tells curl to save the retrieved file in the current directory using the filename from the remote URL.
In this source example, curl requests cover.jpg from www.py4e.com and saves the result in the current directory with the same filename. The -O option is the key state change: curl moves from displaying retrieved content to saving it as a local file. The transferred content may be plain text or binary data; the request-response process is the same.
wget also downloads cover.jpg to the current directory and automatically uses the remote filename. Unlike curl, wget does not need an -O option for this basic saving behavior.
| Retrieval need | curl | wget |
|---|---|---|
| Retrieve one remote file | Use curl with -O when the file should be saved using its remote filename | Downloads and saves using the remote filename automatically |
| Display retrieved content | Without -O, content is printed to the terminal | The source emphasizes wget's automatic file-saving behavior |
| Retrieve many linked resources | Flexible for general-purpose data transfer | Supports recursive downloading with -r |
Mirroring Linked Resources
wget becomes especially useful when the target is not one isolated file. With the -r option, wget recursively downloads an entire website hierarchy. It follows the website's link structure instead of stopping after the first page or file.
The command above is a generated illustration of the -r option. The important behavior comes from the source: when wget encounters a page, it parses the HTML for links and resource references, downloads those connected resources, and continues the process recursively. The result is a local mirror of the website structure.
Recursive downloading can be broad, so wget provides a way to limit how deep it follows links. The -l option controls the number of levels. This lets you choose between following only nearby links and traversing more of the linked website hierarchy.
Tracing a Failed Download
Debugging a download begins with the same request-response cycle used for a successful transfer. First identify the URL sent in the request. Then examine the server response. A 200 OK response indicates successful retrieval. A 404 Not Found response or another error status indicates failure, and the source states that no file is saved in that situation.
Using curl without -O when the goal is to save a file
Without -O, curl prints the retrieved content to the terminal instead of saving it to disk.
Fix:
Use curl -O followed by the URL when the remote filename should be used for the saved file.Assuming wget requires the same saving option as curl
wget automatically saves downloaded files using their original names.
Fix:
Use wget with the URL for the basic retrieval case.Treating a failed server response as a successful download
A 404 or another error status means the retrieval failed, and the source states that no file is saved.
Fix:
Read the reported response status, verify the requested URL, and retry only after correcting the retrieval problem.Using ordinary single-file retrieval when the task requires linked website resources
The command has not been given wget's recursive behavior.
Fix:
Use wget with -r for recursive website downloading, and use -l when the traversal depth should be limited.
Verify retrieval in two stages: interpret the server status and then confirm that the expected local file was saved. A 200 OK response establishes successful retrieval in the source's request-response model; the saved filename establishes where the result was written for commands that save content.
Selecting the Retrieval Tool
One File or a Website?
Choose a command-line approach for each retrieval goal.
Goal 1: save one file with its remote filename: curl can do this when used with -O. wget also does this automatically for a basic remote-file download.
Goal 2: retrieve content without asking curl to save it: curl without -O prints the retrieved content to the terminal.
Goal 3: obtain connected pages and resources: wget is the stronger fit because -r enables recursive downloading, and -l can limit the depth.
Use curl -O or wget for a single saved file, curl without -O when terminal output is intended, and wget -r for recursive website retrieval.
A remote server hosts a binary image named cover.jpg. You want the image saved in the current directory under its remote filename. Which curl option is essential, and what would happen if you omitted it?
Hints
- Focus on the difference between curl and wget when saving content.
- The option uses the letter O.
You need to retrieve a webpage and then follow its links to download connected pages and resources. Which tool and option match this task, and which option can limit how deeply links are followed?
Hints
- Look for the tool that supports recursive downloading.
- The recursive option is -r; the depth option begins with -l.
Key Takeaways
- curl and wget retrieve web content from the command line using HTTP or FTP on Unix-like systems.
- Both tools use an HTTP request-response cycle and can retrieve plain text and binary files.
- curl requires -O to save a file using the remote filename; without -O, curl prints the content to the terminal.
- wget automatically saves ordinary downloads with their original names and supports recursive website downloading with -r.
- A 200 OK response indicates successful retrieval, while a 404 Not Found or another error status indicates failure and no saved file in the source model.
Key Takeaways
- curl and wget make web retrieval possible from a Unix-like command line without relying on a graphical browser.
- The tools send requests, receive server responses, and then display or save the returned content.
- Use curl -O when curl should save a remote file using its URL filename.
- Use wget for automatic single-file saving and especially for recursive website downloading with -r.
- Use the response status and the expected saved file as the first checks when debugging retrieval.