HTTP Protocol Fundamentals
curl and wget are command-line tools for retrieving files from the web on Unix-like systems using HTTP or FTP protocols.
Why Command-Line Retrieval Matters
A graphical browser is not always available or convenient. You may be working on a remote server, automating a task, or retrieving many files at once. curl and wget let you fetch webpages and files directly from a Unix-like command line. Both tools can retrieve content using HTTP or FTP protocols.
The central difference to remember at the beginning is simple: curl prints retrieved content to the terminal unless you use the -O flag to save it, while wget automatically saves files using their original names.
Tracing the HTTP Exchange
When you run curl or wget, the command sends an HTTP GET request to the server named in the URL. The request identifies the content the tool wants. The server locates that content and sends a response back. The command-line tool then either displays the received content or saves it to disk, depending on the command and its options.
The response also provides a status. A 200 OK response means the file was successfully retrieved. A 404 Not Found or another error status indicates that retrieval failed, and the source material states that no file will be saved in that case.
Saving One File with curl
The -O option tells curl to save the retrieved content in the current directory using the filename from the remote URL. In this source example, the remote file is cover.jpg, so curl saves it locally with that same filename. The file is a binary image, but the same retrieval approach also works for plain text files.
Following a curl Download
You need to retrieve the source example's cover.jpg file and keep the remote filename locally.
Choose curl: Use curl as the command-line retrieval tool.
Add -O: The -O option instructs curl to use the filename supplied by the remote URL instead of printing the response to the terminal.
Send the request: curl sends an HTTP GET request for cover.jpg to the remote server.
Save the response: If the server responds with 200 OK, the retrieved file is saved in the current directory as cover.jpg.
The expected local result is a file named cover.jpg in the current directory.
Retrieving Files with wget
wget also sends a request for the remote file and saves the result in the current directory. Like curl -O, it uses the remote filename, so the source example saves cover.jpg as cover.jpg locally. wget can retrieve webpages as well as other remote files.
Choosing Between the Tools
| Task | curl | wget |
|---|---|---|
| Retrieve one file | Use curl with -O to save the remote filename | Use wget; it automatically saves the remote filename |
| Display retrieved content | Without -O, curl prints content to the terminal | wget is described primarily as saving retrieved files |
| Download a webpage or remote file | Can retrieve web content | Can retrieve webpages and remote files |
| Download a website hierarchy | The source identifies curl as flexible for general-purpose data transfer | Use -r for recursive downloading; use -l to limit depth |
Both tools perform the basic download task through the same HTTP request-response cycle. The practical choice depends on the retrieval task. curl is presented as more flexible for general-purpose data transfer, while wget is particularly useful when the goal is to save files automatically or traverse a website recursively.
Following Links Recursively
The -r option enables recursive downloading in wget. After wget retrieves a webpage, it parses the HTML for links and resource references. It then retrieves those connected pages and resources, repeating the process recursively. This can create a local mirror of the website structure and is useful for archiving websites, backing up documentation, or preparing resources for offline viewing.
Diagnosing Download Results
Running curl without -O when the goal is to save a file
Without -O, curl prints the retrieved content to the terminal rather than saving it to disk.
Fix:
Add -O when you want curl to save the file using the remote filename.Assuming an error response produced a valid local download
The source states that a 404 or another error status means retrieval failed and no file will be saved.
Fix:
Check the reported HTTP result and expect a 200 OK response for a successfully retrieved file.Using a single-file command when the task requires a website hierarchy
A single retrieval does not perform the recursive link traversal described for wget.
Fix:
Use wget with -r for recursive downloading, and use -l when you need to limit traversal depth.
Verify both parts of the result: first, confirm that the server response indicates success with 200 OK; second, confirm that the expected filename was saved in the current directory. For curl, also check that you used -O when saving was intended.
Practice the Decision
You need to retrieve one binary image and save it with the filename supplied by the remote URL. Then you need to download a connected website hierarchy while limiting how deeply links are followed. Which tool and option would you choose for each task?
Hints
- For the single image, decide whether curl needs an option to save instead of printing.
- For the website hierarchy, look for wget's recursive option and its depth-control option.
What do you think happens?
What happens if you run curl for a remote file but leave out -O?
Reveal answer
Answer: curl prints the retrieved content to the terminal.
The -O flag is what tells curl to save the file using the remote filename. Recursive website downloading is a wget capability enabled with -r.
Key Takeaways
- curl and wget retrieve web content from the command line through HTTP or FTP.
- Both tools send an HTTP GET request and process the server response.
- Use curl -O when curl should save a file with the remote filename; without -O, curl prints content to the terminal.
- wget automatically saves retrieved files with their original names and supports recursive downloading with -r.
- A 200 OK response indicates successful retrieval, while a 404 Not Found or another error status indicates failure.
Key Takeaways
- curl and wget are command-line tools for retrieving web content without relying on a graphical browser.
- The HTTP process is a request-response cycle: the tool sends an HTTP GET request, the server returns a response, and the tool displays or saves the data.
- curl needs -O to save a remote file with its original filename, while wget saves retrieved files automatically.
- wget -r follows links and resource references recursively, and -l can limit the traversal depth.
- Use the HTTP status and the expected local filename together to verify whether retrieval succeeded.