Concepts / Automating Network Tasks with Shell Scripts

Automating Network Tasks with Shell Scripts

curl and wget are command-line tools for retrieving files from the web on Unix-like systems using HTTP or FTP protocols.

  • Programming

Command-Line Retrieval

A graphical browser is not always available when you need a file from the web. You may be working on a remote server, automating a task, or retrieving many files at once. On Unix-like systems, curl and wget let you retrieve webpages and files directly from the command line using HTTP or FTP protocols.

The important difference for a first curl download is that curl needs the -O flag to save the retrieved file. Without -O, curl prints the retrieved content to the terminal. wget automatically saves a retrieved file using its original filename.

The Request-Response Path

curl and wget follow the same fundamental HTTP request-response pattern. The command sends an HTTP GET request containing the URL. The remote server receives the request, locates the requested content, and sends data back. The tool then either saves that data to disk or displays it on screen, depending on the command and its options.

HTTP GET requestsends data or statussave or displaycurl or wgetcommand with URLRemote serverreceives GET requestHTTP responsecontent or error statusLocal resultsaved file or terminaloutput
What happens between running curl or wget, sending an HTTP request, receiving a response, and saving the retrieved content?

Saving a Remote Filename

Suppose the remote URL ends with cover.jpg. The -O option tells curl to use that remote filename when creating the local file. This applies to binary content such as an image as well as to plain text content: the retrieved data is returned by the server, and curl saves it instead of printing it to the terminal.

Output
A local file named cover.jpg is created in the current directory.
filename is availablesaves ascover.jpgremote filename-Ouse remote filenamecover.jpglocal file
How does the remote filename move from the URL into a newly created local file when curl is used with -O?
  • Running curl with only the URL when the goal is to save the file

    Without -O, curl prints the file contents to the terminal instead of saving them to disk.

    Fix: Use curl -O followed by the URL when you want the remote filename used for a local file.

Choosing the Retrieval Tool

Taskcurlwget
Retrieve one remote fileUse -O when the file should be saved with its remote filenameAutomatically saves the file with its original filename
Display retrieved contentWithout -O, curl prints content to the terminalThe source material emphasizes wget's automatic file saving rather than terminal display
Retrieve a website hierarchyNot identified in the source as the typical strengthUse -r for recursive downloading
General-purpose transferMore flexible for general-purpose data transferPowerful for more complex retrieval tasks, including recursive downloads

For a single file, either tool can perform the basic retrieval. Choose curl when you want its flexible general-purpose data-transfer behavior or when you specifically want curl's -O filename behavior. Choose wget when automatic saving or recursive retrieval is central to the task.

Output
wget downloads cover.jpg and saves it in the current directory using the remote filename.

Recursive Website Retrieval

wget can move beyond one file with its -r option. When recursive downloading is enabled, wget processes a webpage, parses its HTML for links and resource references, downloads the connected pages and resources, and repeats the process. The result is a local folder structure that mirrors the website structure. The -l option can limit how many levels deep wget follows links.

follows hyperlinkfinds referencefollows recursivelysaves locallyadds to mirrorWebpagestarting URLLinked pagedownloaded from a linkNested pagenext linked levelLocal folderstructurewebsite mirrorResourcelinked file or reference
How does wget discover a webpage's linked resources and continue downloading them into a local folder structure?

Tracing Download Failures

  1. Confirm that the command uses the intended URL.
  2. Check whether the tool reports an HTTP error status such as 404 Not Found.
  3. If using curl to save a file, confirm that -O is present; otherwise the content is printed to the terminal.
  4. Check the current directory for the expected filename after a successful retrieval.
  5. Treat a missing expected file as a reason to inspect the reported response and command options again.
receives resultsuccessfailureverify saved contentconfirm no expected fileRun commandcurl or wgetHTTP responsesuccess or error200 OKretrievedExpected filecheck local result404 Not Foundreported failure
How can a learner trace a failed download from command execution through network or HTTP errors to checking whether the correct file exists locally?
  • Assuming that a command succeeded because it ran

    The server may return 404 Not Found or another error status, and the tool reports that failure.

    Fix: Read the reported HTTP result and verify that the expected local file exists.

  • Treating curl and wget as identical in every situation

    Both retrieve files, but wget supports recursive downloading and curl is described as more flexible for general-purpose data transfer.

    Fix: Use the tool whose strength matches the task: curl for flexible transfer and wget for automatic saving or recursive retrieval.

  • Using recursive retrieval without considering depth

    Recursive wget follows links and resource references through multiple levels.

    Fix: Use wget's -l option when the depth of traversal needs to be limited.

Practice and Review

EASY

You need to retrieve a binary image named cover.jpg and save it with the filename supplied by the remote URL. Which command should you choose, and what would happen if you removed the filename-saving option from the curl version?

Hints
  • The source example uses cover.jpg.
  • Recall which curl option tells the tool to use the remote filename.
  • Compare that behavior with wget's automatic filename handling.
MEDIUM

You need a local copy of a website's linked pages and resources, but you do not want wget to follow links without a depth limit. Describe the wget options you would investigate and explain why.

Hints
  • Recursive behavior is controlled by -r.
  • The source identifies -l as the option for limiting traversal depth.
  1. curl and wget retrieve web content from the command line using HTTP or FTP protocols.
  2. Both tools participate in an HTTP request-response cycle and can retrieve plain text and binary files.
  3. curl needs -O to save a file with the remote filename; without it, curl prints content to the terminal.
  4. wget automatically saves files with their original names and supports recursive website downloading with -r.
  5. A successful retrieval should be checked through the reported HTTP result and by verifying the expected local file.

Key Takeaways

  • curl and wget make web retrieval possible from a Unix-like command line without relying on a graphical browser.
  • curl -O saves retrieved content using the filename from the URL, while curl without -O prints the content to the terminal.
  • wget automatically saves remote files and is especially useful for recursive website retrieval with -r.
  • Every retrieval follows a request-response process, so HTTP status results are essential for debugging.
  • After a successful response, verify that the expected file was saved locally with the intended filename.