Concepts / Evaluating a Convnet

Evaluating a Convnet

The original Dogs vs. Cats dataset contains 25,000 color JPEG images.

  • Programming

From Image Collection to Evaluation

Evaluating a convolutional neural network begins before any model code runs. The images must first be arranged so that the training process can find them by both subset and class. In the Dogs vs. Cats task, that means creating separate locations for training, validation, and test images, with separate cat and dog directories inside each location.

The original Dogs vs. Cats dataset contains 25,000 medium-resolution color JPEG images: 12,500 cat images and 12,500 dog images.

The working dataset is smaller than the original collection. It is divided into training, validation, and test subsets, and every subset contains separate cats and dogs directories.

The Three-Way Dataset Split

The dataset is organized into three named subsets: training, validation, and test. This division gives the convnet workflow distinct groups of images to work with instead of one undifferentiated collection. The subset name records where an image belongs in the workflow, while the class directory records whether the image is a cat or a dog.

The selected image ranges determine which original files belong to each subset. The important preparation task is therefore not only choosing images, but also placing every selected file in the directory that matches both its subset and its class.

selected filesselected filesselected filesOriginal dataset25,000 color JPEG imagesTrainingcats and dogsValidationcats and dogsTestcats and dogs
How the original image collection becomes three organized subsets for the convnet workflow.

Required Folder Hierarchy

Each subset needs two class directories. The resulting hierarchy has a training directory containing cats and dogs, a validation directory containing cats and dogs, and a test directory containing cats and dogs. This two-level organization lets the image-loading process identify both the subset and the class from the file location.

Working datasettraincatscatscatsvalidationdogsdogsdogstest
What folder structure is required before the convnet can use the split images by subset and class.

Copying Files into Place

Dataset preparation follows a repeated routing operation. First, a selected original image is identified as belonging to the cat or dog class. Next, its selected range determines whether it belongs in training, validation, or test. Finally, the file is copied into the matching class directory inside that subset.

selected cat filesselected cat filesselected cat filesselected dog filesselected dog filesselected dog filesOriginal cat filesTraining catsOriginal dog filesTraining dogsValidation catsValidation dogsTest catsTest dogs
How selected cat and dog files move from the original collection into the matching subset and class directories.

Tracing One Class Through the Split

Explain the destination of selected cat images as the working dataset is prepared.

Identify the class: The source files are cat images, so they must be routed to cats directories rather than dogs directories.

Apply the selected range: The selected image range determines whether a particular cat file belongs to training, validation, or test.

Copy to the destination: The file is copied into the cats directory inside the subset selected by its range.

Every selected cat file ends in exactly one class directory that also identifies its subset.

Counting Every Destination

Copying files is not the final step. A preparation failure can happen before any model code runs, so the directories should be checked directly. The source procedure obtains the contents of each directory with os.listdir() and counts those contents with len(). The resulting counts can then be compared with the expected count for that directory.

DirectoryExpected image count
Training cats1,000
Training dogs1,000
Validation cats500
Validation dogs500
Test cats500
Test dogs500

Expected counts for the six class directories in the working dataset.

comparecomparecomparecompareTraining catsExpected: 1,000Directory countTraining catsTraining dogsExpected: 1,000Directory countTraining dogsValidation cats anddogsExpected: 500 eachDirectory countsValidation cats and dogsTest cats and dogsExpected: 500 eachDirectory countsTest cats and dogs
How to compare the count obtained from each class directory with its expected count.

Checking the Six Counts

Verify the expected contents of the working dataset.

Count training classes: The training cats directory should contain 1,000 images, and the training dogs directory should also contain 1,000 images.

Count validation classes: The validation cats directory should contain 500 images, and the validation dogs directory should also contain 500 images.

Count test classes: The test cats directory should contain 500 images, and the test dogs directory should also contain 500 images.

Compare results: Compare each directory count with its corresponding expected value rather than checking only one directory.

The preparation check covers six directories: two classes in each of three subsets.

Preparation Mistakes

  • Creating only one class directory inside a subset.

    Every split is required to contain separate cats and dogs directories.

    Fix: Create both class destinations inside training, validation, and test.

  • Copying a file to a subset without matching its class.

    The destination directory identifies the image class.

    Fix: Route cat files to cats directories and dog files to dogs directories.

  • Checking only one directory after copying.

    A preparation failure can affect any subset or class directory.

    Fix: Use directory contents and their lengths to verify all six expected counts.

  • Assuming that successful copying proves the dataset is ready.

    The source identifies counting as the most useful first check for a dataset-preparation failure.

    Fix: Compare every actual directory count with its expected count before training.

Practice the Routing Check

EASY

A working dataset has the correct training cats count and the correct validation counts, but the test dogs directory contains fewer images than expected. Which part of the preparation process should you inspect first, and which expected count should you use for comparison?

Hints
  • Identify the subset and class named by the directory.
  • Use the expected-count table rather than another directory's count.

What do you think happens?

What should the test dogs directory be compared with?

  • 1,000 images
  • 500 images
  • The total number of original dog images
Reveal answer

Answer: 500 images

The expected count is determined by both the subset and the class. Test dogs is one of the two test class directories, and each test class directory is expected to contain 500 images.

What to Remember

  1. The original Dogs vs. Cats dataset contains 25,000 color JPEG images: 12,500 cats and 12,500 dogs.
  2. The working dataset is divided into training, validation, and test subsets.
  3. Every subset must contain separate cats and dogs directories.
  4. Selected image ranges determine which files are copied into each destination.
  5. Use os.listdir() and len() to count directory contents and check all six expected counts: 1,000 for each training class and 500 for each validation and test class.

Key Takeaways

  • A convnet cannot use the Dogs vs. Cats images reliably until the files are arranged by subset and class.
  • The required structure has training, validation, and test directories, each containing cats and dogs directories.
  • Selected image ranges determine where files are copied.
  • Directory counts provide an early check for dataset-preparation errors.
  • The expected counts are 1,000 training cats, 1,000 training dogs, 500 validation cats, 500 validation dogs, 500 test cats, and 500 test dogs.