Concepts / Training a Convnet from Scratch on a Small Dataset

Training a Convnet from Scratch on a Small Dataset

The original Dogs vs. Cats dataset contains 25,000 color JPEG images.

  • Programming

From Archive to Training Set

Training a convolutional network from scratch begins before the model code runs. The image files must first be organized so the training process can find images by both subset and class. In the Dogs vs. Cats example, the original collection is reduced to a smaller working dataset and divided into training, validation, and test subsets.

Dataset preparation is part of the machine-learning workflow, not a separate detail to ignore. A mistake in the folders or file counts can prevent training from working correctly even when the model code is valid.

The Dogs vs. Cats Collection

The original Dogs vs. Cats dataset contains 25,000 medium-resolution color JPEG images. It contains 12,500 cat images and 12,500 dog images. The dataset was made available by Kaggle for a computer-vision competition in 2013, and its compressed size is 543 MB.

containscontainsDogs vs. Cats25,000 color JPEG imagesCats12,500 imagesDogs12,500 images
What are the two image categories in the original collection, and how many images does each contain?

The original collection is much larger than the smaller working dataset used in this exercise. The preparation process selects images and assigns them to separate subsets while preserving the two class categories: cats and dogs.

Three Separate Subsets

The working dataset is divided into training, validation, and test sets. Keeping these subsets separate gives the workflow distinct groups of images for the different stages of developing and assessing the convnet. Each subset must retain both categories, so every subset has a cats directory and a dogs directory.

assigned toassigned toassigned toSelected imagesSmaller working datasetTraining set2,000 imagesValidation set1,000 imagesTest set1,000 images
How are the selected images divided among the three subsets, and how does each subset remain separated from the others?
SubsetCatsDogsTotal
Training1,0001,0002,000
Validation5005001,000
Test5005001,000

Expected image counts in the smaller working dataset

Required Directory Tree

The directory structure mirrors the dataset split. There is a top-level working-dataset directory. Inside it are training, validation, and test directories. Inside each of those are two class directories: cats and dogs.

containscontainscontainscontainscontainscontainscontainscontainscontainsWorking datasettraincats1,000 imagescats500 imagescats500 imagesvalidationdogs1,000 imagesdogs500 imagesdogs500 imagestest
Where are the training, validation, and test images located, and how are the two classes represented inside each split?

Create all six class directories before copying images. This makes the destination of every image explicit: one of the cats or dogs directories inside train, validation, or test.

Moving Images into Place

Each selected image follows a simple path. First, the preparation process identifies whether the file belongs to the cat or dog class. Next, it uses the selected image range to determine whether the file belongs in the training, validation, or test subset. Finally, it copies the file into the matching class directory inside that subset.

identify categoryapply selected rangechoose foldercopy fileOriginal imageSource fileClasscat or dogSubsettrain, validation, or testClass directoryMatching subset and classCopied imageDestination file
How does one selected cat or dog image move from the original collection into the correct split directory?

Tracing a Selected Cat Image

Explain the decisions made before one selected cat image reaches its destination.

Identify the class: The image is treated as a cat image, so its destination must be a cats directory.

Identify the subset: The selected image range determines whether the destination is inside train, validation, or test.

Use the matching directory: The image is copied into the cats directory belonging to the selected subset.

Update the expected count: After copying, the relevant cats directory should contain one more image than it did before.

Every image is classified by category and assigned to one subset before it is copied into a destination directory.

Checking the File Counts

The first useful verification is a directory count. The source procedure uses os.listdir() to obtain the contents of a directory and len() to count those contents. Check all six class directories, not just one of them. A correct preparation should produce 1,000 training cats, 1,000 training dogs, 500 validation cats, 500 validation dogs, 500 test cats, and 500 test dogs.

python
compare within trainingcompare within validationcompare within testtrain/catsExpected: 1,000validation/catsExpected: 500test/catsExpected: 500train/dogsExpected: 1,000validation/dogsExpected: 500test/dogsExpected: 500
How can the observed file count in each class directory be compared with the expected count?

Preparation Mistakes

  • Checking only the training directories

    The preparation is incomplete even though one subset looks correct.

    Fix: Count all six class directories: two inside train, two inside validation, and two inside test.

  • Treating the original dataset count as the working-dataset count

    The working dataset is a smaller selection with the stated per-directory targets.

    Fix: Use the working-dataset targets: 1,000 images for each training class and 500 images for each validation and test class.

  • Copying an image into a class directory without applying the subset selection

    The image may be placed in the wrong subset, making the directory structure inconsistent with the intended split.

    Fix: Apply both decisions: identify the class and use the selected image range to identify the subset.

  • Starting model training before checking the folders

    A dataset-preparation failure can occur before any model code runs.

    Fix: Create the directory structure, copy the files, and verify the six counts first.

Check Your Dataset Plan

EASY

A prepared dataset has 1,000 files in train/cats, 1,000 in train/dogs, 500 in validation/cats, 500 in validation/dogs, 500 in test/cats, and 430 in test/dogs. Which directory requires investigation, and what preparation decision should you review first?

Hints
  • Compare each observed count with the expected count for its subset and class.
  • The test dogs directory has a different count from the other test class directory.
  • Review which selected image range was copied into the test dogs destination.

What do you think happens?

If the original dataset contains 12,500 cat images and 12,500 dog images, should the six directories in the smaller working dataset contain those same counts?

  • Yes, every directory should contain the full original category count
  • No, the smaller working dataset has its own expected counts
  • Only the training directories should contain the full original category count
Reveal answer

Answer: No, the smaller working dataset has its own expected counts.

The original collection contains 25,000 images, while the working dataset targets 1,000 cats and 1,000 dogs for training, plus 500 cats and 500 dogs for each of validation and test.

Preparation Checklist

  1. The original Dogs vs. Cats dataset contains 25,000 medium-resolution color JPEG images: 12,500 cats and 12,500 dogs.
  2. The smaller working dataset is divided into training, validation, and test subsets.
  3. Every subset contains separate cats and dogs directories.
  4. Selected image ranges determine which files are copied into each destination directory.
  5. The expected counts are 1,000 training cats, 1,000 training dogs, 500 validation cats, 500 validation dogs, 500 test cats, and 500 test dogs.
  6. os.listdir() and len() provide a direct way to inspect and count the contents of every class directory.

Key Takeaways

  • The original Dogs vs. Cats dataset has 25,000 color JPEG images split evenly between cats and dogs.
  • The smaller working dataset separates images into training, validation, and test subsets.
  • Each subset requires both a cats directory and a dogs directory.
  • An image reaches its destination through two decisions: its class and its selected subset.
  • Counting all six class directories is the first practical check for a correctly prepared dataset.