Training a Convnet from Scratch on a Small Dataset
The original Dogs vs. Cats dataset contains 25,000 color JPEG images.
From Archive to Training Set
Training a convolutional network from scratch begins before the model code runs. The image files must first be organized so the training process can find images by both subset and class. In the Dogs vs. Cats example, the original collection is reduced to a smaller working dataset and divided into training, validation, and test subsets.
Dataset preparation is part of the machine-learning workflow, not a separate detail to ignore. A mistake in the folders or file counts can prevent training from working correctly even when the model code is valid.
The Dogs vs. Cats Collection
The original Dogs vs. Cats dataset contains 25,000 medium-resolution color JPEG images. It contains 12,500 cat images and 12,500 dog images. The dataset was made available by Kaggle for a computer-vision competition in 2013, and its compressed size is 543 MB.
The original collection is much larger than the smaller working dataset used in this exercise. The preparation process selects images and assigns them to separate subsets while preserving the two class categories: cats and dogs.
Three Separate Subsets
The working dataset is divided into training, validation, and test sets. Keeping these subsets separate gives the workflow distinct groups of images for the different stages of developing and assessing the convnet. Each subset must retain both categories, so every subset has a cats directory and a dogs directory.
| Subset | Cats | Dogs | Total |
|---|---|---|---|
| Training | 1,000 | 1,000 | 2,000 |
| Validation | 500 | 500 | 1,000 |
| Test | 500 | 500 | 1,000 |
Expected image counts in the smaller working dataset
Required Directory Tree
The directory structure mirrors the dataset split. There is a top-level working-dataset directory. Inside it are training, validation, and test directories. Inside each of those are two class directories: cats and dogs.
Create all six class directories before copying images. This makes the destination of every image explicit: one of the cats or dogs directories inside train, validation, or test.
Moving Images into Place
Each selected image follows a simple path. First, the preparation process identifies whether the file belongs to the cat or dog class. Next, it uses the selected image range to determine whether the file belongs in the training, validation, or test subset. Finally, it copies the file into the matching class directory inside that subset.
Tracing a Selected Cat Image
Explain the decisions made before one selected cat image reaches its destination.
Identify the class: The image is treated as a cat image, so its destination must be a cats directory.
Identify the subset: The selected image range determines whether the destination is inside train, validation, or test.
Use the matching directory: The image is copied into the cats directory belonging to the selected subset.
Update the expected count: After copying, the relevant cats directory should contain one more image than it did before.
Every image is classified by category and assigned to one subset before it is copied into a destination directory.
Checking the File Counts
The first useful verification is a directory count. The source procedure uses os.listdir() to obtain the contents of a directory and len() to count those contents. Check all six class directories, not just one of them. A correct preparation should produce 1,000 training cats, 1,000 training dogs, 500 validation cats, 500 validation dogs, 500 test cats, and 500 test dogs.
Preparation Mistakes
Checking only the training directories
The preparation is incomplete even though one subset looks correct.
Fix:
Count all six class directories: two inside train, two inside validation, and two inside test.Treating the original dataset count as the working-dataset count
The working dataset is a smaller selection with the stated per-directory targets.
Fix:
Use the working-dataset targets: 1,000 images for each training class and 500 images for each validation and test class.Copying an image into a class directory without applying the subset selection
The image may be placed in the wrong subset, making the directory structure inconsistent with the intended split.
Fix:
Apply both decisions: identify the class and use the selected image range to identify the subset.Starting model training before checking the folders
A dataset-preparation failure can occur before any model code runs.
Fix:
Create the directory structure, copy the files, and verify the six counts first.
Check Your Dataset Plan
A prepared dataset has 1,000 files in train/cats, 1,000 in train/dogs, 500 in validation/cats, 500 in validation/dogs, 500 in test/cats, and 430 in test/dogs. Which directory requires investigation, and what preparation decision should you review first?
Hints
- Compare each observed count with the expected count for its subset and class.
- The test dogs directory has a different count from the other test class directory.
- Review which selected image range was copied into the test dogs destination.
What do you think happens?
If the original dataset contains 12,500 cat images and 12,500 dog images, should the six directories in the smaller working dataset contain those same counts?
Reveal answer
Answer: No, the smaller working dataset has its own expected counts.
The original collection contains 25,000 images, while the working dataset targets 1,000 cats and 1,000 dogs for training, plus 500 cats and 500 dogs for each of validation and test.
Preparation Checklist
- The original Dogs vs. Cats dataset contains 25,000 medium-resolution color JPEG images: 12,500 cats and 12,500 dogs.
- The smaller working dataset is divided into training, validation, and test subsets.
- Every subset contains separate cats and dogs directories.
- Selected image ranges determine which files are copied into each destination directory.
- The expected counts are 1,000 training cats, 1,000 training dogs, 500 validation cats, 500 validation dogs, 500 test cats, and 500 test dogs.
- os.listdir() and len() provide a direct way to inspect and count the contents of every class directory.
Key Takeaways
- The original Dogs vs. Cats dataset has 25,000 color JPEG images split evenly between cats and dogs.
- The smaller working dataset separates images into training, validation, and test subsets.
- Each subset requires both a cats directory and a dogs directory.
- An image reaches its destination through two decisions: its class and its selected subset.
- Counting all six class directories is the first practical check for a correctly prepared dataset.