Using Data Augmentation
The original Dogs vs. Cats dataset contains 25,000 color JPEG images.
From Images to a Usable Dataset
A computer-vision model cannot use a collection of images effectively unless the images are organized so the training process can find them by subset and class. The Dogs vs. Cats dataset begins as 25,000 medium-resolution color JPEG images: 12,500 cat images and 12,500 dog images. Before training, this collection is arranged into smaller subsets and class-specific directories.
Dataset preparation is part of the machine-learning workflow. A failure in the folders or file counts can occur before any model code runs.
Tracing the Dataset Split
The original collection is divided into three smaller working subsets: training, validation, and test. Each subset has a separate role in the dataset arrangement, while the selected image ranges determine which files belong to each one. This prevents the entire original collection from being treated as one undifferentiated directory.
| Subset | Cat images | Dog images | Total images |
|---|---|---|---|
| Training | 1,000 | 1,000 | 2,000 |
| Validation | 500 | 500 | 1,000 |
| Test | 500 | 500 | 1,000 |
Expected image counts in the smaller working dataset
Reading the Split Counts
Determine how many images should be present in the training subset.
Find the class counts: The training subset is expected to contain 1,000 cat images and 1,000 dog images.
Combine the classes: Adding the two class directories gives 2,000 training images.
The training subset should contain 2,000 images in total, evenly divided between cats and dogs.
Building the Destination Folders
Every split contains two class directories: one for cats and one for dogs. The resulting structure has six class-specific destinations: training cats, training dogs, validation cats, validation dogs, test cats, and test dogs. This organization lets the training process locate an image through both its subset and its class.
The copy process follows the same classification logic as the folder structure. A source cat image is copied into the cat directory belonging to the selected split. A source dog image is copied into the dog directory belonging to that split. The selected image ranges decide the split, and the image's class decides the destination directory within that split.
Checking the Directory Contents
The first practical check is to count the files in every class directory. The source procedure uses os.listdir() to obtain the contents of a directory and len() to count those contents. The purpose is not merely to confirm that folders exist; it is to compare each actual count with the expected count.
| Directory | Expected count |
|---|---|
| Training cats | 1,000 |
| Training dogs | 1,000 |
| Validation cats | 500 |
| Validation dogs | 500 |
| Test cats | 500 |
| Test dogs | 500 |
Counts to verify before model training
A validation cat directory contains 480 images instead of the expected 500. What should you check first?
Hints
- Compare the actual directory count with the expected count.
- Review which image range was assigned to validation.
- Check that cat images were copied into the validation cat directory.
Growing Variation with Transformations
Data augmentation is a technique that applies transformations to existing data to produce additional examples. It artificially increases the size of a dataset rather than requiring an entirely new collection of original images.
The original image is the starting point for the augmented collection. A transformation is then applied to that existing data, producing an additional training example. Transformations can include operations such as rotation, flipping, or cropping. The important distinction is that the transformed example is derived from existing data; it is not another independently collected original image.
| Original data | Transformed data |
|---|---|
| The starting image in the collection | An additional example produced by applying a transformation |
| Part of the source dataset | Derived from existing data |
| Provides the starting point for augmentation | Contributes to an artificially larger dataset |
Following One Image Through Augmentation
Explain what happens when an existing image is used to create additional training examples.
Start with original data: An image from the existing collection is used as the starting point.
Apply a transformation: A transformation such as rotation, flipping, or cropping is applied to the existing image.
Add the resulting example: The transformed result becomes an additional training example in the artificially larger dataset.
The augmented collection contains the original starting data and additional transformed examples derived from it.
Mistakes in Dataset Preparation
Treating the original 25,000-image collection as one training directory.
The required working dataset is divided into three subsets, and each subset has separate class directories.
Fix:
Create the three split destinations and place each image in the class directory belonging to its selected split.Creating split folders without separate cat and dog directories.
The training process must be able to find images by both subset and class.
Fix:
Use separate cats and dogs directories inside training, validation, and test.Assuming that existing folders prove the dataset is correct.
A preparation failure can involve file counts even when the folder structure appears complete.
Fix:
Use directory listing and length counting to verify all six expected counts.Confusing transformed examples with newly collected original data.
Data augmentation applies transformations to existing data.
Fix:
Treat the original image as the starting data and the rotated, flipped, or cropped result as a derived training example.
Practice and Recall
Explain the complete preparation flow in your own words: begin with the 25,000-image Dogs vs. Cats dataset, describe how selected image ranges produce the three subsets, name the two class directories inside each subset, state the six expected directory counts, and explain where data augmentation fits into the training process.
Hints
- Mention training, validation, and test.
- Mention cats and dogs inside every split.
- Use the expected counts of 1,000, 1,000, 500, 500, 500, and 500.
- Distinguish original images from transformed training examples.
What do you think happens?
A directory contains 1,000 training cats, 1,000 training dogs, 500 validation cats, 500 validation dogs, 500 test cats, and 500 test dogs. Is the directory count consistent with the expected dataset split?
Reveal answer
Answer: Yes
The six class-directory counts match the expected counts for the smaller working dataset.
Key Takeaways
- The Dogs vs. Cats dataset contains 25,000 color JPEG images: 12,500 cats and 12,500 dogs.
- The working dataset is divided into training, validation, and test subsets according to selected image ranges.
- Each subset contains separate cats and dogs directories, producing six class-specific destinations.
- The expected counts are 1,000 training cats, 1,000 training dogs, 500 validation cats, 500 validation dogs, 500 test cats, and 500 test dogs.
- Data augmentation applies transformations to existing data to create additional training examples and mitigate overfitting.
Key Takeaways
- The original Dogs vs. Cats dataset contains 25,000 color JPEG images, split evenly between cats and dogs.
- The smaller working dataset uses training, validation, and test subsets, each with separate cat and dog directories.
- Counting files in all six class directories is the first useful check for dataset-preparation errors.
- Data augmentation creates additional examples by transforming existing images.
- Original images are the starting data, while transformed images are derived training examples used to increase variation and mitigate overfitting.