Concepts / Using Data Augmentation

Using Data Augmentation

The original Dogs vs. Cats dataset contains 25,000 color JPEG images.

  • Programming

From Images to a Usable Dataset

A computer-vision model cannot use a collection of images effectively unless the images are organized so the training process can find them by subset and class. The Dogs vs. Cats dataset begins as 25,000 medium-resolution color JPEG images: 12,500 cat images and 12,500 dog images. Before training, this collection is arranged into smaller subsets and class-specific directories.

Dataset preparation is part of the machine-learning workflow. A failure in the folders or file counts can occur before any model code runs.

Tracing the Dataset Split

The original collection is divided into three smaller working subsets: training, validation, and test. Each subset has a separate role in the dataset arrangement, while the selected image ranges determine which files belong to each one. This prevents the entire original collection from being treated as one undifferentiated directory.

selected image rangeselected image rangeselected image rangeOriginal dataset25,000 color JPEG imagesTraining2,000 imagesValidation1,000 imagesTest1,000 images
How is the original image collection organized into the three dataset subsets?
SubsetCat imagesDog imagesTotal images
Training1,0001,0002,000
Validation5005001,000
Test5005001,000

Expected image counts in the smaller working dataset

Reading the Split Counts

Determine how many images should be present in the training subset.

Find the class counts: The training subset is expected to contain 1,000 cat images and 1,000 dog images.

Combine the classes: Adding the two class directories gives 2,000 training images.

The training subset should contain 2,000 images in total, evenly divided between cats and dogs.

Building the Destination Folders

Every split contains two class directories: one for cats and one for dogs. The resulting structure has six class-specific destinations: training cats, training dogs, validation cats, validation dogs, test cats, and test dogs. This organization lets the training process locate an image through both its subset and its class.

DatasetTrainingCats1,000 imagesCats500 imagesCats500 imagesValidationDogs1,000 imagesDogs500 imagesDogs500 imagesTest
Which class directories are contained inside each dataset split?

The copy process follows the same classification logic as the folder structure. A source cat image is copied into the cat directory belonging to the selected split. A source dog image is copied into the dog directory belonging to that split. The selected image ranges decide the split, and the image's class decides the destination directory within that split.

identify subsetidentify subsetcopy by classcopy by classCat imageoriginal collectionSelected splittraining, validation, ortestCat directoryinside selected splitDog imageoriginal collectionSelected splittraining, validation, ortestDog directoryinside selected split
How does an image move from the original collection into the correct split and class directory?

Checking the Directory Contents

The first practical check is to count the files in every class directory. The source procedure uses os.listdir() to obtain the contents of a directory and len() to count those contents. The purpose is not merely to confirm that folders exist; it is to compare each actual count with the expected count.

DirectoryExpected count
Training cats1,000
Training dogs1,000
Validation cats500
Validation dogs500
Test cats500
Test dogs500

Counts to verify before model training

EASY

A validation cat directory contains 480 images instead of the expected 500. What should you check first?

Hints
  • Compare the actual directory count with the expected count.
  • Review which image range was assigned to validation.
  • Check that cat images were copied into the validation cat directory.

Growing Variation with Transformations

Data augmentation is a technique that applies transformations to existing data to produce additional examples. It artificially increases the size of a dataset rather than requiring an entirely new collection of original images.

The original image is the starting point for the augmented collection. A transformation is then applied to that existing data, producing an additional training example. Transformations can include operations such as rotation, flipping, or cropping. The important distinction is that the transformed example is derived from existing data; it is not another independently collected original image.

rotationflippingcroppingOriginal imagestarting dataRotated imagetransformed exampleFlipped imagetransformed exampleCropped imagetransformed example
How can one original image become additional training examples through transformations?
Original dataTransformed data
The starting image in the collectionAn additional example produced by applying a transformation
Part of the source datasetDerived from existing data
Provides the starting point for augmentationContributes to an artificially larger dataset

Following One Image Through Augmentation

Explain what happens when an existing image is used to create additional training examples.

Start with original data: An image from the existing collection is used as the starting point.

Apply a transformation: A transformation such as rotation, flipping, or cropping is applied to the existing image.

Add the resulting example: The transformed result becomes an additional training example in the artificially larger dataset.

The augmented collection contains the original starting data and additional transformed examples derived from it.

Mistakes in Dataset Preparation

  • Treating the original 25,000-image collection as one training directory.

    The required working dataset is divided into three subsets, and each subset has separate class directories.

    Fix: Create the three split destinations and place each image in the class directory belonging to its selected split.

  • Creating split folders without separate cat and dog directories.

    The training process must be able to find images by both subset and class.

    Fix: Use separate cats and dogs directories inside training, validation, and test.

  • Assuming that existing folders prove the dataset is correct.

    A preparation failure can involve file counts even when the folder structure appears complete.

    Fix: Use directory listing and length counting to verify all six expected counts.

  • Confusing transformed examples with newly collected original data.

    Data augmentation applies transformations to existing data.

    Fix: Treat the original image as the starting data and the rotated, flipped, or cropped result as a derived training example.

Practice and Recall

MEDIUM

Explain the complete preparation flow in your own words: begin with the 25,000-image Dogs vs. Cats dataset, describe how selected image ranges produce the three subsets, name the two class directories inside each subset, state the six expected directory counts, and explain where data augmentation fits into the training process.

Hints
  • Mention training, validation, and test.
  • Mention cats and dogs inside every split.
  • Use the expected counts of 1,000, 1,000, 500, 500, 500, and 500.
  • Distinguish original images from transformed training examples.

What do you think happens?

A directory contains 1,000 training cats, 1,000 training dogs, 500 validation cats, 500 validation dogs, 500 test cats, and 500 test dogs. Is the directory count consistent with the expected dataset split?

  • Yes
  • No
Reveal answer

Answer: Yes

The six class-directory counts match the expected counts for the smaller working dataset.

Key Takeaways

  1. The Dogs vs. Cats dataset contains 25,000 color JPEG images: 12,500 cats and 12,500 dogs.
  2. The working dataset is divided into training, validation, and test subsets according to selected image ranges.
  3. Each subset contains separate cats and dogs directories, producing six class-specific destinations.
  4. The expected counts are 1,000 training cats, 1,000 training dogs, 500 validation cats, 500 validation dogs, 500 test cats, and 500 test dogs.
  5. Data augmentation applies transformations to existing data to create additional training examples and mitigate overfitting.

Key Takeaways

  • The original Dogs vs. Cats dataset contains 25,000 color JPEG images, split evenly between cats and dogs.
  • The smaller working dataset uses training, validation, and test subsets, each with separate cat and dog directories.
  • Counting files in all six class directories is the first useful check for dataset-preparation errors.
  • Data augmentation creates additional examples by transforming existing images.
  • Original images are the starting data, while transformed images are derived training examples used to increase variation and mitigate overfitting.