Concepts / Techniques for Improving Convolutional Neural Networks

Techniques for Improving Convolutional Neural Networks

Fine-tuning starts with a pre-trained convnet rather than an uninitialized model.

  • Programming

Starting Point Matters

Fine-tuning changes how a convolutional neural network is prepared for a new task. The model does not begin as an entirely new, uninitialized convnet. It begins with a convnet that has already been trained, and that pre-trained state becomes the starting point for further adjustment.

Fine-tuning is the process of adapting a pre-trained convolutional neural network to a target task by retaining some model parts during a training stage and updating other parts for that task.

fine-tuningUninitialized convnetAdapted convnettarget taskPre-trained convnet
What changes when adaptation begins from a pre-trained convnet instead of an uninitialized model?

The Adaptation Path

The main state transition has three ideas. First, begin with the pre-trained convnet. Next, identify how the model will be used for the target task. During the fine-tuning stage, some parts can retain their existing parameters while other parts have their parameters updated. The result is an adapted convnet whose state reflects the target task.

chosen forguidesproducesPre-trained convnetTarget taskFine-tuning stageAdapted convnet
What happens as a pre-trained convnet is adapted for a target task?

Tracing One Fine-tuning Procedure

Describe the state of a convnet before, during, and after fine-tuning for a target task.

Before adaptation: The model is a pre-trained convnet. It is not treated as an entirely new model with no prior training.

During adaptation: The procedure separates model parts conceptually: some existing parameters are retained, while other parameters are updated for the target task.

After the stage: The convnet has been adapted for the target task. Its final state depends on which parts the particular procedure retained and which parts it updated.

Fine-tuning is a transition from a pre-trained state to an adapted state, with retention and updating determined by the procedure.

Retained and Updated Parts

A useful way to reason about fine-tuning is to divide the convnet into two conceptual groups. One group contains parts whose existing parameters are retained during a stage of training. The other contains parts whose parameters are updated for the target task. This distinction describes the model's state during adaptation.

includesincludesPre-trained convnetRetained partsexisting parametersUpdated partstarget-task parameters
Which conceptual parts of a pre-trained convnet are retained, and which may be updated for the target task?
defines adaptationretainsupdatesTarget taskRetained parametersFine-tuningprocedureUpdated parameters
How does the target-task adaptation affect the parameter groups?

Fixed Extraction and Fine-tuning

used foradapted forPre-trained convnetparameters retainedTarget taskPre-trained convnetselected parameters updatedTarget task
What is the difference between using a pre-trained convnet as a fixed feature extractor and updating parameters for a new task?
  • Describing fine-tuning as training a convnet from scratch

    Fine-tuning starts with a convnet that has already been trained.

    Fix: State that the pre-trained convnet is the starting point for further adjustment.

  • Claiming that every part of the model must be updated

    During a fine-tuning stage, some parts may be retained while others are updated.

    Fix: Identify the retained and updated groups for the specific procedure being discussed.

  • Treating the retained and updated parts as a fixed layer list

    The exact retained and updated parts depend on the procedure.

    Fix: Describe the groups conceptually and specify the actual parts only when the procedure provides them.

  • Confusing reuse with adaptation

    Fine-tuning involves further adjustment for the target task, whereas simple reuse keeps the existing parameters during the described use.

    Fix: Say whether the pre-trained parameters are merely retained or whether selected parts are updated.

QuestionFixed feature extractionFine-tuning
Starting pointA pre-trained convnetA pre-trained convnet
Role of existing parametersRetained during the described useSome may be retained while others are updated
PurposeUse the pre-trained model for a target taskAdapt the pre-trained model to a target task
What must be specifiedThat the existing parameters remain retainedWhich parts are retained and which are updated in the procedure

Practice the State Change

EASY

A pre-trained convnet is being adapted to a new target task. During the described stage, one group of model parts keeps its existing parameters and another group has its parameters updated. Explain why this is fine-tuning, and identify the two conceptual groups.

Hints
  • Begin by identifying the model's starting state.
  • Name the purpose of the adaptation.
  • Separate the retained parameters from the updated parameters.

Practice Answer

Explain the state transition in the scenario.

Starting state: The convnet is pre-trained, so the procedure does not begin with an uninitialized model.

Purpose: The model is being adapted to a target task.

Parameter groups: Some parts retain their existing parameters, while other parts have parameters updated for the target task.

The scenario describes fine-tuning because a pre-trained convnet is further adjusted for a target task, with retention and updating separated according to the procedure.

Key Takeaways

  1. Fine-tuning starts from a pre-trained convnet rather than an uninitialized model.
  2. Its purpose is to adapt the convnet to a target task.
  3. During a fine-tuning stage, some model parts may retain their existing parameters while other parts are updated.
  4. The exact retained and updated parts depend on the procedure being described.
  5. A clear explanation should distinguish simple reuse of pre-trained parameters from further parameter updates for the target task.

Key Takeaways

  • Fine-tuning begins with a pre-trained convolutional neural network.
  • The model is adapted for a target task rather than treated as an entirely new model.
  • Some model parts may be retained and others updated during a fine-tuning stage.
  • The exact division between retained and updated parts depends on the procedure.
  • The most important distinction is between reusing existing parameters and adjusting parameters for the target task.