Concepts / Sequence Processing with Convnets

Sequence Processing with Convnets

Deep learning applications can be organized by their input and output modalities.

  • Programming

Start with the Data Path

When studying a deep learning application, begin with the task's data path rather than its model name. Ask two questions: What kind of data enters the model, and what kind of data leaves it? The answers form an input-output modality mapping. This viewpoint is useful for organizing applications, including tasks involving sequences, without assuming that a particular architecture is required.

A task is not classified only by its input. The output matters equally: image-to-vector, image-to-text, and image-to-image are different mappings.

Trace an Application

inspectthen inspectcombinedo not confuse withApplicationIdentify inputvector, image, timeseries,text, or multiple typesIdentify outputvector, image, text, oranother supported typeName the mappinginput modality to outputmodalityConsider architectureseparatelythe mapping does notspecify the model
What steps let you classify a task by tracing from the input modality to the output modality?
  1. Identify the data entering the model.
  2. Identify the data produced by the model.
  3. Write the input modality followed by the output modality.
  4. Treat the resulting mapping as an organizational label, not as a complete description of the model.

The classification process has two ends. First, determine whether the input is vector data, image data, timeseries data, text, or a combination of modalities. Next, determine whether the output is vector data, image data, text, or another relevant modality. Only after both ends are known can the task be assigned a mapping category.

Compare the Mapping Categories

maps tomaps tomaps tomaps tomaps tomaps tomaps tomaps toVectorinputVectoroutputImageinputTextoutputTimeseriesinputImageoutputTextinputImage and textmultimodal input
How do vector, image, timeseries, text, and multimodal inputs map to different kinds of outputs?
Input modalityOutput modalityMapping category
VectorVectorVector-to-vector
ImageVectorImage-to-vector
ImageTextImage-to-text
ImageImageImage-to-image
TimeseriesVectorTimeseries-to-vector
TextTextText-to-text
TextImageText-to-image
Image and textTextMultimodal image-and-text-to-text
Video and textTextMultimodal video-and-text-to-text

The mapping names the data relationship at the input and output ends.

A vector-to-vector task connects vector data at both ends. An image can instead produce a vector, text, or another image, so the input label alone is insufficient. Timeseries data is listed as mapping to vector data. Text can map to text or images. Multimodal tasks differ because more than one input type participates, such as image and text or video and text before text is produced.

Follow a Sequence Mapping

entersproducesTimeseries datainput modalitySequence-processingtaskmodel operation notspecified by the mappingVector dataoutput modality
How can a sequence-processing task be classified without assuming a particular model architecture?

Classifying a Sequence Task

Suppose a task receives timeseries data and produces vector data. What modality mapping describes it?

Identify the input: The task receives timeseries data, so its input modality is timeseries.

Identify the output: The task produces vector data, so its output modality is vector.

Combine the two ends: Read the input and output together as a timeseries-to-vector mapping.

Separate mapping from architecture: The mapping identifies the data relationship. It does not specify whether a particular architecture, including a convolutional network, is used.

The task is classified as timeseries-to-vector.

This classification is the useful first step when discussing sequence processing with convolutional networks. The source material identifies the timeseries-to-vector relationship, but it does not define the internal operations of a convolutional network. Therefore, do not infer convolution steps, sequence positions, or architectural details from the mapping label alone.

What do you think happens?

A task receives images and produces text. Which mapping describes it?

  • Image-to-vector
  • Image-to-text
  • Text-to-image
  • Vector-to-vector
Reveal answer

Answer: Image-to-text

The input is image data and the output is text data. Both ends are required for the classification.

Keep Generalization Separate

The mapping category describes the relationship between data types, not the full behavior of a model. Knowing that a model maps images to text does not establish how well it will perform on data that differs from its training data. The same caution applies to other mappings, including timeseries-to-vector and multimodal mappings.

Mistakes in Task Classification

  • Classifying a task only by its input.

    An image can map to a vector, text, or another image. Those are different application categories.

    Fix: Always identify both the input modality and the output modality.

  • Treating the mapping as the model architecture.

    The mapping labels the data relationship; it does not by itself specify the model architecture.

    Fix: Describe the input-output mapping first, then discuss architecture separately when the available information supports it.

  • Assuming that a known mapping guarantees generalization.

    The mapping does not establish how well the model performs beyond its training data.

    Fix: Treat the mapping as an organizational tool rather than a complete description of model behavior.

  • Ignoring multimodal inputs.

    More than one input type participates, so the task belongs to a multimodal mapping.

    Fix: Record every input modality before naming the output relationship.

Practice the Two-Ended Test

EASY

Classify each task by identifying its input modality and output modality: a task that receives text and produces images; a task that receives images and produces images; and a task that receives video and text and produces text.

Hints
  • Name the input or inputs first.
  • Name the output second.
  • For more than one input type, use the word multimodal.

Practice Answers

Classify the three tasks using their input and output modalities.

Text and images: Text enters and images leave, so the mapping is text-to-image.

Images and images: Images enter and images leave, so the mapping is image-to-image.

Video and text to text: Video and text both enter before text is produced, so this is a multimodal video-and-text-to-text mapping.

The classifications are text-to-image, image-to-image, and multimodal video-and-text-to-text.

Key Takeaways

  1. Classify a deep learning task by examining both what enters the model and what it produces.
  2. Vector, image, timeseries, text, and multimodal data can participate in different input-output mappings.
  3. A timeseries-to-vector task is identified from its data ends; the mapping alone does not specify the internal architecture.
  4. Image-to-image, image-to-text, and image-to-vector tasks differ because their output modalities differ.
  5. A modality mapping organizes an application but does not guarantee generalization beyond the training data.

Key Takeaways

  • Deep learning applications can be organized as mappings from input modalities to output modalities.
  • Correct classification requires inspecting both ends of the task, not just the input.
  • Timeseries-to-vector is a sequence-related mapping, while image-to-text, text-to-image, and multimodal mappings represent other data relationships.
  • The mapping category does not determine the model architecture or guarantee performance on data beyond training.