Playground / Text Vectorization: Tokens to Vectors

Turn a sentence into ids and vectors

Text Vectorization: Tokens to Vectors

Interactive lab

Try it: Text Vectorization: Tokens to Vectors

The text-to-tensor pipeline: standardize, split into tokens, build a frequency-ranked vocabulary with padding and [UNK], map tokens to integer ids, pad or truncate, then one-hot or multi-hot encode and look the ids up in an embedding table.

How it works

  1. Standardize: lowercase and delete punctuation.
  2. Split on whitespace into tokens.
  3. Build the vocabulary: 0 = padding, 1 = [UNK], then words by frequency (ties: first appearance), capped at max tokens.
  4. Map each token to its id (1 if it did not make the vocabulary), then pad with 0 or truncate to the sequence length.
  5. Encode: one-hot rows per position or one multi-hot row per text; embed: id k selects row k of the embedding table.

Default run (18 steps): Raw text (59 characters): "The movie was great. The plot was great, the acting wasn't!". A model cannot read strings, so it is turned into numbers step by step. … Embedding lookup: id k selects row k of a fixed 8 x 3 table, giving a (8, 3) dense matrix instead of (8, 8) sparse one-hots. Row 0 (padding) is looked up too; a trained Embedding layer would learn these rows.

Simplified: Bounded ASCII text (up to 200 characters, 40 tokens), whitespace splitting only, and a small FIXED embedding table (a real Embedding layer learns its rows). The multi-hot vector keeps the padding slot (always 0) so its positions line up with the ids.

Educational simulation

Loading the simulation…