Text Vectorization: Tokens to Vectors
Try it: Text Vectorization: Tokens to Vectors
The text-to-tensor pipeline: standardize, split into tokens, build a frequency-ranked vocabulary with padding and [UNK], map tokens to integer ids, pad or truncate, then one-hot or multi-hot encode and look the ids up in an embedding table.
How it works
- Standardize: lowercase and delete punctuation.
- Split on whitespace into tokens.
- Build the vocabulary: 0 = padding, 1 = [UNK], then words by frequency (ties: first appearance), capped at max tokens.
- Map each token to its id (1 if it did not make the vocabulary), then pad with 0 or truncate to the sequence length.
- Encode: one-hot rows per position or one multi-hot row per text; embed: id k selects row k of the embedding table.
Default run (18 steps): Raw text (59 characters): "The movie was great. The plot was great, the acting wasn't!". A model cannot read strings, so it is turned into numbers step by step. … Embedding lookup: id k selects row k of a fixed 8 x 3 table, giving a (8, 3) dense matrix instead of (8, 8) sparse one-hots. Row 0 (padding) is looked up too; a trained Embedding layer would learn these rows.
Simplified: Bounded ASCII text (up to 200 characters, 40 tokens), whitespace splitting only, and a small FIXED embedding table (a real Embedding layer learns its rows). The multi-hot vector keeps the padding slot (always 0) so its positions line up with the ids.
Loading the simulation…