Tokenization and the split() Function
The Python split() function preserves punctuation attached to words, treating 'soft!' and 'soft' as different tokens and creating separate dictionary entries.
Why Raw Text Misleads
When you count words in Shakespeare, news articles, or social media posts, the text is rarely uniform. It contains punctuation marks and mixed capitalization. If you count words without cleaning the text first, one word can be divided among several forms. The resulting frequency dictionary can contain inaccurate counts and an artificially large vocabulary.
The central issue is that split() separates text at spaces, not according to the meaning of words. Punctuation attached to a word stays attached, and uppercase and lowercase letters remain different characters.
Punctuation Stays Attached
The Python split() function looks for spaces and treats the text between spaces as individual tokens. It does not understand punctuation as separate from a word. Therefore, in the text soft! soft, the exclamation mark remains attached to the first token. The resulting tokens are soft! and soft, not two copies of soft.
One Word, Two Tokens
Consider the text soft! soft and imagine building a frequency dictionary directly from the tokens produced by split().
Separate at the space: The space divides the text into two tokens.
Preserve attached punctuation: The exclamation mark remains part of the first token, so the tokens are soft! and soft.
Create dictionary entries: Because the two token strings are different, they become separate dictionary entries rather than one shared entry.
The punctuation mark causes soft! and soft to be counted separately.
Capitalization Splits Counts
Python treats uppercase and lowercase letters as distinct characters. A dictionary therefore treats Who and who as different keys. The same principle applies to The, the, and THE: without case normalization, each spelling can receive its own count.
This fragmentation hides the total frequency of a word. For example, if a text contains It at the beginning of sentences and it elsewhere, a dictionary built before case normalization can place those forms under different keys. The count for the word is then spread across entries instead of being combined.
The Cost of Uncleaned Counting
Uncleaned text produces fragmented data. Punctuation variants and capitalization variants divide counts that should belong to the same word. This makes it difficult to answer questions such as how many times a word appears or which words are most common. It also inflates the apparent vocabulary size because several entries may represent one normalized word.
Cleaning Before Parsing
Text cleaning should happen before parsing and frequency counting. The essential strategies described here are removing punctuation and normalizing case. Removing punctuation prevents forms such as sun. and sun from becoming separate tokens. Normalizing case prevents forms such as It and it from being treated as different dictionary keys.
- Start with the raw text.
- Remove punctuation that should not distinguish one word from another.
- Normalize capitalization so equivalent letter sequences use the same case.
- Apply split() to produce tokens.
- Build frequency counts from the cleaned tokens.
The purpose of cleaning is to ensure accurate results for the analysis you want to perform. The source establishes punctuation removal and case normalization as required strategies for accurate word-frequency counting; it does not define one universal cleaning rule for every possible text-analysis task.
Mistakes to Avoid
Assuming split() removes punctuation automatically.
split() looks for spaces and preserves punctuation attached to a token.
Fix:
Remove punctuation before parsing when punctuation should not affect word identity.Treating uppercase and lowercase forms as one dictionary key.
Python treats uppercase and lowercase letters as distinct characters, so the dictionary uses different keys.
Fix:
Normalize case before building frequency counts.Interpreting the number of dictionary entries as the true vocabulary size.
Uncleaned punctuation can create multiple entries for what should be one normalized word.
Fix:
Clean punctuation before counting vocabulary or word frequency.
Check Your Understanding
A text contains the forms It, it, sun., and sun. Before creating a frequency dictionary, identify which cleaning strategies are needed and explain why each strategy matters.
Hints
- Look for differences caused by punctuation.
- Look for differences caused by capitalization.
- Think about whether equivalent forms should share one dictionary entry.
What do you think happens?
If punctuation is preserved and case is not normalized, will It and it necessarily contribute to one shared dictionary entry?
Reveal answer
Answer: No, because capitalization creates distinct keys.
Python treats uppercase and lowercase letters as distinct characters, so the two forms can be stored under separate dictionary keys.
Key Takeaways
- split() separates text at spaces and preserves punctuation attached to words.
- Punctuation can make soft! and soft separate tokens and dictionary entries.
- Capitalization makes forms such as Who and who distinct dictionary keys.
- Uncleaned text fragments frequency counts and inflates apparent vocabulary size.
- Remove punctuation and normalize case before parsing when accurate word-frequency counts are required.
Key Takeaways
- split() uses spaces to separate tokens; it does not separate punctuation from attached words.
- Case-sensitive dictionary keys divide counts across uppercase and lowercase forms.
- Uncleaned text produces inaccurate frequencies and an inflated vocabulary size.
- Cleaning punctuation and normalizing case before parsing combines equivalent word forms.