Escape Sequences in String Literals
Try it: Escape Sequences in String Literals
How Python turns the text of a string literal into characters: the tokenizer first finds where the literal ends (a backslash protects the next character; an unescaped inner quote or a line break in '…' ends it or breaks it), then each escape — \n, \t, \\, \', \", \xhh, \uXXXX, \N{name}, octal — becomes ONE character, a backslash at the end of a line joins lines, unknown escapes keep their backslash, and a raw string keeps everything as typed; len(), repr() and print() then show the result.
How it works
- The tokenizer scans for the closing quote: a backslash always skips the next character, so \' does not end a '…' string; an unescaped copy of the opening quote ends it (anything after is read as code), and a real line break inside '…' or "…" is 'unterminated string literal'. Triple quotes may span lines.
- The body is then decoded left to right: \n newline, \t tab, \\ one backslash, \' and \" quotes, \r, \a, \b, \f, \v, \ooo octal, \xhh (exactly 2 hex digits), \uXXXX, \UXXXXXXXX and \N{name} — each is ONE character.
- A backslash followed by a line break disappears (the source lines are joined); a backslash before a character that is not an escape is kept, with a SyntaxWarning; a truncated \x, \u or \U escape is a SyntaxError that names its positions.
- An r prefix makes a raw string: backslashes are kept as typed (but a raw string still cannot end in an odd backslash).
- len(s) counts the stored characters; repr(s) — what the interpreter echoes — writes invisible characters back as escapes; print(s) shows the characters themselves (line breaks, tab stops).
Default run (16 steps): Python reads a string literal opened by a single quote ('). First the tokenizer looks for where it ends; then the characters between the quotes are decoded. … print(s) shows the characters themselves: each newline starts a new line.
Simplified: One assignment s = <literal> of up to 80 characters and 4 lines, typed from printable ASCII plus a few accented letters and symbols; no f-strings or bytes. The lab's own tokenizer and decoder read the text — nothing is executed. Its value, len, repr, warnings and SyntaxError messages match CPython 3.12 on 2,500+ literals; \N{...} knows only 13 character names, and when the text after an early closing quote is an operator, keyword, f-string or new line the lab stops without saying what Python does next. print() output of invisible characters is shown as [\x..].
Loading the simulation…