Playground / UTF-8: code points to bytes and back

Encode a character into bytes

UTF-8: code points to bytes and back

Interactive lab

Try it: UTF-8: code points to bytes and back

How str.encode('utf-8') turns each character's Unicode code point into 1 to 4 bytes using fixed bit templates, and how bytes.decode('utf-8') reads them back — including the invalid bytes that raise UnicodeDecodeError.

How it works

  1. Every character is a Unicode code point, written U+XXXX (ord() gives its number).
  2. The code point's size picks a template: up to U+007F → 0xxxxxxx; up to U+07FF → 110xxxxx 10xxxxxx; up to U+FFFF → 3 bytes; above → 4 bytes.
  3. The code point's bits fill the x's from left to right, giving the bytes.
  4. Decoding reads a lead byte, which says how many 10xxxxxx continuation bytes follow, and joins their payload bits.
  5. A byte that cannot start a character, a wrong continuation byte, an overlong or surrogate form, or missing bytes at the end raise UnicodeDecodeError (errors='strict').

Default run (14 steps): 'Aé€😀'.encode("utf-8"): 4 characters, each becomes 1–4 bytes. … 4 characters → 10 bytes: b'A\xc3\xa9\xe2\x82\xac\xf0\x9f\x98\x80'. (bytes repr shows ASCII as text and every other byte as \xNN.)

Simplified: Up to 12 characters or 16 bytes; control characters and lone surrogates are removed from the text; only the strict error handler is shown.

Educational simulation

Loading the simulation…