^[control-codes live in libghostty-vt

Parsing

UTF-8 UTF-8 and C1 Controls

What a terminal reading UTF-8 does with bytes that are not valid UTF-8, and with the C1 controls.

C2 80 … C2 9F
DEC STD 070
§3.5.1.2.4 C1 Control Codes, p. 3-19
xterm
C1 (8-Bit) Control Characters

Like nearly every terminal now, the one on this site reads its input as UTF-8, and so does a default xterm in a UTF-8 locale. That leaves two questions the older standards do not answer: what to do with bytes that are not valid UTF-8, and what to do with the C1 controls, which ECMA-48 and DEC STD 070 give as single bytes from 80 to 9F, bytes that in UTF-8 can only continue a character.

C1 controls. DEC STD 070 has a C1 control cancel any sequence it interrupts and take effect itself (§3.5.1.2.4, p. 3-19). In UTF-8 they would be the code points U+0080 to U+009F, and xterm, unless told otherwise with allowC1Printable, ignores those altogether (doparsing), as libghostty-vt does. A program that wants a C1 control sends its 7-bit form, ESC and a letter: CSI is ESC [, not 9B.

Invalid UTF-8. Where a character’s bytes are cut short by another byte, both write U+FFFD for what came before and then read that byte afresh. They part company on the rest (decodeUtf8):

  • A stray continuation byte, 80 to BF where no character has begun, is U+FFFD in libghostty-vt. xterm looks at the byte after it: if that is printable ASCII, it takes the stray byte as the code point of the same number, which makes 80 to 9F a C1 control, and so ignored, and A0 to BF a Latin-1 character, A9 being ©. Otherwise it is U+FFFD.
  • An encoded surrogate, U+D800 to U+DFFF, is one U+FFFD in xterm, and one per byte in libghostty-vt, which follows the WHATWG rule of replacing each piece that cannot begin a valid character.
  • An overlong encoding, or one past U+10FFFF, is replaced the same way by libghostty-vt. xterm, built with 32-bit characters, decodes some of them to values that are not Unicode at all, which no case here can show as text; its utf8Weblike resource makes it follow the WHATWG rule instead.

The cases treat their input as one read: xterm can look past a stray byte only when the next byte arrived with it. They follow a default xterm, so the ones above where the two part company are known differences, and the same goes for the CSI and NEL pages’ cases of the 8-bit forms.

Validation

U8-1: A character cut short

A\xc3B      # C3 starts a two-byte character
|A�B_____|
cursor 1,4

U8-2: Three bytes cut short

A\xe2\x82B  # E2 82 starts a three-byte one
|A�B_____|
cursor 1,4

U8-3: A surrogate

A\xed\xa0\x80B  # U+D800
|A�B_____|
cursor 1,4

U8-4: A stray C1 byte before text is ignored

A\x80B
|AB______|
cursor 1,3

U8-5: A stray Latin-1 byte before text is printed

A\xa9B
|A©B_____|
cursor 1,4

U8-6: A stray byte with nothing after it

A\x80
|A�______|
cursor 1,3

U8-7: A byte UTF-8 never uses

A\xffB
|A�B_____|
cursor 1,4

U8-8: C1 code points are ignored

A\u{9b}2;2H  # U+009B, CSI, as UTF-8
|A2;2H___|
cursor 1,6