UTF-8 UTF-8 and C1 Controls
What a terminal reading UTF-8 does with bytes that are not valid UTF-8, and with the C1 controls.
C2 80 … C2 9F
- DEC STD 070
- §3.5.1.2.4 C1 Control Codes, p. 3-19
- xterm
- C1 (8-Bit) Control Characters
Like nearly every terminal now, the one on this site reads its input as UTF-8, and so does a default xterm in a UTF-8 locale. That leaves two questions the older standards do not answer: what to do with bytes that are not valid UTF-8, and what to do with the C1 controls, which ECMA-48 and DEC STD 070 give as single bytes from 80 to 9F, bytes that in UTF-8 can only continue a character.
C1 controls. DEC STD 070 has a C1 control cancel any sequence it interrupts and take effect itself (§3.5.1.2.4, p. 3-19). In UTF-8 they would be the code points U+0080 to U+009F, and xterm, unless told otherwise with allowC1Printable, ignores those altogether (doparsing), as libghostty-vt does. A program that wants a C1 control sends its 7-bit form, ESC and a letter: CSI is ESC [, not 9B.
Invalid UTF-8. Where a character’s bytes are cut short by another byte, both write U+FFFD for what came before and then read that byte afresh. They part company on the rest (decodeUtf8):
- A stray continuation byte,
80toBFwhere no character has begun, is U+FFFD in libghostty-vt. xterm looks at the byte after it: if that is printable ASCII, it takes the stray byte as the code point of the same number, which makes80to9Fa C1 control, and so ignored, andA0toBFa Latin-1 character,A9being©. Otherwise it is U+FFFD. - An encoded surrogate, U+D800 to U+DFFF, is one U+FFFD in xterm, and one per byte in libghostty-vt, which follows the WHATWG rule of replacing each piece that cannot begin a valid character.
- An overlong encoding, or one past U+10FFFF, is replaced the same way by libghostty-vt. xterm, built with 32-bit characters, decodes some of them to values that are not Unicode at all, which no case here can show as text; its
utf8Weblikeresource makes it follow the WHATWG rule instead.
The cases treat their input as one read: xterm can look past a stray byte only when the next byte arrived with it. They follow a default xterm, so the ones above where the two part company are known differences, and the same goes for the CSI and NEL pages’ cases of the 8-bit forms.
Validation
U8-1: A character cut short
A\xc3B # C3 starts a two-byte character
|A�B_____|
cursor 1,4
U8-2: Three bytes cut short
A\xe2\x82B # E2 82 starts a three-byte one
|A�B_____|
cursor 1,4
U8-3: A surrogate
A\xed\xa0\x80B # U+D800
|A�B_____|
cursor 1,4
U8-4: A stray C1 byte before text is ignored
A\x80B
|AB______|
cursor 1,3
U8-5: A stray Latin-1 byte before text is printed
A\xa9B
|A©B_____|
cursor 1,4
U8-6: A stray byte with nothing after it
A\x80
|A�______|
cursor 1,3
U8-7: A byte UTF-8 never uses
A\xffB
|A�B_____|
cursor 1,4
U8-8: C1 code points are ignored
A\u{9b}2;2H # U+009B, CSI, as UTF-8
|A2;2H___|
cursor 1,6