# UTF-8 and C1 Controls

> What a terminal reading UTF-8 does with bytes that are not valid UTF-8, and with the C1 controls.

- **Sequence:** `C2 80 … C2 9F`
- **DEC STD 070:** [§3.5.1.2.4 C1 Control Codes, p. 3-19](https://archive.org/details/bitsavers_decstandar0VideoSystemsReferenceManualDec91_74264381/page/n123/mode/1up)
- **xterm:** [C1 (8-Bit) Control Characters](https://invisible-island.net/xterm/ctlseqs/ctlseqs.html)

Like nearly every terminal now, the one on this site reads its input as
UTF-8, and so does a default xterm in a UTF-8 locale. That leaves two
questions the older standards do not answer: what to do with bytes that are
not valid UTF-8, and what to do with the C1 controls, which ECMA-48 and
DEC STD 070 give as single bytes from `80` to `9F`, bytes that in UTF-8 can
only continue a character.

**C1 controls.** DEC STD 070 has a C1 control cancel any sequence it
interrupts and take effect itself
([§3.5.1.2.4, p. 3-19](https://archive.org/details/bitsavers_decstandar0VideoSystemsReferenceManualDec91_74264381/page/n123/mode/1up)).
In UTF-8 they would be the code points U+0080 to U+009F, and xterm, unless
told otherwise with `allowC1Printable`, ignores those altogether
([`doparsing`](https://github.com/ThomasDickey/xterm-snapshots/blob/xterm-411a/charproc.c#L3316-L3335)),
as libghostty-vt does. A program that wants a C1 control sends its 7-bit
form, `ESC` and a letter: [CSI](https://control-codes.page/esc/csi/index.html.md) is `ESC [`, not `9B`.

**Invalid UTF-8.** Where a character's bytes are cut short by another byte,
both write U+FFFD for what came before and then read that byte afresh.
They part company on the rest
([`decodeUtf8`](https://github.com/ThomasDickey/xterm-snapshots/blob/xterm-411a/ptydata.c#L69-L245)):

- A *stray continuation byte*, `80` to `BF` where no character has begun,
  is U+FFFD in libghostty-vt. xterm looks at the byte after it: if that is
  printable ASCII, it takes the stray byte as the code point of the same
  number, which makes `80` to `9F` a C1 control, and so ignored, and `A0` to
  `BF` a Latin-1 character, `A9` being `©`. Otherwise it is U+FFFD.
- An encoded *surrogate*, U+D800 to U+DFFF, is one U+FFFD in xterm, and one
  per byte in libghostty-vt, which follows the WHATWG rule of replacing each
  piece that cannot begin a valid character.
- An *overlong* encoding, or one past U+10FFFF, is replaced the same way by
  libghostty-vt. xterm, built with 32-bit characters, decodes some of them
  to values that are not Unicode at all, which no case here can show as
  text; its `utf8Weblike` resource makes it follow the WHATWG rule instead.

The cases treat their input as one read: xterm can look past a stray byte
only when the next byte arrived with it. They follow a default xterm, so the
ones above where the two part company are known differences, and the same
goes for the [CSI](https://control-codes.page/esc/csi/index.html.md) and [NEL](https://control-codes.page/esc/nel/index.html.md) pages' cases of the
8-bit forms.

## Validation

### U8-1: A character cut short

Input, one step per line:

```text
A\xc3B      # C3 starts a two-byte character
```

Expected screen:

```text
|A�B_____|
cursor 1,4
```

### U8-2: Three bytes cut short

Input, one step per line:

```text
A\xe2\x82B  # E2 82 starts a three-byte one
```

Expected screen:

```text
|A�B_____|
cursor 1,4
```

### U8-3: A surrogate

*A known difference: libghostty-vt does not pass this case.*

Input, one step per line:

```text
A\xed\xa0\x80B  # U+D800
```

Expected screen:

```text
|A�B_____|
cursor 1,4
```

### U8-4: A stray C1 byte before text is ignored

*A known difference: libghostty-vt does not pass this case.*

Input, one step per line:

```text
A\x80B
```

Expected screen:

```text
|AB______|
cursor 1,3
```

### U8-5: A stray Latin-1 byte before text is printed

*A known difference: libghostty-vt does not pass this case.*

Input, one step per line:

```text
A\xa9B
```

Expected screen:

```text
|A©B_____|
cursor 1,4
```

### U8-6: A stray byte with nothing after it

Input, one step per line:

```text
A\x80
```

Expected screen:

```text
|A�______|
cursor 1,3
```

### U8-7: A byte UTF-8 never uses

Input, one step per line:

```text
A\xffB
```

Expected screen:

```text
|A�B_____|
cursor 1,4
```

### U8-8: C1 code points are ignored

Input, one step per line:

```text
A\u{9b}2;2H  # U+009B, CSI, as UTF-8
```

Expected screen:

```text
|A2;2H___|
cursor 1,6
```

---

This is the Markdown version of <https://control-codes.page/parser/utf8/>. On that page every validation case runs live in libghostty-vt, the terminal emulation core of Ghostty, compiled to WebAssembly.

The validation cases are written in the notation described in <https://control-codes.page/notation/index.html.md>.
