🌿freegardner

Science

Byteification Fixes AI Letter Counting Blind Spot

08 Oct 2026 · via Nature

Byteification Fixes AI Letter Counting Blind Spot
AI-generated image

Byteification Fixes AI Letter Counting Blind Spot

The Strawberry Problem

Strawberry contains the letter “r” three times. When asked, many large language models answer that the letter appears twice. [1] This is not a trivial arithmetic slip. It is a window into how these systems actually read.

The reason lies in the fundamental unit these models use to process text. Most large language models encode words as “tokens” — chunks that represent sequences of letters rather than individual characters. [1] This encoding delivers excellent performance, but it carries a structural blind spot: the model cannot access the individual characters inside each word, which are encoded as binary sequences called bytes.

Byteification Fixes AI Letter Counting Blind Spot (Image 1)
AI-generated image

This limitation is not new. It has been embedded in the dominant approach to language model design since the method of splitting text into subword units became standard practice. The consequences extend beyond counting letters. Any task that requires character-level awareness — detecting a typo, reversing a string, identifying a rhyme, or counting the letter “i” in “artificial intelligence” — sits at the edge of what these models can do reliably.

Retrofitting the Old Architecture to See Bytes

A technique called byteification now offers a way out that does not require discarding existing models. Writing in Nature, Minixhofer and colleagues report an approach that retrofits token-based large language models to operate at the byte level. [1] Instead of rebuilding from scratch, byteification transforms models that already exist, giving them the ability to read and process individual characters while preserving what they already know.

The authors show that byteified models can achieve competitive performance — meaning they remain capable across the tasks that made token-based models successful in the first place — while retaining the ability to read individual characters. The technique does not replace existing models; it upgrades the ones already in use.

Byteification Fixes AI Letter Counting Blind Spot (Image 2)
AI-generated image

The assessment frames byteification as a retrofit — a modification that keeps the engine and changes what it can perceive.


Sources

  1. DOI: 10.1038/d41586-026-03059-2
  2. Nature — Quote source (original article)

← back to the garden