Retrofitting language models to operate over bytes

原始链接: https://www.nature.com/articles/s41586-026-11111-4

Hacker News new | past | comments | ask | show | jobs | submit login Retrofitting language models to operate over bytes ( nature.com ) 5 points by theanonymousone 1 hour ago | hide | past | favorite | discuss help Consider applying for YC's Winter 2027 batch! Applications are open till November 2. Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact Search:
相关文章

原文

Byteified model architecture details

Our architecture follows the same overall structure as earlier LTLMs, including DTP, BLT and H-Net (‘Related work’ in Methods). It consists of the following components.

Tokenization and embedding

\({\mathcal{T}}\) assigns every input UTF-8 byte in x (we treat x as a sequence over bytes: x ∈ {0, …, 255}n) a corresponding embedding in \({{\mathbb{R}}}^{d}\) from an embedding table containing an entry for every byte. The embedding table over bytes is negligible in size compared with embedding tables over subwords. However, scaling the size and sparsity of the embedding table has been shown to improve performance while having no negative effect on inference speed46. Inspired by the hash embeddings in BLT22,47, we, thus, increased the size of the embedding table. Specifically, we residually added the longest subword embedding (from the original embedding table in the subword-level LLM) that ends at the current byte position to every byte embedding:

$${e}_{i}:= {{\mathcal{T}}}_{\mathrm{Byte}}({x}_{i})+{{\mathcal{T}}}_{\mathrm{SubwordSuffix}}({x}_{:i}),$$

where \({{\mathcal{T}}}_{\mathrm{SubwordSuffix}}\) assigns an embedding to every byte based on the index of the subword token in the vocabulary \({{\mathcal{V}}}_{\mathrm{Subword}}\) with the longest common suffix to the byte sequence up to the current position i. For example, the subword embeddings corresponding to {‘_’, ‘_f’, ‘_fl’, ‘o’, ‘_flow’} could be added to the byte representation of the sequence {‘_’, ‘f’, ‘l’, ‘o’, ‘w’}, assuming none of {‘_flo’, ‘flo’, ‘lo’} are in the original subword vocabulary. Retaining the subword embeddings is not strictly necessary, and we can generally achieve the same performance by increasing the size of the local encoder instead. However, retaining the subword embedding allows us to achieve a better performance–efficiency trade-off by increasing the number of cheap sparsely activated parameters; an alternative to increasing the size and sparsity of the local encoder is using a mixture of experts in the feedforward layer, although we do not investigate this here.

Local encoder

The local encoder \({\mathcal{E}}\) contextualizes the byte-level embeddings through an mLSTM layer33, resulting in the contextualized representations \(\widehat{e}\). We found that mLSTM improves inference speed compared with other linear recurrent neural network (RNN) variants (Extended Data Fig. 3) while attaining competitive performance. We found a single mLSTM layer to be sufficient, as the expressivity of the local encoder was substantially enhanced by the retained subword embeddings.

Boundary predictor

The boundary predictor \({\mathcal{B}}\) predicts a score p ∈ [0, 1] for each byte based on the contextualized representations \(\widehat{e}\). If p is greater than some threshold, a patch boundary is placed after the current byte. In contrast to earlier LTLMs, the boundary predictor of the byteified models is non-causal: it has access to 1 byte of future context, and it is only employed for the prefill, where future information can be used while retaining the ability to generate text. We describe non-causal boundary prediction in detail in section ‘Byteified language model architecture’, where we also discuss how boundary prediction is handled during decoding.

Pooling

We pool byte-level representations into patch representations by selecting the representation of the last byte in every patch as the patch-level representation h. This is equivalent to the pooling done by Hwang et al.2 and does not introduce any extra parameters. Unlike Hwang et al.2, the local models and the global model use the same representation dimensionality, obviating the need for an up-projection. We originally experimented with smaller local dimensions but found the up-projection mechanism to bottleneck performance by restricting the rank of the representations (Extended Data Fig. 1).

Global model

Most of the compute is spent in the deep global model \({\mathcal{M}}\) contextualizing the patch representations h into \(\widehat{h}\). We retain the global model of the original subword-level LLM.

Depooling

The global model is invoked at every patch boundary, providing contextualized representations for every patch. It remains to depool these representations back to representations of bytes. We do so by adding the latest available patch representation in \(\widehat{h}\) at any byte position to a linear projection of the byte representations \(\widehat{e}\), resulting in z. This is like the depooling by H-Net, again forgoing the projection due to equal global and local dimensionality.

Local decoder

The local decoder \({\mathcal{D}}\) contextualizes the depooled byte representations z into \(\widehat{z}\) through another stack of mLSTM layers. We used more mLSTM layers (in practice, four) to increase capacity, because unlike in the encoder, we found it infeasible to meaningfully reincorporate the output subword embedding matrix, which could have potentially allowed us to reduce the number of layers in the decoder in a similar way as for the encoder.

Language-modelling head

The language-modelling head converts the final byte representations \(\widehat{z}\) into scores interpretable as next-byte probabilities (and next-byte-followed-by-boundary probabilities if boundary predictions are fused; compare with ‘Byteified language model architecture’) through a projection to the vocabulary space and softmax. During decoding, the next atom (byte or boundary in the unfused case and byte or byte-and-boundary in the fused case) was generated by looping over the local encoder, local decoder and language-modelling head. If a boundary is predicted, the patch ends and is passed through the global model, after which the next loop over the local encoder, local decoder and language-modelling head starts. We did not find it necessary to employ any mechanism to ensure consistency between the patch representations and decoded bytes. As in standard subword-level language models, the extent of the possible mismatch between the hidden representations and the decoded atoms is dictated by the sampling procedure; any apparent mismatch is analogous to sampling a non-argmax token in a standard subword-level language model.

In practice, Bolmo 1B contains approximately 10M fewer parameters, where M is millions, than OLMo 2 1B (−0.7%), Bolmo 7B contains approximately 330M more than Olmo 3 7B (+4.5%), Blama 8B contains approximately 220M more than Llama 3 8B (+2.7%) and Bwen 8B contains approximately 120M more than Qwen3 8B (+1.5%).

Byteifying procedure

Subword-to-byte distillation

The first stage starts by initializing the parameters of the global model from the subword-level LLM checkpoint, whereas the parameters of the local models and the language-modelling head are initialized randomly. The aim of this stage is to quickly learn weights for the local encoder, local decoder, boundary predictor and language-modelling head that recover the behaviour of the subword model. Efficiency is crucial; the cost of this stage should be minimal to permit fast experimentation and allow increasing the investment into stage 2. To achieve these goals, we designed a stage 1 procedure that allows the model to learn the desired weights without fully backpropagating through the global model. This substantially reduces the time per training step (Extended Data Table 4). The stage 1 loss is minimal if and only if the byte-level model exactly mimics the source subword-level LLM. It is composed of three parts.

Quickly learning a boundary predictor \({\boldsymbol{\mathcal{B}}}_{{\bf{byteify}}}\)

We train the boundary predictor to emulate the boundaries placed by subword tokenization through a binary cross-entropy loss:

$${{\mathcal{L}}}_{{\mathcal{B}}}:= -\sum _{t}({{\mathcal{B}}}_{\mathrm{subword}}{(x)}_{t}\,\log \,{{\mathcal{B}}}_{\mathrm{byteify}}{(\hat{e})}_{t}+(1-{{\mathcal{B}}}_{\mathrm{subword}}{(x)}_{t})\,\log (1-{{\mathcal{B}}}_{\mathrm{byteify}}{(\hat{e})}_{t})),$$

where \({{\mathcal{B}}}_{\mathrm{subword}}(x)\) is 1 for every byte at the last position of a subword patch, otherwise 0. The boundary predictor \({{\mathcal{B}}}_{\mathrm{byteify}}\), which uses future context to tokenize the prefill text, quickly achieves over 99% accuracy.

Quickly learning a local encoder \(\boldsymbol{\mathcal{E}}\)

Assuming our boundary predictor perfectly emulates subword tokenization, our local encoder and pooling mechanism will be a perfect substitute for the subword embedding matrix if they yield the same input to the global model as the subword embedding matrix for every patch. This is the case if all pooled representations \(\mathrm{Pool}({\mathcal{E}}({\mathcal{T}}\,(x)),{{\mathcal{B}}}_{\mathrm{byteify}}(\hat{e}))\) are equal to the corresponding subword embeddings \({{\mathcal{T}}}_{\mathrm{subword}}(x)\). Hwang et al.2 optimized towards this goal by directly minimizing the L2 distance of every pooled representation to the corresponding subword embedding. We took an alternative approach inspired by research on model stitching, which showed that similar representations do not necessarily propagate through subsequent layers in a similar way48. We propagate the pooled representations through n layers of the global model and minimize the L2 distance to the subword representations that result from propagating the subword embeddings through the same n layers:

$$\begin{array}{l}{Y}_{{\mathcal{E}}}:= {{\mathcal{M}}}_{:n}({{\mathcal{T}}}_{\mathrm{subword}}(x)),\\ {\hat{Y}}_{{\mathcal{E}}}:= {{\mathcal{M}}}_{:n}(\mathrm{Pool}({\mathcal{E}}({\mathcal{T}}\,(x)),{{\mathcal{B}}}_{\mathrm{subword}}(x))),\\ {{\mathcal{L}}}_{{\mathcal{E}}}:= \parallel {Y}_{{\mathcal{E}}}-{\hat{Y}}_{{\mathcal{E}}}\parallel .\end{array}$$

where \({Y}_{{\mathcal{E}}}\) are the representations of the original model at layer n, \({\hat{Y}}_{{\mathcal{E}}}\) the retrofitted model representations at the same layer n of the global model, and \({{\mathcal{L}}}_{{\mathcal{E}}}\) the \({{\mathcal{L}}}^{2}\) distance between the two. Notably, we pool the local encoder representations using the true subword boundaries \({{\mathcal{B}}}_{\mathrm{subword}}\) instead of \({{\mathcal{B}}}_{\mathrm{byteify}}\); this is necessary to preserve the alignment of the pooled representations to the representations in \({{\mathcal{T}}}_{\mathrm{subword}}(x)\) along the sequence dimension. \({{\mathcal{M}}}_{:n}\) indicates the global model up to and including the nth layer. The weights of \({\mathcal{M}}\) are kept frozen. If n = 0, this reduces to the setting of ref. 2. Although choosing n > 0 necessitates backpropagating through some parts of the global model, we can minimize the resulting cost by choosing a small n. We found that n = 4 strikes a good balance between performance and efficiency, as it substantially outperformed n = 0 while remaining cheap to compute.

Quickly learning a local decoder \(\boldsymbol{\mathcal{D}}\)

Our local decoder and language-modelling head are optimal if our byte-level LLM assigns the same likelihood as the subword model to every text x. Assuming equal patch boundaries, it is optimal if the likelihood of every patch is equal. As subword-level LLMs implicitly predict output patch boundaries, we cannot easily compute comparable patch likelihoods in byte-level models without output boundary prediction. In this case, we would have to resort to approximations, as in ref. 49. However, as the byteified models do predict output patch boundaries, simply comparing the likelihoods of every patch results in an exact objective (proof in ‘Exactness of the stage 1 objective’ in Supplementary Information):

$$\begin{array}{l}{{\mathcal{L}}}_{{\mathcal{D}},\mathrm{Distill}}:= \sum _{i}f\left(\prod _{j\in T(x,i)}\,\mathrm{LMHead}\,({\hat{z}}_{\mathrm{subword}})[\,j,\mathrm{next}\_\mathrm{byte}(x,j)]\right.\\ \,\left.\parallel \,{\mathrm{LMHead}}_{\mathrm{subword}}({z}_{\mathrm{subword}})[i,\mathrm{next}\_\mathrm{tok}(x,i)]\right),\end{array}$$

where j ∈ T(x, i) indicates all byte indices j that are part of the ith subword patch; this includes the indices of the special <b> symbol if treated as separate or the indices of the 256 special symbols consisting of a byte plus <b> if fused. next_tok( ⋅ ) and next_byte( ⋅ ) map to the index in the vocabulary of the symbol occurring after the current symbol (token or byte), including special symbols. \({z}_{\mathrm{subword}}={\mathcal{M}}({{\mathcal{T}}}_{\mathrm{subword}}(x))\) are the representations of the subword model at the final layer, \({\hat{z}}_{\mathrm{subword}}={\mathcal{D}}(\mathrm{Depool}\,(\hat{e},{z}_{\mathrm{subword}},p))\) is the result of passing these representations through the depooling layer and the local decoder, and LMHeadsubword is the language-modelling head of the source subword-level LLM. We chose for the comparison function f:

$$f(\hat{y}\,\parallel \,y):= -({y}^{1/\tau }\,\log \,{\hat{y}}^{1/\tau }+(1-{y}^{1/\tau })\,\log (1-{\hat{y}}^{1/\tau })),$$

where \(\widehat{y}\) are the predictions and y the targets. In principle, f could be any function that is minimal at \(\widehat{y}=y\); we chose the temperature-modulated binary cross-entropy with temperature τ = 5, as recommended by Minixhofer et al.49. In practice, we conducted the operations involved in the loss computation in log-space to ensure stable numerics. We optionally combine the distillation loss \({{\mathcal{L}}}_{{\mathcal{D}},\mathrm{Distill}}\) with a cross-entropy loss to encourage the system to model the training data well and to start exploiting byte-level information, at the cost of giving up exactness if enabled:

$${{\mathcal{L}}}_{{\mathcal{D}},\mathrm{CE}}:= \sum _{j}-\log \,\mathrm{LMHead}\,({\hat{z}}_{\mathrm{subword}})[\,j,\mathrm{next}\_\mathrm{byte}(x,j)].$$

Putting it together

In principle, the boundary predictor and local encoder on the one hand and the local decoder and language-modelling head on the other could be trained separately (assuming we stop the gradient to the encoder through \({\widehat{z}}_{{\rm{subword}}}\)). Although there may be scenarios where this is beneficial, we chose to train them together for simplicity. The complete stage 1 loss is given by

$${{\mathcal{L}}}_{\mathrm{Stage1}}:= {\lambda }_{{\mathcal{B}}}{{\mathcal{L}}}_{{\mathcal{B}}}+{\lambda }_{{\mathcal{E}}}{{\mathcal{L}}}_{{\mathcal{E}}}+{\lambda }_{{\mathcal{D}},\mathrm{Distill}}{{\mathcal{L}}}_{{\mathcal{D}},\mathrm{Distill}}+{\lambda }_{{\mathcal{D}},\mathrm{CE}}{{\mathcal{L}}}_{{\mathcal{D}},\mathrm{CE}},$$

where \({\lambda }_{{\mathcal{B}}},{\lambda }_{{\mathcal{E}}},{\lambda }_{{\mathcal{D}},\mathrm{Distill}},{\lambda }_{{\mathcal{D}},\mathrm{CE}}\in {\rm{{\mathbb{R}}}}\) are the loss weights, which we set as follows: \({\lambda }_{{\mathcal{B}}}=4,{\lambda }_{{\mathcal{E}}}=1,{\lambda }_{{\mathcal{D}},\mathrm{Distill}}=1\,\mathrm{and}\,{\lambda }_{{\mathcal{D}},\mathrm{CE}}=1\). Stage 1 needs in total one forward pass through all layers and one backward pass through the first n layers of the global model plus forward and backward passes through local encoder, local decoder, boundary predictor and language-modelling head. This makes stage 1 substantially more efficient than training the entire model. It could also be further optimized by quantizing or applying inference-specific optimizations to the global model layers starting from the (n + 1)th layer (which we do not need to backpropagate through). We analysed the difference between inserting stage 1 and directly training the entire model end-to-end with randomly initialized parameters (besides the global model) in ‘Benefits of training in two stages’. Besides performance improvements, stage 1 provides a vehicle for rapid experimentation: we can conduct stage 1 training to rapidly check whether a particular architecture for the local encoder and decoder has sufficient capacity to emulate the input and output embedding matrices, respectively. We use this to guide the architecture search for our byteified models under the hypothesis that byte-level architectures that cannot emulate the subword model after stage 1 will remain inadequate with further stage 2 training.

End-to-end training

In the second stage, we train the entire model end-to-end, retaining only the boundary loss \({{\mathcal{L}}}_{{\mathcal{B}}}\) and the cross-entropy loss. For the cross-entropy loss (previously denoted \({{\mathcal{L}}}_{{\mathcal{D}},\mathrm{CE}}\)), we substitute the depooled representations \({\widehat{z}}_{{\rm{subword}}}\) computed from the subword model representations with the true depooled representations \(\widehat{z}\), and we refer to this new loss as \({{\mathcal{L}}}_{\mathrm{CE}}\) instead:

$${{\mathcal{L}}}_{\mathrm{Stage2}}:= {\lambda }_{{\mathcal{B}}}{{\mathcal{L}}}_{{\mathcal{B}}}+{\lambda }_{\mathrm{CE}}{{\mathcal{L}}}_{\mathrm{CE}}.$$

We now optimize all parameters, including those of the global model \({\mathcal{M}}\). This stage is intended for the model to adjust to the end-to-end setting, as in stage 1 we assumed a local encoder and boundary predictor perfectly emulating the subword model, which, although close, is not true in practice. The global model learns to exploit the new byte-level information in stage 2 and can optionally be trained with higher compression ratios of bytes per patch (section ‘Increasing model efficiency’).

Methods for increasing the compression factor

To increase the compression factor of a byteified model, we fix a compression ratio t for the target average bytes per patch. We then remove subword boundaries (merged subword tokens) of \({{\mathcal{B}}}_{\mathrm{subword}}\) until the desired compression ratio is achieved. We experimented with three merging strategies:

  1. (1)

    BPE. We iteratively merged the most common pair of tokens as in byte pair encoding1. In contrast to conventional BPE, we applied BPE per example instead of over the entire corpus. This was inspired by Feher et al.50, who showed that it is possible to retrofit language models to operate over BPE merges of the tokens in their vocabulary.

  2. (2)

    Entropy. We used an auxiliary 370M-parameter subword-level LLM, trained on 74.3B tokens following a downscaled version of the OLMo 2 training and architecture29, to compute next-token entropies. We then iteratively merged the pair of patches that, when summing their individual entropies, resulted in the lowest entropy among all entropy sums of pairs of patches in the example.

  3. (3)

    Cross-entropy. We used the same auxiliary LLM as for entropy-based merging, but instead of merging the pair of tokens with the lowest total entropy, we iteratively merged the pair of tokens with the lowest total cross-entropy with respect to the next token in the data.

For entropy- and cross-entropy-based merging, the auxiliary LLM was required during training time only to supervise the boundary predictor (as in DTP24). Unlike BLT22, we did not need to retain the auxiliary LLM for inference.

Even though the loss is discontinuous with respect to the parameters of the boundary predictor and we did not employ any technique to backpropagate through the discrete boundary predictions, we observed stable training without loss spikes with all of the above merging methods. An important nuance is that the supervision target compression ratio t was not attained by the model. Despite the boundaries not being learned end to end, the model learns to trade off boundary prediction accuracy with the main next-byte prediction loss, like other multitask models that learn to balance performance on the constituent tasks (for example, ref. 51). An important hyperparameter is, thus, the factor \({\lambda }_{{\mathcal{B}}}\) controlling the importance of the boundary prediction task; we keep \({\lambda }_{{\mathcal{B}}}=4\) from stage 1 training.

Experimental set-up

Data

Our data mix consists of approximately 172B tokens, as tokenized by the Dolma2 Tokenizer (https://huggingface.co/allenai/dolma2-tokenizer) from the Dolma 3 pretraining data mix28, augmented with 75M tokens of CUTE-style data11, sampled so as not to overlap with the CUTE test set, to encourage character understanding (‘Character-level training data’ in Methods). Training ran for less than one epoch on this mix.

Model

We used Qwen3 8B Base30, Llama 3 8B31 and the Olmo 3 7B checkpoint after mid-training and long-context extension28 as the starting points for our byteified models. For the local models, we used stacks of alternating mLSTM33 and feedforward layers of size 1 and 4 for the encoder and decoder, respectively (Extended Data Table 5).

Training

We trained stage 1 on 9.8B tokens (approximately 43B bytes). In this stage, we trained the local encoder, decoder, boundary predictor and language-modelling head, keeping the global model frozen. For stage 2, we trained the entire model on 39.3B tokens (approximately 173B bytes; Extended Data Table 4).

Evaluation

We created the 7B byteification evaluation suite based on the Olmo 3 OlmoBaseEval28, skipping GSM Symbolic and BigCodeBench due to their size, and adding CUTE11 and EXECUTE52 to measure character understanding in English and across other languages, respectively. We created the 1B byteification evaluation suite based on the Base Easy Suite in Olmo 328, again adding CUTE11 to measure character understanding. For the 1B suite, we defined a set of core tasks consisting of ARC53, MMLU54, CSQA55, HellaSwag56, WinoGrande57, SocialIQA58, PiQA59, the Basic Skills benchmark28 and CUTE11 for use in ablations and sweeps (Extended Data Tables 1 and 2). The 7B suite additionally comprises HumanEval60, MBPP61, DS-1000 (ref. 62), DeepSeek LeetCode63, MultiPL-E64, GSM8K65, Minerva MATH66, MedMCQA67, MedQA68, SciQ69, DROP70, Natural Questions71, SQuAD72, CoQA73 and Lambada74, in addition to the Jeopardy and Basic Skills tasks from OlmoBaseEval28. We used PyTorch v.2.8 (ref. 75) and vLLM v.0.11.0 (ref. 76) to generate completions whenever available. Generation capabilities for the byteified models are implemented in our bolmo-core (https://github.com/allenai/bolmo-core) framework.

Character-level training data

To encourage models trained on our data mix to learn information about the characters within a word, we generated approximately 75M tokens (approximately 0.04% of the training data) for tasks requiring character-level understanding using the CUTE repository (https://github.com/Leukas/CUTE). Tasks included spelling out words, reversing words as well as swapping, deleting and substituting characters within a word given words in a source wordlist. We used a list of n = 150,000 words, ensuring zero overlap with the CUTE test words to avoid contamination. These data are purely in English. We did not use any multilingual character-understanding data, but we still observed large improvements on the multilingual EXECUTE benchmark, indicating that some texts requiring character-level understanding can help acquire generalizable knowledge about the characters within words. We observed that byte-level models otherwise do not acquire this knowledge through our short training schedule. However, training for longer, on more diverse data or with larger local models could act as alternative routes to acquiring character-level knowledge.

Statistical tests

To test whether model A statistically significantly outperforms model B, we started from the set of benchmark scores, D = {(a1, b1), (a2, b2), …, (an, bn)}, where (ai, bi) are the scores of model A and model B on benchmark i. The set of benchmarks consists of all n tasks in the respective evaluation suite (the 7B byteification suite or 1B byteification suite with n = 40 and n = 22, respectively; Extended Data Tables 1 and 2). We then drew n pairs from D with replacement for every bootstrap iteration k ∈ {1, 2, …, N} with the number of bootstrap iterations N = 10,000:

$${D}_{k}^{\star }=\{({a}_{{j}_{1}},{b}_{{j}_{1}}),({a}_{{j}_{2}},{b}_{{j}_{2}}),\ldots ,({a}_{{j}_{n}},{b}_{{j}_{n}})\}.$$

We then computed the difference in performance for each bootstrap sample,

$${\delta }_{k}^{\star }=\frac{1}{n}\sum _{(a,b)\in {D}_{k}^{\star }}a-b,$$

resulting in an empirical distribution of performance differences \(\varDelta =\{{\delta }_{1}^{\star },{\delta }_{2}^{\star },\ldots ,{\delta }_{N}^{\star }\}\), which we used to compute an unadjusted P value:

$${p}^{{\rm{unadj}}}=\frac{1}{N}\mathop{\sum }\limits_{k=1}^{N}{\mathbb{1}}\,({\delta }_{k}^{\star }\le 0).$$

We then applied a Holm–Bonferroni correction for multiple comparisons. Given a total of K unadjusted P values {Punadj(A, B), Punadj(A, C), …, Punadj(A, Z)}, we sorted the P values in ascending order and adjusted starting from i = 1 (the smallest P value), sequentially moving to larger P values:

$${p}_{(i)}^{{\rm{adj}}}=\min (\mathop{\max }\limits_{j\le i}((K-j+1){p}_{(j)}^{{\rm{unadj}}}),1).$$

Finally, we undid the sorting operation to arrive at an adjusted P value for every comparison.

Related work

Tokenization

LLMs process information represented as a discrete sequence of symbols called tokens or patches. The process of segmenting the input into this discrete sequence is called tokenization, with different ways to tokenize being used across modalities such as text10, audio77 and images78. The predominant approach for tokenizing text since the inception of LLMs has been subword tokenization1,10: tokenizing text into a discrete sequence of units from a finite vocabulary of subword tokens (usually of size 30k–300k), typically represented as integer IDs. Subword tokenization causes several problems. (1) Information about the characters within each token is lost. Although LLMs have been shown to implicitly learn the constituent characters of their tokens11,79 and although it is possible to explicitly re-introduce character information12,80, they still fall short in tasks requiring character knowledge11,13,81. (2) The implicit reliance of subword tokenization on the future contents of the text (called tokenization bias) causes unexpected behaviour at inference if the prompt ends in the middle of a word or with whitespace18,19,20. (3) The need for a fixed, finite subword vocabulary causes restrictive rigidity: for example, although encoding English efficiently is crucial for pretraining, as the vast majority of current pretraining documents are in English, various downstream tasks have different efficiency requirements across different languages. (4) Tokenization in contemporary LLMs is tied to compute allocation: in a standard LLM, the same amount of compute is spent on processing every token in the prefill, every token contributes equally to the key–value cache (KV cache) size and a fixed amount of compute is spent on sequentially generating any new token. Although there are ways to mitigate this problem post hoc—such as KV cache sparsification82 and multi-token prediction83—directly adapting the tokenization and, thus, the compute allocation based on the input instead might be more effective22,24.

Byte-level LLMs

The shortcomings of subword tokenization have motivated extensive work on a wide range of alternatives, which even include tokenizing text by rendering it into pixels and segmenting these into patches84,85,86. The most common alternative has been tokenizing into a smaller set of finer-grained atomic units, such as UTF-8 bytes. One strand of work directly replaces subword tokens with UTF-8 bytes, keeping other aspects of the architecture mostly the same26,27,49,87. This can solve problems (1) and (2) of subword tokenization and potentially problem (3) with the right choice of fine-grained units88,89. In any case, compute allocation remains a problem, exacerbated by having to process longer sequences. To mitigate this problem, some architectures pool a fixed number of tokens into a single representation with a lightweight local encoder (for example, another transformer network), pass the pooled representations through a deep global model operating over the shortened sequence and then depool the representations back to the original granularity through a local decoder. This approach has been pioneered for autoregressive models by the hourglass transformer90 and later adopted more broadly91,92,93. Recent subsequent work has shown that replacing static pooling with dynamic tokenization improves the performance–efficiency Pareto front24,25. In this case, the token boundaries may be learned end to end, rely on entropy spikes or be externally supervised2,24. We collectively refer to these architectures as LTLMs, as—although operating over bytes—they perform a tokenization step inside the model that aggregates the byte representations into representations over latent patches. Byte-level LTLMs can finally address issues (1) to (4) of subword tokenization. The most recent LTLMs have shown promise by performing on par with subword tokenization when spending the same overall number of FLOPs on training2,22. Although we focused on LTLMs in this work, there are also other strands of promising research relevant to byte-level models, such as MrT594, which uses a soft gating mechanism to reduce sequence lengths at inference, and zip2zip95, which adaptively merges tokens based on the past token context.

Tokenizer transfer and retrofitting

Techniques to alter the architecture of a model with extra training are typically referred to as retrofitting, which often relies on self-distillation82,96. The principal difficulty when this involves a change of tokenizer is finding embeddings for the new tokens; this is usually done using heuristics39,97,98 or training-based methods99. Recently, effective tokenizer transfer methods based on cross-tokenizer distillation have been introduced49,100,101. Here the original model is seen as the teacher, the tokenizer-transferred model is seen as the student and the objective is to match the behaviour of the student to the teacher. Byteification is a special case of tokenizer transfer. Byteification was first done by Pagnoni et al.22 by initializing the LTLM parameters from an existing subword model where possible and training as if from random initialization. Hwang et al.2 later byteified by supervising the boundary prediction to match the subword boundaries and introducing an auxiliary embedding-matching loss. Our key contribution is creating an LTLM that is specifically suited to byteifying. We do so by introducing a new architecture (section ‘Byteified language model architecture’) and a dedicated two-stage procedure that efficiently byteifies by first learning to exactly recover the behaviour of the source subword model (section ‘Byteification training procedure’). Together, these innovations allow us to closely match the performance of state-of-the-art subword-level LLMs with a byteified model.

联系我们 contact @ memedata.com