Skip to content
TrackPodcasts
scienceOct 18, 202520:48pending

TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar

About this episode

🤗 Upvotes: 27 | cs.CL, cs.AI, cs.LG, cs.PL, cs.SE

Authors:
Yinxi Li, Yuntian Deng, Pengyu Nie

Title:
TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar

Arxiv:
http://arxiv.org/abs/2510.14972v1

Abstract:
Large language models (LLMs) for code rely on subword tokenizers, such as byte-pair encoding (BPE), learned from mixed natural language text and programming language code but driven by statistics rather than grammar. As a result, semantically identical code snippets can be tokenized differently depending on superficial factors such as whitespace or identifier naming. To measure the impact of this misalignment, we introduce TokDrift, a framework that applies semantic-preserving rewrite rules to create code variants differing only in tokenization. Across nine code LLMs, including large ones with over 30B parameters, even minor formatting changes can cause substantial shifts in model behavior. Layer-wise analysis shows that the issue originates in early embeddings, where subword segmentation fails to capture grammar token boundaries. Our findings identify misaligned tokenization as a hidden obstacle to reliable code understanding and generation, highlighting the need for grammar-aware tokenization for future code LLMs.

Get every episode summarized

Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar

Daily Paper Cast

0:00
20:48

More episodes

More from Daily Paper Cast

View all episodes →