Skip to content

LlamaTokenizer encodes some runs of repeated characters differently from SentencePiece #7751

Description

@Laurianti

System Information (please complete the following information):

  • OS & Version: Windows 11
  • ML.NET Version: Microsoft.ML.Tokenizers 3.0.0-preview.26457.2, same code on main at 681cfb6
  • .NET Version: .NET 10.0

Describe the bug
LlamaTokenizer encodes some runs of repeated characters differently from SentencePiece. When several pairs of symbols can merge into pieces with the same score, SentencePiece merges the leftmost pair first, while SentencePieceBpeModel merges the rightmost one.

To Reproduce

using Stream stream = File.OpenRead("tokenizer.model"); // https://huggingface.co/hf-internal-testing/llama-tokenizer
LlamaTokenizer tokenizer = LlamaTokenizer.Create(stream);
Console.WriteLine(string.Join(", ", tokenizer.EncodeToIds("____")));
Text LlamaTokenizer SentencePiece 0.2.2
____ 1, 4770, 1649 (▁__, __) 1, 903, 22359 (▁_, ___)
...... 1, 6317, 3045 (▁.., ....) 1, 13035, 636 (▁...., ..)
~~~~ 1, 3695, 30022, 7377 (▁~, ~, ~~) 1, 3695, 7377, 30022 (▁~, ~~, ~)

Hugging Face tokenizers gives the same ids as SentencePiece. The difference is rare in prose but common in source code, where indentation and separator lines such as // ---------------- are runs of one character. Compared with SentencePiece 0.2.2:

Corpus Texts Encoded differently from SentencePiece
WikiText-103, test split 2,891 paragraphs 1 (0.03%)
XNLI, test split, 15 languages 75,150 sentences 39 (0.05%)
Source code in C#, Java, Python, C++ and TypeScript from public repositories 7,560 blocks of 20 lines 1,567 (20.7%)
Total 85,601 1,607

In code, for instance, a line indented by 20 spaces before if starts with ▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁, ▁▁▁, ▁if in SentencePiece and with ▁▁▁, ▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁, ▁if in LlamaTokenizer. With the fix below, no text in the three sets differs.

Expected behavior
The ids SentencePiece produces. Its SymbolPairComparator in bpe_model.cc ranks a pair lower when the scores are equal and its left is greater (i1 == i2 && h1.left > h2.left), so the leftmost pair merges first.

Additional context
SymbolPair.CompareTo in SentencePieceBpeModel.cs breaks ties with other.Left.CompareTo(Left), which dequeues the rightmost pair first; the SymbolPair of CodeGenTokenizer uses Left.CompareTo(other.Left). I have a fix, Left.CompareTo(other.Left), with the three cases above added to LlamaTestData, and would like to contribute it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    untriagedNew issue has not been triaged

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions