System Information (please complete the following information):
- OS & Version: Windows 11
- ML.NET Version: Microsoft.ML.Tokenizers 3.0.0-preview.26457.2, same code on main at 681cfb6
- .NET Version: .NET 10.0
Describe the bug
LlamaTokenizer encodes some runs of repeated characters differently from SentencePiece. When several pairs of symbols can merge into pieces with the same score, SentencePiece merges the leftmost pair first, while SentencePieceBpeModel merges the rightmost one.
To Reproduce
using Stream stream = File.OpenRead("tokenizer.model"); // https://huggingface.co/hf-internal-testing/llama-tokenizer
LlamaTokenizer tokenizer = LlamaTokenizer.Create(stream);
Console.WriteLine(string.Join(", ", tokenizer.EncodeToIds("____")));
| Text |
LlamaTokenizer |
SentencePiece 0.2.2 |
____ |
1, 4770, 1649 (▁__, __) |
1, 903, 22359 (▁_, ___) |
...... |
1, 6317, 3045 (▁.., ....) |
1, 13035, 636 (▁...., ..) |
~~~~ |
1, 3695, 30022, 7377 (▁~, ~, ~~) |
1, 3695, 7377, 30022 (▁~, ~~, ~) |
Hugging Face tokenizers gives the same ids as SentencePiece. The difference is rare in prose but common in source code, where indentation and separator lines such as // ---------------- are runs of one character. Compared with SentencePiece 0.2.2:
| Corpus |
Texts |
Encoded differently from SentencePiece |
| WikiText-103, test split |
2,891 paragraphs |
1 (0.03%) |
| XNLI, test split, 15 languages |
75,150 sentences |
39 (0.05%) |
| Source code in C#, Java, Python, C++ and TypeScript from public repositories |
7,560 blocks of 20 lines |
1,567 (20.7%) |
| Total |
85,601 |
1,607 |
In code, for instance, a line indented by 20 spaces before if starts with ▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁, ▁▁▁, ▁if in SentencePiece and with ▁▁▁, ▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁, ▁if in LlamaTokenizer. With the fix below, no text in the three sets differs.
Expected behavior
The ids SentencePiece produces. Its SymbolPairComparator in bpe_model.cc ranks a pair lower when the scores are equal and its left is greater (i1 == i2 && h1.left > h2.left), so the leftmost pair merges first.
Additional context
SymbolPair.CompareTo in SentencePieceBpeModel.cs breaks ties with other.Left.CompareTo(Left), which dequeues the rightmost pair first; the SymbolPair of CodeGenTokenizer uses Left.CompareTo(other.Left). I have a fix, Left.CompareTo(other.Left), with the three cases above added to LlamaTestData, and would like to contribute it.
System Information (please complete the following information):
Describe the bug
LlamaTokenizerencodes some runs of repeated characters differently from SentencePiece. When several pairs of symbols can merge into pieces with the same score, SentencePiece merges the leftmost pair first, whileSentencePieceBpeModelmerges the rightmost one.To Reproduce
LlamaTokenizer____1, 4770, 1649(▁__,__)1, 903, 22359(▁_,___)......1, 6317, 3045(▁..,....)1, 13035, 636(▁....,..)~~~~1, 3695, 30022, 7377(▁~,~,~~)1, 3695, 7377, 30022(▁~,~~,~)Hugging Face
tokenizersgives the same ids as SentencePiece. The difference is rare in prose but common in source code, where indentation and separator lines such as// ----------------are runs of one character. Compared with SentencePiece 0.2.2:In code, for instance, a line indented by 20 spaces before
ifstarts with▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁,▁▁▁,▁ifin SentencePiece and with▁▁▁,▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁,▁ifinLlamaTokenizer. With the fix below, no text in the three sets differs.Expected behavior
The ids SentencePiece produces. Its
SymbolPairComparatorinbpe_model.ccranks a pair lower when the scores are equal and itsleftis greater (i1 == i2 && h1.left > h2.left), so the leftmost pair merges first.Additional context
SymbolPair.CompareToinSentencePieceBpeModel.csbreaks ties withother.Left.CompareTo(Left), which dequeues the rightmost pair first; theSymbolPairofCodeGenTokenizerusesLeft.CompareTo(other.Left). I have a fix,Left.CompareTo(other.Left), with the three cases above added toLlamaTestData, and would like to contribute it.