Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@
* ✨ Add `--enable-tables` to the CLI for file, standard input and interactive parsing in [#422](https://github.com/executablebooks/markdown-it-py/pull/422)
* 🐛 Fix CLI interactive mode joining input lines with an extra newline, which split every line into its own paragraph and broke hard line breaks, in [#172](https://github.com/executablebooks/markdown-it-py/issues/172)
* 🐛 Fix trimming and splitting with the Python whitespace set instead of the CommonMark one, which dropped U+001C–U+001F and U+0085 from paragraphs, headings, table cells and fence info strings and let distinct reference labels resolve each other, in [#418](https://github.com/executablebooks/markdown-it-py/pull/418), thanks to [@Nexory](https://github.com/Nexory)
* 🐛 Fix image `alt` text dropping backslash escapes and entities, by also joining `text_special` tokens inside an image (ported from markdown-it 15.0.0), in [#445](https://github.com/executablebooks/markdown-it-py/issues/445)
* 📚 Document the Python renderer constructor contract.
* 🔧 Drop the optional dependency on the `commonmark` package and its benchmark, since the upstream package is unmaintained, in [#401](https://github.com/executablebooks/markdown-it-py/pull/401)

Expand Down
64 changes: 34 additions & 30 deletions markdown_it/rules_core/text_join.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,35 +19,39 @@ def text_join(state: StateCore) -> None:
if inline_token.type != "inline":
continue

# convert text_special to text and join all adjacent text nodes
new_tokens: list[Token] = []
children = inline_token.children or []
i = 0
while i < len(children):
child_token = children[i]
if child_token.type == "text_special":
child_token.type = "text"
if (
child_token.type == "text"
and new_tokens
and new_tokens[-1].type == "text"
):
# Collapse a run of adjacent text nodes in a single join, instead
# of pairwise `a + b` concatenation. The pairwise form is O(L*k)
# in the size of the run because each step rebuilds the growing
# prefix; "".join is O(L).
parts = [new_tokens[-1].content, child_token.content]
for child_token in children:
# image `alt` is parsed into its own token tree
if child_token.children:
child_token.children = _join_text_tokens(child_token.children)
inline_token.children = _join_text_tokens(children)


def _join_text_tokens(tokens: list[Token]) -> list[Token]:
"""Convert `text_special` to `text` and join all adjacent text tokens"""
new_tokens: list[Token] = []
i = 0
while i < len(tokens):
token = tokens[i]
if token.type == "text_special":
token.type = "text"
if token.type == "text" and new_tokens and new_tokens[-1].type == "text":
# Collapse a run of adjacent text nodes in a single join, instead
# of pairwise `a + b` concatenation. The pairwise form is O(L*k)
# in the size of the run because each step rebuilds the growing
# prefix; "".join is O(L).
parts = [new_tokens[-1].content, token.content]
i += 1
while i < len(tokens):
next_token = tokens[i]
if next_token.type == "text_special":
next_token.type = "text"
if next_token.type != "text":
break
parts.append(next_token.content)
i += 1
while i < len(children):
next_token = children[i]
if next_token.type == "text_special":
next_token.type = "text"
if next_token.type != "text":
break
parts.append(next_token.content)
i += 1
new_tokens[-1].content = "".join(parts)
else:
new_tokens.append(child_token)
i += 1
inline_token.children = new_tokens
new_tokens[-1].content = "".join(parts)
else:
new_tokens.append(token)
i += 1
return new_tokens
19 changes: 19 additions & 0 deletions tests/test_port/fixtures/issue-fixes.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,3 +54,22 @@ Fix parsing of incorrect numeric character references
<p><a href="&amp;#X22y;"></a> &amp;#X22y;
<a href="&amp;#35y;"></a> &amp;#35y;</p>
.

#445 Image alt keeps backslash escapes and entities
.
![a \* b](/u)

![a &amp; b](/u)

![a &lt; b](/u)

![a \* *b* &amp; c](/u)

![C:\Python26](/u)
.
<p><img src="/u" alt="a * b" /></p>
<p><img src="/u" alt="a &amp; b" /></p>
<p><img src="/u" alt="a &lt; b" /></p>
<p><img src="/u" alt="a * b &amp; c" /></p>
<p><img src="/u" alt="C:\Python26" /></p>
.