log: escape terminal control characters - #1114
Conversation
|
@rubo77 please review |
|
I have this on my radar, just need some proper time reviewing this complex bounds checking and pointer math. Will come back to it |
This comment was marked as low quality.
This comment was marked as low quality.
This comment was marked as low quality.
This comment was marked as low quality.
|
Both locale concerns from @FusionPow were valid. The filter was treating permitted high bytes as UTF-8 regardless of the active locale. Reworked this to instead use the active locale's decoder to identify boundaries. ISO-8859-1 I also tightened the output bounds before expansion and added regressions for UTF-8, ISO-8859-1, EUC-JP and |
|
And @rubo77, I would report a case only where raw control characters demonstrably reach a terminal. Raw bytes in redirected or machine-readable output can be intentional so this should not be reported broadly without reproducing a terminal path. |
|
Updated description; will address any CI failures. |
Also. Appreciated. |
|
Great job @steadytao. LGTM |
Fixes #1112
Summary
Escape terminal control characters in filenames and diagnostics written to standard output, standard error or rsync log files. The output filter preserves printable text in the active locale while escaping:
\#NNNsequences which would otherwise be ambiguousInvalid or incomplete multibyte input falls back to conservative byte filtering. Generic output no longer trusts leading or trailing carriage returns. The three intentional progress displays emit their intended carriage return through a dedicated helper.
Compatibility
Character boundaries are decoded using the active locale. This preserves valid multibyte text including UTF-8 and EUC-JP while still escaping C1 bytes in single-byte locales such as ISO-8859-1. Systems without
wchar.hormbrtowc()use a conservative byte-oriented fallback. Default filtering without--8-bit-outputremains byte-oriented.