Skip to content

Migrate TextFeatures to narwhals, add polars support - #1074

Merged
solegalli merged 2 commits into
narwhals-migrationfrom
narwhals-text-features
Sep 19, 2026
Merged

solegalli merged 2 commits into
narwhals-migrationfrom
narwhals-text-features

Conversation

@solegalli

Copy link
Copy Markdown
Collaborator

Summary

Migrates TextFeatures to narwhals. It now accepts pandas, polars and other dataframes supported by narwhals, and returns the same type it receives.

  • Each of the 20 features is defined once, in TEXT_FEATURES, from a few text statistics (length, regex counts, counts of a set of characters, word counts, ...). Three small classes compute those statistics for each backend:
    • _PandasText: pandas string methods, str.translate to count sets of characters, and Python loops for word counts. Counts shared by several features (length, whitespace, digits, words) are computed once per variable.
    • _PolarsText: polars string methods (count_matches, extract_all), taken from the native namespace, because narwhals has no way to count regex matches.
    • _NarwhalsText: narwhals expressions, for other backends. It counts matches by replacing each match with 2 characters instead of 1 and comparing the lengths.
  • Whitespace is listed explicitly (the characters for which Python's str.isspace() is true). polars' \s does not match \x1c-\x1f, so without the explicit list polars would give different counts from pandas.
  • ends_with_punctuation uses ^[^\n]*[.!?]\n?$ on polars, which gives the same result as pandas' str.match(r".*[.!?]$"). On the narwhals path it also checks for a trailing "\n\n", because Python's $ also matches before a final newline.
  • fit checks the dtype with narwhals: string, object, and categorical columns whose categories are strings are accepted, as before. Polars Categorical/Enum columns are accepted too and cast to string.
  • The user's dataframe is not modified. check_X no longer makes a copy, and the old code wrote the filled text back into it.

Benchmarks

Machine: 10-core Mac, pandas 3.0.3 (python-backed strings, pyarrow not installed), polars 1.43.0, narwhals 2.24.0. Text is random sentences of 0-25 words (about 70 characters on average) with capitals, digits and punctuation. Timings are the median of 5 alternating runs, in ms.

Whole transform(), all 20 features

pandas, old code (on main) vs this PR:

backend dtype rows text columns old new speed-up
pandas str 10,000 1 198 71 2.77x
pandas object 10,000 1 204 77 2.65x
pandas str 10,000 3 598 204 2.94x
pandas object 10,000 3 695 246 2.83x
pandas str 100,000 1 2457 728 3.37x
pandas object 100,000 1 2148 747 2.88x
pandas str 100,000 3 7148 2453 2.91x
pandas object 100,000 3 6757 2505 2.70x
pandas str 500,000 1 11205 3542 3.16x
pandas object 500,000 1 11144 3714 3.00x
pandas str 500,000 3 42293 18414 2.30x
pandas object 500,000 3 30912 10751 2.88x

polars, narwhals expressions (_NarwhalsText) vs polars methods (_PolarsText, used in this PR):

backend rows text columns narwhals polars speed-up
polars 500,000 1 2116 854 2.48x
polars 500,000 3 4425 2007 2.20x
polars 1,000,000 1 3991 1553 2.57x
polars 1,000,000 3 8918 4253 2.10x
polars 2,000,000 1 12012 5260 2.28x
polars 2,000,000 3 48724 24410 2.00x

Per feature (how each implementation was chosen)

pandas, 500k rows, str dtype (ms). pd .str = the old code; py = Python list comprehension on .tolist(); nw = narwhals expression on pandas:

feature pd .str (old) pd .str (shared counts) py nw
char_count 782 390 381 664
word_count 442 394 184 1000
sentence_count 496 470 477 2022
avg_word_length 533 495 232 2038
digit_count 165 168 171 255
letter_count 1000 936 962 1293
lowercase_count 881 960 996 1214
special_char_count 562 696 637 1235
whitespace_count 602 426 450 730
digit_ratio 856 500 - 1433
uppercase_ratio 946 583 - 1535
unique_word_count 1094 1301 456 - (needs pyarrow)
lexical_diversity 1763 1525 - -
has_digits / has_uppercase / starts / ends / is_empty 25-218 22-230 44-259 28-240

Counting sets of characters, 100k rows (ms): regex str.count vs deleting the characters with str.translate and comparing lengths:

count str.count str.translate (pandas) str.translate (Python loop)
a-z 194 69 62
a-zA-Z 197 68 59
whitespace 81 72 62
special 92 63 58
\d 37 63 56

str.translate in pandas is close to the Python loop and keeps pandas' own dtypes (see Behaviour), so pandas uses it. \d and [A-Z] (few matches) stay as regex counts.

polars, 500k rows (ms), single expression:

feature narwhals polars
char_count 222 146
word_count 391 138
sentence_count 711 77
avg_word_length 768 290
digit_count 72 38
letter_count 498 239
lowercase_count 492 237
special_char_count 237 140
whitespace_count 222 136
unique_word_count 590 286
lexical_diversity 1346 552
has_digits / has_uppercase / starts / ends 6-43 6-43 (same expression)

I also tried rewriting the narwhals counts as "delete the runs that don't match, then measure the length" (for example replace_all("[^a-z]+", "")). On polars that took 130-390 ms, still slower than count_matches.

Behaviour

  • pandas: identical outputs (assert_frame_equal, including dtypes) to the old code for object, str and nullable string dtypes and for category. I checked this on 63 tricky strings: empty, None/NaN, only spaces, tabs and newlines, \xa0, \x1c, U+2003 (em space), U+200B (zero-width space), U+0085 (next line), accented and Greek letters (including final sigma), CJK, emoji, Arabic-Indic and full-width digits, İ, ß/ẞ, text ending in .\n and .\n\n, multiple spaces, uppercase-only text. I also checked a non-default and a duplicated index, integer column names, reordered columns, drop_original and two variables.
  • polars and the narwhals path give the same values as pandas on the same strings. Output dtypes are Int64 for counts and flags and Float64 for ratios, like pandas' int64/float64.
  • lexical_diversity stays as unique words / total words.
  • As before, with missing_values="ignore" the text columns in the output have missing values replaced by "". polars does the same.
  • Init error messages now follow the Got {param} instead. convention:
    • variables shows the value instead of the type name.
    • The two features errors ("features must be None or a list of strings" and "Invalid features: {...}. Available features are: [...]") become one message: features must be None or a list with any of [...]. Got {features} instead. The old Invalid features message printed a set, so its order was not deterministic.
  • Docs: the descriptions of char_count, avg_word_length, letter_count, special_char_count, digit_ratio and uppercase_ratio now say what is actually computed. For example, digit_ratio is digits / non-whitespace characters, not / total characters. I also added a "With polars" example and fixed the "ratio of upper- to lowercase" wording. I ran all the examples in the user guide except the 20newsgroups one; the outputs shown match.

Tests

The test file is rewritten to the shared conventions: init error tests with full messages, test_init_param_assignment, make_df for both backends, frame_to_dict, and explicit expected values (ratios written as fractions). It adds:

  • each feature on the old test strings, and on the edge cases above, for both backends

  • the narwhals path, forced on polars data with monkeypatch

  • output dtypes, categorical columns, reordered columns, and the target passed as series/list/array

  • pandas-only tests for integer column names, the pandas index and numeric categories

  • tests/test_text: base narwhals-migration 25 failed, 14 passed (the old code breaks since check_X returns narwhals frames). This branch: 183 passed, 0 failed.

  • With the old test file against the new code, 37 pass and 2 fail. The 2 failures are the changed Invalid features message.

  • No other module imports feature_engine.text.

  • flake8 feature_engine tests is clean. mypy feature_engine shows the same 2 errors as the base (transformation/log.py), none new.

Needs decision

  1. avg_word_length is len(text.strip()) / word_count, so it counts the spaces between words: "Hello World!" gives 6.0, while the average word length is 5.5. I kept the old behaviour and documented it. Should it become (characters - whitespace) / words?
  2. pyarrow-backed pandas strings (the default str dtype when pyarrow is installed): not benchmarked or tested, because pyarrow isn't installed here. There, pandas' .str.count/.str.contains run pyarrow's RE2 engine. In RE2, \d is ASCII-only and $ matches only at the very end, so results already differed from python-backed strings on non-ASCII digits and on text ending in "\n", and they still do. str.translate and the Python word loops fall back to Python elements there, so their speed on pyarrow strings may differ from the table above.
  3. The new features / variables error messages (see Behaviour).

Pre-existing issues, not fixed

  • The "Combining with sklearn's bag-of-words" example in the user guide fails. It passes the column as 'text', so ColumnTransformer sends a Series and check_X raises TypeError. With ['text'] it runs, but the output keeps the raw text column, which StandardScaler can't handle. The fix is TextFeatures(variables=['text'], drop_original=True) with ['text']. I didn't change it because I couldn't re-run it to update the accuracies: it needs the 20newsgroups download.
  • object columns holding non-string values pass the dtype check in fit. The old code then failed in transform for most features, and gave 0 for some. Now the word features raise AttributeError.
  • get_feature_names_out ignores input_features and doesn't check it against feature_names_in_, unlike the shared mixin.

solegalli and others added 2 commits September 19, 2026 11:32
TextFeatures now accepts pandas, polars and other dataframes supported by
narwhals, and returns the same type it receives. All features are defined once
in terms of a few text statistics, computed with pandas string methods and
Python loops for pandas, polars string methods for polars, and narwhals
expressions for other backends. Pandas outputs are identical to before and
the transform is about 3x faster.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…mes_out in TextFeatures

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@solegalli
solegalli merged commit 91d01de into narwhals-migration Sep 19, 2026
4 of 10 checks passed
@solegalli
solegalli deleted the narwhals-text-features branch September 19, 2026 11:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant