Migrate TextFeatures to narwhals, add polars support - #1074
Merged
Merged
Conversation
TextFeatures now accepts pandas, polars and other dataframes supported by narwhals, and returns the same type it receives. All features are defined once in terms of a few text statistics, computed with pandas string methods and Python loops for pandas, polars string methods for polars, and narwhals expressions for other backends. Pandas outputs are identical to before and the transform is about 3x faster. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…mes_out in TextFeatures Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Migrates
TextFeaturesto narwhals. It now accepts pandas, polars and other dataframes supported by narwhals, and returns the same type it receives.TEXT_FEATURES, from a few text statistics (length, regex counts, counts of a set of characters, word counts, ...). Three small classes compute those statistics for each backend:_PandasText: pandas string methods,str.translateto count sets of characters, and Python loops for word counts. Counts shared by several features (length, whitespace, digits, words) are computed once per variable._PolarsText: polars string methods (count_matches,extract_all), taken from the native namespace, because narwhals has no way to count regex matches._NarwhalsText: narwhals expressions, for other backends. It counts matches by replacing each match with 2 characters instead of 1 and comparing the lengths.str.isspace()is true). polars'\sdoes not match\x1c-\x1f, so without the explicit list polars would give different counts from pandas.ends_with_punctuationuses^[^\n]*[.!?]\n?$on polars, which gives the same result as pandas'str.match(r".*[.!?]$"). On the narwhals path it also checks for a trailing"\n\n", because Python's$also matches before a final newline.fitchecks the dtype with narwhals: string, object, and categorical columns whose categories are strings are accepted, as before. Polars Categorical/Enum columns are accepted too and cast to string.check_Xno longer makes a copy, and the old code wrote the filled text back into it.Benchmarks
Machine: 10-core Mac, pandas 3.0.3 (python-backed strings, pyarrow not installed), polars 1.43.0, narwhals 2.24.0. Text is random sentences of 0-25 words (about 70 characters on average) with capitals, digits and punctuation. Timings are the median of 5 alternating runs, in ms.
Whole
transform(), all 20 featurespandas, old code (on
main) vs this PR:polars, narwhals expressions (
_NarwhalsText) vs polars methods (_PolarsText, used in this PR):Per feature (how each implementation was chosen)
pandas, 500k rows,
strdtype (ms).pd .str= the old code;py= Python list comprehension on.tolist();nw= narwhals expression on pandas:.str(old).str(shared counts)Counting sets of characters, 100k rows (ms): regex
str.countvs deleting the characters withstr.translateand comparing lengths:str.countstr.translate(pandas)str.translate(Python loop)\dstr.translatein pandas is close to the Python loop and keeps pandas' own dtypes (see Behaviour), so pandas uses it.\dand[A-Z](few matches) stay as regex counts.polars, 500k rows (ms), single expression:
I also tried rewriting the narwhals counts as "delete the runs that don't match, then measure the length" (for example
replace_all("[^a-z]+", "")). On polars that took 130-390 ms, still slower thancount_matches.Behaviour
assert_frame_equal, including dtypes) to the old code forobject,strand nullablestringdtypes and forcategory. I checked this on 63 tricky strings: empty, None/NaN, only spaces, tabs and newlines,\xa0,\x1c, U+2003 (em space), U+200B (zero-width space), U+0085 (next line), accented and Greek letters (including final sigma), CJK, emoji, Arabic-Indic and full-width digits,İ,ß/ẞ, text ending in.\nand.\n\n, multiple spaces, uppercase-only text. I also checked a non-default and a duplicated index, integer column names, reordered columns,drop_originaland two variables.lexical_diversitystays as unique words / total words.missing_values="ignore"the text columns in the output have missing values replaced by"". polars does the same.Got {param} instead.convention:variablesshows the value instead of the type name.featureserrors ("features must be None or a list of strings" and "Invalid features: {...}. Available features are: [...]") become one message:features must be None or a list with any of [...]. Got {features} instead.The oldInvalid featuresmessage printed a set, so its order was not deterministic.char_count,avg_word_length,letter_count,special_char_count,digit_ratioanduppercase_rationow say what is actually computed. For example,digit_ratiois digits / non-whitespace characters, not / total characters. I also added a "With polars" example and fixed the "ratio of upper- to lowercase" wording. I ran all the examples in the user guide except the 20newsgroups one; the outputs shown match.Tests
The test file is rewritten to the shared conventions: init error tests with full messages,
test_init_param_assignment,make_dffor both backends,frame_to_dict, and explicit expected values (ratios written as fractions). It adds:each feature on the old test strings, and on the edge cases above, for both backends
the narwhals path, forced on polars data with
monkeypatchoutput dtypes, categorical columns, reordered columns, and the target passed as series/list/array
pandas-only tests for integer column names, the pandas index and numeric categories
tests/test_text: basenarwhals-migration25 failed, 14 passed (the old code breaks sincecheck_Xreturns narwhals frames). This branch: 183 passed, 0 failed.With the old test file against the new code, 37 pass and 2 fail. The 2 failures are the changed
Invalid featuresmessage.No other module imports
feature_engine.text.flake8 feature_engine testsis clean.mypy feature_engineshows the same 2 errors as the base (transformation/log.py), none new.Needs decision
avg_word_lengthislen(text.strip()) / word_count, so it counts the spaces between words:"Hello World!"gives 6.0, while the average word length is 5.5. I kept the old behaviour and documented it. Should it become(characters - whitespace) / words?strdtype when pyarrow is installed): not benchmarked or tested, because pyarrow isn't installed here. There, pandas'.str.count/.str.containsrun pyarrow's RE2 engine. In RE2,\dis ASCII-only and$matches only at the very end, so results already differed from python-backed strings on non-ASCII digits and on text ending in"\n", and they still do.str.translateand the Python word loops fall back to Python elements there, so their speed on pyarrow strings may differ from the table above.features/variableserror messages (see Behaviour).Pre-existing issues, not fixed
'text', soColumnTransformersends a Series andcheck_XraisesTypeError. With['text']it runs, but the output keeps the raw text column, whichStandardScalercan't handle. The fix isTextFeatures(variables=['text'], drop_original=True)with['text']. I didn't change it because I couldn't re-run it to update the accuracies: it needs the 20newsgroups download.objectcolumns holding non-string values pass the dtype check infit. The old code then failed intransformfor most features, and gave 0 for some. Now the word features raiseAttributeError.get_feature_names_outignoresinput_featuresand doesn't check it againstfeature_names_in_, unlike the shared mixin.