Skip to content

Migrate DropCorrelatedFeatures to narwhals, add polars support - #1087

Open
solegalli wants to merge 3 commits into
narwhals-migrationfrom
narwhals-drop-correlated-features
Open

solegalli wants to merge 3 commits into
narwhals-migrationfrom
narwhals-drop-correlated-features

Conversation

@solegalli

Copy link
Copy Markdown
Collaborator

Stacked on #1070 (selection base classes). Its commits show in this diff until #1070 merges. This PR's own commit is the last one.

Summary

  • DropCorrelatedFeatures.fit() validates X with check_X(X) and passes the native dataframe to _select_numerical_variables, _check_contains_na / _check_contains_inf, find_correlated_features and _get_feature_names_in. The correlation work happens in find_correlated_features from Migrate the selection base classes and helpers to narwhals, add polars support #1070. transform is inherited from BaseSelector.
  • On the base ref, fit rebinds X = check_X(X), so pandas input reached find_correlated_features as a narwhals frame and took the non-pandas paths. Passing the native frame lets pandas use its own paths again. In particular, spearman with missing values, kendall and callables go back to DataFrame.corr(), and the error for an unknown method is pandas' ValueError again, as on main.
  • import pandas is removed. The type hints are now IntoDataFrame / IntoSeries.
  • The missing_values init check now checks isinstance(missing_values, str) before the membership test, per AGENTS.md. The allowed values and the message are unchanged.
  • Docstring changes: a polars example; fit parameters are backend-neutral; the intro explains that each pair is compared on the rows where both have values. Attribute descriptions are fixed: features_to_drop_ is a list, not a set; correlated_feature_sets_ holds sets; correlated_feature_dict_ referred to a nonexistent correlated_feature_groups.
  • User guide: a "With polars" section, with output from a real run. Two typos are fixed: "Spearman, Kendall, or Spearman", and "variables 6, 7 and 8" should be 6, 7 and 9. The existing pandas outputs in the guide were re-run and still match.

Benchmarks

The selector itself only wires the steps together, so the comparison is between the whole fit() / fit_transform() on the base ref (a narwhals frame handed to the helpers) and on this PR (the native frame). Times are the median of 5 alternating runs in ms (1 run for kendall and for spearman with NaN). Data has correlated columns, and "NaN" means 5% missing values.

pandas

method NaN rows × cols base fit this PR fit base fit_transform this PR fit_transform
pearson no 10k × 50 2.5 0.7 2.6 0.9
pearson no 100k × 50 9.8 7.0 10.7 7.9
pearson no 500k × 10 11.1 8.6 12.2 9.3
pearson no 500k × 50 38.0 31.2 41.7 35.0
pearson yes 10k × 50 3.9 2.0 4.0 2.1
pearson yes 100k × 50 21.6 18.9 22.6 19.8
pearson yes 500k × 50 105.0 98.4 107.8 100.7
spearman no 10k × 50 31.7 29.4 32.0 29.7
spearman no 100k × 50 366.9 380.4 363.3 379.0
spearman no 500k × 10 424.7 436.8 424.1 436.0
spearman no 500k × 50 2123 2211 2125 2210
spearman yes 10k × 10 465.9 51.2 467.4 51.3
spearman yes 100k × 10 4045 663 4025 668
kendall no / yes 10k × 10 62.1 / 58.2 60.8 / 56.6 61.3 / 61.2 67.4 / 56.7

polars (the same code path in both versions: the helpers convert with nw.from_native either way)

method NaN rows × cols base fit this PR fit
pearson no 500k × 50 43.9 44.0
pearson no 1M × 50 85.7 83.6
pearson yes 500k × 50 104.0 109.3
pearson yes 1M × 50 229.0 228.1
spearman no 500k × 50 304.3 309.1
spearman no 1M × 50 685.7 676.4
spearman yes 500k × 10 2623 2773
kendall no / yes 100k × 50 19134 / 18707 19387 / 19135

The native frame wins or ties everywhere, except pandas spearman without NaN. There, the base ref's accidental path (narwhals rank on the pandas backend) is 3-4% faster than the scipy rankdata that find_correlated_features picks for pandas. That choice belongs to the base file, so this PR doesn't change it (see Needs decision). Spearman with NaN is 6-9x faster on pandas because DataFrame.corr() is used again.

Behaviour

  • Outputs were recorded on main (pandas) for 64 cases. The cases cover the 4 methods (pearson, spearman, kendall, a callable) at 4 thresholds, with and without NaN, plus mixed dtypes (categorical, int, constant, negatively correlated), reordered columns, and missing_values="raise" with NaN and with inf. Also covered: variables subsets, confirm_variables, unknown and non-numerical variables, fewer than 2 variables, an unknown method, thresholds 0.0 and 1.0, transform with reordered columns and with a different number of columns, empty/all-NaN/two-row inputs, integer column names with a non-default index, and a named columns index.
  • pandas: 63 of 64 identical to main, including attributes, get_support, get_feature_names_out, transform values, dtypes, column order and index. The exception is the empty-dataframe error message, which comes from check_X on narwhals-migration and not from this PR.
  • polars gives the same values as pandas in 60 of 62 cases. The two differences are both errors:
  • Compared with the base ref on pandas, only those same two errors change, and both go back to main's behaviour.

Tests

  • tests/test_selection/test_drop_correlated_features.py is rewritten to the conventions:
    • # init parameters: one error test per message (threshold, missing_values, confirm_variables) with wrong values and wrong types, and test_init_param_assignment.
    • # fit and transform: every test runs on pandas and polars via make_df, with data as dicts and frame_to_dict comparisons.
    • A small handcrafted dataset makes pearson, spearman, kendall and a callable give different drops, so the methods are actually told apart. The make_classification data gives the same result for all methods.
    • There are tests for NaN handled pairwise (all 4 methods), negative correlation, non-numerical columns ignored and kept, variables subsets, confirm_variables, and the NaN, inf, <2 variables, non-numerical and non-dataframe errors.
    • pandas-only tests: integer column names (assert_frame_equal), and the unknown-method error, whose message comes from pandas.
  • tests/test_selection, run one file at a time: 137 failing on the base ref, 136 on this branch. No new failures. The fixed one is the old test_raises_error_when_method_not_permitted, which failed on the base ref because pandas went through the narwhals path; its replacement passes.
  • tests/parametrize_with_checks_selection_v16.py: 289 failing before and after, the same set (numpy-array input checks).
  • flake8 feature_engine tests is clean. mypy feature_engine shows the same 2 errors as the base ref (datetime_subtraction.py, log.py).

Needs decision

  • Unknown method with polars. method is not validated in __init__. pandas raises at fit with pandas' ValueError, while polars raises TypeError: 'str' object is not callable from find_correlated_features. I kept the current behaviour. Proposal: validate in __init__ (isinstance(method, str) and method in ["pearson", "spearman", "kendall"] or callable(method)), with a message ending "Got {method} instead.". That moves the error from fit to init and changes the pandas message, so it's your call. SmartCorrelatedSelection has the same parameter and would need the same change.
  • method type hint. find_correlated_features(method: str) in the base file rejects callables for mypy, so the selector keeps method: str although callables are allowed. Widening it to Union[str, Callable] needs a change in base_selection_functions.py.
  • Spearman without NaN on pandas (base file): narwhals rank on pandas was 3-4% faster than scipy rankdata in the fit-level numbers above. Migrate the selection base classes and helpers to narwhals, add polars support #1070 measured the opposite at the function level, so it may be noise. It's worth a re-check there, not here.

Pre-existing issues, not fixed

  • pandas dataframes that mix integer and string column names raise TypeError: '<' not supported between instances of 'str' and 'int' in fit, because the variables are sorted alphabetically. This also happens on main. Fixing it (for example, sorting by str) could change which feature is kept in existing pipelines.

solegalli and others added 3 commits September 19, 2026 11:44
…s support

BaseSelector.transform() returns the retained features in the train set
order, in the same library as the input (pandas X[features], narwhals
select otherwise). BaseRecursiveSelector.fit() trains the estimators on
native frames and returns (nw_X, y). The helpers in
base_selection_functions no longer import pandas: correlations are
computed with numpy (np.corrcoef, or matrix products for pairwise
complete observations when there are missing values), and feature
importances are pandas Series for pandas input and dicts otherwise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
fit() validates X with check_X and hands the native dataframe to the variable,
missing-value and correlation helpers, so pandas keeps its fast paths in
find_correlated_features and polars is supported. The tests are rewritten to
run on pandas and polars, and the user guide gets a polars example.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant