Conversation
…s support BaseSelector.transform() returns the retained features in the train set order, in the same library as the input (pandas X[features], narwhals select otherwise). BaseRecursiveSelector.fit() trains the estimators on native frames and returns (nw_X, y). The helpers in base_selection_functions no longer import pandas: correlations are computed with numpy (np.corrcoef, or matrix products for pairwise complete observations when there are missing values), and feature importances are pandas Series for pandas input and dicts otherwise. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
fit() no longer resets the index of the input in place: pandas input gets the probes in a new frame (DataFrame(copy=False) + concat, faster than narwhals), other backends append them with a narwhals horizontal concat. Probe values are identical for a given random_state. probe_features_ is a dataframe in the input library; feature importances are a pandas Series for pandas input and a dict otherwise, as in the base selectors. Tests rewritten to run on pandas and polars; user guide and docstring get a polars example. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #1070 (selection base classes). Its commit shows in the diff until #1070 is merged.
Summary
ProbeFeatureSelectionnow accepts pandas, polars and other dataframes supported by narwhals, and returns the same library it receives.fit():nw_X, y = check_X_y(X, y); the nativeXgoes to_select_numerical_variablesand_get_feature_names_in._generate_probe_features(n_obs)returns a dict{name: numpy array}. The numpy calls and their order did not change, so the probe values for a givenrandom_stateare the same as before.DataFrame(probes, copy=False)+concat([X[variables_].reset_index(drop=True), probes], axis=1), called throughnw.get_native_namespace(X)like_importance_seriesdoes, so there is no pandas import.nw.from_dict(probes, backend=...)+nw.concat(..., how="horizontal").fit()no longer runsX.reset_index(drop=True, inplace=True). It was safe only whilecheck_Xcopied the input. Now thatcheck_Xdoes not copy, it would reset the user's index.collective=Falsebuildsfeature_importances_/feature_importances_std_with_importance_series, as the base does. They are a pandas Series for pandas input and a dict for other input._get_features_to_drop(probes)takes the probe names and reads the importances by key, so it works with both the Series and the dict. For a single probe it still compares against the 1-element array, which keeps the old behaviour for the degenerate cases described below.probe_features_is a dataframe in the input's library.'binomial', which raises an error. It now says'binary', the valid name.Benchmarks
The machine was heavily loaded (load average about 45 on 10 cores), so the numbers are noisy. Each value is the median of 21-25 alternating repeats, in ms. Cross-validation takes almost all of
fit(), so I timed the part the selector does itself: building the probe frame and appending it toX, with the probe arrays already generated.pandas:
DataFrame(copy=False)+concatis the fastest at every size.The two narwhals columns come from a separate run, so compare them with the other columns only roughly.
fit()with the cross-validation replaced by a stub (the selector's own work). Old =origin/main, whosecheck_Xstill copiedX. Each cell lists 3 runs, each the median of 9 fits, in ms:polars: appending the probes costs well under 1 ms at 500k rows and more. Generating the probes with numpy costs 10-1000 ms, and polars-native
hstackand narwhalsconcatare within noise of each other:hstackI kept narwhals
concatfor the non-pandas path. It also covers the other backends, whilehstackwould need a third, polars-only branch to save at most about 0.3 ms.Behaviour
pandas: identical to
origin/mainin 112 recorded cases. The cases cover:'all'and lists, n_probes 1 and 3, the 3 thresholdsvariables,confirm_variables, a splits generator,groupsdistribution=[]andn_probes=0In each case these match exactly:
probe_features_(values, dtypes, RangeIndex),feature_importances_and_std_(values, index),features_to_drop_,variables_, the transform output, the error messages, and the user'sXstaying untouched (same index).polars: same values as pandas, bit for bit, in the 108 applicable cases.
probe_features_is a polars DataFrame, and the importances are dicts (see below).All outputs in the user guide and the docstring were re-run: unchanged. I added the missing
dtype: float64line to one output.Tests
tests/test_selection/test_probe_feature_selection.pyis rewritten to the conventions:match=re.escape, plustest_init_param_assignment.make_df, using thedata_classificationfixture that is already inconftest.py. The expected importances are kept from the old tests.fitresets the index in place)._get_features_to_dropunit tests, with both the Series and the dict container.tests/test_selectionfailing tests: 137 onorigin/narwhals-selection-base, 128 on this branch, no new failures. The 9 tests that now pass are the 5 old probe tests and 4test_check_estimator_selectorscases for this class.tests/parametrize_with_checks_selection_v16.py -k Probe: the same 21 failures before and after. They are caused by numpy input tocheck_X, which fails for every migrated transformer.flake8 feature_engine testsis clean.mypy feature_engineshows the same 2 errors as the base (intransformation/log.py).Needs decision
feature_importances_/feature_importances_std_are dicts. This follows the base container from Migrate the selection base classes and helpers to narwhals, add polars support #1070 and its open question.Pre-existing issues, not fixed
fit()callsnp.random.seed(random_state)and so resets numpy's global random generator. This makes unseeded estimators such asRandomForestClassifier()deterministic, and the user guide and tests rely on it. Switching tonp.random.RandomState(random_state)would give the same probe values, but results with unseeded estimators would change, so I left it.distribution=[]orn_probes <= 0pass__init__, and thenfit()fails with numpy's "The truth value of an empty array is ambiguous". The fix would be to validate both in__init__, which changes when and how the error is raised.collective=Truefails in scikit-learn, because the frame then mixes int and str column names (the probes).collective=Falseworks.n_probes=Trueis accepted as 1, becauseboolis a subclass ofint.