Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
89 changes: 67 additions & 22 deletions docs/user_guide/selection/SelectByShuffling.rst
Original file line number Diff line number Diff line change
Expand Up @@ -104,7 +104,7 @@ the entire dataset, without shuffling, using cross-validation.

.. code:: python

0.488702767247119
0.48870212980353145

In the following sections, we'll explore some of the additional useful data stored by
:class:`SelectByShuffling()`.
Expand All @@ -125,15 +125,15 @@ each feature:
.. code:: python

{'age': -0.0054698043007869734,
'sex': 0.03325633986510784,
'bmi': 0.184158237207512,
'sex': 0.03325633986510779,
'bmi': 0.1841582372075119,
'bp': 0.10089894421748086,
's1': 0.49324432634948095,
's2': 0.21163252880660438,
's3': 0.02006839198785859,
's4': 0.011098050006761673,
's5': 0.4828781996541602,
's6': 0.003963360084439538}
's1': 0.4932443263494815,
's2': 0.2116325288066046,
's3': 0.020068391987858536,
's4': 0.011098050006761617,
's5': 0.4828781996541603,
's6': 0.003963360084439482}

:class:`SelectByShuffling()` stores the standard deviation of the performance change:

Expand All @@ -146,16 +146,16 @@ shuffling:

.. code:: python

{'age': 0.012788500580799392,
{'age': 0.012788500580799462,
'sex': 0.040792331972680645,
'bmi': 0.042212436355346106,
'bp': 0.05397012536801143,
's1': 0.35198797776358015,
's2': 0.167636042355086,
's3': 0.03455158514716544,
's4': 0.007755675852874145,
's5': 0.1449579162698361,
's6': 0.011193022434166025}
'bmi': 0.04221243635534631,
'bp': 0.053970125368011726,
's1': 0.3519879777635807,
's2': 0.1676360423550863,
's3': 0.03455158514716549,
's4': 0.007755675852874141,
's5': 0.14495791626983628,
's6': 0.011193022434166125}

We can plot the performance change together with the standard deviation to get a better
idea of how shuffling features affect the model performance:
Expand All @@ -181,8 +181,8 @@ feature:

.. figure:: ../../images/shuffle-features-std.png

With this set up, features that elicited a mean performance drop greater than the mean
performance of all features, will be removed. If, for any reason, this threshold is too
With this set up, features that elicited a performance drop smaller than the mean
performance drop of all features will be removed. If, for any reason, this threshold is too
conservative or too permissive, by analysing the former barplot, you can get a better
idea of how these features affect the predictions of the model, and select a different
threshold.
Expand All @@ -198,7 +198,7 @@ threshold:
tr.features_to_drop_

The following features were deemed as non-important, because their performance drift is
greater than the mean performance drift of all features:
smaller than the mean performance drift of all features:

.. code:: python

Expand All @@ -221,7 +221,52 @@ In the following output, we see the dataframe with the selected features:
3 -0.011595 0.012191 0.024991 0.022688
4 -0.036385 0.003935 0.015596 -0.031988


Using polars
~~~~~~~~~~~~

:class:`SelectByShuffling()` also works with polars dataframes, and returns a polars
dataframe when the input is a polars dataframe. With the same `random_state`, the
features are shuffled in the same way as with pandas, so the performance drifts and the
selected features are the same.

.. code:: python

import polars as pl

X_pl = pl.DataFrame(X.to_dict(orient="list"))
y_pl = pl.Series("target", y.to_list())

tr = SelectByShuffling(
estimator=LinearRegression(),
scoring="r2",
cv=3,
random_state=0,
)

Xt = tr.fit_transform(X_pl, y_pl)

print(tr.features_to_drop_)
print(Xt.head())

In the following output, we see the features to drop and the polars dataframe with the
selected features:

.. code:: text

['age', 'sex', 'bp', 's3', 's4', 's6']
shape: (5, 4)
┌───────────┬───────────┬───────────┬───────────┐
│ bmi ┆ s1 ┆ s2 ┆ s5 │
│ --- ┆ --- ┆ --- ┆ --- │
│ f64 ┆ f64 ┆ f64 ┆ f64 │
╞═══════════╪═══════════╪═══════════╪═══════════╡
│ 0.061696 ┆ -0.044223 ┆ -0.034821 ┆ 0.019907 │
│ -0.051474 ┆ -0.008449 ┆ -0.019163 ┆ -0.068332 │
│ 0.044451 ┆ -0.045599 ┆ -0.034194 ┆ 0.002861 │
│ -0.011595 ┆ 0.012191 ┆ 0.024991 ┆ 0.022688 │
│ -0.036385 ┆ 0.003935 ┆ 0.015596 ┆ -0.031988 │
└───────────┴───────────┴───────────┴───────────┘

Additional resources
--------------------

Expand Down
96 changes: 48 additions & 48 deletions feature_engine/selection/base_recursive_selector.py
Original file line number Diff line number Diff line change
@@ -1,22 +1,23 @@
from types import GeneratorType
from typing import List, Union
from typing import List, Tuple, Union

import pandas as pd
import narwhals as nw
import numpy as np
from narwhals.typing import IntoDataFrame, IntoSeries
from sklearn.inspection import permutation_importance
from sklearn.model_selection import cross_validate

from feature_engine._check_init_parameters.check_variables import (
_check_variables_input_value,
)
from feature_engine.dataframe_checks import check_X_y
from feature_engine.selection.base_selection_functions import get_feature_importances
from feature_engine.selection.base_selection_functions import (
_importance_series,
_select_numerical_variables,
get_feature_importances,
)
from feature_engine.selection.base_selector import BaseSelector
from feature_engine.tags import _return_tags
from feature_engine.variable_handling import (
check_numerical_variables,
find_numerical_variables,
retain_variables_if_in_df,
)

Variables = Union[None, int, str, List[Union[str, int]]]

Expand Down Expand Up @@ -81,10 +82,13 @@ class BaseRecursiveSelector(BaseSelector):
Performance of the model trained using the original dataset.

feature_importances_:
Pandas Series with the feature importance (comes from step 2)
The feature importance (comes from step 2). A pandas Series with the
features as index when X is a pandas dataframe, and a dictionary with the
features as keys otherwise.

feature_importances_std_:
Pandas Series with the standard deviation of the feature importance.
The standard deviation of the feature importance, as a pandas Series or a
dictionary, like `feature_importances_`.

features_to_drop_:
List with the features to remove from the dataset.
Expand Down Expand Up @@ -116,7 +120,9 @@ def __init__(
):

if not isinstance(threshold, (int, float)):
raise ValueError("threshold can only be integer or float")
raise ValueError(
f"threshold must be an integer or a float. Got {threshold} instead."
)

super().__init__(confirm_variables)
self.variables = _check_variables_input_value(variables)
Expand All @@ -126,84 +132,78 @@ def __init__(
self.cv = cv
self.groups = groups

def fit(self, X: pd.DataFrame, y: pd.Series):
def fit(self, X: IntoDataFrame, y: IntoSeries) -> Tuple[nw.DataFrame, IntoSeries]:
"""
Find initial model performance. Sort features by importance.

Parameters
----------
X: pandas dataframe of shape = [n_samples, n_features]
X: dataframe of shape = [n_samples, n_features]
The input dataframe

y: array-like of shape (n_samples)
Target variable. Required to train the estimator.
"""

# check input dataframe
X, y = check_X_y(X, y)
Returns
-------
nw_X: narwhals dataframe
The input dataframe, as a narwhals dataframe.

if self.variables is None:
self.variables_ = find_numerical_variables(X)
else:
if self.confirm_variables is True:
variables_ = retain_variables_if_in_df(X, self.variables)
self.variables_ = check_numerical_variables(X, variables_)
else:
self.variables_ = check_numerical_variables(X, self.variables)
y: Series or numpy array
The target, checked.
"""
nw_X, y = check_X_y(X, y)

self.variables_ = _select_numerical_variables(
X, self.variables, self.confirm_variables
)

self._cv = list(self.cv) if isinstance(self.cv, GeneratorType) else self.cv

# check that there are more than 1 variable to select from
self._check_variable_number()

# save input features
self._get_feature_names_in(X)

X_model = nw_X.select(nw.col(*self.variables_)).to_native()

# train model with all features and cross-validation
model = cross_validate(
estimator=self.estimator,
X=X[self.variables_],
X=X_model,
y=y,
cv=self._cv,
groups=self.groups,
scoring=self.scoring,
return_estimator=True,
)

# store initial model performance
self.initial_model_performance_ = model["test_score"].mean()

# Initialize a dataframe that will contain the list of the feature/coeff
# importance for each cross validation fold
feature_importances_cv = pd.DataFrame()

# Populate the feature_importances_cv dataframe with columns containing
# the feature importance values for each model returned by the cross
# validation.
# There are as many columns as folds.
for i in range(len(model["estimator"])):
m = model["estimator"][i]

# one row of feature importance per cross-validation fold
importances = []
for m in model["estimator"]:
if hasattr(m, "feature_importances_") or hasattr(m, "coef_"):
feature_importances_cv[i] = get_feature_importances(m)
importances.append(get_feature_importances(m))
else:
r = permutation_importance(
m,
X[self.variables_],
X_model,
y,
n_repeats=1,
random_state=10,
)
feature_importances_cv[i] = r.importances_mean
importances.append(r.importances_mean)
importances_arr = np.array(importances)

# Add the variables as index to feature_importances_cv
feature_importances_cv.index = self.variables_

# Aggregate the feature importance returned in each fold
self.feature_importances_ = feature_importances_cv.mean(axis=1)
self.feature_importances_std_ = feature_importances_cv.std(axis=1)
self.feature_importances_ = _importance_series(
X, self.variables_, importances_arr.mean(axis=0)
)
self.feature_importances_std_ = _importance_series(
X, self.variables_, importances_arr.std(axis=0, ddof=1)
)

return X, y
return nw_X, y

def _more_tags(self):
tags_dict = _return_tags()
Expand Down
Loading