Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
72 changes: 67 additions & 5 deletions docs/user_guide/selection/DropConstantFeatures.rst
Original file line number Diff line number Diff line change
Expand Up @@ -71,8 +71,8 @@ Next, we load the Titanic dataset and separate it into a training set and a test
)

Now, we set up the :class:`DropConstantFeatures()` to remove features that show the same
value in more than 70% of the observations. We do this through the parameter `tol`. The
default value for this parameter is zero, in which case it will remove constant features.
value in 70% or more of the observations. We do this through the parameter `tol`. The
default value for this parameter is 1, in which case it will remove constant features.

.. code:: python

Expand All @@ -93,7 +93,7 @@ The variables to drop are stored in the attribute `features_to_drop_`:

transformer.features_to_drop_

These are the 4 features that show the same value in more than 70% of the rows:
These are the 4 features that show the same value in 70% or more of the rows:

.. code:: python

Expand All @@ -115,7 +115,7 @@ We obtain the following proportions:
C 0.195415
Q 0.090611
Missing 0.002183
Name: embarked, dtype: float64
Name: proportion, dtype: float64


Based on the previous results, 71% of the passengers embarked in S.
Expand All @@ -138,7 +138,7 @@ We obtain the following proportions:
5 0.003275
6 0.002183
9 0.001092
Name: parch, dtype: float64
Name: proportion, dtype: float64

Based on the previous results, 77% of the passengers had 0 parent or child. Because of this,
these features were deemed quasi-constant and will be removed in the next step.
Expand Down Expand Up @@ -215,6 +215,68 @@ This and other feature selection methods may not necessarily avoid overfitting,
contribute to simplifying our machine learning pipelines and creating more interpretable
machine learning models.

Missing values
--------------

By default, :class:`DropConstantFeatures()` raises an error if the variables contain
missing values. With `missing_values="include"`, missing values are counted as one more
value of the variable. With `missing_values="ignore"`, missing values are not counted as a
value, but the proportion of the most frequent value is still calculated over all the rows.
In addition, with `tol=1` and `missing_values="ignore"`, variables that show a single value
besides the missing data are dropped.

With polars
-----------

:class:`DropConstantFeatures()` also works with polars dataframes, and returns a polars
dataframe. Both null and NaN are treated as missing values:

.. code:: python

import polars as pl
from feature_engine.selection import DropConstantFeatures

X = pl.DataFrame({
"city": ["London", "London", "London", "London", "Paris"],
"rooms": [3, 2, None, 4, 3],
"garden": [True, True, True, True, True],
"floor": [1.0, 1.0, None, None, 1.0],
})

dcf = DropConstantFeatures(tol=0.8, missing_values="ignore")
Xt = dcf.fit_transform(X)

print(dcf.features_to_drop_)

The variables `city` and `garden` show the same value in 80% or more of the rows. The
variable `floor` shows the value 1.0 in 3 of 5 rows, 60%, because the missing values count
towards the total number of rows:

.. code:: python

['city', 'garden']

The transformed dataframe is a polars dataframe:

.. code:: python

print(Xt)

.. code:: text

shape: (5, 2)
┌───────┬───────┐
│ rooms ┆ floor │
│ --- ┆ --- │
│ i64 ┆ f64 │
╞═══════╪═══════╡
│ 3 ┆ 1.0 │
│ 2 ┆ 1.0 │
│ null ┆ null │
│ 4 ┆ null │
│ 3 ┆ 1.0 │
└───────┴───────┘

Additional resources
--------------------

Expand Down
96 changes: 48 additions & 48 deletions feature_engine/selection/base_recursive_selector.py
Original file line number Diff line number Diff line change
@@ -1,22 +1,23 @@
from types import GeneratorType
from typing import List, Union
from typing import List, Tuple, Union

import pandas as pd
import narwhals as nw
import numpy as np
from narwhals.typing import IntoDataFrame, IntoSeries
from sklearn.inspection import permutation_importance
from sklearn.model_selection import cross_validate

from feature_engine._check_init_parameters.check_variables import (
_check_variables_input_value,
)
from feature_engine.dataframe_checks import check_X_y
from feature_engine.selection.base_selection_functions import get_feature_importances
from feature_engine.selection.base_selection_functions import (
_importance_series,
_select_numerical_variables,
get_feature_importances,
)
from feature_engine.selection.base_selector import BaseSelector
from feature_engine.tags import _return_tags
from feature_engine.variable_handling import (
check_numerical_variables,
find_numerical_variables,
retain_variables_if_in_df,
)

Variables = Union[None, int, str, List[Union[str, int]]]

Expand Down Expand Up @@ -81,10 +82,13 @@ class BaseRecursiveSelector(BaseSelector):
Performance of the model trained using the original dataset.

feature_importances_:
Pandas Series with the feature importance (comes from step 2)
The feature importance (comes from step 2). A pandas Series with the
features as index when X is a pandas dataframe, and a dictionary with the
features as keys otherwise.

feature_importances_std_:
Pandas Series with the standard deviation of the feature importance.
The standard deviation of the feature importance, as a pandas Series or a
dictionary, like `feature_importances_`.

features_to_drop_:
List with the features to remove from the dataset.
Expand Down Expand Up @@ -116,7 +120,9 @@ def __init__(
):

if not isinstance(threshold, (int, float)):
raise ValueError("threshold can only be integer or float")
raise ValueError(
f"threshold must be an integer or a float. Got {threshold} instead."
)

super().__init__(confirm_variables)
self.variables = _check_variables_input_value(variables)
Expand All @@ -126,84 +132,78 @@ def __init__(
self.cv = cv
self.groups = groups

def fit(self, X: pd.DataFrame, y: pd.Series):
def fit(self, X: IntoDataFrame, y: IntoSeries) -> Tuple[nw.DataFrame, IntoSeries]:
"""
Find initial model performance. Sort features by importance.

Parameters
----------
X: pandas dataframe of shape = [n_samples, n_features]
X: dataframe of shape = [n_samples, n_features]
The input dataframe

y: array-like of shape (n_samples)
Target variable. Required to train the estimator.
"""

# check input dataframe
X, y = check_X_y(X, y)
Returns
-------
nw_X: narwhals dataframe
The input dataframe, as a narwhals dataframe.

if self.variables is None:
self.variables_ = find_numerical_variables(X)
else:
if self.confirm_variables is True:
variables_ = retain_variables_if_in_df(X, self.variables)
self.variables_ = check_numerical_variables(X, variables_)
else:
self.variables_ = check_numerical_variables(X, self.variables)
y: Series or numpy array
The target, checked.
"""
nw_X, y = check_X_y(X, y)

self.variables_ = _select_numerical_variables(
X, self.variables, self.confirm_variables
)

self._cv = list(self.cv) if isinstance(self.cv, GeneratorType) else self.cv

# check that there are more than 1 variable to select from
self._check_variable_number()

# save input features
self._get_feature_names_in(X)

X_model = nw_X.select(nw.col(*self.variables_)).to_native()

# train model with all features and cross-validation
model = cross_validate(
estimator=self.estimator,
X=X[self.variables_],
X=X_model,
y=y,
cv=self._cv,
groups=self.groups,
scoring=self.scoring,
return_estimator=True,
)

# store initial model performance
self.initial_model_performance_ = model["test_score"].mean()

# Initialize a dataframe that will contain the list of the feature/coeff
# importance for each cross validation fold
feature_importances_cv = pd.DataFrame()

# Populate the feature_importances_cv dataframe with columns containing
# the feature importance values for each model returned by the cross
# validation.
# There are as many columns as folds.
for i in range(len(model["estimator"])):
m = model["estimator"][i]

# one row of feature importance per cross-validation fold
importances = []
for m in model["estimator"]:
if hasattr(m, "feature_importances_") or hasattr(m, "coef_"):
feature_importances_cv[i] = get_feature_importances(m)
importances.append(get_feature_importances(m))
else:
r = permutation_importance(
m,
X[self.variables_],
X_model,
y,
n_repeats=1,
random_state=10,
)
feature_importances_cv[i] = r.importances_mean
importances.append(r.importances_mean)
importances_arr = np.array(importances)

# Add the variables as index to feature_importances_cv
feature_importances_cv.index = self.variables_

# Aggregate the feature importance returned in each fold
self.feature_importances_ = feature_importances_cv.mean(axis=1)
self.feature_importances_std_ = feature_importances_cv.std(axis=1)
self.feature_importances_ = _importance_series(
X, self.variables_, importances_arr.mean(axis=0)
)
self.feature_importances_std_ = _importance_series(
X, self.variables_, importances_arr.std(axis=0, ddof=1)
)

return X, y
return nw_X, y

def _more_tags(self):
tags_dict = _return_tags()
Expand Down
Loading