Spec-Zone.ru › scikit-learn

Примечание

Перейти к концу для загрузки полного примера кода. или запустить этот пример в браузере через JupyterLite или Binder

Основанный на модели и последовательный отбор признаков

В этом примере проиллюстрированы и сравниваются два подхода к отбору признаков: SelectFromModel, основанный на важности признаков, и SequentialFeatureSelector, основанный на жадном подходе.

Мы используем набор данных по диабету, который содержит 10 признаков, собранных у 442 пациентов с диабетом.

Авторы: Manoj Kumar, Maria Telenczuk, Nicolas Hug.

Лицензия: BSD 3-х пунктов

# Authors: The scikit-learn developers
# SPDX-License-Identifier: BSD-3-Clause

Загрузка данных

Сначала мы загружаем набор данных по диабету, доступный в scikit-learn, и выводим его описание:

from sklearn.datasets import load_diabetes

diabetes = load_diabetes()
X, y = diabetes.data, diabetes.target
print(diabetes.DESCR)
.. _diabetes_dataset:

Diabetes dataset
----------------

Ten baseline variables, age, sex, body mass index, average blood
pressure, and six blood serum measurements were obtained for each of n =
442 diabetes patients, as well as the response of interest, a
quantitative measure of disease progression one year after baseline.

**Data Set Characteristics:**

:Number of Instances: 442

:Number of Attributes: First 10 columns are numeric predictive values

:Target: Column 11 is a quantitative measure of disease progression one year after baseline

:Attribute Information:
    - age     age in years
    - sex
    - bmi     body mass index
    - bp      average blood pressure
    - s1      tc, total serum cholesterol
    - s2      ldl, low-density lipoproteins
    - s3      hdl, high-density lipoproteins
    - s4      tch, total cholesterol / HDL
    - s5      ltg, possibly log of serum triglycerides level
    - s6      glu, blood sugar level

Note: Each of these 10 feature variables have been mean centered and scaled by the standard deviation times the square root of `n_samples` (i.e. the sum of squares of each column totals 1).

Source URL:
https://www4.stat.ncsu.edu/~boos/var.select/diabetes.html

For more information see:
Bradley Efron, Trevor Hastie, Iain Johnstone and Robert Tibshirani (2004) "Least Angle Regression," Annals of Statistics (with discussion), 407-499.
(https://web.stanford.edu/~hastie/Papers/LARS/LeastAngle_2002.pdf)

Важность признаков по коэффициентам

Чтобы получить представление о важности признаков, мы будем использовать оценщик RidgeCV. Признаки с наибольшим абсолютным coef_ значением считаются наиболее важными. Мы можем наблюдать коэффициенты непосредственно, не нуждаясь в их масштабировании (или масштабировании данных), потому что из описания выше мы знаем, что признаки уже были стандартизированы. Для более полного примера по интерпретации коэффициентов линейных моделей вы можете обратиться к Распространенные ошибки в интерпретации коэффициентов линейных моделей. # noqa: E501

import matplotlib.pyplot as plt
import numpy as np

from sklearn.linear_model import RidgeCV

ridge = RidgeCV(alphas=np.logspace(-6, 6, num=5)).fit(X, y)
importance = np.abs(ridge.coef_)
feature_names = np.array(diabetes.feature_names)
plt.bar(height=importance, x=feature_names)
plt.title("Feature importances via coefficients")
plt.show()
Feature importances via coefficients

Выбор признаков на основе важности

Теперь мы хотим выбрать два наиболее важных признака по коэффициентам. Для этого предназначен SelectFromModel. SelectFromModel принимает параметр threshold и выбирает признаки, важность которых (определяемая коэффициентами) превышает этот порог.

Поскольку мы хотим выбрать только 2 признака, мы установим этот порог немного выше коэффициента третьего по важности признака.

from time import time

from sklearn.feature_selection import SelectFromModel

threshold = np.sort(importance)[-3] + 0.01

tic = time()
sfm = SelectFromModel(ridge, threshold=threshold).fit(X, y)
toc = time()
print(f"Features selected by SelectFromModel: {feature_names[sfm.get_support()]}")
print(f"Done in {toc - tic:.3f}s")
Features selected by SelectFromModel: ['s1' 's5']
Done in 0.002s

Выбор признаков с помощью последовательного отбора признаков

Другой способ выбора признаков – использование SequentialFeatureSelector (SFS). SFS – жадный процесс, где на каждой итерации мы выбираем лучший новый признак для добавления к нашим выбранным признакам на основе результата перекрестной проверки. То есть, мы начинаем с 0 признаков и выбираем лучший одиночный признак с наивысшим результатом. Процедура повторяется до тех пор, пока мы не достигнем желаемого количества выбранных признаков.

Мы также можем пойти в обратном направлении (обратный SFS), т.е. начать с всех признаков и жадно выбирать признаки для удаления по одному.

from sklearn.feature_selection import SequentialFeatureSelector

tic_fwd = time()
sfs_forward = SequentialFeatureSelector(
    ridge, n_features_to_select=2, direction="forward"
).fit(X, y)
toc_fwd = time()

tic_bwd = time()
sfs_backward = SequentialFeatureSelector(
    ridge, n_features_to_select=2, direction="backward"
).fit(X, y)
toc_bwd = time()

print(
    "Features selected by forward sequential selection: "
    f"{feature_names[sfs_forward.get_support()]}"
)
print(f"Done in {toc_fwd - tic_fwd:.3f}s")
print(
    "Features selected by backward sequential selection: "
    f"{feature_names[sfs_backward.get_support()]}"
)
print(f"Done in {toc_bwd - tic_bwd:.3f}s")
Features selected by forward sequential selection: ['bmi' 's5']
Done in 0.212s
Features selected by backward sequential selection: ['bmi' 's5']
Done in 0.604s

Интересно, что прямой и обратный отбор выбрали один и тот же набор признаков. Как правило, это не так, и оба метода дадут разные результаты.

Мы также отмечаем, что признаки, выбранные SFS, отличаются от тех, которые были выбраны по важности признаков: SFS выбирает bmi вместо s1. Это, однако, кажется разумным, так как bmi соответствует третьему по важности признаку по коэффициентам. Это довольно примечательно, учитывая, что SFS вообще не использует коэффициенты.

Завершая, следует отметить, что SelectFromModel значительно быстрее, чем SFS. Действительно, SelectFromModel требует обучения модели только один раз, тогда как SFS требует перекрестной проверки многих разных моделей для каждой итерации. Однако SFS работает с любой моделью, в то время как SelectFromModel требует, чтобы базовый оценщик предоставлял атрибут coef_ или атрибут feature_importances_. Прямой SFS быстрее обратного SFS, поскольку он требует только n_features_to_select = 2 итераций, в то время как обратный SFS требует n_features - n_features_to_select = 8 итераций.

Использование отрицательных значений толерантности

SequentialFeatureSelector может использоваться для удаления признаков, присутствующих в наборе данных, и возврата меньшего подмножества исходных признаков с direction="backward" и отрицательным значением tol.

Мы начинаем с загрузки набора данных по раку молочной железы, содержащего 30 различных признаков и 569 образцов.

import numpy as np

from sklearn.datasets import load_breast_cancer

breast_cancer_data = load_breast_cancer()
X, y = breast_cancer_data.data, breast_cancer_data.target
feature_names = np.array(breast_cancer_data.feature_names)
print(breast_cancer_data.DESCR)
.. _breast_cancer_dataset:

Breast cancer wisconsin (diagnostic) dataset
--------------------------------------------

**Data Set Characteristics:**

:Number of Instances: 569

:Number of Attributes: 30 numeric, predictive attributes and the class

:Attribute Information:
    - radius (mean of distances from center to points on the perimeter)
    - texture (standard deviation of gray-scale values)
    - perimeter
    - area
    - smoothness (local variation in radius lengths)
    - compactness (perimeter^2 / area - 1.0)
    - concavity (severity of concave portions of the contour)
    - concave points (number of concave portions of the contour)
    - symmetry
    - fractal dimension ("coastline approximation" - 1)

    The mean, standard error, and "worst" or largest (mean of the three
    worst/largest values) of these features were computed for each image,
    resulting in 30 features.  For instance, field 0 is Mean Radius, field
    10 is Radius SE, field 20 is Worst Radius.

    - class:
            - WDBC-Malignant
            - WDBC-Benign

:Summary Statistics:

===================================== ====== ======
                                        Min    Max
===================================== ====== ======
radius (mean):                        6.981  28.11
texture (mean):                       9.71   39.28
perimeter (mean):                     43.79  188.5
area (mean):                          143.5  2501.0
smoothness (mean):                    0.053  0.163
compactness (mean):                   0.019  0.345
concavity (mean):                     0.0    0.427
concave points (mean):                0.0    0.201
symmetry (mean):                      0.106  0.304
fractal dimension (mean):             0.05   0.097
radius (standard error):              0.112  2.873
texture (standard error):             0.36   4.885
perimeter (standard error):           0.757  21.98
area (standard error):                6.802  542.2
smoothness (standard error):          0.002  0.031
compactness (standard error):         0.002  0.135
concavity (standard error):           0.0    0.396
concave points (standard error):      0.0    0.053
symmetry (standard error):            0.008  0.079
fractal dimension (standard error):   0.001  0.03
radius (worst):                       7.93   36.04
texture (worst):                      12.02  49.54
perimeter (worst):                    50.41  251.2
area (worst):                         185.2  4254.0
smoothness (worst):                   0.071  0.223
compactness (worst):                  0.027  1.058
concavity (worst):                    0.0    1.252
concave points (worst):               0.0    0.291
symmetry (worst):                     0.156  0.664
fractal dimension (worst):            0.055  0.208
===================================== ====== ======

:Missing Attribute Values: None

:Class Distribution: 212 - Malignant, 357 - Benign

:Creator:  Dr. William H. Wolberg, W. Nick Street, Olvi L. Mangasarian

:Donor: Nick Street

:Date: November, 1995

This is a copy of UCI ML Breast Cancer Wisconsin (Diagnostic) datasets.
https://goo.gl/U2Uwz2

Features are computed from a digitized image of a fine needle
aspirate (FNA) of a breast mass.  They describe
characteristics of the cell nuclei present in the image.

Separating plane described above was obtained using
Multisurface Method-Tree (MSM-T) [K. P. Bennett, "Decision Tree
Construction Via Linear Programming." Proceedings of the 4th
Midwest Artificial Intelligence and Cognitive Science Society,
pp. 97-101, 1992], a classification method which uses linear
programming to construct a decision tree.  Relevant features
were selected using an exhaustive search in the space of 1-4
features and 1-3 separating planes.

The actual linear program used to obtain the separating plane
in the 3-dimensional space is that described in:
[K. P. Bennett and O. L. Mangasarian: "Robust Linear
Programming Discrimination of Two Linearly Inseparable Sets",
Optimization Methods and Software 1, 1992, 23-34].

This database is also available through the UW CS ftp server:

ftp ftp.cs.wisc.edu
cd math-prog/cpo-dataset/machine-learn/WDBC/

.. dropdown:: References

  - W.N. Street, W.H. Wolberg and O.L. Mangasarian. Nuclear feature extraction
    for breast tumor diagnosis. IS&T/SPIE 1993 International Symposium on
    Electronic Imaging: Science and Technology, volume 1905, pages 861-870,
    San Jose, CA, 1993.
  - O.L. Mangasarian, W.N. Street and W.H. Wolberg. Breast cancer diagnosis and
    prognosis via linear programming. Operations Research, 43(4), pages 570-577,
    July-August 1995.
  - W.H. Wolberg, W.N. Street, and O.L. Mangasarian. Machine learning techniques
    to diagnose breast cancer from fine-needle aspirates. Cancer Letters 77 (1994)
    163-171.

Мы будем использовать оценщик LogisticRegression с SequentialFeatureSelector для выполнения отбора признаков.

from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

for tol in [-1e-2, -1e-3, -1e-4]:
    start = time()
    feature_selector = SequentialFeatureSelector(
        LogisticRegression(),
        n_features_to_select="auto",
        direction="backward",
        scoring="roc_auc",
        tol=tol,
        n_jobs=2,
    )
    model = make_pipeline(StandardScaler(), feature_selector, LogisticRegression())
    model.fit(X, y)
    end = time()
    print(f"\ntol: {tol}")
    print(f"Features selected: {feature_names[model[1].get_support()]}")
    print(f"ROC AUC score: {roc_auc_score(y, model.predict_proba(X)[:, 1]):.3f}")
    print(f"Done in {end - start:.3f}s")
tol: -0.01
Features selected: ['worst perimeter']
ROC AUC score: 0.975
Done in 11.479s

tol: -0.001
Features selected: ['radius error' 'fractal dimension error' 'worst texture'
 'worst perimeter' 'worst concave points']
ROC AUC score: 0.997
Done in 11.204s

tol: -0.0001
Features selected: ['mean compactness' 'mean concavity' 'mean concave points' 'radius error'
 'area error' 'concave points error' 'symmetry error'
 'fractal dimension error' 'worst texture' 'worst perimeter' 'worst area'
 'worst concave points' 'worst symmetry']
ROC AUC score: 0.998
Done in 9.760s

Мы можем заметить, что количество выбранных признаков имеет тенденцию увеличиваться, поскольку отрицательные значения tol приближаются к нулю. Время, затрачиваемое на отбор признаков, также уменьшается по мере приближения значений tol к нулю.

Общее время выполнения сценария: (0 минут 33.356 секунды)

Launch binder
Launch JupyterLite

Download Jupyter notebook: plot_select_from_model_diabetes.ipynb

Download Python source code: plot_select_from_model_diabetes.py

Download zipped: plot_select_from_model_diabetes.zip

Связанные примеры

Pipeline ANOVA SVM

Основные моменты выпуска scikit-learn 0.24

Настройка порога функции принятия решений задним числом

Удаление признаков с помощью RFE с перекрестной проверкой

© 2007–2025 The scikit-learn developers
Licensed under the 3-clause BSD License.
https://scikit-learn.org/1.6/auto_examples/feature_selection/plot_select_from_model_diabetes.html

Spec-Zone.ru

Настройки Оффлайн Что нового Помощь О нас
Spec-Zone .ru
спецификации, руководства, описания, API