Browse documentation
2026-09-28AutoDataOpen in dashboard

Feature Selection

Rank a finished run's columns in the Feature Eng tab, or drop weak columns during a run: methods, parameters and API.

Feature selection ranks the columns of a dataset by how useful they are for predicting your targets, and can drop the weakest. It appears in three places:

WhereWhat it doesChanges data?
The Feature Eng tabRanks a completed run's columns without re-running anything.No
Feature Selection, before imputation (Common tier)A pipeline stage between type conversion and missing-data handling that keeps only the selected columns.Yes
Column selection, after scaling (Enterprise, Experimental tier)Target-aware selection on the fully prepared columns.Yes

Targets and excluded columns are always kept, and if selection fails the data passes through unchanged.

Methods

MethodIdKindRanks or prunes by
Mutual Information (default)mutual_informationsupervisedInformation each column carries about the target.
SHAP (LightGBM)shap_importancesupervisedContribution to a fast tree model's predictions.
Recursive Feature EliminationrfesupervisedRepeatedly removing the weakest column.
Variance Thresholdvariance_thresholdunsupervisedRemoves near-constant columns (threshold default 0.01).
Correlation Pruningcorrelation_pruningunsupervisedRemoves one of each highly correlated pair (threshold default 0.95).
PCApcaunsupervisedReplaces the features with principal components.
Permutation importance (API)permutation_importancesupervisedHow much a model worsens when the column is shuffled.
Leakage guard (API)leakage_guardsupervisedFlags columns that predict the target almost perfectly.

Supervised methods work on a sample of up to 50,000 rows; with several targets, a column useful for any of them survives.

The Feature Eng tab

  1. Session: pick a completed run, or choose Enter a session ID manually and paste the full session ID (the shortened ID shown in Runs is rejected).
  2. Target column(s): click the column chips (the run's own targets are pre-selected). The columns are those of the run's type-converted output. If they cannot be loaded, type them comma-separated.
  3. Methods: toggle one or more; Top K 1 – 200, default 20.
  4. Analyze: the Ranking table shows rank, feature, an importance bar and a score from 0 to 1. With several methods, each method's scores are scaled to 0 – 1 and merged. Warnings appear above the table.

If you like the ranking, re-run the pipeline with Feature Selection switched on.

API

Both /api/v1/feature-selection/... and /api/v1/feature-engineering/... work (also without /v1), with a session cookie or an API key.

GET /api/v1/feature-selection/columns?session_id=<full id>
-> {"success": true, "columns": [...], "row_count": 5000}

POST /api/v1/feature-selection/recommend
{"session_id": "<full id>", "target_columns": ["churn"],
 "methods": ["mutual_information", "shap_importance"], "top_k": 15}
-> {"success": true, "ranking": [{"column": "tenure", "score": 1.0, "rank": 1}],
    "per_method": {...}, "leakage": {...}, "warnings": [...],
    "methods": [...], "top_k": 15, "row_count": 5000}
FieldTypeDefaultNotes
session_idstringrequiredFull ID of a completed run you own or that is shared with you.
target_columnslistrequiredMust be columns of the ranked output.
methodslist["mutual_information"]Any ids above; unknown ids are ignored with a warning.
top_kinteger20Length of the merged ranking.

Errors: 400 for missing fields or unknown targets, 403 for another user's run, 404 for an unknown session or a run with no stored output.

Feature Selection as a pipeline stage

Tick Feature Selection, before imputation in the Common tier and open Advanced parameters for Method and Keep top N (1 – 1000, default 20; Variance and Correlation prune on their threshold instead). The stage adds a progress step and writes feature_selection_output.csv and feature_selection_report.json.

SettingJSON keyMultipart fieldDefault
On / offenable_feature_selectionenableFeatureSelectionfalse
Methodfeature_selection_config.methodfeatureSelectionConfigmutual_information from the dashboard; variance_threshold when omitted
Keep top Nfeature_selection_config.top_k(same object)20
Threshold (API)feature_selection_config.threshold(same object)variance 0.01, correlation 0.95

In JSON, method may also be a list (each runs on the previous result) or auto (leakage guard, variance, correlation, then mutual information). On /api/v1/process switch the stage on with tools.feature_selection and pass feature_selection_config.

Column selection, after scaling (Enterprise)

Enterprise accounts can add target-aware selection on the prepared columns: leakage guard, variance, correlation, mutual information, permutation importance, recursive elimination and SHAP, run in the order chosen, with an optional top N, and suspected leakage flagged or removed. Key selection_config (multipart selectionConfig). For other accounts the Dimensionality & Similarity Manager removes only constant and near-duplicate columns.