Feature Selection
Rank a finished run's columns in the Feature Eng tab, or drop weak columns during a run: methods, parameters and API.
Feature selection ranks the columns of a dataset by how useful they are for predicting your targets, and can drop the weakest. It appears in three places:
| Where | What it does | Changes data? |
|---|---|---|
| The Feature Eng tab | Ranks a completed run's columns without re-running anything. | No |
| Feature Selection, before imputation (Common tier) | A pipeline stage between type conversion and missing-data handling that keeps only the selected columns. | Yes |
| Column selection, after scaling (Enterprise, Experimental tier) | Target-aware selection on the fully prepared columns. | Yes |
Targets and excluded columns are always kept, and if selection fails the data passes through unchanged.
Methods
| Method | Id | Kind | Ranks or prunes by |
|---|---|---|---|
| Mutual Information (default) | mutual_information | supervised | Information each column carries about the target. |
| SHAP (LightGBM) | shap_importance | supervised | Contribution to a fast tree model's predictions. |
| Recursive Feature Elimination | rfe | supervised | Repeatedly removing the weakest column. |
| Variance Threshold | variance_threshold | unsupervised | Removes near-constant columns (threshold default 0.01). |
| Correlation Pruning | correlation_pruning | unsupervised | Removes one of each highly correlated pair (threshold default 0.95). |
| PCA | pca | unsupervised | Replaces the features with principal components. |
| Permutation importance (API) | permutation_importance | supervised | How much a model worsens when the column is shuffled. |
| Leakage guard (API) | leakage_guard | supervised | Flags columns that predict the target almost perfectly. |
Supervised methods work on a sample of up to 50,000 rows; with several targets, a column useful for any of them survives.
The Feature Eng tab
- Session: pick a completed run, or choose Enter a session ID manually and paste the full session ID (the shortened ID shown in Runs is rejected).
- Target column(s): click the column chips (the run's own targets are pre-selected). The columns are those of the run's type-converted output. If they cannot be loaded, type them comma-separated.
- Methods: toggle one or more; Top K 1 – 200, default 20.
- Analyze: the Ranking table shows rank, feature, an importance bar and a score from 0 to 1. With several methods, each method's scores are scaled to 0 – 1 and merged. Warnings appear above the table.
If you like the ranking, re-run the pipeline with Feature Selection switched on.
API
Both /api/v1/feature-selection/... and /api/v1/feature-engineering/... work (also without /v1), with a session cookie or an API key.
GET /api/v1/feature-selection/columns?session_id=<full id>
-> {"success": true, "columns": [...], "row_count": 5000}
POST /api/v1/feature-selection/recommend
{"session_id": "<full id>", "target_columns": ["churn"],
"methods": ["mutual_information", "shap_importance"], "top_k": 15}
-> {"success": true, "ranking": [{"column": "tenure", "score": 1.0, "rank": 1}],
"per_method": {...}, "leakage": {...}, "warnings": [...],
"methods": [...], "top_k": 15, "row_count": 5000}
| Field | Type | Default | Notes |
|---|---|---|---|
session_id | string | required | Full ID of a completed run you own or that is shared with you. |
target_columns | list | required | Must be columns of the ranked output. |
methods | list | ["mutual_information"] | Any ids above; unknown ids are ignored with a warning. |
top_k | integer | 20 | Length of the merged ranking. |
Errors: 400 for missing fields or unknown targets, 403 for another user's run, 404 for an unknown session or a run with no stored output.
Feature Selection as a pipeline stage
Tick Feature Selection, before imputation in the Common tier and open Advanced parameters for Method and Keep top N (1 – 1000, default 20; Variance and Correlation prune on their threshold instead). The stage adds a progress step and writes feature_selection_output.csv and feature_selection_report.json.
| Setting | JSON key | Multipart field | Default |
|---|---|---|---|
| On / off | enable_feature_selection | enableFeatureSelection | false |
| Method | feature_selection_config.method | featureSelectionConfig | mutual_information from the dashboard; variance_threshold when omitted |
| Keep top N | feature_selection_config.top_k | (same object) | 20 |
| Threshold (API) | feature_selection_config.threshold | (same object) | variance 0.01, correlation 0.95 |
In JSON, method may also be a list (each runs on the previous result) or auto (leakage guard, variance, correlation, then mutual information). On /api/v1/process switch the stage on with tools.feature_selection and pass feature_selection_config.
Column selection, after scaling (Enterprise)
Enterprise accounts can add target-aware selection on the prepared columns: leakage guard, variance, correlation, mutual information, permutation importance, recursive elimination and SHAP, run in the order chosen, with an optional top N, and suspected leakage flagged or removed. Key selection_config (multipart selectionConfig). For other accounts the Dimensionality & Similarity Manager removes only constant and near-duplicate columns.