Browse documentation
2026-09-30AutoDataOpen in dashboard

Pipeline Modules

Every module of the AutoData pipeline in run order: what it does to your data, when it runs, the files it writes and the settings it takes.

A training run passes your data through the modules below, in this order. Each entry says what the module does to the data, when it runs, what it writes and which settings reach it. Where each control sits is in the Configuration reference. The detailed logic of each module is Enterprise documentation, available in the dashboard.

Run order at a glance

#ModuleDefaultMain output
1Data Completion & Verification (DCV)offcompleated_dataset.csv
2Load and nested JSON flatteningonnested_flattening_report.csv
3Preliminary cleaningonpreliminary_cleaning_report.csv
4Column profiling and automatic settingsalwaysauto_selection.json
5Pivot and validation policy (Enterprise)offpivot and validation reports
6Anomaly Detection (AD)offanormaly_fixed_output
7Data Type Conversion (DTC)ondtc_output
8Feature Selection (before imputation)offfeature_selection_output.csv
9Missing Data Handler (MDH)onmdh_output
10Custom Data Scaling (CDS)oncds_output
11Quality evaluationalwaysquality_report.json
12Dimensionality & Similarity Manager (DSM)ondsm_output
13Data Synthetic Generator (DSG) and media synthesisondsg_output, multimodal_output
14Model-ready exportonmodel_ready
15Result Evaluatoroffresult_evaluation_report.json
16Pipeline reportonpipeline_report.pdf

Data files are named <input>_<date>_<time>_<step>.<ext> by default (for example sales_20260922_013000_dsg_output.parquet) or <step>.<ext> with plain naming. Progress is shown as step N of M: ten steps for a plain run, plus one each for completion, feature selection and the PDF report when on.

Data Completion & Verification (DCV)

Fills empty cells and checks existing values against outside sources before anything else. Validation either overwrites wrong values or marks them in an added column.

  • Runs when: completion or validation is switched on. Skipped when LLM use is off for the run or the deployment.
  • Writes: compleated_dataset.csv (always CSV) and data_gathering_report.json.
SettingJSON key · multipart fieldDefaultOptions / rangeEffect
Completion / Validationcompletion_params.completion_enabled, validation_enabledofftrue / falseTurn each half on.
Modescompletion_mode, validation_modefastfast, accurateAccurate cross-checks more sources.
Validation Actionvalidation_actionoverwrite (dashboard)overwrite, markReplace or flag wrong values.
Columnsfilling_columns, validation_columnsallcolumn namesLimit each half.
Minimum confidenceconfidence_threshold0.30 – 1Lower-confidence values are not written.
(none)api_usage_per_cell0.01numberReserved; not available yet.

Load and nested JSON flattening

Reads the file or connector result. Cells holding JSON objects become one column per leaf (details.guest.email); lists become columns as well, never rows, so one input row is always one output row. The flattening chosen at upload is replayed at inference.

  • Runs when: always; flattening when on and nested cells exist.
  • Writes: nested_flattening_report.csv.
SettingJSON key · multipart fieldDefaultOptions / rangeEffect
Nested flatteningenable_nested_flattening · enableNestedFlatteningontrue / falseOff leaves JSON as text.
Columns to flattenflatten_columns · flattenColumnsdetectlist, or []Only these; [] none.

Dataset metadata (Enterprise)

A declared data dictionary: per-column role, type, semantic type, valid range, ordinal order, unit, description, preferred imputation and scaler, and keep, plus dataset description, domain and missing-value sentinels. Declarations override detection in every stage and are replayed at inference. Out-of-range values become missing and are imputed; rows are never dropped and targets never modified. Key dataset_metadata (multipart datasetMetadata); writes metadata_report.json. Check a dictionary with POST /api/v1/dataset-metadata/validate.

Preliminary cleaning

Removes exact duplicate rows (duplicates of a rare class are kept), turns tokens such as na, n/a and null into empty cells, and drops nearly-empty, constant and ID-like columns. Target, excluded and media columns are protected; free text is not mistaken for an identifier.

  • Runs when: on by default.
  • Writes: preliminary_cleaning_report.csv.
SettingJSON key · multipart fieldDefaultOptions / rangeEffect
Enable Preliminary Cleaningenable_preliminary_cleaning · enablePreliminaryCleaningontrue / falseOff keeps duplicates and every column.

Column profiling and automatic settings

Works out what each column means. The Synthesis Mode, Data Similarity Threshold, Z-Score Limit and MDH Mode are chosen for the dataset when you do not supply them; a value you send always wins. On very large datasets, stages learn from a representative subset of rows while every row is still transformed and delivered.

  • Runs when: always; no settings.
  • Writes: semantic_profile.json, auto_selection.json.

A Pipeline tab upload sends the Synthesis Mode only when you pick GAN, so leaving it on Gaussian Copula lets AutoData choose.

Pivot and validation policy (Enterprise)

The pivot (as-of) column records when each row came into existence: dates are measured from it, day-count features can be added, and columns recorded after it are kept with a warning, reduced to the time difference, or removed. The validation method decides how rows are divided when preparation choices are scored.

  • Runs when: an Enterprise account sends pivot_config or validation_policy.
  • Writes: pivot_report.json, validation_report.json.
SettingJSON key · multipart fieldDefaultOptions / rangeEffect
Pivot columnpivot_config.columnnonea columnAs-of time.
After-pivot columnspivot_config.leakage_actionwarnwarn, relative, dropKeep and warn / time difference / remove.
Day-count features, suggestionrelative_deltas, auto_suggeston with a pivot, offtrue / falseAdd features; propose a pivot.
Split strategyvalidation_policy.split_strategyrandomrandom, stratified, time_ordered, groupedHow rows are divided.
Protocoleval_protocolholdoutholdout, kfold, walk_forwardFaster vs more reliable scoring.
Folds / windowseval_folds, walk_forward_splits3, 3whole numbersFor kfold / walk_forward.

Anomaly Detection (AD)

Repairs values whose format blocks numeric processing: currencies (to USD), percentages, dates and times, whitespace, booleans, written numbers, ordinal categories, misspellings, units and types, and can extract phone area codes, e-mail domains and emoji meanings. Identifier columns are left alone. With LLM use off, local checks still run.

  • Runs when: at least one of the fourteen anomaly_params operations is on and the stage is enabled (enable_anomaly_detection). Eight operations are on by default; see the Configuration reference.
  • Writes: anormaly_fixed_output, early_anomaly_report.json.

Data Type Conversion (DTC)

Turns every column into numbers: numeric strings become numbers, categories are encoded, dates become elapsed days plus calendar parts, text is vectorised per Text Tokenization, and image or audio files become feature columns. A column nothing can encode is carried through unchanged and reported. What DTC learns is stored so inference encodes new rows identically.

  • Runs when: on by default.
  • Writes: dtc_output.
SettingJSON key · multipart fieldDefaultOptions / rangeEffect
Text Tokenizationtext_mode · textMode20 Drop Text, 1 Neural, 2 TF-IDF, 3 AutoHow text becomes features.
Datetime expansiondatetime_mode · datetimeModebasicnone, minimal, basic, fullColumns per date.
Audio / Image Modeaudio_processing_mode, image_processing_mode00 – 3Media encoding (No Processing, Basic, Advanced, Mel-Spectrogram Tokenization / Neural Embeddings).
Media columnsaudio_columns, image_columnsnonecolumn namesColumns of file names.

Text mode 0 does not remove text. Although labelled Drop Text, free-text columns are carried through unencoded; the model-ready export moves them to model_ready_excluded_columns.csv.

Feature Selection (before imputation)

Drops the least useful columns right after DTC so later stages work on fewer columns. Targets, excluded columns and text and media features are kept; on any error the data passes through unchanged.

  • Runs when: switched on (off by default).
  • Writes: feature_selection_output.csv, feature_selection_report.json.

Methods and parameters: see Feature selection.

Missing Data Handler (MDH)

Two modes. Imputation (0, default) fills every missing value and keeps every row, with the method chosen for your data and replayed at inference. Imputation/dropping (2, Enterprise) removes the rows and columns that are too empty to fill reliably, including rows with no target value, and fills the rest; every remaining column can still be filled at inference, and inference never removes a row. 2D-Removal (1) removes rows and columns with too many missing values. After this stage no output cell is empty.

  • Runs when: on by default.
  • Writes: mdh_output.
SettingJSON key · multipart fieldDefaultOptions / rangeEffect
MDH Modemdh_mode · mdhMode00, 1Fill gaps or remove sparse rows and columns.
Column decisions (Enterprise)imputation_overridesnoneper columnPin a column's method.

Custom Data Scaling (CDS)

Puts numeric features on a common scale. Targets, flags, excluded columns and text and media features are not scaled; row order matches the input.

  • Runs when: on by default.
  • Writes: cds_output.
SettingJSON key · multipart fieldDefaultOptions / rangeEffect
Scaling Methodcds_scaler · cdsScalerautoauto, standard, minmax, robust, maxabs, yeojohnson, quantile_normal, quantile_uniform, log1p, noneauto picks the best-scoring scaler.
Column decisions (Enterprise)scaling_overridesnoneper columnPin a column's scaler.

Quality evaluation

A quick usefulness score of the real prepared rows for predicting the target, plus the automatic choices and row accounting. Never changes data.

  • Runs when: always.
  • Writes: quality_report.json.

Dimensionality & Similarity Manager (DSM)

Column selection: removes constant columns and, of any pair more similar than the threshold, keeps one. It never removes or splits rows. Targets, excluded columns and text and media features are kept. Enterprise accounts can add target-aware selection with selection_config.

  • Runs when: on by default.
  • Writes: dsm_output (the real prepared rows: evaluate models on this file), column_drops.json.
SettingJSON key · multipart fieldDefaultOptions / rangeEffect
Data Similarity Thresholdsimilarity_p · similarityP0.99between 0 and 1Lower removes more columns.

Data Synthetic Generator (DSG)

Delivers the dataset at the requested size.

  • Runs when: on and Target Rows set.
  • Writes: dsg_output; with media, multimodal_output.
SettingJSON key · multipart fieldDefaultOptions / rangeEffect
Target Rowsoutput_number · outputNumber10000positive number, or emptyFinal size; empty skips DSG.
Synthesis Modedsg_modecopulacopula, ganGenerator type.
Z-Score Limitzscore_limit3greater than 0Caps extreme values while the generator learns.
Class balancingdsg_y_oversample_threshold, dsg_y_num_binsoff, 100 – 1 exclusive; 2 – 100Minimum share per class or bin; output can exceed Target Rows.
Allow Stratified Downsamplingdsg_allow_downsampleofftrue / falseClass-preserving subsample when Target Rows is below the real rows.
Force Pure Synthetic Outputforce_syntheticofftrue / falseAll rows synthetic; Pipeline tab uploads only.
(none)dsg_rare_oversample_thresholdnone0 – 0.5Reserved; not available yet.
SituationDelivered
Target Rows larger than real rowsAll real rows first, then synthetic rows up to Target Rows, with a warning saying how many are real.
Target Rows smaller, downsampling offThe full real dataset, without synthesis, with a warning.
Target Rows smaller, downsampling onA class-preserving subsample of the real rows.
Force Pure Synthetic OutputOnly synthetic rows.

Synthetic text, image and audio features are generated statistically by default and with a learned generator in Accurate mode. For tokenized text, multimodal_output currently writes merged token cells as bracketed strings; train on dsg_output or the model-ready export.

Model-ready export

An all-numeric file from the final dataset: token cells expanded to one column per position, booleans as numbers, no empty cells. Columns that cannot be numeric go to model_ready_excluded_columns.csv. Optional X/y and train/test splits; the test set is taken from real rows only by default.

  • Runs when: Model-Ready Dataset is on (default).
  • Writes: model_ready or its split files, model_ready_manifest.json.
SettingJSON key · multipart fieldDefaultOptions / rangeEffect
Split layoutmodel_ready_splitnonenone, xy, train_test, xy_train_testFile layout.
Test sizemodel_ready_test_size0.20.05 – 0.5Held-out fraction.
Real rows onlymodel_ready_real_test_onlyontrue / falseOff adds a leakage warning.

Result Evaluator

Trains several standard models on the real prepared data and records their metrics and the run time. Never changes data.

  • Runs when: switched on (off by default).
  • Writes: result_evaluation_report.json and .csv.
SettingJSON key · multipart fieldDefaultOptions / rangeEffect
Result Evaluatorenable_result_evaluatorofftrue / falseRun the benchmark.

Pipeline report

A PDF of the stages, timings, shapes and sample rows.

  • Runs when: Pipeline Report PDF is on (default).
  • Writes: pipeline_report.pdf.

Outside a normal training run

Data Outlier Removal (retraining only)

Outlier-row removal runs only when retraining a session (run_dor, default on; dor_eps and dor_min_samples default to the parent session's values). A training run accepts enable_dor but ignores it and writes no outlier file.

Drift

Training records each final feature's distribution; inference and retraining report a drift score per column and overall with a severity of stable, moderate or significant.

LLM use and data residency

Every run writes llm_audit.json naming the stages that contacted a model and the host. llm_enabled: false stops all model calls for a run; strict_llm: true fails the run instead of falling back. Both are read from automation payloads and /api/v1/process, not from a Pipeline tab upload.

Paired values

Right after loading, a column whose cells hold two numbers, such as blood pressure 140/90 or a resolution 1920x1080, is split into two numeric columns, for example Blood Pressure_systolic and Blood Pressure_diastolic. Dates, times and codes are left alone. A target that is split becomes two targets with the same task type. Declare a column's type as compound in the dataset metadata to choose the separator or part names, or as another type to keep it whole. Inference, retraining, streams and LM Readiness split new rows the same way.