Pipeline Modules
Every module of the AutoData pipeline in run order: what it does to your data, when it runs, the files it writes and the settings it takes.
A training run passes your data through the modules below, in this order. Each entry says what the module does to the data, when it runs, what it writes and which settings reach it. Where each control sits is in the Configuration reference. The detailed logic of each module is Enterprise documentation, available in the dashboard.
Run order at a glance
| # | Module | Default | Main output |
|---|---|---|---|
| 1 | Data Completion & Verification (DCV) | off | compleated_dataset.csv |
| 2 | Load and nested JSON flattening | on | nested_flattening_report.csv |
| 3 | Preliminary cleaning | on | preliminary_cleaning_report.csv |
| 4 | Column profiling and automatic settings | always | auto_selection.json |
| 5 | Pivot and validation policy (Enterprise) | off | pivot and validation reports |
| 6 | Anomaly Detection (AD) | off | anormaly_fixed_output |
| 7 | Data Type Conversion (DTC) | on | dtc_output |
| 8 | Feature Selection (before imputation) | off | feature_selection_output.csv |
| 9 | Missing Data Handler (MDH) | on | mdh_output |
| 10 | Custom Data Scaling (CDS) | on | cds_output |
| 11 | Quality evaluation | always | quality_report.json |
| 12 | Dimensionality & Similarity Manager (DSM) | on | dsm_output |
| 13 | Data Synthetic Generator (DSG) and media synthesis | on | dsg_output, multimodal_output |
| 14 | Model-ready export | on | model_ready |
| 15 | Result Evaluator | off | result_evaluation_report.json |
| 16 | Pipeline report | on | pipeline_report.pdf |
Data files are named <input>_<date>_<time>_<step>.<ext> by default (for example sales_20260922_013000_dsg_output.parquet) or <step>.<ext> with plain naming. Progress is shown as step N of M: ten steps for a plain run, plus one each for completion, feature selection and the PDF report when on.
Data Completion & Verification (DCV)
Fills empty cells and checks existing values against outside sources before anything else. Validation either overwrites wrong values or marks them in an added column.
- Runs when: completion or validation is switched on. Skipped when LLM use is off for the run or the deployment.
- Writes:
compleated_dataset.csv(always CSV) anddata_gathering_report.json.
| Setting | JSON key · multipart field | Default | Options / range | Effect |
|---|---|---|---|---|
| Completion / Validation | completion_params.completion_enabled, validation_enabled | off | true / false | Turn each half on. |
| Modes | completion_mode, validation_mode | fast | fast, accurate | Accurate cross-checks more sources. |
| Validation Action | validation_action | overwrite (dashboard) | overwrite, mark | Replace or flag wrong values. |
| Columns | filling_columns, validation_columns | all | column names | Limit each half. |
| Minimum confidence | confidence_threshold | 0.3 | 0 – 1 | Lower-confidence values are not written. |
| (none) | api_usage_per_cell | 0.01 | number | Reserved; not available yet. |
Load and nested JSON flattening
Reads the file or connector result. Cells holding JSON objects become one column per leaf (details.guest.email); lists become columns as well, never rows, so one input row is always one output row. The flattening chosen at upload is replayed at inference.
- Runs when: always; flattening when on and nested cells exist.
- Writes:
nested_flattening_report.csv.
| Setting | JSON key · multipart field | Default | Options / range | Effect |
|---|---|---|---|---|
| Nested flattening | enable_nested_flattening · enableNestedFlattening | on | true / false | Off leaves JSON as text. |
| Columns to flatten | flatten_columns · flattenColumns | detect | list, or [] | Only these; [] none. |
Dataset metadata (Enterprise)
A declared data dictionary: per-column role, type, semantic type, valid range, ordinal order, unit, description, preferred imputation and scaler, and keep, plus dataset description, domain and missing-value sentinels. Declarations override detection in every stage and are replayed at inference. Out-of-range values become missing and are imputed; rows are never dropped and targets never modified. Key dataset_metadata (multipart datasetMetadata); writes metadata_report.json. Check a dictionary with POST /api/v1/dataset-metadata/validate.
Preliminary cleaning
Removes exact duplicate rows (duplicates of a rare class are kept), turns tokens such as na, n/a and null into empty cells, and drops nearly-empty, constant and ID-like columns. Target, excluded and media columns are protected; free text is not mistaken for an identifier.
- Runs when: on by default.
- Writes:
preliminary_cleaning_report.csv.
| Setting | JSON key · multipart field | Default | Options / range | Effect |
|---|---|---|---|---|
| Enable Preliminary Cleaning | enable_preliminary_cleaning · enablePreliminaryCleaning | on | true / false | Off keeps duplicates and every column. |
Column profiling and automatic settings
Works out what each column means. The Synthesis Mode, Data Similarity Threshold, Z-Score Limit and MDH Mode are chosen for the dataset when you do not supply them; a value you send always wins. On very large datasets, stages learn from a representative subset of rows while every row is still transformed and delivered.
- Runs when: always; no settings.
- Writes:
semantic_profile.json,auto_selection.json.
A Pipeline tab upload sends the Synthesis Mode only when you pick GAN, so leaving it on Gaussian Copula lets AutoData choose.
Pivot and validation policy (Enterprise)
The pivot (as-of) column records when each row came into existence: dates are measured from it, day-count features can be added, and columns recorded after it are kept with a warning, reduced to the time difference, or removed. The validation method decides how rows are divided when preparation choices are scored.
- Runs when: an Enterprise account sends
pivot_configorvalidation_policy. - Writes:
pivot_report.json,validation_report.json.
| Setting | JSON key · multipart field | Default | Options / range | Effect |
|---|---|---|---|---|
| Pivot column | pivot_config.column | none | a column | As-of time. |
| After-pivot columns | pivot_config.leakage_action | warn | warn, relative, drop | Keep and warn / time difference / remove. |
| Day-count features, suggestion | relative_deltas, auto_suggest | on with a pivot, off | true / false | Add features; propose a pivot. |
| Split strategy | validation_policy.split_strategy | random | random, stratified, time_ordered, grouped | How rows are divided. |
| Protocol | eval_protocol | holdout | holdout, kfold, walk_forward | Faster vs more reliable scoring. |
| Folds / windows | eval_folds, walk_forward_splits | 3, 3 | whole numbers | For kfold / walk_forward. |
Anomaly Detection (AD)
Repairs values whose format blocks numeric processing: currencies (to USD), percentages, dates and times, whitespace, booleans, written numbers, ordinal categories, misspellings, units and types, and can extract phone area codes, e-mail domains and emoji meanings. Identifier columns are left alone. With LLM use off, local checks still run.
- Runs when: at least one of the fourteen
anomaly_paramsoperations is on and the stage is enabled (enable_anomaly_detection). Eight operations are on by default; see the Configuration reference. - Writes:
anormaly_fixed_output,early_anomaly_report.json.
Data Type Conversion (DTC)
Turns every column into numbers: numeric strings become numbers, categories are encoded, dates become elapsed days plus calendar parts, text is vectorised per Text Tokenization, and image or audio files become feature columns. A column nothing can encode is carried through unchanged and reported. What DTC learns is stored so inference encodes new rows identically.
- Runs when: on by default.
- Writes:
dtc_output.
| Setting | JSON key · multipart field | Default | Options / range | Effect |
|---|---|---|---|---|
| Text Tokenization | text_mode · textMode | 2 | 0 Drop Text, 1 Neural, 2 TF-IDF, 3 Auto | How text becomes features. |
| Datetime expansion | datetime_mode · datetimeMode | basic | none, minimal, basic, full | Columns per date. |
| Audio / Image Mode | audio_processing_mode, image_processing_mode | 0 | 0 – 3 | Media encoding (No Processing, Basic, Advanced, Mel-Spectrogram Tokenization / Neural Embeddings). |
| Media columns | audio_columns, image_columns | none | column names | Columns of file names. |
Text mode 0 does not remove text. Although labelled Drop Text, free-text columns are carried through unencoded; the model-ready export moves them to model_ready_excluded_columns.csv.
Feature Selection (before imputation)
Drops the least useful columns right after DTC so later stages work on fewer columns. Targets, excluded columns and text and media features are kept; on any error the data passes through unchanged.
- Runs when: switched on (off by default).
- Writes:
feature_selection_output.csv,feature_selection_report.json.
Methods and parameters: see Feature selection.
Missing Data Handler (MDH)
Two modes. Imputation (0, default) fills every missing value and keeps every row, with the method chosen for your data and replayed at inference. Imputation/dropping (2, Enterprise) removes the rows and columns that are too empty to fill reliably, including rows with no target value, and fills the rest; every remaining column can still be filled at inference, and inference never removes a row. 2D-Removal (1) removes rows and columns with too many missing values. After this stage no output cell is empty.
- Runs when: on by default.
- Writes:
mdh_output.
| Setting | JSON key · multipart field | Default | Options / range | Effect |
|---|---|---|---|---|
| MDH Mode | mdh_mode · mdhMode | 0 | 0, 1 | Fill gaps or remove sparse rows and columns. |
| Column decisions (Enterprise) | imputation_overrides | none | per column | Pin a column's method. |
Custom Data Scaling (CDS)
Puts numeric features on a common scale. Targets, flags, excluded columns and text and media features are not scaled; row order matches the input.
- Runs when: on by default.
- Writes:
cds_output.
| Setting | JSON key · multipart field | Default | Options / range | Effect |
|---|---|---|---|---|
| Scaling Method | cds_scaler · cdsScaler | auto | auto, standard, minmax, robust, maxabs, yeojohnson, quantile_normal, quantile_uniform, log1p, none | auto picks the best-scoring scaler. |
| Column decisions (Enterprise) | scaling_overrides | none | per column | Pin a column's scaler. |
Quality evaluation
A quick usefulness score of the real prepared rows for predicting the target, plus the automatic choices and row accounting. Never changes data.
- Runs when: always.
- Writes:
quality_report.json.
Dimensionality & Similarity Manager (DSM)
Column selection: removes constant columns and, of any pair more similar than the threshold, keeps one. It never removes or splits rows. Targets, excluded columns and text and media features are kept. Enterprise accounts can add target-aware selection with selection_config.
- Runs when: on by default.
- Writes:
dsm_output(the real prepared rows: evaluate models on this file),column_drops.json.
| Setting | JSON key · multipart field | Default | Options / range | Effect |
|---|---|---|---|---|
| Data Similarity Threshold | similarity_p · similarityP | 0.99 | between 0 and 1 | Lower removes more columns. |
Data Synthetic Generator (DSG)
Delivers the dataset at the requested size.
- Runs when: on and Target Rows set.
- Writes:
dsg_output; with media,multimodal_output.
| Setting | JSON key · multipart field | Default | Options / range | Effect |
|---|---|---|---|---|
| Target Rows | output_number · outputNumber | 10000 | positive number, or empty | Final size; empty skips DSG. |
| Synthesis Mode | dsg_mode | copula | copula, gan | Generator type. |
| Z-Score Limit | zscore_limit | 3 | greater than 0 | Caps extreme values while the generator learns. |
| Class balancing | dsg_y_oversample_threshold, dsg_y_num_bins | off, 10 | 0 – 1 exclusive; 2 – 100 | Minimum share per class or bin; output can exceed Target Rows. |
| Allow Stratified Downsampling | dsg_allow_downsample | off | true / false | Class-preserving subsample when Target Rows is below the real rows. |
| Force Pure Synthetic Output | force_synthetic | off | true / false | All rows synthetic; Pipeline tab uploads only. |
| (none) | dsg_rare_oversample_threshold | none | 0 – 0.5 | Reserved; not available yet. |
| Situation | Delivered |
|---|---|
| Target Rows larger than real rows | All real rows first, then synthetic rows up to Target Rows, with a warning saying how many are real. |
| Target Rows smaller, downsampling off | The full real dataset, without synthesis, with a warning. |
| Target Rows smaller, downsampling on | A class-preserving subsample of the real rows. |
| Force Pure Synthetic Output | Only synthetic rows. |
Synthetic text, image and audio features are generated statistically by default and with a learned generator in Accurate mode. For tokenized text, multimodal_output currently writes merged token cells as bracketed strings; train on dsg_output or the model-ready export.
Model-ready export
An all-numeric file from the final dataset: token cells expanded to one column per position, booleans as numbers, no empty cells. Columns that cannot be numeric go to model_ready_excluded_columns.csv. Optional X/y and train/test splits; the test set is taken from real rows only by default.
- Runs when: Model-Ready Dataset is on (default).
- Writes:
model_readyor its split files,model_ready_manifest.json.
| Setting | JSON key · multipart field | Default | Options / range | Effect |
|---|---|---|---|---|
| Split layout | model_ready_split | none | none, xy, train_test, xy_train_test | File layout. |
| Test size | model_ready_test_size | 0.2 | 0.05 – 0.5 | Held-out fraction. |
| Real rows only | model_ready_real_test_only | on | true / false | Off adds a leakage warning. |
Result Evaluator
Trains several standard models on the real prepared data and records their metrics and the run time. Never changes data.
- Runs when: switched on (off by default).
- Writes:
result_evaluation_report.jsonand .csv.
| Setting | JSON key · multipart field | Default | Options / range | Effect |
|---|---|---|---|---|
| Result Evaluator | enable_result_evaluator | off | true / false | Run the benchmark. |
Pipeline report
A PDF of the stages, timings, shapes and sample rows.
- Runs when: Pipeline Report PDF is on (default).
- Writes:
pipeline_report.pdf.
Outside a normal training run
Data Outlier Removal (retraining only)
Outlier-row removal runs only when retraining a session (run_dor, default on; dor_eps and dor_min_samples default to the parent session's values). A training run accepts enable_dor but ignores it and writes no outlier file.
Drift
Training records each final feature's distribution; inference and retraining report a drift score per column and overall with a severity of stable, moderate or significant.
LLM use and data residency
Every run writes llm_audit.json naming the stages that contacted a model and the host. llm_enabled: false stops all model calls for a run; strict_llm: true fails the run instead of falling back. Both are read from automation payloads and /api/v1/process, not from a Pipeline tab upload.
Paired values
Right after loading, a column whose cells hold two numbers, such as blood pressure 140/90 or a resolution 1920x1080, is split into two numeric columns, for example Blood Pressure_systolic and Blood Pressure_diastolic. Dates, times and codes are left alone. A target that is split becomes two targets with the same task type. Declare a column's type as compound in the dataset metadata to choose the separator or part names, or as another type to keep it whole. Inference, retraining, streams and LM Readiness split new rows the same way.