Configuration Reference
Every setting a pipeline run accepts: its control, JSON and multipart names, default, range and effect, plus speed-mode knobs and the /api/v1/process layout.
Where settings live
Every way of starting a run (a file upload on the Pipeline tab, a connector run, a folder or cloud listener, an SFTP inbox, a scheduled run, a trigger) opens the same Configure window, so a setting means the same thing and has the same default everywhere. Its controls are grouped into three tiers: Common (most runs touch them), Advanced (for datasets that misbehave) and Experimental (rarely right; all off by default).
| Transport | Spelling | Used by |
|---|---|---|
| JSON | snake_case, for example text_mode | Stored pipeline configs of schedules, listeners and SFTP inboxes, trigger target_pipeline_config, POST /api/v1/process-connector. camelCase is also accepted; snake_case wins. |
| Multipart form | camelCase, for example textMode | The Pipeline tab upload, POST /process-data. Lists and objects are sent as JSON strings. |
config field | its own layout | POST /api/v1/process and the Python SDK; see the last section. |
Settings marked (Enterprise) are shown only to Enterprise accounts and are ignored by the server for every other account, without an error. Speed-preset knobs, Enterprise internals and module logic are outside the scope of this page; see Pipeline modules for what each stage does with these settings.
Targets, columns and size
| Setting (control) | JSON key | Multipart field | Default | Options / range | What it does |
|---|---|---|---|---|---|
| Target Columns (Y Variables) | y_columns | yColumns | none (required) | column names | The columns to predict; several targets of mixed task types are allowed. Replay runs take them from the parent session. |
| Excluded Columns | excluded_columns | excludedColumns | none | column names | Columns that are not features. They stay in the outputs untouched: not encoded, imputed, scaled or removed. |
| Task Type Selector | task_types | taskTypes | regression on the Pipeline tab, classification on automation | c or r per target, in target order | Classification or regression. /api/v1/process treats every target as classification. |
| Target Rows / Output Sample Size | output_number (also output_rows) | outputNumber | 10000 | positive whole number, or empty | Size of the final dataset. Empty skips synthesis. Zero, negative or unreadable values fall back to 10000. |
| Speed / accuracy | speed_mode | mode | balanced | rapid, balanced, accurate | Speed preset; see Speed modes. |
In JSON a bare mode is the run mode (full_pipeline, inference, retraining), never the speed preset.
Data cleaning (Anomaly Detection)
The Data cleaning section lists fourteen operations; the stage runs when at least one is ticked. In JSON the stage switch is enable_anomaly_detection (default false; multipart enableAnomalyDetection) and the operations are the anomaly_params object (multipart anomalyParams). Missing keys take the defaults below; a request with no anomaly_params uses the operations saved on your account.
| Control | anomaly_params key | Default | Effect |
|---|---|---|---|
| Currency conversion | currency_conversion | on | Currency amounts become USD numbers. |
| Percentage to decimal | numeric_range_conversion | on | "85%" becomes 0.85. |
| Date/time standardisation | date_time_conversion | on | One date and time format per column. |
| Whitespace cleaning | whitespace_cleaning | on | Extra spaces removed. |
| Boolean normalisation | boolean_normalization | on | Yes/No, 1/0, true/false unified. |
| Written numbers to digits | textual_numeric_conversion | on | "five" becomes 5. |
| Ordinal detection | ordinal_detection | on | Ordered categories such as low, medium, high are recognised. |
| Spelling error detection | spelling_error_detection | on | Misspelt category values corrected. |
| Unit standardisation | unit_standardization | off | "5kg" becomes 5 with the unit recorded. |
| Data type consistency | data_type_consistency | off | One type per column. |
| Phone area code | phone_area_code_extraction | off | Adds an area-code column. |
| Email domain | email_domain_extraction | off | Adds an e-mail domain column. |
| Emoji and symbols | special_character_analysis | off | Emoji converted to text. |
| (API and automation only) | temporal_feature_extraction | off | Time features derived from dates. |
Instructions, completion and run options
| Setting (control) | JSON key | Multipart field | Default | Options / range | What it does |
|---|---|---|---|---|---|
| Instructions for the pipeline | prompt_text | promptText (also projectDescription) | empty | text | Plain-language context for the cleaning stages. Experimental. |
| Save artifacts for inference / retraining | retraining | retraining | false | true / false | A label on the run only: every run keeps its replay artifacts, and any run whose artifacts are still stored can be a parent. |
| Session name | session_name | sessionName | none | text | Name shown in Runs. No Pipeline tab field; rename from Runs. |
| (no control) | text_cleaning | textCleaning | true | true / false | Reserved; not available yet. |
| (no control) | llm_enabled | not read | true | true / false | false: no stage contacts a language model for this run; completion and validation are skipped. JSON surfaces and advanced_params on /api/v1/process. |
| (no control) | strict_llm | not read | false | true / false | true: a stage that could not reach its model fails the run instead of falling back. |
Dataset completion & validation
JSON completion_params, multipart completionParams. Not read by /api/v1/process.
| Field | Default | Options | Effect |
|---|---|---|---|
completion_enabled, validation_enabled | false, false | true / false | Fill empty cells; check existing values. The stage runs when either is on. |
completion_mode, validation_mode | fast, fast | fast, accurate | Accurate cross-checks more sources and is slower. |
validation_action | overwrite from the dashboard, mark when omitted | overwrite, mark | Replace wrong values, or flag them in a new column. |
filling_columns, validation_columns | all | column names | Limit each half to these columns. |
confidence_threshold | 0.3 | 0 – 1 | Values found with lower confidence are not written. |
api_usage_per_cell | 0.01 | number | Reserved; not available yet. |
Stages and outputs (Output & Tool Selector)
One dialog chooses which stages run, which files they write, the export format, file names, the model-ready export and the PDF report. It travels as output_preferences (multipart outputPreferences). Automation payloads also accept enable_dtc, enable_mdh, enable_cds, enable_dsm, enable_dsg (default true) and enable_anomaly_detection (default false); a key in output_preferences wins. Unsent keys use your saved account defaults.
| Selector option | output_preferences key(s) | Default | Effect | |
|---|---|---|---|---|
| Data Completion & Verification | completion_use_tool, compleated_dataset_output | off | Run DCV; keep compleated_dataset.csv. | |
| Anomaly Detection | anormaly_detection_use_tool, anormaly_fixed_output | off | Derived from Data cleaning; keep anormaly_fixed_output. | |
| DTC, MDH, CDS, DSM, DSG | <stage>_use_tool, <stage>_output | on, on | Run the stage; keep its file. | |
| (hidden) DOR | dor_use_tool, dor_output | on | Applies to retraining only; outlier removal does not run in training runs. | |
| Write every output as | output_format | csv | csv, parquet, feather, orc, xlsx, json, jsonl, ndjson | Format of every data output. An output Excel cannot hold, or a failed write, falls back to CSV and the run says so. Inference/retraining results, compleated_dataset and feature_selection_output are always CSV. |
| File names | output_name_pattern | prefixed | prefixed, plain | prefixed: sales_20260922_013000_dsg_output.csv; plain: dsg_output.csv. Automated runs omit the input name. |
| Multimodal Output | multimodal_output | off in the dashboard | true / false | Write the merged text/media frame. |
| Pipeline Report PDF | pipeline_report_output | on | true / false | Write pipeline_report.pdf. Automation may send enable_pdf_report, which wins. |
| Model-Ready Dataset | model_ready_output | on | true / false | Write the all-numeric export. |
| Split layout | model_ready_split | none | none, xy, train_test, xy_train_test | Single file, X and y, train and test, or all four. |
| Test size | model_ready_test_size | 0.2 | 0.05 – 0.5 | Fraction held out. |
| Test set from real rows only | model_ready_real_test_only | on | true / false | Off splits all rows naively and adds a leakage warning. |
Common settings
| Setting (control) | JSON key | Multipart field | Default | Options / range | What it does |
|---|---|---|---|---|---|
| Feature Selection, before imputation | enable_feature_selection | enableFeatureSelection | off | true / false | Drops the least useful columns right after DTC; targets and excluded columns are kept. See Feature selection. |
| Method / Keep top N | feature_selection_config {method, top_k, threshold} | featureSelectionConfig | mutual_information, 20 | mutual_information, shap_importance, rfe, variance_threshold, correlation_pruning, pca (API also permutation_importance, leakage_guard, a list, auto); top N 1 – 1000 | How columns are scored and how many are kept. Variance and correlation use threshold (0.01 / 0.95) instead of top N. Omitting method in JSON means variance_threshold. |
| Text Tokenization | text_mode | textMode | 2 | 0 Drop Text, 1 Neural Tokenization, 2 TF-IDF, 3 Auto | How free-text columns become numbers. Mode 0 does not remove text: those columns are carried through unencoded. |
| Output Sample Size | output_number | outputNumber | 10000 | as Target Rows | Same value as Target Rows. |
| Audio / Image column detection | (dashboard only) | (dashboard only) | auto | auto, manual | Detect media columns, or pick them yourself. |
| Select Audio / Image Columns | audio_columns, image_columns | audioColumns, imageColumns | none | column names | Columns holding media file names. Automation never detects media on its own. |
| Enable Preliminary Cleaning | enable_preliminary_cleaning | enablePreliminaryCleaning | on | true / false | Removes exact duplicate rows, turns tokens such as na and null into empty cells, drops nearly-empty, constant and ID-like columns. Target, excluded and media columns are protected. |
| Target model family (Enterprise) | model_families | modelFamilies | linear + tree_gbdt (one pre-selected on automation forms) | linear, tree_gbdt, mlp, sequence_neural; up to 4 | Prepares the data for a model family; with several, prepares once per family and recommends one. |
Advanced settings
| Setting (control) | JSON key | Multipart field | Default | Options / range | What it does |
|---|---|---|---|---|---|
| MDH Mode | mdh_mode | mdhMode | 0 | 0 Imputation, 2 Imputation/dropping (Enterprise), 1 2D-Removal | Imputation fills every gap and keeps every row; Imputation/dropping removes what is too empty to fill reliably and fills the rest; 2D-Removal removes rows and columns with too many gaps. |
| Audio Mode | audio_processing_mode | audioMode | 0 | 0 No Processing, 1 Basic Features, 2 Advanced Processing, 3 Mel-Spectrogram Tokenization | Encoding of audio files. Out-of-range values mean 0. |
| Image Mode | image_processing_mode | imageMode | 0 | 0 No Processing, 1 Basic Features, 2 Advanced Processing, 3 Neural Embeddings | Encoding of image files. |
| Scaling Method | cds_scaler | cdsScaler | auto | auto, standard, minmax, robust, maxabs, yeojohnson, quantile_normal, quantile_uniform, log1p, none | auto picks the best-scoring scaler; anything else forces one. |
| Data Similarity Threshold | similarity_p | similarityP | 0.99 | greater than 0, less than 1 | DSM keeps one column of any pair more similar than this; lower removes more columns. DSM never splits rows. Values above 1 are read as percentages; invalid values fall back to 0.99. |
| Z-Score Limit | zscore_limit | zscoreLimit | 3 | greater than 0 | Extreme values are capped at this many standard deviations while the generator learns; lower gives fewer extreme synthetic values. |
| Synthesis Mode | dsg_mode | dsgMode | copula | copula, gan | GAN captures complex relationships but is much slower. A Pipeline tab upload sends only an explicit GAN, so leaving Copula there lets AutoData choose. |
| Class Balancing: Minimum Class Share | dsg_y_oversample_threshold | dsgYOversampleThreshold | empty (off) | greater than 0, less than 1 | Raises every target class (or bin of a continuous target) to at least this share with synthetic rows; the output can exceed Target Rows. Out of range returns 400. |
| Class Balancing: bins | dsg_y_num_bins | dsgYNumBins | 10 | 2 – 100 | Equal-frequency bins for a continuous target. |
| Allow Stratified Downsampling | dsg_allow_downsample | dsgAllowDownsample | off | true / false | When Target Rows is below the processed rows: a class-preserving subsample (on) or the full processed data without synthesis (off). |
| Pivot (as-of) column (Enterprise) | pivot_config | pivotConfig | off | column; leakage_action warn / relative / drop; relative_deltas; auto_suggest | When each row came into existence; dates are measured from it and later-recorded columns are handled as chosen. |
| Validation method (Enterprise) | validation_policy | validationPolicy | off | split_strategy random / stratified / time_ordered / grouped; eval_protocol holdout / kfold / walk_forward; order_column, group_column, eval_folds (3), walk_forward_splits (3) | How rows are divided when preparation choices are scored. |
| Dataset metadata (Enterprise) | dataset_metadata | datasetMetadata | none | JSON dictionary (version 1) or CSV data dictionary | Declared roles, types, ranges, order and missing-value sentinels override detection in every stage and are replayed at inference. Unreadable metadata returns 400. |
Experimental settings
| Setting (control) | JSON key | Multipart field | Default | Options / range | What it does |
|---|---|---|---|---|---|
| Force Pure Synthetic Output | force_synthetic | forceSynthetic | off | true / false | All rows synthetic. Honoured only on Pipeline tab uploads; other surfaces accept and ignore it. |
| Result Evaluator | enable_result_evaluator | enableResultEvaluator | off | true / false | Benchmarks several models on the prepared data; writes result_evaluation_report.json. |
| Column selection, after scaling (Enterprise) | selection_config | selectionConfig | off | methods (leakage_guard, variance_threshold, correlation_pruning, mutual_information, permutation_importance, rfe, shap_importance), top_k, threshold, leakage_action warn / drop | Target-aware selection inside DSM. |
| Probe regularization (Enterprise) | probe_config | probeConfig | l2 | penalty l2 / l1 / elasticnet / none; C; l1_ratio | Regularisation of the model that scores preparation choices, not of your model. |
| Column decisions (Enterprise) | imputation_overrides, scaling_overrides | imputationOverrides, scalingOverrides | none | per column, e.g. {"age": "Median Imputation"} | Pin a column's imputation or scaling; a pin that cannot be honoured is refused and recorded. |
Settings without a control
| Setting (control) | JSON key | Multipart field | Default | Options / range | What it does |
|---|---|---|---|---|---|
| Datetime expansion | datetime_mode | datetimeMode | basic (or the speed mode's value) | none, minimal, basic, full | Columns per date: none drops dates, minimal keeps elapsed days, basic adds month, day of week, year and hour, full adds every component. Not read by /api/v1/process. |
| Nested flattening | enable_nested_flattening | enableNestedFlattening | on | true / false | Expand JSON objects in cells into columns; lists stay as text. |
| Columns to flatten | flatten_columns | flattenColumns | absent (detect) | list, or [] | Only these columns; [] expands none. |
| Rare-class oversampling | dsg_rare_oversample_threshold | dsgRareOversampleThreshold | none | 0 – 0.5 | Reserved; not available yet. Use Class Balancing. |
| Idempotency key | header Idempotency-Key | idempotencyKey | none | string | /process-data only: a repeated key returns the existing session. |
| Webhooks | webhook_ids | webhookIds | none | webhook ids | Webhooks notified when the run finishes or fails. |
| Automatic write-back | auto_sink | not read | none | object | Automation and /api/v1/process: where results are written on completion. |
| Media files | n/a | imageFiles, soundFiles | none | files | /process-data only: the files named by the media columns. |
| Schema / ranges | n/a | schema, ranges | none | JSON | Reserved; not available yet. |
Speed modes
rapid finishes sooner with rougher choices, balanced (default) is the standard behaviour, accurate takes longer to fit your data more closely. The price does not change with the mode, only time. Modes do not apply to inference or retraining. Each mode sets sixteen knobs, all overridable:
| Knob | Override key | Range | Rapid / Balanced / Accurate | What it changes |
|---|---|---|---|---|
| One-hot / free-text cardinality limit | dtc_OHElimit | 2 – 100000 | 40 / 100 / 250 | Which columns count as free text rather than categories. |
| Text-detection sample rows | dtc_sample_rows_detect_text | 10 – 100000 | 100 / 200 / 400 | Rows read when detecting text columns. |
| Auto text-mode sample rows | dtc_sample_rows_text_mode | 10 – 100000 | 100 / 300 / 600 | Rows read when Auto picks TF-IDF or neural. |
| Neural tokenizer vocabulary size | dtc_seq_max_features | 100 – 1000000 | 20000 / 50000 / 50000 | Vocabulary cap, text mode 1. |
| TF-IDF vocabulary size | dtc_tfidf_max_features | 100 – 1000000 | 8000 / 20000 / 20000 | TF-IDF feature cap, text mode 2. |
| Datetime expansion | dtc_datetime_mode | none, minimal, basic, full | minimal / basic / full | Columns per date. |
| Synthetic text method | text_synth_method | statistical, vae | statistical / statistical / vae | Generator for synthetic text. |
| Anomaly detection: skip LLM checks | anomaly_force_no_llm | true / false | true / false / false | LLM-backed anomaly checks off. |
| Anomaly detection: sample scale | anomaly_sample_scale | 0.1 – 10 | 0.5 / 1.0 / 2.0 | Rows each anomaly check inspects. |
| Missing-data threshold trials | mdh_nooftrials | 1 – 100 | 2 / 5 / 10 | Thresholds tried by 2D-Removal. |
| Skip expensive imputation strategies | imputer_skip_expensive | true / false | true / false / false | Slowest imputation methods left out. |
| Force all imputation strategies | imputer_force_all | true / false | false / false / true | Every imputation method tried. |
| Scalers evaluated | cds_eval_top_n | 1 – 9 | 1 / 3 / 9 | Candidate scalers tried under auto. |
| Synthetic-generation fit budget | dsg_iess_budget_multiplier | 0.1 – 10 | 0.5 / 1.0 / 2.0 | How much data the generator fits on. |
| Synthetic image method | image_synth_method | statistical, vae | statistical / statistical / vae | Generator for synthetic image features. |
| Synthetic audio method | audio_synth_method | statistical, vae | statistical / statistical / vae | Generator for synthetic audio features. |
Out-of-range values are clamped; unreadable ones are ignored.
| Key (JSON / multipart) | Meaning |
|---|---|
mode_overrides / modeOverrides | Your values for the knobs, keyed {"dtc.OHElimit": 250} or {"dtc_OHElimit": 250}. On /api/v1/process put the flat keys directly in advanced_params. |
mode_decisions / modeDecisions | Per setting, user (keep yours, or the Balanced value if you set none) or mode (preset wins). Keys are groups text, datetime, missing_data, scaling, anomaly, synthetic, media, or single knobs such as dtc.OHElimit. |
touched_fields / touchedFields | Which settings you changed yourself; carried with the run. |
Without a decision, a value you set wins; otherwise the preset. The Pipeline tab asks in its Confirm processing run dialog (Keep my settings / Use <mode>); automation forms ask once in the Configure window and store the answer; /api/v1/process returns 409 mode_confirmation_required with mode_contested and policy_signature. Re-send with advanced_params.mode_ack set to the signature to keep your values, or with advanced_params.mode_decisions. The SDK raises ModeConfirmationRequired.
Names on POST /api/v1/process
This route takes a multipart file and a config JSON field:
{
"target_columns": ["churn"],
"output_rows": 10000,
"session_name": "churn-weekly",
"tools": {"anomaly": false, "dtc": true, "mdh": true, "dor": false,
"cds": true, "dsm": true, "dsg": true,
"feature_selection": false, "preliminary_cleaning": true,
"nested_flattening": true, "result_evaluator": false},
"feature_selection_config": {"method": "mutual_information", "top_k": 20},
"anomaly_params": {"currency_conversion": true},
"advanced_params": {
"excluded_columns": ["id"], "text_mode": 2, "mdh_mode": 0,
"cds_scaler": "auto", "similarity_p": 0.99, "zscore_limit": 3,
"dsg_mode": "copula", "dsg_y_oversample_threshold": 0.3,
"dsg_y_num_bins": 10, "dsg_allow_downsample": false,
"audio_mode": 0, "audio_columns": [], "image_mode": 0, "image_columns": [],
"output_format": "parquet", "output_name_pattern": "prefixed",
"llm_enabled": true, "strict_llm": false,
"mode": "balanced", "mode_decisions": {}, "touched_fields": []
},
"webhook_ids": [], "auto_sink": null
}
On this route advanced_params.mode is the speed mode, media modes are audio_mode / image_mode, every target is classification, the model-ready split comes from your saved defaults, run_dsg: false skips synthesis, and completion_params, datetime_mode and force_synthetic are not read. Enterprise keys go in advanced_params.