Browse documentation
2026-09-28AutoDataOpen in dashboard

Configuration Reference

Every setting a pipeline run accepts: its control, JSON and multipart names, default, range and effect, plus speed-mode knobs and the /api/v1/process layout.

Where settings live

Every way of starting a run (a file upload on the Pipeline tab, a connector run, a folder or cloud listener, an SFTP inbox, a scheduled run, a trigger) opens the same Configure window, so a setting means the same thing and has the same default everywhere. Its controls are grouped into three tiers: Common (most runs touch them), Advanced (for datasets that misbehave) and Experimental (rarely right; all off by default).

TransportSpellingUsed by
JSONsnake_case, for example text_modeStored pipeline configs of schedules, listeners and SFTP inboxes, trigger target_pipeline_config, POST /api/v1/process-connector. camelCase is also accepted; snake_case wins.
Multipart formcamelCase, for example textModeThe Pipeline tab upload, POST /process-data. Lists and objects are sent as JSON strings.
config fieldits own layoutPOST /api/v1/process and the Python SDK; see the last section.

Settings marked (Enterprise) are shown only to Enterprise accounts and are ignored by the server for every other account, without an error. Speed-preset knobs, Enterprise internals and module logic are outside the scope of this page; see Pipeline modules for what each stage does with these settings.

Targets, columns and size

Setting (control)JSON keyMultipart fieldDefaultOptions / rangeWhat it does
Target Columns (Y Variables)y_columnsyColumnsnone (required)column namesThe columns to predict; several targets of mixed task types are allowed. Replay runs take them from the parent session.
Excluded Columnsexcluded_columnsexcludedColumnsnonecolumn namesColumns that are not features. They stay in the outputs untouched: not encoded, imputed, scaled or removed.
Task Type Selectortask_typestaskTypesregression on the Pipeline tab, classification on automationc or r per target, in target orderClassification or regression. /api/v1/process treats every target as classification.
Target Rows / Output Sample Sizeoutput_number (also output_rows)outputNumber10000positive whole number, or emptySize of the final dataset. Empty skips synthesis. Zero, negative or unreadable values fall back to 10000.
Speed / accuracyspeed_modemodebalancedrapid, balanced, accurateSpeed preset; see Speed modes.

In JSON a bare mode is the run mode (full_pipeline, inference, retraining), never the speed preset.

Data cleaning (Anomaly Detection)

The Data cleaning section lists fourteen operations; the stage runs when at least one is ticked. In JSON the stage switch is enable_anomaly_detection (default false; multipart enableAnomalyDetection) and the operations are the anomaly_params object (multipart anomalyParams). Missing keys take the defaults below; a request with no anomaly_params uses the operations saved on your account.

Controlanomaly_params keyDefaultEffect
Currency conversioncurrency_conversiononCurrency amounts become USD numbers.
Percentage to decimalnumeric_range_conversionon"85%" becomes 0.85.
Date/time standardisationdate_time_conversiononOne date and time format per column.
Whitespace cleaningwhitespace_cleaningonExtra spaces removed.
Boolean normalisationboolean_normalizationonYes/No, 1/0, true/false unified.
Written numbers to digitstextual_numeric_conversionon"five" becomes 5.
Ordinal detectionordinal_detectiononOrdered categories such as low, medium, high are recognised.
Spelling error detectionspelling_error_detectiononMisspelt category values corrected.
Unit standardisationunit_standardizationoff"5kg" becomes 5 with the unit recorded.
Data type consistencydata_type_consistencyoffOne type per column.
Phone area codephone_area_code_extractionoffAdds an area-code column.
Email domainemail_domain_extractionoffAdds an e-mail domain column.
Emoji and symbolsspecial_character_analysisoffEmoji converted to text.
(API and automation only)temporal_feature_extractionoffTime features derived from dates.

Instructions, completion and run options

Setting (control)JSON keyMultipart fieldDefaultOptions / rangeWhat it does
Instructions for the pipelineprompt_textpromptText (also projectDescription)emptytextPlain-language context for the cleaning stages. Experimental.
Save artifacts for inference / retrainingretrainingretrainingfalsetrue / falseA label on the run only: every run keeps its replay artifacts, and any run whose artifacts are still stored can be a parent.
Session namesession_namesessionNamenonetextName shown in Runs. No Pipeline tab field; rename from Runs.
(no control)text_cleaningtextCleaningtruetrue / falseReserved; not available yet.
(no control)llm_enablednot readtruetrue / falsefalse: no stage contacts a language model for this run; completion and validation are skipped. JSON surfaces and advanced_params on /api/v1/process.
(no control)strict_llmnot readfalsetrue / falsetrue: a stage that could not reach its model fails the run instead of falling back.

Dataset completion & validation

JSON completion_params, multipart completionParams. Not read by /api/v1/process.

FieldDefaultOptionsEffect
completion_enabled, validation_enabledfalse, falsetrue / falseFill empty cells; check existing values. The stage runs when either is on.
completion_mode, validation_modefast, fastfast, accurateAccurate cross-checks more sources and is slower.
validation_actionoverwrite from the dashboard, mark when omittedoverwrite, markReplace wrong values, or flag them in a new column.
filling_columns, validation_columnsallcolumn namesLimit each half to these columns.
confidence_threshold0.30 – 1Values found with lower confidence are not written.
api_usage_per_cell0.01numberReserved; not available yet.

Stages and outputs (Output & Tool Selector)

One dialog chooses which stages run, which files they write, the export format, file names, the model-ready export and the PDF report. It travels as output_preferences (multipart outputPreferences). Automation payloads also accept enable_dtc, enable_mdh, enable_cds, enable_dsm, enable_dsg (default true) and enable_anomaly_detection (default false); a key in output_preferences wins. Unsent keys use your saved account defaults.

Selector optionoutput_preferences key(s)DefaultEffect
Data Completion & Verificationcompletion_use_tool, compleated_dataset_outputoffRun DCV; keep compleated_dataset.csv.
Anomaly Detectionanormaly_detection_use_tool, anormaly_fixed_outputoffDerived from Data cleaning; keep anormaly_fixed_output.
DTC, MDH, CDS, DSM, DSG<stage>_use_tool, <stage>_outputon, onRun the stage; keep its file.
(hidden) DORdor_use_tool, dor_outputonApplies to retraining only; outlier removal does not run in training runs.
Write every output asoutput_formatcsvcsv, parquet, feather, orc, xlsx, json, jsonl, ndjsonFormat of every data output. An output Excel cannot hold, or a failed write, falls back to CSV and the run says so. Inference/retraining results, compleated_dataset and feature_selection_output are always CSV.
File namesoutput_name_patternprefixedprefixed, plainprefixed: sales_20260922_013000_dsg_output.csv; plain: dsg_output.csv. Automated runs omit the input name.
Multimodal Outputmultimodal_outputoff in the dashboardtrue / falseWrite the merged text/media frame.
Pipeline Report PDFpipeline_report_outputontrue / falseWrite pipeline_report.pdf. Automation may send enable_pdf_report, which wins.
Model-Ready Datasetmodel_ready_outputontrue / falseWrite the all-numeric export.
Split layoutmodel_ready_splitnonenone, xy, train_test, xy_train_testSingle file, X and y, train and test, or all four.
Test sizemodel_ready_test_size0.20.05 – 0.5Fraction held out.
Test set from real rows onlymodel_ready_real_test_onlyontrue / falseOff splits all rows naively and adds a leakage warning.

Common settings

Setting (control)JSON keyMultipart fieldDefaultOptions / rangeWhat it does
Feature Selection, before imputationenable_feature_selectionenableFeatureSelectionofftrue / falseDrops the least useful columns right after DTC; targets and excluded columns are kept. See Feature selection.
Method / Keep top Nfeature_selection_config {method, top_k, threshold}featureSelectionConfigmutual_information, 20mutual_information, shap_importance, rfe, variance_threshold, correlation_pruning, pca (API also permutation_importance, leakage_guard, a list, auto); top N 1 – 1000How columns are scored and how many are kept. Variance and correlation use threshold (0.01 / 0.95) instead of top N. Omitting method in JSON means variance_threshold.
Text Tokenizationtext_modetextMode20 Drop Text, 1 Neural Tokenization, 2 TF-IDF, 3 AutoHow free-text columns become numbers. Mode 0 does not remove text: those columns are carried through unencoded.
Output Sample Sizeoutput_numberoutputNumber10000as Target RowsSame value as Target Rows.
Audio / Image column detection(dashboard only)(dashboard only)autoauto, manualDetect media columns, or pick them yourself.
Select Audio / Image Columnsaudio_columns, image_columnsaudioColumns, imageColumnsnonecolumn namesColumns holding media file names. Automation never detects media on its own.
Enable Preliminary Cleaningenable_preliminary_cleaningenablePreliminaryCleaningontrue / falseRemoves exact duplicate rows, turns tokens such as na and null into empty cells, drops nearly-empty, constant and ID-like columns. Target, excluded and media columns are protected.
Target model family (Enterprise)model_familiesmodelFamilieslinear + tree_gbdt (one pre-selected on automation forms)linear, tree_gbdt, mlp, sequence_neural; up to 4Prepares the data for a model family; with several, prepares once per family and recommends one.

Advanced settings

Setting (control)JSON keyMultipart fieldDefaultOptions / rangeWhat it does
MDH Modemdh_modemdhMode00 Imputation, 2 Imputation/dropping (Enterprise), 1 2D-RemovalImputation fills every gap and keeps every row; Imputation/dropping removes what is too empty to fill reliably and fills the rest; 2D-Removal removes rows and columns with too many gaps.
Audio Modeaudio_processing_modeaudioMode00 No Processing, 1 Basic Features, 2 Advanced Processing, 3 Mel-Spectrogram TokenizationEncoding of audio files. Out-of-range values mean 0.
Image Modeimage_processing_modeimageMode00 No Processing, 1 Basic Features, 2 Advanced Processing, 3 Neural EmbeddingsEncoding of image files.
Scaling Methodcds_scalercdsScalerautoauto, standard, minmax, robust, maxabs, yeojohnson, quantile_normal, quantile_uniform, log1p, noneauto picks the best-scoring scaler; anything else forces one.
Data Similarity Thresholdsimilarity_psimilarityP0.99greater than 0, less than 1DSM keeps one column of any pair more similar than this; lower removes more columns. DSM never splits rows. Values above 1 are read as percentages; invalid values fall back to 0.99.
Z-Score Limitzscore_limitzscoreLimit3greater than 0Extreme values are capped at this many standard deviations while the generator learns; lower gives fewer extreme synthetic values.
Synthesis Modedsg_modedsgModecopulacopula, ganGAN captures complex relationships but is much slower. A Pipeline tab upload sends only an explicit GAN, so leaving Copula there lets AutoData choose.
Class Balancing: Minimum Class Sharedsg_y_oversample_thresholddsgYOversampleThresholdempty (off)greater than 0, less than 1Raises every target class (or bin of a continuous target) to at least this share with synthetic rows; the output can exceed Target Rows. Out of range returns 400.
Class Balancing: binsdsg_y_num_binsdsgYNumBins102 – 100Equal-frequency bins for a continuous target.
Allow Stratified Downsamplingdsg_allow_downsampledsgAllowDownsampleofftrue / falseWhen Target Rows is below the processed rows: a class-preserving subsample (on) or the full processed data without synthesis (off).
Pivot (as-of) column (Enterprise)pivot_configpivotConfigoffcolumn; leakage_action warn / relative / drop; relative_deltas; auto_suggestWhen each row came into existence; dates are measured from it and later-recorded columns are handled as chosen.
Validation method (Enterprise)validation_policyvalidationPolicyoffsplit_strategy random / stratified / time_ordered / grouped; eval_protocol holdout / kfold / walk_forward; order_column, group_column, eval_folds (3), walk_forward_splits (3)How rows are divided when preparation choices are scored.
Dataset metadata (Enterprise)dataset_metadatadatasetMetadatanoneJSON dictionary (version 1) or CSV data dictionaryDeclared roles, types, ranges, order and missing-value sentinels override detection in every stage and are replayed at inference. Unreadable metadata returns 400.

Experimental settings

Setting (control)JSON keyMultipart fieldDefaultOptions / rangeWhat it does
Force Pure Synthetic Outputforce_syntheticforceSyntheticofftrue / falseAll rows synthetic. Honoured only on Pipeline tab uploads; other surfaces accept and ignore it.
Result Evaluatorenable_result_evaluatorenableResultEvaluatorofftrue / falseBenchmarks several models on the prepared data; writes result_evaluation_report.json.
Column selection, after scaling (Enterprise)selection_configselectionConfigoffmethods (leakage_guard, variance_threshold, correlation_pruning, mutual_information, permutation_importance, rfe, shap_importance), top_k, threshold, leakage_action warn / dropTarget-aware selection inside DSM.
Probe regularization (Enterprise)probe_configprobeConfigl2penalty l2 / l1 / elasticnet / none; C; l1_ratioRegularisation of the model that scores preparation choices, not of your model.
Column decisions (Enterprise)imputation_overrides, scaling_overridesimputationOverrides, scalingOverridesnoneper column, e.g. {"age": "Median Imputation"}Pin a column's imputation or scaling; a pin that cannot be honoured is refused and recorded.

Settings without a control

Setting (control)JSON keyMultipart fieldDefaultOptions / rangeWhat it does
Datetime expansiondatetime_modedatetimeModebasic (or the speed mode's value)none, minimal, basic, fullColumns per date: none drops dates, minimal keeps elapsed days, basic adds month, day of week, year and hour, full adds every component. Not read by /api/v1/process.
Nested flatteningenable_nested_flatteningenableNestedFlatteningontrue / falseExpand JSON objects in cells into columns; lists stay as text.
Columns to flattenflatten_columnsflattenColumnsabsent (detect)list, or []Only these columns; [] expands none.
Rare-class oversamplingdsg_rare_oversample_thresholddsgRareOversampleThresholdnone0 – 0.5Reserved; not available yet. Use Class Balancing.
Idempotency keyheader Idempotency-KeyidempotencyKeynonestring/process-data only: a repeated key returns the existing session.
Webhookswebhook_idswebhookIdsnonewebhook idsWebhooks notified when the run finishes or fails.
Automatic write-backauto_sinknot readnoneobjectAutomation and /api/v1/process: where results are written on completion.
Media filesn/aimageFiles, soundFilesnonefiles/process-data only: the files named by the media columns.
Schema / rangesn/aschema, rangesnoneJSONReserved; not available yet.

Speed modes

rapid finishes sooner with rougher choices, balanced (default) is the standard behaviour, accurate takes longer to fit your data more closely. The price does not change with the mode, only time. Modes do not apply to inference or retraining. Each mode sets sixteen knobs, all overridable:

KnobOverride keyRangeRapid / Balanced / AccurateWhat it changes
One-hot / free-text cardinality limitdtc_OHElimit2 – 10000040 / 100 / 250Which columns count as free text rather than categories.
Text-detection sample rowsdtc_sample_rows_detect_text10 – 100000100 / 200 / 400Rows read when detecting text columns.
Auto text-mode sample rowsdtc_sample_rows_text_mode10 – 100000100 / 300 / 600Rows read when Auto picks TF-IDF or neural.
Neural tokenizer vocabulary sizedtc_seq_max_features100 – 100000020000 / 50000 / 50000Vocabulary cap, text mode 1.
TF-IDF vocabulary sizedtc_tfidf_max_features100 – 10000008000 / 20000 / 20000TF-IDF feature cap, text mode 2.
Datetime expansiondtc_datetime_modenone, minimal, basic, fullminimal / basic / fullColumns per date.
Synthetic text methodtext_synth_methodstatistical, vaestatistical / statistical / vaeGenerator for synthetic text.
Anomaly detection: skip LLM checksanomaly_force_no_llmtrue / falsetrue / false / falseLLM-backed anomaly checks off.
Anomaly detection: sample scaleanomaly_sample_scale0.1 – 100.5 / 1.0 / 2.0Rows each anomaly check inspects.
Missing-data threshold trialsmdh_nooftrials1 – 1002 / 5 / 10Thresholds tried by 2D-Removal.
Skip expensive imputation strategiesimputer_skip_expensivetrue / falsetrue / false / falseSlowest imputation methods left out.
Force all imputation strategiesimputer_force_alltrue / falsefalse / false / trueEvery imputation method tried.
Scalers evaluatedcds_eval_top_n1 – 91 / 3 / 9Candidate scalers tried under auto.
Synthetic-generation fit budgetdsg_iess_budget_multiplier0.1 – 100.5 / 1.0 / 2.0How much data the generator fits on.
Synthetic image methodimage_synth_methodstatistical, vaestatistical / statistical / vaeGenerator for synthetic image features.
Synthetic audio methodaudio_synth_methodstatistical, vaestatistical / statistical / vaeGenerator for synthetic audio features.

Out-of-range values are clamped; unreadable ones are ignored.

Key (JSON / multipart)Meaning
mode_overrides / modeOverridesYour values for the knobs, keyed {"dtc.OHElimit": 250} or {"dtc_OHElimit": 250}. On /api/v1/process put the flat keys directly in advanced_params.
mode_decisions / modeDecisionsPer setting, user (keep yours, or the Balanced value if you set none) or mode (preset wins). Keys are groups text, datetime, missing_data, scaling, anomaly, synthetic, media, or single knobs such as dtc.OHElimit.
touched_fields / touchedFieldsWhich settings you changed yourself; carried with the run.

Without a decision, a value you set wins; otherwise the preset. The Pipeline tab asks in its Confirm processing run dialog (Keep my settings / Use <mode>); automation forms ask once in the Configure window and store the answer; /api/v1/process returns 409 mode_confirmation_required with mode_contested and policy_signature. Re-send with advanced_params.mode_ack set to the signature to keep your values, or with advanced_params.mode_decisions. The SDK raises ModeConfirmationRequired.

Names on POST /api/v1/process

This route takes a multipart file and a config JSON field:

{
  "target_columns": ["churn"],
  "output_rows": 10000,
  "session_name": "churn-weekly",
  "tools": {"anomaly": false, "dtc": true, "mdh": true, "dor": false,
            "cds": true, "dsm": true, "dsg": true,
            "feature_selection": false, "preliminary_cleaning": true,
            "nested_flattening": true, "result_evaluator": false},
  "feature_selection_config": {"method": "mutual_information", "top_k": 20},
  "anomaly_params": {"currency_conversion": true},
  "advanced_params": {
    "excluded_columns": ["id"], "text_mode": 2, "mdh_mode": 0,
    "cds_scaler": "auto", "similarity_p": 0.99, "zscore_limit": 3,
    "dsg_mode": "copula", "dsg_y_oversample_threshold": 0.3,
    "dsg_y_num_bins": 10, "dsg_allow_downsample": false,
    "audio_mode": 0, "audio_columns": [], "image_mode": 0, "image_columns": [],
    "output_format": "parquet", "output_name_pattern": "prefixed",
    "llm_enabled": true, "strict_llm": false,
    "mode": "balanced", "mode_decisions": {}, "touched_fields": []
  },
  "webhook_ids": [], "auto_sink": null
}

On this route advanced_params.mode is the speed mode, media modes are audio_mode / image_mode, every target is classification, the model-ready split comes from your saved defaults, run_dsg: false skips synthesis, and completion_params, datetime_mode and force_synthetic are not read. Enterprise keys go in advanced_params.