Browse documentation
2026-09-28AutoDataOpen in dashboard

AutoData Overview

What AutoData does, the pipeline stages in order, what each run delivers and where to go next.

From raw data to model-ready data

In machine learning and analytics, up to 60% of project time is spent cleaning, preparing and structuring data. AutoData automates that work. It reads a file or a live source, turns every column into clean, numeric, scaled features, fills gaps, removes redundant columns and tops the dataset up with realistic synthetic rows. A finished run can then score or refresh new data with exactly the same transforms, on demand, on a schedule or as data arrives.

The product at a glance

AreaWhat you can do
IngestUpload CSV, Excel, Parquet, JSON, JSON Lines/NDJSON, Feather or ORC, or read from 38 source types: SQL databases, warehouses (Snowflake, BigQuery, Databricks, Redshift, Synapse, Fabric), document stores, cloud storage and drives, SaaS apps, industrial sources (OPC UA, PI, MQTT, InfluxDB) and streams (Kafka, Kinesis). Nested JSON is flattened automatically.
PrepareA configurable pipeline runs on dedicated workers with live progress; see below.
DeliverDownload any stage output in CSV, Parquet, Feather, ORC, Excel or JSON, a model-ready file with optional train/test split, and a PDF report; or write results back to a destination.
ReuseInference scores new data with a finished run's transforms; retraining refreshes it; streaming prepares records continuously.
AutomateScheduled runs, folder and cloud listeners, SFTP inboxes, triggers and incremental sync, all with the same configuration window.
MonitorRun history and reports, quality alerts including drift, webhooks, e-mail notifications and column lineage.
IntegrateA REST API and a Python SDK with API keys, per-key limits and permissions.

The AutoData pipeline

Stages run in this order; every stage except loading can be switched off. Pipeline modules documents each in full.

Step 1: Data Completion & Verification (DCV), optional

Fills empty cells and checks existing values against outside sources, then overwrites wrong values or marks them in a new column. Off by default.

Step 2: Preliminary cleaning

Removes exact duplicate rows, normalises missing-value tokens and drops nearly-empty, constant and ID-like columns, protecting targets.

Step 3: Anomaly Detection (AD)

Repairs formats that block numeric processing: currencies to USD, percentages to decimals, unified dates, whitespace, booleans, written numbers, ordinal categories, misspellings, units and types, with optional extraction of phone area codes, e-mail domains and emoji. Fourteen operations, eight on by default.

Step 4: Data Type Conversion (DTC)

Encodes categories, turns dates into time-based features, vectorises text (TF-IDF by default) and encodes images and audio. Columns nothing can encode are carried through, never silently dropped.

Step 5: Feature Selection, optional

Keeps the most useful columns before imputation.

Step 6: Missing Data Handler (MDH)

Imputation (default) fills every missing value and keeps every row, with the method chosen for your data; Imputation/dropping (Enterprise) removes the rows and columns that are too empty to fill reliably and fills the rest; 2D-Removal instead removes rows and columns with too many gaps.

Step 7: Custom Data Scaling (CDS)

Puts features on a common scale, with the scaler chosen for your data or forced by you.

Step 8: Dimensionality & Similarity Manager (DSM)

Removes constant columns and near-duplicate columns, for a leaner feature set. It does not split rows.

Step 9: Data Synthetic Generator (DSG)

Delivers the requested number of rows: all real rows plus synthetic rows that follow the real data's distribution, with optional class balancing.

What you receive

  • Final dataset (dsg_output): numeric, scaled and at the size you asked for. Real rows come first; to evaluate a model on real data use dsm_output or the model-ready test split.
  • Model-ready export: an all-numeric file, optionally split into X/y and train/test, with the test set taken from real rows.
  • Stage outputs: one file per stage you keep, for an auditable lineage of transformations.
  • Pipeline report PDF: the steps, timings and sample rows.
  • Reports: a quality report and, when enabled, a model benchmark.

Accounts

Trial accounts get three processing runs with uploads up to 1 MB. Paid and Unrestricted accounts have full access. Enterprise adds per-column lineage, preparation for a chosen model family with a recommendation, and extra configuration (pivot column, validation method, column selection, dataset metadata), billed at a higher rate.

Where to go next