Browse documentation
2026-09-28AutoDataOpen in dashboard

Dashboard User Guide

A complete walkthrough of the AutoData dashboard: the Home screen, uploading or connecting data, configuring, running and downloading results.

Overview

The AutoData dashboard is where you prepare datasets for machine learning, automate that preparation and watch every run. After signing in you land on the Home screen (/dashboard); every tab has its own address, /dashboard/<tab>.

Home

PartWhat it shows
GreetingYour first name, a New pipeline run button and a Documentation button.
Run in progressWhile a run is processing: its step (step N of M) and message. Click to return to it.
Get startedA four-step checklist (upload a dataset, run the pipeline, download your results, create an API key), Try with demo CSV and Dismiss.
StatsTotal runs, Completed, Failed and Rows processed.
DestinationsEvery tab, grouped as Build (Pipeline, Inference, Feature Eng), Automate (Connections, Listeners, Triggers, Sync, Scheduled, Streaming), Monitor (Runs, Quality Alerts, Lineage) and Manage (API Keys, Collaboration, Documentation). Profile is in the header menu.
Recent runsYour five latest runs with status; View all opens Runs.

Step 1: Source

Open Pipeline. The Source card toggles between File Upload and Connector.

File upload

Click Select file or drag a file onto the drop zone. Accepted: CSV, Excel (XLSX, XLS), Parquet, JSON, JSON Lines (JSONL, NDJSON), Feather and ORC. A progress bar shows speed and time remaining.

Once read, the page shows the column count, Dataset Information with the exact price of the file with your settings (every row is measured; a larger file shows Measuring your data… P% and the start button stays disabled until the price is ready), any dataset warning, and a Dataset preview: a profile of each column (kind, null %, unique values, flags), sample rows and a dry-run preview. Cells holding JSON objects are expanded into columns; the Which columns were treated as nested panel lets you keep a column as text.

Connector

Choose a source type, pick a saved credential (managed in the Connections tab) or enter the fields, then test the connection, discover tables and pick one or write a query. See Connector data sources.

Upload limits

AccountData fileImage / audio filesRuns
Trial1 MB1 MB each per run3 free runs (a run counts once it completes)
All other accounts1 GB in the dashboard (the server accepts requests up to 5 GB)1 GB eachCharged to your balance

Step 2: Configure

Targets

Under Available columns, click a column chip to make it a target: the first click makes it a regression target, the second a classification target, the third clears it. Show All lists every column.

Run configuration

The Run configuration card holds the Processing Mode (rapid, balanced or accurate), a summary (mode, targets, stages, rows), Templates (apply, save, rename or delete a saved configuration) and Pipeline Configuration → Configure, which opens the configuration window.

The configuration window

The same window configures every kind of run. It contains:

  • Target Columns and Excluded Columns. Excluded columns stay in the outputs untouched: they are not encoded, imputed, scaled or removed.
  • Target Rows: the size of the final dataset (default 10,000). Leave it empty to skip synthetic generation.
  • Speed / accuracy.
  • Task Type Selector: classification or regression per target.
  • Data cleaning: the anomaly operations in three groups (Cleaning, Interpretation, Feature extraction). The stage runs when at least one is ticked.
  • Instructions for the pipeline (optional, experimental).
  • Output & Tool Selector: which stages run and which files are written, the export format (CSV, Parquet, Feather, ORC, Excel, JSON, JSON Lines, NDJSON), file names, the model-ready dataset and its split, the multimodal output and the pipeline report PDF.
  • Save artifacts so these runs can later be used for inference or retraining.
  • The Common, Advanced and Experimental tiers.
  • Dataset completion & validation.
  • Decide everything: fills in the target (when it can tell which column it is), classification or regression for each target, and Target Rows from the file's row count. Targets you already chose are kept, as are your outputs, cleaning operations, media settings, speed mode and excluded columns. Available for a profiled file upload.
  • Reset config: puts every setting back, including targets, excluded columns and dataset metadata, then re-applies your saved defaults. It asks first, and both buttons can be undone until you close the pop-up.
  • Save as my defaults.
TierSettings
CommonFeature Selection (method, keep top N), Text Tokenization (Drop Text, Auto, TF-IDF default, Neural), Output Sample Size, audio and image column detection and selection, Enable Preliminary Cleaning, Target model family (Enterprise).
AdvancedMDH Mode (Imputation default, Imputation/dropping on Enterprise, or 2D-Removal), Audio Mode, Image Mode, Scaling Method (Auto default), Data Similarity Threshold (0.99), Z-Score Limit (3), Synthesis Mode (Gaussian Copula or GAN), Class Balancing, and, for Enterprise, Pivot column, Validation method and Dataset metadata.
ExperimentalForce Pure Synthetic Output, Result Evaluator, and, for Enterprise, Column selection, Probe regularization and Column decisions.

Every setting, with its default, range and API name, is in the Configuration reference; what each stage does is in Pipeline modules.

Text Tokenization “Drop Text” does not remove text columns: they are carried through unencoded. Choose TF-IDF, Neural or Auto to turn text into features.

Images and audio

Your table is a manifest: the media column holds file names, and you add the files in the Image Files and Sound Files zones. A zone is enabled once an Image Mode or Audio Mode is set. Click Validate Image Files / Validate Audio Files before starting. Images: jpg, jpeg, png, gif, bmp, webp, tiff, svg. Audio: mp3, wav, ogg, flac, aac, m4a, wma.

Notify on completion

If you have webhooks, tick the ones to call when the run finishes or fails.

Step 3: Process

Click Start Data Preparation. The Confirm processing run dialog shows the estimated time for each processing mode, the exact price (the same for every mode; if it is above your balance you are warned but can still start) and, if the chosen mode would change settings you configured, one line per setting with Keep my settings or Use <mode>. Click Confirm $X and start: you are charged exactly that price. If the file or a setting that affects the price changed since, the run is not started and the new price is shown for you to confirm.

Progress appears as stage (N of M) with a stage timeline. A plain run has ten steps: Loading & Validation, Preprocessing, Anomaly Detection, Data Type Conversion, Missing Data Handling, Train/Test Split (internal; your output is not split), Column Dimension Standardizer, Dimensionality & Similarity Manager, Data Synthetic Generator and Multimodal Token Merger. Data Gathering, Feature Selection and PDF Report add a step each when on. Click CANCEL PROCESSING to stop; the worker stops at the next checkpoint.

Step 4: Results

  • Run Report: rows and columns in and out, duration and warnings.
  • Data Quality Report: rows before and after, null %, and the columns removed.
  • Generated Files: download each file, Download all (zip), or Map Columns to a destination schema.
  • The database button per file, and the Write to a destination panel for a chosen stage, to send an output to one of your saved connections.
  • Process Another File to start over.

The final file, dsg_output, may be mostly synthetic (real rows come first). Evaluate models on dsm_output or the model-ready test split. Every run stays in Runs, where you can download, rename, retry, compare, share and read its report.

Handy links

  • /dashboard/pipeline?demo=1 loads a 200-row demo CSV (the Home screen's Try with demo CSV). It does not pick targets or start a run.
  • /dashboard/pipeline?clone=<session_id> copies an earlier run's configuration.

Inference and retraining

The Inference tab replays a completed run on rows it has never seen:

  • Run inference: encode, impute and scale the new rows with the run's stored transforms, so the same columns come out on the same scale.
  • Retrain a session: refresh the pipeline with the new rows while keeping its output schema. Outlier removal and synthetic generation are optional here.

Targets and settings come from the parent run, so there is nothing to reconfigure. A retraining output is itself a valid parent. See Inference and retraining.

Automating a run mode

Scheduled runs, triggers, folder and cloud listeners and SFTP inboxes share the configuration window and three run modes, chosen in “What should this run?”: Prepare a new model, Run inference or Retrain a session. The replay modes need a base session. Because nobody watches these runs, the window asks once whether the speed mode may change your settings and applies that answer every time.

Account limits

TrialPaidUnrestrictedEnterprise
Processing runs3charged to your balancecharged to your balancecharged to your balance
Data file size (dashboard)1 MB1 GB1 GB1 GB
Connectorsallallallall
API keys131550
Enterprise settingsnononoyes

Tips

  • Before uploading, Estimate price gives an estimated range from rows, columns and, optionally, the kinds of columns and the share of missing values; it charges and binds nothing. Once a file is uploaded, the exact price is shown under the start button. When the run finishes, its summary shows the amount charged.
  • Use Templates to reuse a configuration, or Save as my defaults to make it your starting point.
  • Anomaly Detection takes noticeably longer on large datasets.