Browse documentation
2026-09-29AutoDataOpen in dashboard

Inference and Retraining

Reuse a finished session on new data: score new rows with exactly the transforms your model was trained on, or refresh its training set without changing its schema.

Why replay instead of rerun?

A model trained on AutoData output expects its input prepared the same way every time. That means the same category codes, the same imputed values, the same scaling, and the same columns in the same order. If you ran the full pipeline again on new data, all of that would be re-fitted, and the model would receive a dataset it has never seen.

Inference and retraining replay a completed session (the parent) instead. They load the transforms the parent fitted and apply them to the new rows. Nothing is re-fitted, so the output always has the parent's schema.

The two run modes

InferenceRetraining
Use it toPrepare new rows for scoring by an existing modelBuild a refreshed training set from new rows
StagesAnomaly Detection (if the parent used it), then DTC, MDH and CDS, all replayedThe same replay, followed by outlier row removal and synthetic row generation (each can be turned off)
Rows outOne row for every input rowInput rows, minus any removed outliers, plus synthetic rows up to the target size
Main outputfeatures.csv, plus y_columns.csv when targets are presentfeatures.csv and y_columns.csv
Billed atAbout a fifth of preparing the same dataAbout two fifths of preparing the same data

Feature selection never runs in a replay, and the output is not split into train and test sets. A refreshed set therefore always has the parent's columns, and you can use it in place of the set it updates.

Running it from the dashboard

  1. Open the Inference tab.
  2. Choose where the rows come from:
    • File Upload accepts .csv, .xlsx, .xls, .parquet, .json, .jsonl, .ndjson, .feather and .orc.
    • Connector reads a table or a query from a saved connection.
  3. Under Select Training Session, pick the parent. The list shows:
    • your completed sessions whose stored transforms are still available;
    • sessions shared with you with Can edit access.
    Retraining results can be parents themselves, so you can keep a pipeline current indefinitely. Inference results cannot.
  4. Choose Inference or Retraining.
  5. For a parent that was trained on image or audio, add media files if they are needed.
  6. Click Run Inference or Start Retraining, then download the files from the results.

The Pipeline tab always trains a new session. Replays are run from the Inference tab.

Retraining options

OptionAPI fieldDefaultEffect
Remove Outlier Rowsrun_doronDrops rows that lie far outside the rest of the data.
Add Synthetic Rowsrun_dsgonAdds synthetic rows until the set reaches the target size. Real rows are always kept. If the data already has at least that many rows, nothing is added.
Output Sample Sizeoutput_sample_sizethe parent's output size, or 10,000 if it had noneThe target total row count.
(API only)dor_epsthe parent's value, otherwise derived from the dataOutlier distance tolerance. Larger values keep more rows.
(API only)dor_min_samplesthe parent's value, otherwise 5How many close neighbours a row needs in order to count as normal.

What the new data must look like

  • Columns: every column that was an input feature at training time must be present. If one is missing, the run stops with a SCHEMA_MISMATCH error that lists it. Columns that training itself dropped may be absent.
  • Extra columns are ignored.
  • Column names that differ only in case or whitespace, or that carry a byte-order mark or non-breaking space, are matched to the original names automatically.
  • Targets always come from the parent session. They are optional in the new data. If they are missing, you get features.csv only, along with a NO_TARGET_COLUMNS warning.
  • Anomaly Detection uses the parent's setting. To turn it on or off for a single run, send the enableAnomalyDetection override.

Drift

Every replay compares the new data with the parent's training data and reports a drift score. A higher score means the data has changed more.

Drift scoreReading
below 0.10stable: the new data is still close to the original
0.10 to 0.25moderate: a shift that is worth looking at
0.25 and abovesignificant: the run adds an INPUT_DRIFT warning that names the columns that shifted most. Consider retraining.

If no comparison is possible, the drift result is unavailable and includes a reason. Retraining reports drift for the input and for the final output separately. To be alerted automatically, create a quality alert on drift_score.

Images and audio

If the parent processed image or audio columns, the data file stays a table whose media column holds file names or http(s) URLs. You do not choose media modes or columns for a replay. They come from the parent.

AutoData looks for each referenced file in these places:

  • files uploaded with this run (the imageFiles and soundFiles parts);
  • for listeners, files that arrived beside the data file;
  • the parent session's own media;
  • the URL in the cell.

Files that cannot be found are reported with a MEDIA_UNRESOLVED warning. They are never silently encoded as zeros. Trial accounts can upload up to 1 MB of each media type per run.

Output files

FileWritten byContents
features.csvboth modesThe prepared features, in the parent's column order. In a retraining run, real rows come first.
y_columns.csvboth modesThe prepared targets, in the same row order. Only written when the new data includes targets.
features_complete.csv and features_imputed.csvinferenceThe rows that needed no imputation, and the rows that were imputed
inference_metadata.json and retraining_metadata.jsonthe matching modeRow counts, schema validation, drift, media staging and warnings. Retraining metadata also includes a before-and-after comparison.
transformation_log.json, dor_report.json, dsg_report.jsonas applicableStage logs

Replay outputs are always CSV and always keep these names, regardless of your export format or naming settings.

Automating replays

Scheduled runs, triggers, folder and cloud listeners, SFTP inboxes and one-off connector runs all take a run mode: full_pipeline (the default), inference or retraining. Send it as run_mode, together with the base_session_id to replay. That session must be one of your own completed sessions. A retraining automation can also set its outlier-removal, synthetic-row and target-row options.

Automated runs cannot upload files. A media session run this way must therefore name its media by http(s) URL, or the files must still be present in the original session. Listeners also pick up media files that land beside the data file.

From the API and the Python SDK

POST /api/v1/inference and POST /api/v1/retraining accept a session cookie or an API key.

  • They are synchronous. The response comes back when the run finishes, up to about 25 minutes. If the run takes longer, you get a TIMEOUT response containing the session_id; poll /progress/<id> for that session.
  • To upload a file, send multipart with file and original_session_id.
  • To read from a connector, send JSON with:
    • original_session_id;
    • source_type;
    • table or custom_query;
    • credential_id, inline_secrets, or both. inline_secrets holds connection fields sent with the request. They override any saved credential and are not stored.

The response includes:

  • new_session_id;
  • output_files;
  • schema_validation;
  • warnings;
  • metadata, which includes the drift result;
  • amount_charged_usd and balance_usd, in US dollars (the older credits_charged and credits_remaining are deprecated).
from datatoolpack import AutoDataClient

with AutoDataClient(api_key="dtpk_xxxxxxxxxxxx") as client:
    scored = client.infer(
        original_session_id="3f2a...",
        file_path="new_rows.csv",
        download_path="./scored",
    )
    refreshed = client.retrain(
        original_session_id="3f2a...",
        file_path="new_training_rows.csv",
        output_rows=20000,
        download_path="./refreshed",
    )

You can also send POST /api/v1/process with run_mode and base_session_id in its config. This queues a run that you poll. The legacy paths /inference and /retraining behave the same as the v1 paths. /inference-retraining has been removed. See the API Integration Guide.

Billing

A replay is priced on what is in the new data, the same way as a training run (rows, columns and their kinds, missing values and the stages that replay), then by run type:

  • inference costs about a fifth of preparing the same data;
  • retraining costs about two fifths.

Image and audio bytes are charged per GB, the same way as in a training run. Each replay is charged the exact price of the data it processes, measured when it starts.

Enterprise accounts (Enterprise)

Enterprise accounts are billed at 2x, and the run-type factor applies on top: Enterprise inference costs about two fifths, and retraining about four fifths, of what preparing the same data costs on another plan.

When a replay is refused

  • The session is not available to you or has not finished. It must be yours, or shared with you with edit access, and it must have completed.
  • The session's stored transforms have been cleaned up (MISSING_ARTIFACTS). Replay a newer session instead.
  • A required input column is missing (SCHEMA_MISMATCH).
  • Retraining found no target columns in the parent (MISSING_Y_COLUMNS).

LM Readiness sessions

LM Readiness sessions carry an LM Readiness label in the session picker. Inference applies the session's classic pipeline and saved LM transformations to the new rows without refitting; retraining prepares the new rows again with the session's own LM Readiness settings. The results show the readiness report with Pipeline outputs and LM inputs. The new rows need the session's mapped role columns and encoded source columns; up to 200,000 rows per run.