Inference and Retraining
Reuse a finished session on new data: score new rows with exactly the transforms your model was trained on, or refresh its training set without changing its schema.
Why replay instead of rerun?
A model trained on AutoData output expects its input prepared the same way every time. That means the same category codes, the same imputed values, the same scaling, and the same columns in the same order. If you ran the full pipeline again on new data, all of that would be re-fitted, and the model would receive a dataset it has never seen.
Inference and retraining replay a completed session (the parent) instead. They load the transforms the parent fitted and apply them to the new rows. Nothing is re-fitted, so the output always has the parent's schema.
The two run modes
| Inference | Retraining | |
|---|---|---|
| Use it to | Prepare new rows for scoring by an existing model | Build a refreshed training set from new rows |
| Stages | Anomaly Detection (if the parent used it), then DTC, MDH and CDS, all replayed | The same replay, followed by outlier row removal and synthetic row generation (each can be turned off) |
| Rows out | One row for every input row | Input rows, minus any removed outliers, plus synthetic rows up to the target size |
| Main output | features.csv, plus y_columns.csv when targets are present | features.csv and y_columns.csv |
| Billed at | About a fifth of preparing the same data | About two fifths of preparing the same data |
Feature selection never runs in a replay, and the output is not split into train and test sets. A refreshed set therefore always has the parent's columns, and you can use it in place of the set it updates.
Running it from the dashboard
- Open the Inference tab.
- Choose where the rows come from:
- File Upload accepts
.csv,.xlsx,.xls,.parquet,.json,.jsonl,.ndjson,.featherand.orc. - Connector reads a table or a query from a saved connection.
- File Upload accepts
- Under Select Training Session, pick the parent. The list shows:
- your completed sessions whose stored transforms are still available;
- sessions shared with you with Can edit access.
- Choose Inference or Retraining.
- For a parent that was trained on image or audio, add media files if they are needed.
- Click Run Inference or Start Retraining, then download the files from the results.
The Pipeline tab always trains a new session. Replays are run from the Inference tab.
Retraining options
| Option | API field | Default | Effect |
|---|---|---|---|
| Remove Outlier Rows | run_dor | on | Drops rows that lie far outside the rest of the data. |
| Add Synthetic Rows | run_dsg | on | Adds synthetic rows until the set reaches the target size. Real rows are always kept. If the data already has at least that many rows, nothing is added. |
| Output Sample Size | output_sample_size | the parent's output size, or 10,000 if it had none | The target total row count. |
| (API only) | dor_eps | the parent's value, otherwise derived from the data | Outlier distance tolerance. Larger values keep more rows. |
| (API only) | dor_min_samples | the parent's value, otherwise 5 | How many close neighbours a row needs in order to count as normal. |
What the new data must look like
- Columns: every column that was an input feature at training time must be present. If one is missing, the run stops with a
SCHEMA_MISMATCHerror that lists it. Columns that training itself dropped may be absent. - Extra columns are ignored.
- Column names that differ only in case or whitespace, or that carry a byte-order mark or non-breaking space, are matched to the original names automatically.
- Targets always come from the parent session. They are optional in the new data. If they are missing, you get
features.csvonly, along with aNO_TARGET_COLUMNSwarning. - Anomaly Detection uses the parent's setting. To turn it on or off for a single run, send the
enableAnomalyDetectionoverride.
Drift
Every replay compares the new data with the parent's training data and reports a drift score. A higher score means the data has changed more.
| Drift score | Reading |
|---|---|
| below 0.10 | stable: the new data is still close to the original |
| 0.10 to 0.25 | moderate: a shift that is worth looking at |
| 0.25 and above | significant: the run adds an INPUT_DRIFT warning that names the columns that shifted most. Consider retraining. |
If no comparison is possible, the drift result is unavailable and includes a reason. Retraining reports drift for the input and for the final output separately. To be alerted automatically, create a quality alert on drift_score.
Images and audio
If the parent processed image or audio columns, the data file stays a table whose media column holds file names or http(s) URLs. You do not choose media modes or columns for a replay. They come from the parent.
AutoData looks for each referenced file in these places:
- files uploaded with this run (the
imageFilesandsoundFilesparts); - for listeners, files that arrived beside the data file;
- the parent session's own media;
- the URL in the cell.
Files that cannot be found are reported with a MEDIA_UNRESOLVED warning. They are never silently encoded as zeros. Trial accounts can upload up to 1 MB of each media type per run.
Output files
| File | Written by | Contents |
|---|---|---|
features.csv | both modes | The prepared features, in the parent's column order. In a retraining run, real rows come first. |
y_columns.csv | both modes | The prepared targets, in the same row order. Only written when the new data includes targets. |
features_complete.csv and features_imputed.csv | inference | The rows that needed no imputation, and the rows that were imputed |
inference_metadata.json and retraining_metadata.json | the matching mode | Row counts, schema validation, drift, media staging and warnings. Retraining metadata also includes a before-and-after comparison. |
transformation_log.json, dor_report.json, dsg_report.json | as applicable | Stage logs |
Replay outputs are always CSV and always keep these names, regardless of your export format or naming settings.
Automating replays
Scheduled runs, triggers, folder and cloud listeners, SFTP inboxes and one-off connector runs all take a run mode: full_pipeline (the default), inference or retraining. Send it as run_mode, together with the base_session_id to replay. That session must be one of your own completed sessions. A retraining automation can also set its outlier-removal, synthetic-row and target-row options.
Automated runs cannot upload files. A media session run this way must therefore name its media by http(s) URL, or the files must still be present in the original session. Listeners also pick up media files that land beside the data file.
From the API and the Python SDK
POST /api/v1/inference and POST /api/v1/retraining accept a session cookie or an API key.
- They are synchronous. The response comes back when the run finishes, up to about 25 minutes. If the run takes longer, you get a
TIMEOUTresponse containing thesession_id; poll/progress/<id>for that session. - To upload a file, send multipart with
fileandoriginal_session_id. - To read from a connector, send JSON with:
original_session_id;source_type;tableorcustom_query;credential_id,inline_secrets, or both.inline_secretsholds connection fields sent with the request. They override any saved credential and are not stored.
The response includes:
new_session_id;output_files;schema_validation;warnings;metadata, which includes the drift result;amount_charged_usdandbalance_usd, in US dollars (the oldercredits_chargedandcredits_remainingare deprecated).
from datatoolpack import AutoDataClient
with AutoDataClient(api_key="dtpk_xxxxxxxxxxxx") as client:
scored = client.infer(
original_session_id="3f2a...",
file_path="new_rows.csv",
download_path="./scored",
)
refreshed = client.retrain(
original_session_id="3f2a...",
file_path="new_training_rows.csv",
output_rows=20000,
download_path="./refreshed",
)
You can also send POST /api/v1/process with run_mode and base_session_id in its config. This queues a run that you poll. The legacy paths /inference and /retraining behave the same as the v1 paths. /inference-retraining has been removed. See the API Integration Guide.
Billing
A replay is priced on what is in the new data, the same way as a training run (rows, columns and their kinds, missing values and the stages that replay), then by run type:
- inference costs about a fifth of preparing the same data;
- retraining costs about two fifths.
Image and audio bytes are charged per GB, the same way as in a training run. Each replay is charged the exact price of the data it processes, measured when it starts.
Enterprise accounts (Enterprise)
Enterprise accounts are billed at 2x, and the run-type factor applies on top: Enterprise inference costs about two fifths, and retraining about four fifths, of what preparing the same data costs on another plan.
When a replay is refused
- The session is not available to you or has not finished. It must be yours, or shared with you with edit access, and it must have completed.
- The session's stored transforms have been cleaned up (
MISSING_ARTIFACTS). Replay a newer session instead. - A required input column is missing (
SCHEMA_MISMATCH). - Retraining found no target columns in the parent (
MISSING_Y_COLUMNS).
LM Readiness sessions
LM Readiness sessions carry an LM Readiness label in the session picker. Inference applies the session's classic pipeline and saved LM transformations to the new rows without refitting; retraining prepares the new rows again with the session's own LM Readiness settings. The results show the readiness report with Pipeline outputs and LM inputs. The new rows need the session's mapped role columns and encoded source columns; up to 200,000 rows per run.