Runs and Results
The Runs tab, renaming, tags, compare, retry and templates; every output file a run produces, how files are named and formatted, how to download them, and how long they are kept.
Runs
Each pipeline execution is a run, also called a session, and has a session ID. The dashboard's Runs tab lists all of your own runs, newest first:
- training runs;
- inference and retraining runs;
- runs started by automations.
Runs that other people share with you appear on the Collaboration tab instead.
At the top of the tab are counts of total, completed, failed and cancelled runs. You can search by name, file or session ID, filter by status, and filter by tag. Each row shows:
- the run name and the first part of its session ID, with a button that copies the full ID;
- the run type: Pipeline, Inference or Retraining;
- the input file and its row count, column count and size;
- the start time;
- the status.
Row actions
| Action | What it does |
|---|---|
| Rename | Hover over the name and click the pencil. Names can be up to 255 characters. API: PUT /api/v1/sessions/<id>/name with {"session_name": "..."}. |
| Tags and note | Up to 10 tags of up to 32 characters each, and a private note of up to 500 characters. API: PUT /api/v1/sessions/<id>/meta. |
| Zip | Downloads every output as autodata_<first 8 characters of the ID>.zip. Available on completed runs. |
| Clone config | Opens the Pipeline tab filled in with this run's configuration, so you can start a new run on a new file. The link form is /dashboard/pipeline?clone=<session_id>. |
| Retry | Shown on failed runs. It resubmits the run from its last finished stage, using the same session ID. The retry needs the original upload, and uploads are removed when a run ends. So in practice it works only after an interruption such as a server restart. Otherwise, clone the configuration and upload the file again. API: POST /api/v1/sessions/<id>/retry. |
| Show files | Lists the run's output files, each with a download button. |
| Compare | Tick two runs, then click Compare. You see rows and columns in and out, duration, output files and warnings side by side, with differences highlighted. |
Pipeline templates
The Templates menu on the Pipeline tab can save the current configuration under a name of up to 80 characters. You can then apply, rename or delete saved templates. Each account can hold up to 50 templates, and each template's configuration can be up to 64 KB. Templates include media settings. They do not include a run mode.
API: GET and POST /api/pipeline-templates; PUT and DELETE /api/pipeline-templates/<id>.
Starting, following and cancelling a run
- Confirmation. Before a dashboard run starts, a dialog shows a recap and the exact price of the uploaded file with your settings.
- Time is shown for each speed mode, relative to Rapid.
- The price is the same in every speed mode. Confirm $X and start starts the run, and you are charged exactly that amount. If the file or a setting that affects the price changed since, the run is not started and the new price is shown for you to confirm. If the price is above your balance you are warned, but the run can still start.
- For every setting you configured that the chosen speed mode would change, you pick Keep my settings or Use <mode>.
- The API equivalent is
POST /api/v1/estimatewith thepricing_profile_idfrom validation. It never charges.
- Waiting for approval. Scheduled runs, triggers and folder listeners can have a Maximum price per run ($). A run whose price is above it is held instead of running and listed under Waiting for approval on the Runs tab, with its price and the limit. Approve starts it at that price, and it is charged exactly that; Discard drops it and nothing is charged. You are notified by a
job.heldwebhook event and, if you receive failure emails, by email. API:GET /api/pending-approvals,POST /api/pending-approvals/{id}/approve,POST /api/pending-approvals/{id}/discard(409 if already decided, 410 if the held data is gone). - Duplicate protection. Send an
Idempotency-Keyheader, or anidempotencyKeyform field, with/process-data. If a later request repeats that key, you get back the existing session ID instead of a second run. - Progress. The Pipeline tab shows a stage timeline with a duration for each stage, a progress bar for long stages, and the warnings raised so far. A sidebar chip, Run in progress · Step N/M, follows the run from any tab. Keep the dashboard open until a run finishes the first time, because a dashboard run is added to your history when its result is first fetched.
- Cancel. Stops the run, which is then recorded as Cancelled. Files from stages that already finished may remain. API:
POST /api/v1/cancel/<session_id>.
Output files
You choose which files are written under Output Preferences. These are saved per account, and each run can override them. A stage's file is written only when its preference is on and the stage actually ran.
| File (step) | Default | Contents |
|---|---|---|
anormaly_fixed_output | on | Data after anomaly repair |
dtc_output | on | Data after type conversion and encoding |
mdh_output | on | Data after missing values have been handled |
cds_output | on | Scaled data |
dsm_output | on | Data after feature selection. Contains all real rows, fully processed. |
dsg_output | on | The processed data plus synthetic rows, up to your requested row count. Real rows come first. |
multimodal_output | on | The final merged frame |
model_ready | on | Flattened, all-numeric data with no empty cells, plus model_ready_manifest.json. See below. |
compleated_dataset | off | The Data Completion and Verification result. Always CSV. |
pipeline_report.pdf | on | The PDF pipeline report |
quality_report.json | always | A baseline model score on the real, processed rows |
result_evaluation_report.json and .csv | off | A benchmark of several standard models. Written only when the Result Evaluator is on. Included in ZIP downloads. |
step_metrics.json and .csv | always | Timing and row counts for each stage |
Which file to evaluate on
dsg_output is not an evaluation set. If the row count you requested is larger than your data, most of its rows can be synthetic, and a model scored on it will look better than it is.
Score on one of these instead:
dsm_output, which contains the real rows, fully processed;- the first
real_rowsrows ofdsg_output; - the model-ready test split, which is drawn from real rows only by default.
Every result reports dsg_provenance, which gives the real and synthetic row counts.
Model-ready export
| Setting | Values | Default |
|---|---|---|
model_ready_split | none, xy (X and y files), train_test, xy_train_test (X_train, X_test, y_train and y_test) | none |
model_ready_test_size | 0.05 to 0.5 | 0.2 |
model_ready_real_test_only | Draw the test split from real rows only | on |
Columns that cannot be made numeric are written to model_ready_excluded_columns, in the same row order as the export. If a split cannot be made, the export falls back to a simpler layout and adds a warning.
File names
output_name_pattern controls how data files and model-ready files are named:
prefixed(the default) gives[input]_[YYYYMMDD_HHMMSS]_[step].[ext], for examplesales_20260922_013000_dsg_output.parquet. Automated runs leave out the input part.plaingives justdsg_output.parquet.
Reports, manifests and inference or retraining outputs always keep their fixed names.
Export format
output_format can be csv (the default), parquet, xlsx, json, jsonl, ndjson, feather or orc. The format never makes a run fail:
- A file that is too large for Excel, which allows 1,048,576 rows and 16,384 columns, is written as CSV instead.
- A file that cannot be written in the chosen format is also written as CSV instead.
- Columns of mixed type in Parquet, Feather or ORC are written as text.
Every such change is listed in output_format_warnings. Inference and retraining outputs are always CSV.
Downloading
| Route | Returns |
|---|---|
GET /download/<session_id>/<filename> | One file |
GET /api/sessions/<session_id>/files | The list of files, with their sizes and download URLs |
GET /api/v1/sessions/<session_id>/download-all | A ZIP of every output. Your input file and empty files are left out. |
POST /api/v1/download-archive/<session_id> | A ZIP of selected files. Send {"files": ["dsg_output", "model_ready"]}; names can be full file names, stems or step names. |
Outputs are always stored on disk. The older in-memory output routes under /api/v1/sessions/<id>/outputs are kept for compatibility, but current runs do not fill them.
How long files are kept
| What | Kept for |
|---|---|
| Your uploaded input file | Removed when the run ends |
| Output data files, the PDF, model-ready files and most reports | 7 days |
| The run's stored configuration and transforms, which inference and retraining need, plus uploaded images and audio | 1 year |
| Run history, names, tags and notes | Not removed |
Download what you need within the week, or set the run to write its results to a database automatically. Self-hosted installations can change these periods.
Statuses and errors
| Status | Meaning |
|---|---|
| Processing | The run is in progress. A run stuck in this state for more than 2 hours is marked failed. |
| Completed | Every enabled stage finished. |
| Failed | A stage stopped the run. The result carries error_info with a code, stage, column, message and hint. The code is one of TOO_FEW_ROWS, TARGET_INVALID, COLUMN_NOT_FOUND, OUT_OF_MEMORY, NO_DATA, WORKER_DIED, TIMEOUT, DATA_TYPE_CONFLICT, FILE_UNREADABLE or PROCESSING_FAILED. |
| Cancelled | You stopped the run. |
To list runs, call GET /api/v1/sessions. It accepts these filters: q, status, tag, date_from and date_to. Each run in the response includes is_retrainable, which tells you whether its transforms are still stored. To get one run's details, call GET /api/v1/sessions/<id>. For the configuration used to clone a run, call GET /api/v1/sessions/<id>/config.