Reports and Insights
The dataset preview on upload, the Run Report and what each plan sees, the Data Quality Report, the PDF pipeline report and the optional Result Evaluator.
Overview
AutoData reports on your data at several points:
- a dataset preview when you upload;
- a Run Report and a Data Quality Report when a run finishes;
- a PDF pipeline report you can download;
- an optional benchmark from the Result Evaluator.
Every report describes outcomes: what changed, by how much, and what you should check.
Dataset preview
After you upload a file, the Pipeline tab shows three sections:
- Column profile covers up to 300 columns. For each column it shows the kind (number, text, date, boolean or other), the percentage of empty cells, the number of distinct values, and a zero variance flag. The figures are computed from the first 100,000 rows.
- Sample rows shows the first 10 rows of the first 12 columns.
- Dry-run preview gives each column an advisory verdict: Kept, Dropped (the column is empty, or holds the same value in every row) or Needs review (many values are missing, the format was not recognised, or the column looks like an identifier). The run itself decides using the full data.
Run Report
The Run Report appears when a run completes, and the Compare view on the Runs tab uses it too. You can also fetch it with GET /api/v1/sessions/<session_id>/report.
| Field | Meaning |
|---|---|
summary.rows_in, summary.rows_out | Rows in the input and in the final dataset |
summary.cols_in, summary.cols_out | Columns in the input and in the final dataset |
summary.duration_seconds | Total run time |
summary.output_files | Each output file with a plain-language label, such as Final dataset or Model-ready dataset |
warnings | Outcome-level notes, each with a code, message and stage |
What each plan sees
| Plan | Tier | Contents |
|---|---|---|
| Trial, Paid, Unrestricted, Admin | summary | The summary and warnings above |
| Enterprise | full | The summary, plus details: which columns each stage removed and why, what was imputed, the final schema, row accounting for each stage, and the per-column decision ledger |
Data Quality Report
The Data Quality Report shows:
- rows before and after the run;
- columns before and after;
- the percentage change in rows, flagged when more than 20% of rows were lost;
- alerts;
- which columns contained empty cells, with the percentage for each.
Two further panels appear when they apply. Columns Removed lists columns dropped because they carried no usable signal. A media warning appears when some image or audio files could not be read.
Every run also writes quality_report.json. It contains a quick baseline model score (accuracy for classification, R² for regression) computed on the real, processed rows before any synthetic rows are added. That makes it an honest measure of how learnable your data is. If the score cannot be computed, status is skipped and a reason is given, for example that the dataset does not have exactly one target or has too few rows.
PDF pipeline report
pipeline_report.pdf is written when the Pipeline report output preference is on, which it is by default. For API and automation runs, use enable_pdf_report. Every plan can get the report. It contains:
- An executive summary, followed by the data collection and validation analysis (only when Data Completion ran), the input data analysis and an input sample.
- A stage execution summary with timings, and an output sample.
- Performance and technical diagnostics: a system overview, the status of advanced features, the columns that were removed, and pipeline metadata.
- Four-row data samples after each stage. A stage that did not run is marked "Not generated".
If the PDF cannot be generated, the run still succeeds without it. Like other outputs, the PDF is kept for 7 days.
Result Evaluator
The Result Evaluator trains a few standard models on the prepared data for each target and reports their scores:
- Regression: R², MAE, MSE, RMSE and MAPE.
- Classification: accuracy, macro F1, macro precision, macro recall and ROC AUC.
It also names the best model for each target. It is off by default. Turn it on with enableResultEvaluator in the dashboard, or with enable_result_evaluator or tools.result_evaluator through the API. It is available on every plan.
The results go to result_evaluation_report.json and result_evaluation_report.csv. You get these files through a ZIP download of the run.
A note on evaluation
dsg_output can consist mostly of synthetic rows, so scores computed on it overstate real-world accuracy. For your own evaluation, use dsm_output or the model-ready test split, which comes from real rows only.