API Integration Guide
Get started with the AutoData REST API: authentication, base URL, your first run, polling, errors, rate limits, balance and idempotency.
Overview
The AutoData REST API lets you submit data-preparation runs, follow their progress and download the prepared files without the browser dashboard. This guide gets you from an API key to a finished run. Every endpoint, field and error is listed in the REST API reference; the machine-readable spec is served at /openapi.yaml.
Authentication
Create an API key in the dashboard (API Keys tab) or on your API Account page. Keys start with dtpk_ and are shown once, when created or rotated. Send the key as a Bearer token:
Authorization: Bearer dtpk_xxxxxxxxxxxx
The same key is accepted as X-API-Key: dtpk_... wherever the Bearer form is, including the core job routes (/api/v1/process, /status, /result, /download, /download-archive, /cancel, /keys, /usage). Your 12-digit account passcode can be sent as X-Passcode instead of a key; per-key limits do not apply then.
Each key can carry a daily and a lifetime spending limit, in US dollars, and a list of stages it may run (anomaly, DTC, MDH, CDS, DSM, DSG). A run that needs a stage the key may not use is refused with 403; a key over its spending limit gets 429. Keys can have an expiry date and can be rotated without changing their settings. Keys are created, rotated, edited and revoked in the dashboard; through the API you can list your keys (GET /api/v1/keys) and read a key's usage (GET /api/v1/usage).
Base URL
https://autodata.datatoolpack.com
The stable scripting surface lives under /api/v1/. Timestamps in responses are UTC, written without a zone suffix.
Your first run
A run is three steps: submit, poll, collect.
# 1. Submit: upload a file and name the target column
curl -X POST https://autodata.datatoolpack.com/api/v1/process \
-H "Authorization: Bearer dtpk_xxxxxxxxxxxx" \
-F "file=@data.csv" \
-F 'config={"target_columns": ["churn"], "output_rows": 10000}'
# -> {"success": true, "session_id": "3f2a...", "status": "queued", ...}
# 2. Poll until status is "completed" (or "error" / "failed" / "cancelled")
curl https://autodata.datatoolpack.com/api/v1/status/SESSION_ID \
-H "Authorization: Bearer dtpk_xxxxxxxxxxxx"
# 3. List the outputs, then download them
curl https://autodata.datatoolpack.com/api/v1/result/SESSION_ID \
-H "Authorization: Bearer dtpk_xxxxxxxxxxxx"
curl -X POST https://autodata.datatoolpack.com/api/v1/download-archive/SESSION_ID \
-H "Authorization: Bearer dtpk_xxxxxxxxxxxx" -o outputs.zip
Accepted upload formats: CSV, XLSX, XLS, Parquet, JSON, JSONL, NDJSON, Feather and ORC, up to 5 GB (1 MB on Trial). config also takes tools (which stages run) and advanced_params (stage settings, the export format and file naming). Every pipeline setting the dashboard offers can be sent as well, at the top level of config or inside advanced_params, in snake_case or camelCase: for example task_types (c or r per target, or auto; when omitted, every target is treated as classification), session_name, anomaly_params, webhook_ids and output_preferences. output_preferences selects the output files of this one run; without it, the files written follow your saved output preferences. advanced_params.output_format / output_name_pattern set the format and naming for one run.
To read rows from a database or cloud source instead of uploading, use POST /api/v1/process-connector with a saved connection. To score new rows with a finished run, use POST /api/v1/inference; to refresh its training set, POST /api/v1/retraining. Both answer when the run is done.
Polling
{
"session_id": "3f2a...",
"status": "processing",
"message": "Imputing missing values...",
"current_step": 6,
"total_steps": 11,
"progress_percent": 55,
"duration_seconds": 41.2
}
total_steps is the number of stages the run reports: 10 for a standard run, plus one each for Data Completion & Verification, feature selection and the PDF report when they are enabled. Poll every few seconds; there is no need to send heartbeats. A failed run carries error_info with a stable code, the stage and a hint. Instead of polling, you can register a webhook for job.completed and job.failed (see the reference).
Speed presets and the 409 answer
advanced_params.mode (or speed_mode at the top level of config) picks a speed preset: rapid, balanced (default) or accurate. A top-level mode on /api/v1/process is the run mode (full_pipeline, inference, retraining), not the preset. If the preset would change a parameter you set yourself, the submission is answered with 409 and "error": "mode_confirmation_required", listing the contested settings and a policy_signature. Re-send with advanced_params.mode_ack set to that signature to keep your values, or with advanced_params.mode_decisions to answer per parameter. POST /api/v1/mode-preview shows the changes without submitting anything.
Idempotency
A submission that timed out may still have started a run. To retry safely, send an Idempotency-Key header with POST /api/v1/process: any unique string, such as a UUID, one per run you intend to start. A retry with the same key from the same caller returns the first request's session_id with "idempotent_replay": true instead of starting a second run. Reuse a key only for retries of the same submission, and use a new key for each new run.
curl -X POST https://autodata.datatoolpack.com/api/v1/process \
-H "Authorization: Bearer dtpk_xxxxxxxxxxxx" \
-H "Idempotency-Key: 7d3f6c1e-0b0a-4c57-9a57-2f1f3f0e9b11" \
-F "file=@data.csv" \
-F 'config={"target_columns": ["churn"]}'
The dashboard upload route (/process-data) accepts the same header (or an idempotencyKey field). Idempotency is best-effort and not guaranteed across a server restart. The Python SDK sends a key with every process() call, so its automatic retries are safe. For other submissions, check GET /api/v1/sessions for a run you already started before submitting again.
Errors
Errors return JSON with an error message, sometimes with hint or code:
{"error": "target_columns is required"}
| Status | Meaning |
|---|---|
400 | Missing or invalid field, unsupported file type, invalid config JSON, an out-of-range setting, or a replay that cannot run |
401 | Missing or invalid API key |
403 | The session is not yours, the key may not run a requested stage, a Trial limit, or an Enterprise-only endpoint |
404 | Session, file, connection or other resource not found; also any unknown /api/ path |
409 | mode_confirmation_required (see above), or a confirmed price that no longer applies (price_changed, file_changed, quote_expired; estimate again) |
410 | Removed endpoint ("code": "ENDPOINT_REMOVED") or an expired share link |
429 | The key reached its daily or lifetime spending limit, or a login / trigger rate limit |
Rate limits
There is no general per-request rate limit on the job API. The limits that exist are: 10 login attempts per 5 minutes per IP, 10 fires per minute per inbound trigger, 20 user searches per minute, and each key's own daily and lifetime spending limits, checked when a run is submitted.
Plans and balance
API access is available on every plan. Balances, prices and charges are in US dollars. New accounts start with this balance and key allowance; an administrator can raise both.
| Plan | Starting balance | API keys | Billing |
|---|---|---|---|
| Trial | $0.10 | 1 | Standard price; 3 runs, 1 MB files |
| Paid | $10 | 3 | Standard price |
| Unrestricted | $10 or custom | 15 | Standard price |
| Enterprise | $10 or custom | 50 | 2x the standard price |
A run is charged when it finishes, in dollars rounded to the cent; a run priced above $0 is charged at least $0.01. A run is priced on what is in the data, not on the size of the file: the rows (the price per row falls as a run grows, with a minimum price per row), the columns and their kinds, missing values and the stages that run; images and audio are charged per GB. Inference costs about a fifth and retraining about two fifths of preparing the same data, because both reuse fitted transforms. See Account and billing for what raises the price.
To know the price before a run, validate the file with /api/v1/validate. Every row is measured; the response carries a pricing_profile_id and a pricing_status (ready, or measuring for a larger file: poll GET /api/pricing/profile/{id} until it is ready). POST /api/v1/estimate with that pricing_profile_id and your run settings returns the exact price (price_usd) and a quote_id. Pass it as config.quote_id on /api/v1/process with the same file and the run is charged exactly that price. If the file is not the one that was priced, the settings no longer price the same, or the quote is older than 7 days, you get 409 (file_changed, price_changed, quote_expired) and nothing starts. Without a file, input_rows and columns (and optionally column_types and missing_rate) give an estimated range (low_usd, high_usd, point_usd) that binds nothing. A run started without a quote_id is charged the exact price of its data, measured when it starts. Every /api/v1/result response reports amount_charged_usd and balance_usd; a low balance never blocks a run, and a negative balance is reported in balance_warning.
Money fields end in _usd. The older *_credits fields and other credit-named fields (credits_charged, credits_remaining, credit_warning) are deprecated: still returned, equal to the dollar amount × 1,000,000, and will be removed in a future version.
Python SDK
pip install datatoolpack
from datatoolpack import AutoDataClient
with AutoDataClient(api_key="dtpk_xxxxxxxxxxxx") as client:
result = client.process(
file_path="data.csv",
target_columns=["churn"],
output_rows=10000,
download_path="./outputs",
)
scored = client.infer(
original_session_id=result["session_id"],
file_path="new_rows.csv",
download_path="./scored",
)
The client also reads AUTODATA_API_KEY from the environment. process() waits and downloads by default; pass wait=False to get the session_id back immediately. It takes the same settings as the route, for example task_types, session_name, anomaly_params and output_config (the per-run output selection).
Enterprise (Enterprise)
Enterprise is not a separate submission path: an Enterprise run is started through the same endpoints, with extra settings such as the model families to prepare the data for and per-column method pins. Those settings are ignored for other plans. Read-only endpoints under /api/v1/enterprise/ return the per-column record, the family comparison, lineage and an audit report for Enterprise accounts, and answer 403 with "required_role": "Enterprise" for others. See the REST API reference.