REST API Reference
Every AutoData REST endpoint: methods, paths, authentication, request fields, responses and errors.
About this reference
Every customer-facing HTTP endpoint of AutoData: method and path (with alias paths that reach the same handler), accepted authentication, request fields with types and defaults, response shape and errors. For a walkthrough, start with the API integration guide. Base URL: https://autodata.datatoolpack.com. Admin-only routes are not listed.
Authentication labels
| Label | Accepts |
|---|---|
| Any | Session cookie, Authorization: Bearer dtpk_..., X-API-Key: dtpk_... or X-Passcode |
| Key | Authorization: Bearer dtpk_..., X-API-Key: dtpk_... or X-Passcode (no cookie) |
| Session | Signed-in browser session only |
| Public | No authentication |
| Enterprise | Any method, Enterprise accounts only (403 otherwise) |
X-API-Key is accepted wherever the Bearer form is, including the core job routes (process, status, result, cancel, download, download-archive, keys, usage). API keys are created, rotated, edited and revoked in the dashboard, where the profile and the passcode are changed too; through the API a key can list the account's keys and read its own usage.
Cookie-authenticated requests that change something (POST, PUT, PATCH, DELETE) outside /api/v1/, /trigger/ and the login routes must send X-CSRF-Token from GET /auth/csrf-token; Bearer requests are exempt. Errors are JSON with error and sometimes hint or code. Failed runs carry error_info {code, stage, column, message, hint}. Timestamps are UTC without a zone suffix.
Status codes
| Code | Meaning |
|---|---|
| 200 / 201 / 202 | Success / created or replay queued / inbound trigger accepted |
| 400 | Invalid or missing field, bad file type, bad JSON, out-of-range value |
| 401 | Not authenticated or bad signature |
| 403 | Not your resource, wrong plan, Trial limit, key stage permission, CSRF |
| 404 | Not found (also unknown /api/ paths) |
| 405 | Wrong method on a real path (Allow header) |
| 409 | mode_confirmation_required; a confirmed price that no longer applies (price_changed, file_changed, quote_expired); a held run already approved or discarded; duplicate watermark; email in use |
| 410 | Removed endpoint, expired share link, disabled trigger, held run whose data is gone |
| 413 | Upload or body too large |
| 429 | Rate or spending limit |
| 500 / 502 / 503 | Server error / a trigger's run failed to start / degraded health or Spark off |
Limits and billing
- Login: 10 attempts per 5 minutes per IP. Inbound trigger: 10 fires per minute. User search: 20 per minute.
- Per-key daily and lifetime spending limits, in US dollars (429), and stage permissions (403) are checked when a run is started with a key through
/api/v1/process,/api/v1/process-connectoror/process-data. - Trial: 3 runs, 1 MB per data file, 1 MB of images and 1 MB of audio per request. Paid plans: 5 GB per request.
- Money is US dollars everywhere: balances, charges, prices and key limits. A run is charged when it finishes, in dollars rounded to the cent; a run priced above $0 is charged at least $0.01. A run is priced on what is in the data, not on the file size: rows (the price per row falls past 1M, 5M, 10M, 50M, 100M and 1B rows, with a minimum price per row), columns and their kinds, missing values, differences in numeric scale, the number of features and the stages that run; synthetic rows are charged per doubling of the row count, image and audio bytes per GB, dataset completion and validation per cell. Inference costs about a fifth and retraining about two fifths of preparing the same data; Enterprise 2x. See Account and billing.
- Money fields end in
_usd. The older*_creditsfields and other credit-named fields (such ascredits_charged,credits_remaining,credit_warning,daily_credit_limit) are deprecated: still returned, equal to the dollar amount × 1,000,000, and will be removed in a future version. Existing*_dollarsfields are unchanged and hold the same amounts as their_usdcounterparts. - An uploaded file is measured in full and priced exactly:
/api/v1/estimatewith itspricing_profile_idreturns one price and aquote_id, and a run started with thatquote_idis charged exactly that price (409price_changed,file_changedorquote_expiredwhen it no longer applies). A description without a file gets an estimated range, which binds and caps nothing. Runs without aquote_id, and automated runs, are charged the exact price of the data they process, measured when they start; automations can setmax_price_dollarsto hold pricier runs for approval. Streams are billed hourly for rows scored, at the inference rate. A low balance never blocks a run; the balance can go negative.
Authentication and session
POST /auth/login
Public. Body {"passcode": "123456789012"}. Returns {success, message, user: {id, name}} and sets the session cookie (24 hours). 400 no body, 401 Invalid passcode, 429.
POST /auth/email-login
Public. Body email, password. Returns {success, message, user: {id, name, email}}. 400, 401 Invalid email or password, 429.
POST /auth/signup
Public. Body email, name, surname (each up to 100 characters), password (8+ characters with upper case, lower case and a digit). Creates a Trial account (3 runs, a $0.10 balance, one API key), signs it in and returns the login passcode once. 400 validation or existing email.
POST /auth/logout
Public. Clears the session. {success, message}.
GET /auth/status
Session. {authenticated, admin_authenticated, account_authenticated, user: {id, name, role, role_name, balance_usd, spent_usd, total_usd}} (total_usd is balance + spent; the older credits field is deprecated). Roles: 1 Trial, 2 Paid, 3 Unrestricted, 4 Admin, 5 Enterprise.
POST /auth/account-login, POST /auth/account-logout, GET /auth/account-status
Separate sign-in for the account pages (email and password). Status returns {authenticated, user_id}.
GET /auth/csrf-token
Public. {csrf_token, csrf_enabled}.
GET /auth/trial-status
Any. Also /api/v1/account/trial-status. {success, is_trial, trial_uploads_remaining, role, role_name, total_runs, completed_runs, failed_runs, total_data_mb, limits}.
GET /auth/user/history
Any. Alias of GET /api/v1/sessions (also /api/auth/user/history).
Upload and dashboard runs
POST /validate-csv
Any. Also /api/v1/validate. Multipart file (CSV, XLSX, XLS, Parquet, JSON, JSONL, NDJSON, Feather, ORC), optional enableNestedFlattening (default true) and flattenColumns (JSON array; absent = detect, [] = none). Returns {columns, shape, missing_values, file_size_mb, validation_level, dataset_info: {rows, columns}, nested_flattening, column_scan, has_audio_column, has_image_column, audio_file_count, image_file_count, pricing_profile_id, pricing_status, warning?, note?}. For the preview, under 50 MB the file is read in full; 50 to 200 MB gives an estimated row count such as "~45,000"; above 200 MB only the header is read. For the price, every row is measured: pricing_status is ready when the file was measured during the request (up to 50 MB) or measuring for a larger file; poll GET /api/pricing/profile/{id} (also /api/v1/pricing/profile/{id}), which returns {status: measuring | ready | failed, progress_pct, exact, rows, columns}. No run is started. 400 No file provided / Invalid file; 403 Trial file over 1 MB.
POST /estimate-job
Any. Body rows, cols, file_size_mb, tools {anomaly, dtc, mdh, cds, dsm, dsg}, output_number, dsg_mode, mode, params, touched_fields, decisions, form_params, pricing_profile_id (from /validate-csv), media_bytes {image, audio}, completion. Returns quote: for a measured file the exact price {quote_id, exact: true, price_usd, breakdown [{item, usd}], drivers, rate_card_version} (its low, high and point figures all equal that price), otherwise an estimated range (low_usd, high_usd, point_usd); while the file is being measured quote is null with pricing_status: "measuring" and pricing_progress_pct. Also per-preset estimates (estimated_seconds, estimated_usd; the price is the same for every preset), price_note (the explanation text), balance_usd, can_afford (whether the price, or the top of a range, fits your balance; a warning only), and what the chosen preset would change (mode_groups, mode_adjustments, mode_contested, policy_signature). Deprecated, still returned: price_credits, low_credits, high_credits, point_credits, estimated_credits, credits_note, credits_remaining.
POST /process-data
Any (the dashboard route; scripts should prefer /api/v1/process). A run started here with an API key is owned by and billed to the key's account, and the key's spending limit and stage permissions apply. Multipart. Booleans are "true"/"false"; lists and objects are JSON strings.
| Field | Type | Default | Meaning |
|---|---|---|---|
file | file, required | The data file | |
yColumns | list, required | Target columns, e.g. ["churn"] | |
taskTypes | list | c per target | c classification, r regression |
outputNumber | integer | 10000 | Target row count; empty skips synthetic generation |
excludedColumns | list | [] | Columns left out |
sessionName, projectDescription | string | Name and free-text context | |
idempotencyKey (or header Idempotency-Key) | string | A repeat with the same key returns the first run's session instead of starting another (best-effort) | |
shared, retraining | boolean | false | Recorded with the run |
webhookIds | list | [] = all subscribed | Webhooks notified when the run ends |
schema, ranges | object | Optional declared schema and ranges, stored with the run | |
quote_id | string | The exact price confirmed through /estimate-job; the run is charged exactly that. 409 (price_changed, file_changed, quote_expired) when the file or a price-affecting setting changed, or after 7 days | |
enableAnomalyDetection | boolean | false | Anomaly stage |
enableDTC, enableMDH, enableCDS, enableDSM, enableDSG | boolean | true | Stage switches |
enableDOR | boolean | false | Accepted; outlier-row removal runs only in retraining |
enableFeatureSelection, enableResultEvaluator | boolean | false | Optional stages |
enablePreliminaryCleaning, enableNestedFlattening | boolean | true | Cleaning and nested JSON expansion |
flattenColumns | list | detect | Columns to flatten |
outputPreferences | object | saved | Output files, stage overrides, output_format, output_name_pattern, model-ready split |
textMode | 0 to 3 | 2 | 0 none, 1 neural tokens, 2 TF-IDF, 3 automatic |
textCleaning | boolean | true | Reserved; not available yet |
datetimeMode | string | basic | none, minimal, basic, full |
audioMode, imageMode | 0 to 3 | 0 | Media encoding; 0 off |
audioColumns, imageColumns | list | [] | Media filename columns |
imageFiles, soundFiles | file, repeatable | Media files (jpg, jpeg, png, gif, bmp, webp, tiff, svg / mp3, wav, ogg, flac, aac, m4a, wma) | |
anomalyParams | object | saved | The 14 anomaly toggles |
completionParams | object | {} | Data Completion & Verification settings |
mdhMode | integer | 0 | 0 imputation, 1 2D removal, 2 imputation/dropping (Enterprise) |
cdsScaler | string | auto | Scaler override |
similarityP | number | 0.99 | Column-pruning similarity threshold (legacy dataSimilarityThreshold) |
zscoreLimit | number | 3 | Outlier bound for synthetic values |
dsgMode | string | copula | Only gan is acted on |
forceSynthetic, dsgAllowDownsample | boolean | false | Force synthetic rows; allow shrinking to a smaller target |
dsgYOversampleThreshold | number | off | Class balancing target, between 0 and 1 |
dsgYNumBins | integer | 10 | 2 to 100 |
dsgRareOversampleThreshold | number | off | 0 to 0.5 |
featureSelectionConfig | object | {} | {method, top_k, threshold} |
mode, modeOverrides, modeDecisions, touchedFields | string / object / list | balanced | Speed preset and your answers to it |
modelFamilies, imputationOverrides, scalingOverrides, validationPolicy, pivotConfig, selectionConfig, probeConfig, datasetMetadata | object | Enterprise only; removed for other plans |
Returns {success, session_id, status: "processing", message, heartbeat_interval_seconds, idle_timeout_seconds, dataset_info, warning?}. 400 missing file, bad yColumns, bad file type, targets not in the file, bad balancing value; 403 Trial limits; 500.
POST /upload-images, POST /upload-sounds
Any. Multipart files (repeatable). Returns {success, session_id, files: [{filename, size, size_mb, path}], count}.
POST /api/v1/media/validate
Any (also /validate-media-files). Checks media file names against a dataset; nothing is uploaded or started. Multipart csv, mediaType (audio | image), uploadedFilenames (repeatable), optional audioColumns / imageColumns. Returns {valid, column_found, uploaded_files_without_match, total_uploaded, total_matched, message}.
Job API (v1)
Run modes: full_pipeline (default, trains a new session), inference and retraining (replay a completed session given as base_session_id, taking its targets and settings). Send run_mode; a top-level mode counts as the run mode only when its value is one of these three.
POST /api/v1/process
Key. Multipart file, config (JSON string), optional repeatable imageFiles / soundFiles. Optional header Idempotency-Key (see below).
config accepts the full configuration vocabulary. The request shape {"target_columns", "output_rows", "tools": {...}, "advanced_params": {...}} keeps its meaning, and every pipeline setting the dashboard offers (the /process-data fields above) may also be sent at the top level of config or inside advanced_params, in snake_case or camelCase. The tables list the snake_case names.
| config key | Type | Default | Meaning |
|---|---|---|---|
target_columns | string or list | required for training | Targets |
task_types | list or string | c per target | c classification or r regression, one per target; one string applies to every target; auto detects the type per target. A wrong length or value answers 400 |
output_rows | integer | 10000 | Target rows; "" skips synthetic generation |
run_mode, base_session_id | string | full_pipeline | Run mode. A top-level mode on this route is the run mode too, and is never read as the speed preset |
speed_mode, mode_overrides | string / object | balanced | Speed preset (rapid, balanced, accurate; advanced_params.mode is accepted as well) and your own values for the parameters a preset controls |
tools | object | anomaly (false), dtc, mdh, cds, dsm, dsg (true), dor (false, retraining only), feature_selection (false), preliminary_cleaning (true), nested_flattening (true), result_evaluator (false) | |
enable_anomaly_detection, enable_dtc, enable_mdh, enable_dor, enable_cds, enable_dsm, enable_dsg, enable_feature_selection, enable_preliminary_cleaning, enable_nested_flattening, enable_result_evaluator | boolean | as tools | The same stage switches by their own names |
anomaly_params | object | saved cleaning defaults for API runs (/api/v1/preferences/api-anomaly) | Data-cleaning switches for this run (the 14 toggles) |
completion_params | object | {} | Data Completion & Verification settings |
output_preferences | object | saved output preferences | Output selection for this run: the object GET /api/v1/preferences/outputs returns (*_output file switches, *_use_tool stage switches, model_ready_split none | xy | train_test | xy_train_test, model_ready_test_size, model_ready_real_test_only, output_format, output_name_pattern). model_ready_split, model_ready_test_size and model_ready_real_test_only are also accepted directly. What the request names itself (tools, enable_*, output_format) takes precedence over this object |
feature_selection_config, flatten_columns | object / list | Also accepted in advanced_params. flatten_columns: absent = detect, [] = flatten nothing | |
quote_id | string | An exact quote from /api/v1/estimate (also in advanced_params); the run is charged exactly that price, provided it is the same file (checked by content) and the settings price the same, else 409 | |
session_name, prompt_text, webhook_ids | string / string / list | Run name, free-text context, and the webhooks notified when the run ends | |
force_synthetic, retraining | boolean | false | Force synthetic rows; recorded with the run |
auto_sink, run_dsg, run_dor, dor_eps, dor_min_samples | various | Replay runs only | |
model_families, imputation_overrides, scaling_overrides, dataset_metadata | various | Enterprise only |
| advanced_params key | Default | Meaning |
|---|---|---|
excluded_columns | [] | Columns left out |
text_mode | 2 | 0 to 3, as textMode |
datetime_mode | basic | none, minimal, basic, full |
mdh_mode | 0 | 0 imputation, 1 2D removal, 2 imputation/dropping (Enterprise) |
cds_scaler | auto | Scaler override |
similarity_p | 0.99 | Column-pruning threshold |
zscore_limit | 3.0 | Synthetic outlier bound |
dsg_mode | copula | copula or gan |
dsg_allow_downsample, dsg_y_oversample_threshold, dsg_y_num_bins, dsg_rare_oversample_threshold | false, off, 10, off | Synthetic sizing and balancing |
image_mode, audio_mode, image_columns, audio_columns | 0, 0, [], [] | Media |
output_format | saved (csv) | csv, parquet, xlsx, json, jsonl, ndjson, feather, orc |
output_name_pattern | prefixed | prefixed or plain |
llm_enabled, strict_llm | true, false | Whether stages may call a language model |
mode, mode_ack, mode_decisions, touched_fields | balanced | Speed preset and your answers |
dtc_OHElimit, dtc_sample_rows_detect_text, dtc_sample_rows_text_mode, dtc_seq_max_features, dtc_tfidf_max_features, dtc_datetime_mode, text_synth_method, image_synth_method, audio_synth_method, anomaly_force_no_llm, anomaly_sample_scale, mdh_nooftrials, imputer_skip_expensive, imputer_force_all, cds_eval_top_n, dsg_iess_budget_multiplier | set by the preset | Preset-controlled parameters you can pin; labels and ranges come from /api/v1/mode-preview |
validation_policy, pivot_config, selection_config, probe_config | Enterprise only |
Training: 200 {success, session_id, status: "queued", message, config, mode, mode_adjustments, policy_signature}. Replay: 201 {success, session_id, mode, base_session_id, rows_read}. 409 mode_confirmation_required with mode_contested and policy_signature, or 409 {error, code, quote_id} with code price_changed, file_changed or quote_expired (estimate again and resend). 400 with a message for a missing target, a task_types list that does not match the targets, or a non-numeric zscore_limit, similarity_p, text_mode or mdh_mode. 401, 403, 429, 500.
Idempotency-Key. Send any unique string in the Idempotency-Key header to make a submission safe to retry. A retry with the same key from the same caller returns the first request's session, {success, session_id, status: "queued", message, idempotent_replay: true}, and no second run starts. A request that was answered with an error does not keep its key.
POST /api/v1/process-connector
Any (also /dashboard/api/runs). JSON: source_type (required), credential_id or inline_secrets, table or custom_query, incremental + watermark_column (read only new rows; the position advances when the run completes), source_key, type_casts ({"col": "int|float|string|bool|datetime|category"}), y_columns (required for training), run_mode, base_session_id, speed_mode, output_number, plus every pipeline setting in snake_case (enable_*, text_mode, mdh_mode, similarity_p, output_preferences, output_format, llm_enabled, ...) and webhook_ids, auto_sink. Returns {success, run: {session_id}, session_id, message}; replay 201 with rows_read; session_id: null when an incremental read found nothing new. 400, 403, 404 Credential not found, 429.
POST /api/v1/mode-preview
Any. Body mode, advanced_params, touched_fields, decisions. Returns {success, mode, mode_groups, mode_adjustments, mode_contested, policy_signature, knob_specs}.
POST /api/v1/pipeline-config/decide
Any (also /api/pipeline-config/decide). Proposes a configuration for a dataset (the dashboard's "Decide everything"); nothing is stored or started. Body columns (required: the column names), y_columns, excluded_columns, rows, profile_columns (the per-column profile from validation) and, for Enterprise accounts, dataset_metadata. Returns {success, tier, config, decided, needs}: the proposed settings, what was chosen for you and what is still yours to choose; Enterprise accounts also get reasons and warnings. 400 when columns is missing or empty.
POST /api/v1/estimate
Any. Runs and charges nothing. Body: pricing_profile_id (from validation) for the exact price, plus tools {anomaly, dtc, mdh, cds, dsm, dsg}, output_rows, text_mode, media_bytes {image, audio}, completion (dataset completion and validation settings, priced per cell) and session_type (api, inference, retraining). Returns {quote_id, exact: true, price_usd, breakdown [{item, usd}], drivers, rate_card_version, session_type, pricing_status: "ready"}; while the file is still being measured, {pricing_status: "measuring", progress_pct}. Send quote_id as config.quote_id on /api/v1/process with the same file and the run is charged exactly price_usd; a different file, settings that price differently, or a quote older than 7 days answer 409 and nothing starts.
Without a profile, a description: input_rows and columns (required), column_types (counts of numeric, boolean, categorical_low, categorical_high, datetime, text, identifier), missing_rate (0 to 1, or a percentage). Returns an estimated range {quote_id, exact: false, low_usd, high_usd, point_usd, breakdown, rate_card_version}, which binds and caps nothing. The deprecated price_credits, low_credits, high_credits and point_credits are still returned, equal to the dollar amount × 1,000,000, and will be removed in a future version. 400 when input_rows or columns is missing, or the profile is unknown, expired or could not be measured.
POST /api/pricing/estimate
Public (no account; the pricing page calculator). The same description fields, plus mode (preparation, inference, retraining) and enterprise (boolean). Returns the range only, with no quote_id; it binds nothing.
Status, results and downloads
GET /progress/{session_id}
Any. {status, step, current_step, total_steps, message, stage_key, stage_label, duration_seconds, estimated_remaining, estimated_total, timestamp, queue_info, warnings, substep, last_heartbeat, error_info?}. Status: queued, processing, completed, error, failed, cancelling, cancelled. total_steps is 10 plus one each for Data Completion & Verification, feature selection and the PDF report when enabled. Error codes include COLUMN_NOT_FOUND, DATA_TYPE_CONFLICT, NO_DATA, OUT_OF_MEMORY, TARGET_INVALID, TIMEOUT, TOO_FEW_ROWS, WORKER_DIED, PROCESSING_FAILED.
POST /heartbeat/{session_id}
Any. {ok, ts}. Optional: runs are not cancelled when heartbeats stop.
GET /result/{session_id}
Any. The full result: status, message, files [{name, path, size, description, storage}], row_count, output_shape, output_format, output_format_warnings, warning, warnings, dsg_provenance, model_ready, quality_report, column_verdicts, metadata_report, audio_decode_report, image_encode_report, dsm_column_drops, mode_policy, processing_time_seconds, completed_at, and on a completed run billing {amount_charged_usd, quote_low_usd, quote_high_usd, list_usd, capped, rate_card_version}: the amount charged, in US dollars; for a run started with a confirmed price, low and high both equal that price. The deprecated credits_charged, quote_low_credits, quote_high_credits and list_credits are still returned. The first fetch of a finished run records history, settles billing and sends webhooks and notification emails.
GET /download/{session_id}/{filename}
Any. Serves a disk output or a memory-held output. Share recipients may download the shared files. 400, 403, 404 (File not found or an empty file).
POST /cancel/{session_id}
Any. {ok, status: "cancelled"} or {ok, message: "No active task for this session"}. The id active cancels your in-flight synchronous inference or retraining call.
GET /api/v1/status/{session_id}
Key. {session_id, status, message, current_step, total_steps, progress_percent, duration_seconds, error_info?, substep?, run_warnings?}.
GET /api/v1/result/{session_id}
Key. {session_id, status, files [{name, url, size, description}], row_count, duration_seconds, amount_charged_usd, balance_usd} plus, when relevant, warning, message, dsg_provenance, audio_decode_report, image_encode_report, dsm_column_drops, warnings, media_staging, balance_warning (text in dollars, set when the balance is negative), llm. The deprecated credits_charged, credits_remaining and credit_warning are still returned. Before completion: {session_id, status, message}.
GET /api/v1/download/{session_id}/{filename}
Key. One disk output as an attachment. 400, 403, 404.
POST /api/v1/download-archive/{session_id}
Key. Optional body {"files": ["dsg_output", "dsm_output.csv"]} (exact name, stem or stage name); omit for all outputs. Returns a ZIP; the input is never included.
POST /api/v1/cancel/{session_id}
Key. {ok, cancelled}.
Inference and retraining
POST /api/v1/inference
Any (also /inference). Synchronous. Multipart file + original_session_id (required) + optional enableAnomalyDetection, imageFiles, soundFiles; or JSON {original_session_id, source_type, credential_id or inline_secrets, table or custom_query}. Returns {status: "success", new_session_id, output_files, transformations_applied, schema_validation, warnings, metadata, amount_charged_usd, balance_usd}, plus balance_warning when the balance is negative (credits_charged, credits_remaining and credit_warning are deprecated). A failed run is 400 {status: "error", message, error} (SESSION_NOT_FOUND, MISSING_ARTIFACTS, SCHEMA_MISMATCH, FILE_READ_ERROR, JOB_SUBMIT_FAILED, WORKER_ERROR); after 25 minutes error: TIMEOUT with the session id, and the run continues. 404 parent not found or not yours. A run started with an API key is recorded against that key's usage; the same holds for retraining.
POST /api/v1/retraining
Any (also /retraining). Inference fields plus run_dor (true), dor_eps (automatic), dor_min_samples (5), run_dsg (true), output_sample_size; run_dsm is accepted and ignored. Outputs features.csv and y_columns.csv.
POST /api/v1/inference-retraining
Removed: 410 ENDPOINT_REMOVED.
Sessions
GET /api/v1/sessions
Any. Query selectable_only, q, status, tag, date_from, date_to. {success, history: [{id, session_id, session_name, retraining_enabled, is_retrainable, original_filename, total_rows, total_columns, file_size_bytes, file_size_mb, processing_date, status, created_at, session_type, tags, note}], total}.
GET /api/v1/sessions/{id}
Any (also /api/session/{id}/details, /session/{id}/details). History fields plus results_available, media, processing_parameters, steps, total_duration_seconds, enterprise; share recipients get a viewer view with permission and capabilities.
PUT /api/v1/sessions/{id}/name, PUT /api/v1/sessions/{id}/retraining
Any. {"session_name": "..."} (255 characters); {"retraining_enabled": true}.
GET /api/v1/sessions/{id}/config, PUT /api/v1/sessions/{id}/meta, GET /api/v1/sessions/{id}/download-all
Any (also under /api/session/{id}/...). Config returns {session_id, config, run_mode} in /process-data field names. Meta takes tags (up to 10, 32 characters each) and note (500). Download-all returns every output as a ZIP. 404 when not yours.
GET /api/v1/sessions/{id}/report
Any. {session_id, tier, status, summary, warnings, details?}; details (full tier) for Enterprise accounts.
POST /api/v1/sessions/{id}/retry
Any (also /dashboard/api/sessions/{id}/retry). Re-runs from the last completed stage with the configuration the run was started with (targets, task types, outputs, cleaning and speed preset).
GET / DELETE /api/v1/sessions/{id}/outputs, GET /api/v1/sessions/{id}/outputs/{name}
Any (also /api/memory-outputs/...). Memory-held outputs (about 4 hours); download with ?format= (csv default; json returns rows inline).
GET /api/v1/sessions/{id}/files
Any (also /api/sessions/{id}/files). The downloadable files of a run (shared files only for recipients): {success, files: [{name, size_mb, url}]}. 404 when the session is neither yours nor shared with you.
POST /api/v1/sessions/{id}/write-output
Any (also /dashboard/api/sessions/{id}/write-output). Body connector_type (any connector type except ga4; defaults to the saved connection's type, else sql), credential_id or inline_secrets, table_name (required), if_exists (replace, append, fail, merge; left out, the destination's default: replace where offered, else append), merge_keys (required for merge), output_stage (dsg), file_format (file destinations: csv, parquet, xlsx, json, jsonl, ndjson, feather, orc), partition_by (one target per distinct value of a column, up to 1,000), partition_format (%Y-%m-%d), options (destination settings such as time_column or partition_key_column). Returns {success, rows_written, table, message}, with partitions for a partitioned write.
A mode the destination does not offer is a 400 that lists the supported ones; fail on an existing target is a 409. Salesforce, HubSpot, NetSuite, SharePoint lists, Stripe, PI Web API and OPC UA are written to only through a saved connection with Allow writing switched on (allow_write); otherwise 403. When a destination accepts some rows and declines others, the answer reports rows_written, rows_failed and errors. See Write Output & Auto-Sink for each destination's modes, options and row limits.
Sharing and comments
Also without /v1. Permissions: read, comment, edit.
POST /api/v1/shared-sessions
Any. session_id (required), shared_with_id (omit for a public link), permission (read), expires_at (ISO-8601, future), note, shared_files. 201 {success, shared_session}.
GET /api/v1/shared-sessions, GET /api/v1/shared-sessions/received, PUT / DELETE /api/v1/shared-sessions/{share_id}
Any. List outbound and received shares; change or revoke. Changing or revoking a link applies to everyone who joined it.
GET /api/v1/shared-sessions/access/{token}, GET .../access/{token}/download/{filename}
Public. The link's session, files and viewer state; download a shared file. 404 unknown, 410 expired.
POST /api/v1/shared-sessions/access/{token}/join
Any. Adds the link's grant to your account.
GET / POST /api/v1/shared-sessions/by-session/{id}/comments, DELETE .../comments/{comment_id}
Any. Comments up to 2000 characters; posting needs comment permission.
GET /api/v1/users/search
Any (also /api/users/search). Finds a share recipient: ?q= exact email or a known collaborator's name. Returns {users: [{id, name, email_hint}]} with masked emails; pass id as shared_with_id. 429 above 20 searches per minute.
Connections
Connector types (38): sql, snowflake, bigquery, mongodb, s3, gcs, azure_blob, adls_gen2, databricks, delta, fabric, synapse, kafka, kinesis, mqtt, opcua, oracle, sap_hana, clickhouse, redshift, timescaledb, cassandra, dynamodb, elasticsearch, salesforce, hubspot, stripe, google_sheets, influxdb, ga4, netsuite, pi_web_api, sharepoint, http, google_drive, dropbox, onedrive, box. Also without /v1.
GET / POST /api/v1/credentials, PUT / DELETE /api/v1/credentials/{id}
Any. Create with name, connector_type, secrets (all required; stored encrypted, never returned). PUT takes name, secrets, merge_secrets and clear_secrets: it replaces the secrets unless merge_secrets: true, and clear_secrets lists stored field names to remove. The list shows non-secret hints, the last test result, stored_fields (the names of the fields that are set, never their values) and write: what the connection can do as a destination (modes, target_kind, formats, options, guarded, max_rows, enabled).
POST /api/v1/connectors/{type}/test | discover | preview
Any. Body credential_id or inline_secrets; preview also table or a read-only custom_query (a single SELECT or WITH). When the cause of a failure is recognised, the answer adds error_category (auth, permission, not_found, tls, timeout, network) and a plain-language error_hint.
Data tools
POST /api/v1/mapping/suggest, POST /api/v1/mapping/apply
Any. Suggest: source_columns, target_columns, threshold (0.4). Apply: session_id, mappings, output_stage (dsg), drop_unmapped (false); returns a preview, nothing is saved.
GET /api/v1/feature-engineering/columns, POST /api/v1/feature-engineering/recommend
Any (also feature-selection spellings and without /v1). Recommend: session_id, target_columns, methods (mutual_information default; also variance_threshold, correlation_pruning, pca, rfe, shap_importance, permutation_importance, leakage_guard), top_k (20). Returns {ranking, per_method, leakage, warnings, ...}.
POST /api/v1/pivot/suggest (Enterprise)
Enterprise. session_id, or columns + rows; returns the suggested as-of column and candidates.
POST /api/v1/dataset-metadata/validate (Enterprise)
Enterprise. A metadata file (2 MB) or JSON {metadata, columns, y_columns}; returns {metadata, warnings, summary}.
Automation
Every automation object takes mode (full_pipeline, inference, retraining) and base_session_id. pipeline_config uses the process-connector keys.
Each automated run is charged the exact price of the data it processes, measured when the run starts. Scheduled runs, triggers and listeners can set an optional max_price_dollars (Maximum price per run) in pipeline_config (a trigger: target_pipeline_config); a run priced above it is held for approval instead of running.
/api/v1/pending-approvals
Any (also /api/pending-approvals). GET /api/v1/pending-approvals lists your held runs: {approvals: [{id, session_id, source, name, price_usd, cap_usd, status, created_at, decided_at}]} (price_dollars and cap_dollars hold the same amounts; price_credits and cap_credits are deprecated). POST /api/v1/pending-approvals/{id}/approve starts the run at the price shown and charges exactly that ({success, session_id, price_usd, approval}); POST /api/v1/pending-approvals/{id}/discard drops it. 404 unknown or not yours; 409 already decided; 410 the held data is gone. A held run sends a job.held webhook event {session_id, source, name, price_dollars, cap_dollars, message} and, if you receive failure emails, an email.
/api/v1/scheduled-runs
Any (also /api/scheduled-runs). GET, POST (201), PUT/DELETE /{id}, POST /{id}/run-now, POST /{id}/toggle. Fields: name (required), description, schedule_type (interval | cron), schedule_value (30m, 1h, 2d, or a cron expression in UTC), connector_type, credential_id, source_table, sql_query, pipeline_config (including y_columns; without them the last column is used; and the optional max_price_dollars), is_active. A held run sets last_run_status to held.
/api/v1/triggers and POST /trigger/inbound/{secret}
Any (also /api/triggers). CRUD plus POST /{id}/test (dry run). trigger_type: inbound_webhook (optional config.hmac_key, 16+ characters), watermark_advance (credential_id, table, watermark_column, min_advance), pipeline_completion (source_template_id or source_session_name), quality_alert (not available yet). target_pipeline_config names the source, targets, run mode and settings, and the optional max_price_dollars. Five consecutive failures disable a trigger. The inbound URL takes a CSV or JSON body up to 64 KB (or reads the target's connector), an X-Signature: t=...,v1=... header when an HMAC key is set (5-minute window), and answers 202 {accepted, session_id}, 400, 401, 404, 410, 413, 429 or 502.
/api/v1/listeners/folder
Any. source_type (local, s3, gcs, azure_blob, adls_gen2, sftp), folder_path, bucket_name, prefix, credential_id, name, y_columns (required for full_pipeline), pipeline_config (including the optional max_price_dollars), poll_interval_seconds (60), allowed_extensions, max_file_size_mb (500), cooldown_seconds (60), max_files_per_hour (10), auto_sink_config, enabled.
/api/v1/sftp
Any. GET /info, GET/POST /credentials (POST returns the SFTP password once), PUT /credentials/{id}/pipeline, DELETE /credentials/{id}. One inbox per account.
/api/v1/sync/watermarks
Any (also /api/sync/watermarks). GET /list lists your watermarks and POST /{id}/reset resets one. On /api/v1/sync/watermarks itself a watermark is addressed by connector_type + source_key: GET (query parameters) reads it, {success, watermark} with watermark: null when none is stored; PUT creates or updates it (table_name, watermark_column, watermark_value, watermark_type, row_count, last_sync_at); POST creates it (201; 409 when one already exists); DELETE (query parameters) removes it (404 when none is stored).
/api/v1/quality-alerts
Any. name, metric (row_loss_pct, null_pct, column_drop_count, duration_seconds, drift_score), operator (>, <, >=, <=, ==), threshold, stage (anomaly, dtc, mdh, cds, dsm, dsg, drift; empty = all), severity (warning), enabled. GET /events lists the last 100 alerts. webhook_ids is stored but alerts are not sent to webhooks.
/api/v1/streams (Enterprise)
Enterprise. Continuous inference between Kafka, Kinesis and MQTT: CRUD, /{id}/start|stop|restart, /{id}/status, /{id}/schema, /test-connection. Billed hourly for the rows scored, at the inference rate. See Streaming for the fields.
/api/v1/webhooks
Any (also /webhooks). url (public addresses only), name, events (error, warning, info, success, job.completed, job.failed, quality.breach; default ["error"]), secret, active. POST /{id}/test sends a test event; GET /{id}/deliveries lists the last 50. Runs send job.completed and job.failed with {event, data: {session_id, status, original_filename, timestamp, files | error}, timestamp, webhook_id}, signed as X-Webhook-Signature: sha256=<hex> when a secret is set. One attempt, 10-second timeout.
Preferences
Any. GET, POST and POST /reset on /api/v1/preferences/outputs, /advanced, /anomaly, /completion and /api-anomaly (legacy: /api/output-preferences, /api/advanced-parameters, /api/anomaly-parameters, /api/completion-parameters, and /api/output-config in the older disk/memory shape). Output preferences hold the *_output file switches, *_use_tool stage switches, output_format (csv), output_name_pattern (prefixed) and the model-ready split.
/api/v1/preferences/anomaly and /api/v1/preferences/api-anomaly
Any. Two separate stores of data-cleaning switches, each with GET, POST and POST /reset, each returning {success, parameters}. /api/v1/preferences/anomaly holds the settings of runs started in the dashboard. /api/v1/preferences/api-anomaly holds the defaults of runs started through the API or the SDK: POST /api/v1/process uses them when a request sends no anomaly_params. Changing one does not change the other.
Account and keys
GET /api/v1/keys, GET /api/v1/usage
Key. /api/v1/keys: your active keys (never secrets), each with daily_limit_usd, lifetime_limit_usd, daily_spent_usd and lifetime_spent_usd; null means no limit. /api/v1/usage: the calling key's spend, {today_usd, daily_limit_usd, daily_remaining_usd, lifetime_spent_usd, lifetime_limit_usd, lifetime_remaining_usd, avg_daily_usd_7d, avg_daily_usd_30d, chart_data: {labels, usd}}. All amounts are US dollars. The credit-named fields (daily_credit_limit, lifetime_credit_limit, daily_credits_used, lifetime_credits_used) are deprecated and still returned. Creating, rotating, editing and revoking keys, team keys included, is done in the dashboard.
GET /api/v1/account/profile, GET /api/v1/account/usage
Any. profile: balance_usd, spent_usd and total_usd (balance + spent), role and limits. usage: usage by stage across all keys, with spent_usd per stage and summary.total_spent_usd. Deprecated, still returned: api_credits_remaining, api_credits_used, api_credits_total, credits_used, total_credits_used.
/api/account/*
Session. Profile (GET/PUT), regenerate passcode, usage statistics, API profile, API sessions, usage by tool, personal keys (create with name, expires_in_days; update daily_limit_usd, lifetime_limit_usd and stage permissions; rotate; revoke) and team keys (with per-key spending limits and permissions, billed to you). Profiles carry balance_usd, spent_usd and total_usd; API sessions and usage by tool carry spent_usd; key usage carries the same _usd fields as /api/v1/usage. The older daily_credit_limit and lifetime_credit_limit are still accepted, as credits. Key allowance: Trial 1, Paid 3, Unrestricted 15, Enterprise 50.
GET /api/v1/account/contracts
Any (also /api/account/contracts). Your pricing contracts with usage against each allowance: {contracts: [...]}, an empty list when the account has none.
GET /api/v1/account/trial-status
Any (also /auth/trial-status). Trial state and the free runs remaining: {success, is_trial, trial_uploads_remaining, role, role_name, total_runs, completed_runs, failed_runs, total_data_mb, limits}.
GET / PUT /api/v1/notification-preferences
Any (also /api/notification-preferences). Email on completion or failure: {email_on_complete, email_on_fail, configured}. PUT takes either field and changes only the ones sent; 400 when a value is not a boolean.
GET / POST /api/v1/pipeline-templates, PUT / DELETE /api/v1/pipeline-templates/{template_id}
Any (also /api/pipeline-templates). Saved configurations, up to 50 per account. GET returns {templates: [...]}. POST takes name, config and optional description and answers 201 {template}; PUT changes any of the three; DELETE answers {ok: true}. 400 for an invalid body or when the limit is reached; 404 for a template that is not yours.
System
GET /health (public; 200 or 503), GET /api/v1/workers/status (any), GET /openapi.yaml and GET /api (public), and /api/v1/spark/status|read|transform for self-hosted installs with Spark enabled (503 otherwise).
Enterprise endpoints (Enterprise)
Read-only run evidence under /api/v1/enterprise/: families, per-run column record, family comparison, dataset metadata, lineage and diff, and an audit report. Other plans get 403 with "required_role": "Enterprise".