Browse documentation
2026-09-28AutoDataOpen in dashboard

REST API Reference

Every AutoData REST endpoint: methods, paths, authentication, request fields, responses and errors.

About this reference

Every customer-facing HTTP endpoint of AutoData: method and path (with alias paths that reach the same handler), accepted authentication, request fields with types and defaults, response shape and errors. For a walkthrough, start with the API integration guide. Base URL: https://autodata.datatoolpack.com. Admin-only routes are not listed.

Authentication labels

LabelAccepts
AnySession cookie, Authorization: Bearer dtpk_..., X-API-Key: dtpk_... or X-Passcode
KeyAuthorization: Bearer dtpk_..., X-API-Key: dtpk_... or X-Passcode (no cookie)
SessionSigned-in browser session only
PublicNo authentication
EnterpriseAny method, Enterprise accounts only (403 otherwise)

X-API-Key is accepted wherever the Bearer form is, including the core job routes (process, status, result, cancel, download, download-archive, keys, usage). API keys are created, rotated, edited and revoked in the dashboard, where the profile and the passcode are changed too; through the API a key can list the account's keys and read its own usage.

Cookie-authenticated requests that change something (POST, PUT, PATCH, DELETE) outside /api/v1/, /trigger/ and the login routes must send X-CSRF-Token from GET /auth/csrf-token; Bearer requests are exempt. Errors are JSON with error and sometimes hint or code. Failed runs carry error_info {code, stage, column, message, hint}. Timestamps are UTC without a zone suffix.

Status codes

CodeMeaning
200 / 201 / 202Success / created or replay queued / inbound trigger accepted
400Invalid or missing field, bad file type, bad JSON, out-of-range value
401Not authenticated or bad signature
403Not your resource, wrong plan, Trial limit, key stage permission, CSRF
404Not found (also unknown /api/ paths)
405Wrong method on a real path (Allow header)
409mode_confirmation_required; a confirmed price that no longer applies (price_changed, file_changed, quote_expired); a held run already approved or discarded; duplicate watermark; email in use
410Removed endpoint, expired share link, disabled trigger, held run whose data is gone
413Upload or body too large
429Rate or spending limit
500 / 502 / 503Server error / a trigger's run failed to start / degraded health or Spark off

Limits and billing

  • Login: 10 attempts per 5 minutes per IP. Inbound trigger: 10 fires per minute. User search: 20 per minute.
  • Per-key daily and lifetime spending limits, in US dollars (429), and stage permissions (403) are checked when a run is started with a key through /api/v1/process, /api/v1/process-connector or /process-data.
  • Trial: 3 runs, 1 MB per data file, 1 MB of images and 1 MB of audio per request. Paid plans: 5 GB per request.
  • Money is US dollars everywhere: balances, charges, prices and key limits. A run is charged when it finishes, in dollars rounded to the cent; a run priced above $0 is charged at least $0.01. A run is priced on what is in the data, not on the file size: rows (the price per row falls past 1M, 5M, 10M, 50M, 100M and 1B rows, with a minimum price per row), columns and their kinds, missing values, differences in numeric scale, the number of features and the stages that run; synthetic rows are charged per doubling of the row count, image and audio bytes per GB, dataset completion and validation per cell. Inference costs about a fifth and retraining about two fifths of preparing the same data; Enterprise 2x. See Account and billing.
  • Money fields end in _usd. The older *_credits fields and other credit-named fields (such as credits_charged, credits_remaining, credit_warning, daily_credit_limit) are deprecated: still returned, equal to the dollar amount × 1,000,000, and will be removed in a future version. Existing *_dollars fields are unchanged and hold the same amounts as their _usd counterparts.
  • An uploaded file is measured in full and priced exactly: /api/v1/estimate with its pricing_profile_id returns one price and a quote_id, and a run started with that quote_id is charged exactly that price (409 price_changed, file_changed or quote_expired when it no longer applies). A description without a file gets an estimated range, which binds and caps nothing. Runs without a quote_id, and automated runs, are charged the exact price of the data they process, measured when they start; automations can set max_price_dollars to hold pricier runs for approval. Streams are billed hourly for rows scored, at the inference rate. A low balance never blocks a run; the balance can go negative.

Authentication and session

POST /auth/login

Public. Body {"passcode": "123456789012"}. Returns {success, message, user: {id, name}} and sets the session cookie (24 hours). 400 no body, 401 Invalid passcode, 429.

POST /auth/email-login

Public. Body email, password. Returns {success, message, user: {id, name, email}}. 400, 401 Invalid email or password, 429.

POST /auth/signup

Public. Body email, name, surname (each up to 100 characters), password (8+ characters with upper case, lower case and a digit). Creates a Trial account (3 runs, a $0.10 balance, one API key), signs it in and returns the login passcode once. 400 validation or existing email.

POST /auth/logout

Public. Clears the session. {success, message}.

GET /auth/status

Session. {authenticated, admin_authenticated, account_authenticated, user: {id, name, role, role_name, balance_usd, spent_usd, total_usd}} (total_usd is balance + spent; the older credits field is deprecated). Roles: 1 Trial, 2 Paid, 3 Unrestricted, 4 Admin, 5 Enterprise.

POST /auth/account-login, POST /auth/account-logout, GET /auth/account-status

Separate sign-in for the account pages (email and password). Status returns {authenticated, user_id}.

GET /auth/csrf-token

Public. {csrf_token, csrf_enabled}.

GET /auth/trial-status

Any. Also /api/v1/account/trial-status. {success, is_trial, trial_uploads_remaining, role, role_name, total_runs, completed_runs, failed_runs, total_data_mb, limits}.

GET /auth/user/history

Any. Alias of GET /api/v1/sessions (also /api/auth/user/history).

Upload and dashboard runs

POST /validate-csv

Any. Also /api/v1/validate. Multipart file (CSV, XLSX, XLS, Parquet, JSON, JSONL, NDJSON, Feather, ORC), optional enableNestedFlattening (default true) and flattenColumns (JSON array; absent = detect, [] = none). Returns {columns, shape, missing_values, file_size_mb, validation_level, dataset_info: {rows, columns}, nested_flattening, column_scan, has_audio_column, has_image_column, audio_file_count, image_file_count, pricing_profile_id, pricing_status, warning?, note?}. For the preview, under 50 MB the file is read in full; 50 to 200 MB gives an estimated row count such as "~45,000"; above 200 MB only the header is read. For the price, every row is measured: pricing_status is ready when the file was measured during the request (up to 50 MB) or measuring for a larger file; poll GET /api/pricing/profile/{id} (also /api/v1/pricing/profile/{id}), which returns {status: measuring | ready | failed, progress_pct, exact, rows, columns}. No run is started. 400 No file provided / Invalid file; 403 Trial file over 1 MB.

POST /estimate-job

Any. Body rows, cols, file_size_mb, tools {anomaly, dtc, mdh, cds, dsm, dsg}, output_number, dsg_mode, mode, params, touched_fields, decisions, form_params, pricing_profile_id (from /validate-csv), media_bytes {image, audio}, completion. Returns quote: for a measured file the exact price {quote_id, exact: true, price_usd, breakdown [{item, usd}], drivers, rate_card_version} (its low, high and point figures all equal that price), otherwise an estimated range (low_usd, high_usd, point_usd); while the file is being measured quote is null with pricing_status: "measuring" and pricing_progress_pct. Also per-preset estimates (estimated_seconds, estimated_usd; the price is the same for every preset), price_note (the explanation text), balance_usd, can_afford (whether the price, or the top of a range, fits your balance; a warning only), and what the chosen preset would change (mode_groups, mode_adjustments, mode_contested, policy_signature). Deprecated, still returned: price_credits, low_credits, high_credits, point_credits, estimated_credits, credits_note, credits_remaining.

POST /process-data

Any (the dashboard route; scripts should prefer /api/v1/process). A run started here with an API key is owned by and billed to the key's account, and the key's spending limit and stage permissions apply. Multipart. Booleans are "true"/"false"; lists and objects are JSON strings.

FieldTypeDefaultMeaning
filefile, requiredThe data file
yColumnslist, requiredTarget columns, e.g. ["churn"]
taskTypeslistc per targetc classification, r regression
outputNumberinteger10000Target row count; empty skips synthetic generation
excludedColumnslist[]Columns left out
sessionName, projectDescriptionstringName and free-text context
idempotencyKey (or header Idempotency-Key)stringA repeat with the same key returns the first run's session instead of starting another (best-effort)
shared, retrainingbooleanfalseRecorded with the run
webhookIdslist[] = all subscribedWebhooks notified when the run ends
schema, rangesobjectOptional declared schema and ranges, stored with the run
quote_idstringThe exact price confirmed through /estimate-job; the run is charged exactly that. 409 (price_changed, file_changed, quote_expired) when the file or a price-affecting setting changed, or after 7 days
enableAnomalyDetectionbooleanfalseAnomaly stage
enableDTC, enableMDH, enableCDS, enableDSM, enableDSGbooleantrueStage switches
enableDORbooleanfalseAccepted; outlier-row removal runs only in retraining
enableFeatureSelection, enableResultEvaluatorbooleanfalseOptional stages
enablePreliminaryCleaning, enableNestedFlatteningbooleantrueCleaning and nested JSON expansion
flattenColumnslistdetectColumns to flatten
outputPreferencesobjectsavedOutput files, stage overrides, output_format, output_name_pattern, model-ready split
textMode0 to 320 none, 1 neural tokens, 2 TF-IDF, 3 automatic
textCleaningbooleantrueReserved; not available yet
datetimeModestringbasicnone, minimal, basic, full
audioMode, imageMode0 to 30Media encoding; 0 off
audioColumns, imageColumnslist[]Media filename columns
imageFiles, soundFilesfile, repeatableMedia files (jpg, jpeg, png, gif, bmp, webp, tiff, svg / mp3, wav, ogg, flac, aac, m4a, wma)
anomalyParamsobjectsavedThe 14 anomaly toggles
completionParamsobject{}Data Completion & Verification settings
mdhModeinteger00 imputation, 1 2D removal, 2 imputation/dropping (Enterprise)
cdsScalerstringautoScaler override
similarityPnumber0.99Column-pruning similarity threshold (legacy dataSimilarityThreshold)
zscoreLimitnumber3Outlier bound for synthetic values
dsgModestringcopulaOnly gan is acted on
forceSynthetic, dsgAllowDownsamplebooleanfalseForce synthetic rows; allow shrinking to a smaller target
dsgYOversampleThresholdnumberoffClass balancing target, between 0 and 1
dsgYNumBinsinteger102 to 100
dsgRareOversampleThresholdnumberoff0 to 0.5
featureSelectionConfigobject{}{method, top_k, threshold}
mode, modeOverrides, modeDecisions, touchedFieldsstring / object / listbalancedSpeed preset and your answers to it
modelFamilies, imputationOverrides, scalingOverrides, validationPolicy, pivotConfig, selectionConfig, probeConfig, datasetMetadataobjectEnterprise only; removed for other plans

Returns {success, session_id, status: "processing", message, heartbeat_interval_seconds, idle_timeout_seconds, dataset_info, warning?}. 400 missing file, bad yColumns, bad file type, targets not in the file, bad balancing value; 403 Trial limits; 500.

POST /upload-images, POST /upload-sounds

Any. Multipart files (repeatable). Returns {success, session_id, files: [{filename, size, size_mb, path}], count}.

POST /api/v1/media/validate

Any (also /validate-media-files). Checks media file names against a dataset; nothing is uploaded or started. Multipart csv, mediaType (audio | image), uploadedFilenames (repeatable), optional audioColumns / imageColumns. Returns {valid, column_found, uploaded_files_without_match, total_uploaded, total_matched, message}.

Job API (v1)

Run modes: full_pipeline (default, trains a new session), inference and retraining (replay a completed session given as base_session_id, taking its targets and settings). Send run_mode; a top-level mode counts as the run mode only when its value is one of these three.

POST /api/v1/process

Key. Multipart file, config (JSON string), optional repeatable imageFiles / soundFiles. Optional header Idempotency-Key (see below).

config accepts the full configuration vocabulary. The request shape {"target_columns", "output_rows", "tools": {...}, "advanced_params": {...}} keeps its meaning, and every pipeline setting the dashboard offers (the /process-data fields above) may also be sent at the top level of config or inside advanced_params, in snake_case or camelCase. The tables list the snake_case names.

config keyTypeDefaultMeaning
target_columnsstring or listrequired for trainingTargets
task_typeslist or stringc per targetc classification or r regression, one per target; one string applies to every target; auto detects the type per target. A wrong length or value answers 400
output_rowsinteger10000Target rows; "" skips synthetic generation
run_mode, base_session_idstringfull_pipelineRun mode. A top-level mode on this route is the run mode too, and is never read as the speed preset
speed_mode, mode_overridesstring / objectbalancedSpeed preset (rapid, balanced, accurate; advanced_params.mode is accepted as well) and your own values for the parameters a preset controls
toolsobjectanomaly (false), dtc, mdh, cds, dsm, dsg (true), dor (false, retraining only), feature_selection (false), preliminary_cleaning (true), nested_flattening (true), result_evaluator (false)
enable_anomaly_detection, enable_dtc, enable_mdh, enable_dor, enable_cds, enable_dsm, enable_dsg, enable_feature_selection, enable_preliminary_cleaning, enable_nested_flattening, enable_result_evaluatorbooleanas toolsThe same stage switches by their own names
anomaly_paramsobjectsaved cleaning defaults for API runs (/api/v1/preferences/api-anomaly)Data-cleaning switches for this run (the 14 toggles)
completion_paramsobject{}Data Completion & Verification settings
output_preferencesobjectsaved output preferencesOutput selection for this run: the object GET /api/v1/preferences/outputs returns (*_output file switches, *_use_tool stage switches, model_ready_split none | xy | train_test | xy_train_test, model_ready_test_size, model_ready_real_test_only, output_format, output_name_pattern). model_ready_split, model_ready_test_size and model_ready_real_test_only are also accepted directly. What the request names itself (tools, enable_*, output_format) takes precedence over this object
feature_selection_config, flatten_columnsobject / listAlso accepted in advanced_params. flatten_columns: absent = detect, [] = flatten nothing
quote_idstringAn exact quote from /api/v1/estimate (also in advanced_params); the run is charged exactly that price, provided it is the same file (checked by content) and the settings price the same, else 409
session_name, prompt_text, webhook_idsstring / string / listRun name, free-text context, and the webhooks notified when the run ends
force_synthetic, retrainingbooleanfalseForce synthetic rows; recorded with the run
auto_sink, run_dsg, run_dor, dor_eps, dor_min_samplesvariousReplay runs only
model_families, imputation_overrides, scaling_overrides, dataset_metadatavariousEnterprise only
advanced_params keyDefaultMeaning
excluded_columns[]Columns left out
text_mode20 to 3, as textMode
datetime_modebasicnone, minimal, basic, full
mdh_mode00 imputation, 1 2D removal, 2 imputation/dropping (Enterprise)
cds_scalerautoScaler override
similarity_p0.99Column-pruning threshold
zscore_limit3.0Synthetic outlier bound
dsg_modecopulacopula or gan
dsg_allow_downsample, dsg_y_oversample_threshold, dsg_y_num_bins, dsg_rare_oversample_thresholdfalse, off, 10, offSynthetic sizing and balancing
image_mode, audio_mode, image_columns, audio_columns0, 0, [], []Media
output_formatsaved (csv)csv, parquet, xlsx, json, jsonl, ndjson, feather, orc
output_name_patternprefixedprefixed or plain
llm_enabled, strict_llmtrue, falseWhether stages may call a language model
mode, mode_ack, mode_decisions, touched_fieldsbalancedSpeed preset and your answers
dtc_OHElimit, dtc_sample_rows_detect_text, dtc_sample_rows_text_mode, dtc_seq_max_features, dtc_tfidf_max_features, dtc_datetime_mode, text_synth_method, image_synth_method, audio_synth_method, anomaly_force_no_llm, anomaly_sample_scale, mdh_nooftrials, imputer_skip_expensive, imputer_force_all, cds_eval_top_n, dsg_iess_budget_multiplierset by the presetPreset-controlled parameters you can pin; labels and ranges come from /api/v1/mode-preview
validation_policy, pivot_config, selection_config, probe_configEnterprise only

Training: 200 {success, session_id, status: "queued", message, config, mode, mode_adjustments, policy_signature}. Replay: 201 {success, session_id, mode, base_session_id, rows_read}. 409 mode_confirmation_required with mode_contested and policy_signature, or 409 {error, code, quote_id} with code price_changed, file_changed or quote_expired (estimate again and resend). 400 with a message for a missing target, a task_types list that does not match the targets, or a non-numeric zscore_limit, similarity_p, text_mode or mdh_mode. 401, 403, 429, 500.

Idempotency-Key. Send any unique string in the Idempotency-Key header to make a submission safe to retry. A retry with the same key from the same caller returns the first request's session, {success, session_id, status: "queued", message, idempotent_replay: true}, and no second run starts. A request that was answered with an error does not keep its key.

POST /api/v1/process-connector

Any (also /dashboard/api/runs). JSON: source_type (required), credential_id or inline_secrets, table or custom_query, incremental + watermark_column (read only new rows; the position advances when the run completes), source_key, type_casts ({"col": "int|float|string|bool|datetime|category"}), y_columns (required for training), run_mode, base_session_id, speed_mode, output_number, plus every pipeline setting in snake_case (enable_*, text_mode, mdh_mode, similarity_p, output_preferences, output_format, llm_enabled, ...) and webhook_ids, auto_sink. Returns {success, run: {session_id}, session_id, message}; replay 201 with rows_read; session_id: null when an incremental read found nothing new. 400, 403, 404 Credential not found, 429.

POST /api/v1/mode-preview

Any. Body mode, advanced_params, touched_fields, decisions. Returns {success, mode, mode_groups, mode_adjustments, mode_contested, policy_signature, knob_specs}.

POST /api/v1/pipeline-config/decide

Any (also /api/pipeline-config/decide). Proposes a configuration for a dataset (the dashboard's "Decide everything"); nothing is stored or started. Body columns (required: the column names), y_columns, excluded_columns, rows, profile_columns (the per-column profile from validation) and, for Enterprise accounts, dataset_metadata. Returns {success, tier, config, decided, needs}: the proposed settings, what was chosen for you and what is still yours to choose; Enterprise accounts also get reasons and warnings. 400 when columns is missing or empty.

POST /api/v1/estimate

Any. Runs and charges nothing. Body: pricing_profile_id (from validation) for the exact price, plus tools {anomaly, dtc, mdh, cds, dsm, dsg}, output_rows, text_mode, media_bytes {image, audio}, completion (dataset completion and validation settings, priced per cell) and session_type (api, inference, retraining). Returns {quote_id, exact: true, price_usd, breakdown [{item, usd}], drivers, rate_card_version, session_type, pricing_status: "ready"}; while the file is still being measured, {pricing_status: "measuring", progress_pct}. Send quote_id as config.quote_id on /api/v1/process with the same file and the run is charged exactly price_usd; a different file, settings that price differently, or a quote older than 7 days answer 409 and nothing starts.

Without a profile, a description: input_rows and columns (required), column_types (counts of numeric, boolean, categorical_low, categorical_high, datetime, text, identifier), missing_rate (0 to 1, or a percentage). Returns an estimated range {quote_id, exact: false, low_usd, high_usd, point_usd, breakdown, rate_card_version}, which binds and caps nothing. The deprecated price_credits, low_credits, high_credits and point_credits are still returned, equal to the dollar amount × 1,000,000, and will be removed in a future version. 400 when input_rows or columns is missing, or the profile is unknown, expired or could not be measured.

POST /api/pricing/estimate

Public (no account; the pricing page calculator). The same description fields, plus mode (preparation, inference, retraining) and enterprise (boolean). Returns the range only, with no quote_id; it binds nothing.

Status, results and downloads

GET /progress/{session_id}

Any. {status, step, current_step, total_steps, message, stage_key, stage_label, duration_seconds, estimated_remaining, estimated_total, timestamp, queue_info, warnings, substep, last_heartbeat, error_info?}. Status: queued, processing, completed, error, failed, cancelling, cancelled. total_steps is 10 plus one each for Data Completion & Verification, feature selection and the PDF report when enabled. Error codes include COLUMN_NOT_FOUND, DATA_TYPE_CONFLICT, NO_DATA, OUT_OF_MEMORY, TARGET_INVALID, TIMEOUT, TOO_FEW_ROWS, WORKER_DIED, PROCESSING_FAILED.

POST /heartbeat/{session_id}

Any. {ok, ts}. Optional: runs are not cancelled when heartbeats stop.

GET /result/{session_id}

Any. The full result: status, message, files [{name, path, size, description, storage}], row_count, output_shape, output_format, output_format_warnings, warning, warnings, dsg_provenance, model_ready, quality_report, column_verdicts, metadata_report, audio_decode_report, image_encode_report, dsm_column_drops, mode_policy, processing_time_seconds, completed_at, and on a completed run billing {amount_charged_usd, quote_low_usd, quote_high_usd, list_usd, capped, rate_card_version}: the amount charged, in US dollars; for a run started with a confirmed price, low and high both equal that price. The deprecated credits_charged, quote_low_credits, quote_high_credits and list_credits are still returned. The first fetch of a finished run records history, settles billing and sends webhooks and notification emails.

GET /download/{session_id}/{filename}

Any. Serves a disk output or a memory-held output. Share recipients may download the shared files. 400, 403, 404 (File not found or an empty file).

POST /cancel/{session_id}

Any. {ok, status: "cancelled"} or {ok, message: "No active task for this session"}. The id active cancels your in-flight synchronous inference or retraining call.

GET /api/v1/status/{session_id}

Key. {session_id, status, message, current_step, total_steps, progress_percent, duration_seconds, error_info?, substep?, run_warnings?}.

GET /api/v1/result/{session_id}

Key. {session_id, status, files [{name, url, size, description}], row_count, duration_seconds, amount_charged_usd, balance_usd} plus, when relevant, warning, message, dsg_provenance, audio_decode_report, image_encode_report, dsm_column_drops, warnings, media_staging, balance_warning (text in dollars, set when the balance is negative), llm. The deprecated credits_charged, credits_remaining and credit_warning are still returned. Before completion: {session_id, status, message}.

GET /api/v1/download/{session_id}/{filename}

Key. One disk output as an attachment. 400, 403, 404.

POST /api/v1/download-archive/{session_id}

Key. Optional body {"files": ["dsg_output", "dsm_output.csv"]} (exact name, stem or stage name); omit for all outputs. Returns a ZIP; the input is never included.

POST /api/v1/cancel/{session_id}

Key. {ok, cancelled}.

Inference and retraining

POST /api/v1/inference

Any (also /inference). Synchronous. Multipart file + original_session_id (required) + optional enableAnomalyDetection, imageFiles, soundFiles; or JSON {original_session_id, source_type, credential_id or inline_secrets, table or custom_query}. Returns {status: "success", new_session_id, output_files, transformations_applied, schema_validation, warnings, metadata, amount_charged_usd, balance_usd}, plus balance_warning when the balance is negative (credits_charged, credits_remaining and credit_warning are deprecated). A failed run is 400 {status: "error", message, error} (SESSION_NOT_FOUND, MISSING_ARTIFACTS, SCHEMA_MISMATCH, FILE_READ_ERROR, JOB_SUBMIT_FAILED, WORKER_ERROR); after 25 minutes error: TIMEOUT with the session id, and the run continues. 404 parent not found or not yours. A run started with an API key is recorded against that key's usage; the same holds for retraining.

POST /api/v1/retraining

Any (also /retraining). Inference fields plus run_dor (true), dor_eps (automatic), dor_min_samples (5), run_dsg (true), output_sample_size; run_dsm is accepted and ignored. Outputs features.csv and y_columns.csv.

POST /api/v1/inference-retraining

Removed: 410 ENDPOINT_REMOVED.

Sessions

GET /api/v1/sessions

Any. Query selectable_only, q, status, tag, date_from, date_to. {success, history: [{id, session_id, session_name, retraining_enabled, is_retrainable, original_filename, total_rows, total_columns, file_size_bytes, file_size_mb, processing_date, status, created_at, session_type, tags, note}], total}.

GET /api/v1/sessions/{id}

Any (also /api/session/{id}/details, /session/{id}/details). History fields plus results_available, media, processing_parameters, steps, total_duration_seconds, enterprise; share recipients get a viewer view with permission and capabilities.

PUT /api/v1/sessions/{id}/name, PUT /api/v1/sessions/{id}/retraining

Any. {"session_name": "..."} (255 characters); {"retraining_enabled": true}.

GET /api/v1/sessions/{id}/config, PUT /api/v1/sessions/{id}/meta, GET /api/v1/sessions/{id}/download-all

Any (also under /api/session/{id}/...). Config returns {session_id, config, run_mode} in /process-data field names. Meta takes tags (up to 10, 32 characters each) and note (500). Download-all returns every output as a ZIP. 404 when not yours.

GET /api/v1/sessions/{id}/report

Any. {session_id, tier, status, summary, warnings, details?}; details (full tier) for Enterprise accounts.

POST /api/v1/sessions/{id}/retry

Any (also /dashboard/api/sessions/{id}/retry). Re-runs from the last completed stage with the configuration the run was started with (targets, task types, outputs, cleaning and speed preset).

GET / DELETE /api/v1/sessions/{id}/outputs, GET /api/v1/sessions/{id}/outputs/{name}

Any (also /api/memory-outputs/...). Memory-held outputs (about 4 hours); download with ?format= (csv default; json returns rows inline).

GET /api/v1/sessions/{id}/files

Any (also /api/sessions/{id}/files). The downloadable files of a run (shared files only for recipients): {success, files: [{name, size_mb, url}]}. 404 when the session is neither yours nor shared with you.

POST /api/v1/sessions/{id}/write-output

Any (also /dashboard/api/sessions/{id}/write-output). Body connector_type (any connector type except ga4; defaults to the saved connection's type, else sql), credential_id or inline_secrets, table_name (required), if_exists (replace, append, fail, merge; left out, the destination's default: replace where offered, else append), merge_keys (required for merge), output_stage (dsg), file_format (file destinations: csv, parquet, xlsx, json, jsonl, ndjson, feather, orc), partition_by (one target per distinct value of a column, up to 1,000), partition_format (%Y-%m-%d), options (destination settings such as time_column or partition_key_column). Returns {success, rows_written, table, message}, with partitions for a partitioned write.

A mode the destination does not offer is a 400 that lists the supported ones; fail on an existing target is a 409. Salesforce, HubSpot, NetSuite, SharePoint lists, Stripe, PI Web API and OPC UA are written to only through a saved connection with Allow writing switched on (allow_write); otherwise 403. When a destination accepts some rows and declines others, the answer reports rows_written, rows_failed and errors. See Write Output & Auto-Sink for each destination's modes, options and row limits.

Sharing and comments

Also without /v1. Permissions: read, comment, edit.

POST /api/v1/shared-sessions

Any. session_id (required), shared_with_id (omit for a public link), permission (read), expires_at (ISO-8601, future), note, shared_files. 201 {success, shared_session}.

GET /api/v1/shared-sessions, GET /api/v1/shared-sessions/received, PUT / DELETE /api/v1/shared-sessions/{share_id}

Any. List outbound and received shares; change or revoke. Changing or revoking a link applies to everyone who joined it.

GET /api/v1/shared-sessions/access/{token}, GET .../access/{token}/download/{filename}

Public. The link's session, files and viewer state; download a shared file. 404 unknown, 410 expired.

POST /api/v1/shared-sessions/access/{token}/join

Any. Adds the link's grant to your account.

GET / POST /api/v1/shared-sessions/by-session/{id}/comments, DELETE .../comments/{comment_id}

Any. Comments up to 2000 characters; posting needs comment permission.

GET /api/v1/users/search

Any (also /api/users/search). Finds a share recipient: ?q= exact email or a known collaborator's name. Returns {users: [{id, name, email_hint}]} with masked emails; pass id as shared_with_id. 429 above 20 searches per minute.

Connections

Connector types (38): sql, snowflake, bigquery, mongodb, s3, gcs, azure_blob, adls_gen2, databricks, delta, fabric, synapse, kafka, kinesis, mqtt, opcua, oracle, sap_hana, clickhouse, redshift, timescaledb, cassandra, dynamodb, elasticsearch, salesforce, hubspot, stripe, google_sheets, influxdb, ga4, netsuite, pi_web_api, sharepoint, http, google_drive, dropbox, onedrive, box. Also without /v1.

GET / POST /api/v1/credentials, PUT / DELETE /api/v1/credentials/{id}

Any. Create with name, connector_type, secrets (all required; stored encrypted, never returned). PUT takes name, secrets, merge_secrets and clear_secrets: it replaces the secrets unless merge_secrets: true, and clear_secrets lists stored field names to remove. The list shows non-secret hints, the last test result, stored_fields (the names of the fields that are set, never their values) and write: what the connection can do as a destination (modes, target_kind, formats, options, guarded, max_rows, enabled).

POST /api/v1/connectors/{type}/test | discover | preview

Any. Body credential_id or inline_secrets; preview also table or a read-only custom_query (a single SELECT or WITH). When the cause of a failure is recognised, the answer adds error_category (auth, permission, not_found, tls, timeout, network) and a plain-language error_hint.

Data tools

POST /api/v1/mapping/suggest, POST /api/v1/mapping/apply

Any. Suggest: source_columns, target_columns, threshold (0.4). Apply: session_id, mappings, output_stage (dsg), drop_unmapped (false); returns a preview, nothing is saved.

GET /api/v1/feature-engineering/columns, POST /api/v1/feature-engineering/recommend

Any (also feature-selection spellings and without /v1). Recommend: session_id, target_columns, methods (mutual_information default; also variance_threshold, correlation_pruning, pca, rfe, shap_importance, permutation_importance, leakage_guard), top_k (20). Returns {ranking, per_method, leakage, warnings, ...}.

POST /api/v1/pivot/suggest (Enterprise)

Enterprise. session_id, or columns + rows; returns the suggested as-of column and candidates.

POST /api/v1/dataset-metadata/validate (Enterprise)

Enterprise. A metadata file (2 MB) or JSON {metadata, columns, y_columns}; returns {metadata, warnings, summary}.

Automation

Every automation object takes mode (full_pipeline, inference, retraining) and base_session_id. pipeline_config uses the process-connector keys.

Each automated run is charged the exact price of the data it processes, measured when the run starts. Scheduled runs, triggers and listeners can set an optional max_price_dollars (Maximum price per run) in pipeline_config (a trigger: target_pipeline_config); a run priced above it is held for approval instead of running.

/api/v1/pending-approvals

Any (also /api/pending-approvals). GET /api/v1/pending-approvals lists your held runs: {approvals: [{id, session_id, source, name, price_usd, cap_usd, status, created_at, decided_at}]} (price_dollars and cap_dollars hold the same amounts; price_credits and cap_credits are deprecated). POST /api/v1/pending-approvals/{id}/approve starts the run at the price shown and charges exactly that ({success, session_id, price_usd, approval}); POST /api/v1/pending-approvals/{id}/discard drops it. 404 unknown or not yours; 409 already decided; 410 the held data is gone. A held run sends a job.held webhook event {session_id, source, name, price_dollars, cap_dollars, message} and, if you receive failure emails, an email.

/api/v1/scheduled-runs

Any (also /api/scheduled-runs). GET, POST (201), PUT/DELETE /{id}, POST /{id}/run-now, POST /{id}/toggle. Fields: name (required), description, schedule_type (interval | cron), schedule_value (30m, 1h, 2d, or a cron expression in UTC), connector_type, credential_id, source_table, sql_query, pipeline_config (including y_columns; without them the last column is used; and the optional max_price_dollars), is_active. A held run sets last_run_status to held.

/api/v1/triggers and POST /trigger/inbound/{secret}

Any (also /api/triggers). CRUD plus POST /{id}/test (dry run). trigger_type: inbound_webhook (optional config.hmac_key, 16+ characters), watermark_advance (credential_id, table, watermark_column, min_advance), pipeline_completion (source_template_id or source_session_name), quality_alert (not available yet). target_pipeline_config names the source, targets, run mode and settings, and the optional max_price_dollars. Five consecutive failures disable a trigger. The inbound URL takes a CSV or JSON body up to 64 KB (or reads the target's connector), an X-Signature: t=...,v1=... header when an HMAC key is set (5-minute window), and answers 202 {accepted, session_id}, 400, 401, 404, 410, 413, 429 or 502.

/api/v1/listeners/folder

Any. source_type (local, s3, gcs, azure_blob, adls_gen2, sftp), folder_path, bucket_name, prefix, credential_id, name, y_columns (required for full_pipeline), pipeline_config (including the optional max_price_dollars), poll_interval_seconds (60), allowed_extensions, max_file_size_mb (500), cooldown_seconds (60), max_files_per_hour (10), auto_sink_config, enabled.

/api/v1/sftp

Any. GET /info, GET/POST /credentials (POST returns the SFTP password once), PUT /credentials/{id}/pipeline, DELETE /credentials/{id}. One inbox per account.

/api/v1/sync/watermarks

Any (also /api/sync/watermarks). GET /list lists your watermarks and POST /{id}/reset resets one. On /api/v1/sync/watermarks itself a watermark is addressed by connector_type + source_key: GET (query parameters) reads it, {success, watermark} with watermark: null when none is stored; PUT creates or updates it (table_name, watermark_column, watermark_value, watermark_type, row_count, last_sync_at); POST creates it (201; 409 when one already exists); DELETE (query parameters) removes it (404 when none is stored).

/api/v1/quality-alerts

Any. name, metric (row_loss_pct, null_pct, column_drop_count, duration_seconds, drift_score), operator (>, <, >=, <=, ==), threshold, stage (anomaly, dtc, mdh, cds, dsm, dsg, drift; empty = all), severity (warning), enabled. GET /events lists the last 100 alerts. webhook_ids is stored but alerts are not sent to webhooks.

/api/v1/streams (Enterprise)

Enterprise. Continuous inference between Kafka, Kinesis and MQTT: CRUD, /{id}/start|stop|restart, /{id}/status, /{id}/schema, /test-connection. Billed hourly for the rows scored, at the inference rate. See Streaming for the fields.

/api/v1/webhooks

Any (also /webhooks). url (public addresses only), name, events (error, warning, info, success, job.completed, job.failed, quality.breach; default ["error"]), secret, active. POST /{id}/test sends a test event; GET /{id}/deliveries lists the last 50. Runs send job.completed and job.failed with {event, data: {session_id, status, original_filename, timestamp, files | error}, timestamp, webhook_id}, signed as X-Webhook-Signature: sha256=<hex> when a secret is set. One attempt, 10-second timeout.

Preferences

Any. GET, POST and POST /reset on /api/v1/preferences/outputs, /advanced, /anomaly, /completion and /api-anomaly (legacy: /api/output-preferences, /api/advanced-parameters, /api/anomaly-parameters, /api/completion-parameters, and /api/output-config in the older disk/memory shape). Output preferences hold the *_output file switches, *_use_tool stage switches, output_format (csv), output_name_pattern (prefixed) and the model-ready split.

/api/v1/preferences/anomaly and /api/v1/preferences/api-anomaly

Any. Two separate stores of data-cleaning switches, each with GET, POST and POST /reset, each returning {success, parameters}. /api/v1/preferences/anomaly holds the settings of runs started in the dashboard. /api/v1/preferences/api-anomaly holds the defaults of runs started through the API or the SDK: POST /api/v1/process uses them when a request sends no anomaly_params. Changing one does not change the other.

Account and keys

GET /api/v1/keys, GET /api/v1/usage

Key. /api/v1/keys: your active keys (never secrets), each with daily_limit_usd, lifetime_limit_usd, daily_spent_usd and lifetime_spent_usd; null means no limit. /api/v1/usage: the calling key's spend, {today_usd, daily_limit_usd, daily_remaining_usd, lifetime_spent_usd, lifetime_limit_usd, lifetime_remaining_usd, avg_daily_usd_7d, avg_daily_usd_30d, chart_data: {labels, usd}}. All amounts are US dollars. The credit-named fields (daily_credit_limit, lifetime_credit_limit, daily_credits_used, lifetime_credits_used) are deprecated and still returned. Creating, rotating, editing and revoking keys, team keys included, is done in the dashboard.

GET /api/v1/account/profile, GET /api/v1/account/usage

Any. profile: balance_usd, spent_usd and total_usd (balance + spent), role and limits. usage: usage by stage across all keys, with spent_usd per stage and summary.total_spent_usd. Deprecated, still returned: api_credits_remaining, api_credits_used, api_credits_total, credits_used, total_credits_used.

/api/account/*

Session. Profile (GET/PUT), regenerate passcode, usage statistics, API profile, API sessions, usage by tool, personal keys (create with name, expires_in_days; update daily_limit_usd, lifetime_limit_usd and stage permissions; rotate; revoke) and team keys (with per-key spending limits and permissions, billed to you). Profiles carry balance_usd, spent_usd and total_usd; API sessions and usage by tool carry spent_usd; key usage carries the same _usd fields as /api/v1/usage. The older daily_credit_limit and lifetime_credit_limit are still accepted, as credits. Key allowance: Trial 1, Paid 3, Unrestricted 15, Enterprise 50.

GET /api/v1/account/contracts

Any (also /api/account/contracts). Your pricing contracts with usage against each allowance: {contracts: [...]}, an empty list when the account has none.

GET /api/v1/account/trial-status

Any (also /auth/trial-status). Trial state and the free runs remaining: {success, is_trial, trial_uploads_remaining, role, role_name, total_runs, completed_runs, failed_runs, total_data_mb, limits}.

GET / PUT /api/v1/notification-preferences

Any (also /api/notification-preferences). Email on completion or failure: {email_on_complete, email_on_fail, configured}. PUT takes either field and changes only the ones sent; 400 when a value is not a boolean.

GET / POST /api/v1/pipeline-templates, PUT / DELETE /api/v1/pipeline-templates/{template_id}

Any (also /api/pipeline-templates). Saved configurations, up to 50 per account. GET returns {templates: [...]}. POST takes name, config and optional description and answers 201 {template}; PUT changes any of the three; DELETE answers {ok: true}. 400 for an invalid body or when the limit is reached; 404 for a template that is not yours.

System

GET /health (public; 200 or 503), GET /api/v1/workers/status (any), GET /openapi.yaml and GET /api (public), and /api/v1/spark/status|read|transform for self-hosted installs with Spark enabled (503 otherwise).

Enterprise endpoints (Enterprise)

Read-only run evidence under /api/v1/enterprise/: families, per-run column record, family comparison, dataset metadata, lineage and diff, and an audit report. Other plans get 403 with "required_role": "Enterprise".