Python SDK
Reference for the datatoolpack Python library: install, authentication, retries and timeouts, pipeline options, speed modes, media uploads, and every AutoDataClient method with its parameters, endpoint, return value and errors.
The official datatoolpack Python library wraps the AutoData REST API. Use it to run the pipeline on files or connectors, follow jobs, and download results as files or DataFrames. It can also replay a session for inference or retraining, and manage connectors, automations, webhooks, sharing and account settings.
Install
pip install datatoolpack
Package datatoolpack, version 0.15.0. Requires Python 3.8 or later. The only required dependency is requests. The DataFrame helpers (get_dataframe, get_dataframes, return_dataframes=True) also need pandas; install it yourself (pip install pandas, plus pyarrow for Parquet, Feather and ORC outputs). The package exports AutoDataClient, AutoDataError, ModeConfirmationRequired and __version__.
Authentication and client setup
Every call is authenticated. With an API key the client sends Authorization: Bearer dtpk_...; with a passcode it sends X-Passcode: <passcode>. Create API keys on the API Account page. A key can carry its own spending limits (in US dollars) and per-stage permissions, so a call can fail with 429 (limit reached) or 403 (stage not allowed for this key).
from datatoolpack import AutoDataClient
# 1. API key (recommended)
client = AutoDataClient(api_key="dtpk_your_key")
# 2. Passcode
client = AutoDataClient(passcode="123456789012")
# 3. Environment variables: AUTODATA_API_KEY or AUTODATA_PASSCODE,
# optional AUTODATA_BASE_URL
client = AutoDataClient()
# Context manager closes the HTTP session for you
with AutoDataClient(api_key="dtpk_your_key") as client:
print(client.worker_status())
AutoDataClient(api_key=None, passcode=None, base_url=None, timeout=300, max_retries=3)
| Parameter | Type | Default | Meaning |
|---|---|---|---|
api_key | str | $AUTODATA_API_KEY | API key. Must start with dtpk_, otherwise ValueError. Wins over passcode when both are given. |
passcode | str | $AUTODATA_PASSCODE | Account passcode, used only when no API key is given. Passcode calls have no per-key limits. get_usage() needs an API key and returns {} with a passcode. |
base_url | str | $AUTODATA_BASE_URL, else https://autodata.datatoolpack.com | Server address. A trailing slash is removed. Every endpoint is called under <base_url>/api/v1. |
timeout | int (seconds) | 300 | Timeout for each HTTP request. Downloads, ZIP archives and infer/retrain use at least 900 s. |
max_retries | int | 3 | How many times a request is retried after a transient failure. See Errors, retries and timeouts. |
ValueError is raised when neither an API key nor a passcode is available, or when the key does not start with dtpk_. The client keeps one requests.Session; call close() or use with to release it. repr(client) masks the credentials.
Errors, retries and timeouts
| Exception | When | Attributes |
|---|---|---|
AutoDataError | The server answered with a non-2xx status, a job failed, was cancelled, stalled or timed out while waiting, or a download was empty, truncated or not a valid ZIP. The message is the server's error field when there is one. | status_code (int or None for errors that are not HTTP errors) |
ModeConfirmationRequired (subclass of AutoDataError) | process() only: the speed mode would change a parameter you set yourself, and you did not say how to resolve it. HTTP 409. | mode, adjustments (every change the mode makes), contested (only the ones that override your values), policy_signature, status_code (409). str(e) lists each contested parameter with your value and the mode's value. |
ValueError | Raised by the client before any request: unsupported file extension, bad mode/base_session_id combination, missing required argument, unknown preference kind, both file_path and source. | - |
FileNotFoundError | process() input file or a media file does not exist. | - |
requests.ConnectionError / requests.Timeout | The network still fails after all retries. | - |
from datatoolpack import AutoDataClient, AutoDataError, ModeConfirmationRequired
client = AutoDataClient(api_key="dtpk_your_key")
try:
result = client.process("data.csv", target_columns="label")
except ModeConfirmationRequired as e:
print(e) # lists every contested parameter
print(e.policy_signature)
except AutoDataError as e:
if e.status_code == 429:
print("Spending or rate limit reached:", e)
elif e.status_code == 403:
print("This key may not run one of the requested stages:", e)
else:
raise
Retries. Every request is retried automatically on HTTP 429, 502, 503 and 504 and on connection errors and timeouts, up to max_retries times. Before each retry the client waits max(Retry-After, 2**attempt) seconds (1, 2, 4, ...). Retry-After is honoured as a number of seconds or as an HTTP date, and a single wait is capped at 300 s. When the last attempt still returns a retryable status, it is raised as AutoDataError.
Retries also apply to job submission.
process()sends anIdempotency-Keyheader with every submission (one is generated per call unless you passidempotency_key), so a retry returns the first run instead of starting a second one, and a retried upload re-sends the file from its first byte. Forprocess_from_connector,inferandretrain, a submit that times out or gets a 502/504 may still have created the job on the server. Checklist_sessions()before you submit again.
Timeouts. timeout covers one HTTP request, not a whole job. wait_for_completion() has its own timeout (whole wait, off by default) and stall_timeout (default 3600 s with no change in status, percent or message). infer() and retrain() are synchronous: the server holds the request open for up to about 25 minutes while the job runs, so the client waits at least 900 s per request. Raise timeout= on those calls for large files.
Quick start
from datatoolpack import AutoDataClient
with AutoDataClient(api_key="dtpk_your_key") as client:
# Check how the server reads the file (no job, no charge)
info = client.validate("customers.csv")
print(info["columns"], info["dataset_info"])
# Get the exact price of this file with these settings
if info.get("pricing_status") == "measuring":
client.wait_for_price(info["pricing_profile_id"], timeout=600)
quote = client.estimate(pricing_profile_id=info["pricing_profile_id"],
tools={"anomaly": True}, output_rows=100_000, text_mode=2)
print("price: $", quote["price_usd"])
# Run the pipeline, wait, download, and load the outputs as DataFrames
result = client.process(
"customers.csv",
target_columns="churn",
output_rows=100_000,
tools={"anomaly": True},
advanced_params={"excluded_columns": ["customer_id"], "datetime_mode": "basic",
"text_mode": 2, "quote_id": quote["quote_id"]}, # charged exactly that price
return_dataframes=True,
)
print(result["files"]) # [{name, url, size, description}, ...]
train = result["dataframes"]["dsg_output"]
# Later: score new rows with the same fitted transforms
scored = client.infer(result["session_id"], "new_customers.csv", return_dataframes=True)
features = scored["dataframes"]["features"]
By default process() blocks until the job ends, prints progress to stdout, and extracts every output into ./auto_data_outputs/<session_id>/. Pressing Ctrl+C while it waits cancels the job on the server and re-raises KeyboardInterrupt.
Pipeline options: tools and advanced_params
process() and process_from_connector() take two option dicts. tools switches stages on and off. advanced_params tunes them. Both methods accept the same keys and forward every key you pass.
tools
| Key | Default | Effect | process / process_from_connector |
|---|---|---|---|
anomaly | False | Anomaly detection and format repair. Which checks run comes from anomaly_params when you pass it; a file run without it uses your saved cleaning defaults for API runs (get_preferences('api_anomaly')). | both |
dtc | True | Data Type Conversion: encoding, text, date and media features. | both |
mdh | True | Missing Data Handler. See mdh_mode. | both |
dor | False | Accepted, but outlier-row removal does not run in a normal pipeline. It runs only in retraining (retrain(run_dor=True)). | both |
cds | True | Column scaling. | both |
dsm | True | Column pruning (drops near-duplicate and low-information columns; rows are never split). | both |
dsg | True | Synthetic row generation up to output_rows. | both |
feature_selection | False | Optional feature-selection stage between DTC and MDH, configured by advanced_params['feature_selection_config']. | both |
preliminary_cleaning | True | Removes duplicates, normalises missing-value tokens, drops empty, constant and ID-like columns. | both |
nested_flattening | True | Expands JSON objects held in cells into their own columns (details becomes details.guest.email). Lists stay as JSON text; one input row is always one output row. | both |
result_evaluator | False | Adds a multi-model benchmark report. | both |
Both methods send only the
toolskeys you set. A key you leave out takes the default above.
advanced_params
| Key | Type / values | Default | Effect | process / process_from_connector |
|---|---|---|---|---|
excluded_columns | list[str] | [] | Columns removed before processing. | both |
text_mode | 0 off, 1 neural tokenization, 2 TF-IDF, 3 auto | 2 | How free-text columns are turned into features. | both |
text_cleaning | bool | True | Reserved; not available yet. | both |
mdh_mode | 0 Imputation, 1 2D-Removal, 2 Imputation/dropping | 0 | 0 fills missing values and keeps every row. 1 removes the rows and columns with too many missing values instead. 2 (Enterprise) removes only what is too empty to fill and fills the rest. | both |
zscore_limit | float | 3.0 | Outlier limit for generated synthetic values (in standard deviations). | both |
datetime_mode | none, minimal, basic, full | basic | How many feature columns each date column becomes. | both |
dsg_mode | copula, gan | copula | Synthetic-data generator. | both |
similarity_p | float, 0 to 1 | 0.99 | Correlation above which DSM drops one of two near-duplicate columns. A value of 1 or more turns this pruning off. | both |
task_types | list of 'c'/'r', one per target; one value for all targets; or 'auto' | 'c' for every target | Classification or regression per target. 'auto' lets the pipeline detect the type of each target. Also a keyword argument of both methods. | both |
anomaly_params | dict of per-check toggles | saved cleaning defaults for API runs | Which data-cleaning checks run in this run. Also a keyword argument of both methods. | both |
feature_selection_config | {method, top_k, threshold} | top_k 20 | Configures the feature-selection stage when tools['feature_selection'] is on. method is a name, a list, or 'auto'. | both |
llm_enabled | bool | True | False guarantees no stage sends data to a language model. The result carries an llm audit block. | both |
strict_llm | bool | False | True fails the run instead of silently falling back when a model is unreachable. | both |
mode | rapid, balanced, accurate | balanced | Speed preset. See Speed modes. | both |
speed-mode knobs (for example dtc_OHElimit) | see Speed modes | set by the mode | Your own value for a knob the mode would otherwise choose. | both |
mode_decisions | {'section.knob': 'user' or 'mode'} | {} | Per-knob answer when the mode disagrees with your value. | both |
touched_fields | list[str] | [] | Names of the knobs you set on purpose. | both |
mode_ack | str | - | The policy_signature from a 409, meaning keep all your values. confirm_mode_changes=True sends it for you. | both |
cds_scaler | auto, standard, minmax, robust, maxabs, yeojohnson, quantile_normal, quantile_uniform, log1p, none | auto | Forces one scaler for every column. | both |
output_format | csv, parquet, xlsx, json, jsonl, ndjson, feather, orc | account preference, else csv | File format of the outputs. | both |
output_name_pattern | prefixed, plain | account preference, else prefixed | prefixed names files <input>_<timestamp>_<step>.<ext>. plain names them <step>.<ext>. | both |
flatten_columns | list[str] or None | None (detect) | Columns to expand with nested flattening. [] expands none. | both |
dsg_allow_downsample | bool | False | When output_rows is below the processed row count, take a class-preserving subsample instead of skipping synthesis. | both |
dsg_y_oversample_threshold | float, 0 < v < 1 | off | Raises every target class to at least this share of the generated rows. | both |
dsg_y_num_bins | int, 2 to 100 | 10 | Bins used to apply that rule to a continuous target. | both |
dsg_rare_oversample_threshold | float, 0 to 0.5 | off | Oversamples rare categories in generated rows. | both |
image_mode / audio_mode | 0 none, 1 basic features, 2 advanced processing, 3 neural embeddings (image) or Mel-spectrogram tokenization (audio) | 0 | How media columns are encoded. See Image and sound files. | both |
image_columns / audio_columns | list[str] | [] | The columns that hold media file names. | both |
model_families, imputation_overrides, scaling_overrides, pivot_config, validation_policy, selection_config, probe_config, dataset_metadata | see Enterprise options | off | Enterprise preparation policy. Removed by the server for other accounts. | both |
advanced_paramsaccepts any pipeline setting by the name the dashboard's configuration uses, and both methods forward every key you pass. Insideadvanced_params,modeis the speed preset. The run mode (full_pipeline,inference,retraining) is the separatemodeargument ofprocess()andprocess_from_connector().
Speed modes
advanced_params['mode'] picks a preset: rapid, balanced (default) or accurate. A preset sets the knobs below. To keep a value yourself, pass the knob's field name in advanced_params. If the preset disagrees with a value you set, process() raises ModeConfirmationRequired (HTTP 409) instead of running with changed parameters. There are three ways to answer: pass confirm_mode_changes=True to keep all your values, pass advanced_params['mode_decisions'] to decide per knob ('user' keeps yours, 'mode' takes the preset's), or pass advanced_params['mode_ack'] with the policy_signature from the error. Call preview_mode() to see the changes without running anything. Replays (infer, retrain, replay modes) reuse the parent session's settings, so they never ask.
| Decision key (mode_decisions) | Field name (advanced_params) | Type / range | What it controls |
|---|---|---|---|
dtc.OHElimit | dtc_OHElimit | int 2 to 100000 | One-hot / free-text cardinality limit: which object columns are treated as free text. |
dtc.sample_rows_detect_text | dtc_sample_rows_detect_text | int 10 to 100000 | Rows read when detecting free-text columns. |
dtc.sample_rows_text_mode | dtc_sample_rows_text_mode | int 10 to 100000 | Rows read when text mode 3 (auto) chooses a method. |
dtc.seq_max_features | dtc_seq_max_features | int 100 to 1000000 | Vocabulary size for neural text tokenization (text mode 1). |
dtc.tfidf_max_features | dtc_tfidf_max_features | int 100 to 1000000 | Number of TF-IDF features (text mode 2). |
dtc.datetime_mode | dtc_datetime_mode | none, minimal, basic, full | Datetime expansion (same as datetime_mode). |
anomaly.force_no_llm | anomaly_force_no_llm | bool | Turns the language-model anomaly checks off; local checks still run. |
anomaly.sample_scale | anomaly_sample_scale | float 0.1 to 10 | Scales how many rows each anomaly check inspects. |
mdh.nooftrials | mdh_nooftrials | int 1 to 100 | Missing-data threshold trials (2D-Removal mode). |
imputer.skip_expensive | imputer_skip_expensive | bool | Leaves out the slowest imputation strategies. |
imputer.force_all | imputer_force_all | bool | Evaluates every imputation strategy, even on large data. |
cds.eval_top_n | cds_eval_top_n | int 1 to 9 | How many candidate scalers are tried. |
dsg.iess_budget_multiplier | dsg_iess_budget_multiplier | float 0.1 to 10 | Scales the fitting budget for synthetic generation. |
multimodal.text_synth_method | text_synth_method | statistical, vae | Generator for synthetic text. |
multimodal.image_synth_method | image_synth_method | statistical, vae | Generator for synthetic image features. |
multimodal.audio_synth_method | audio_synth_method | statistical, vae | Generator for synthetic audio features. |
from datatoolpack import ModeConfirmationRequired
params = {"mode": "rapid", "dtc_OHElimit": 250}
preview = client.preview_mode("rapid", params)
for change in preview["mode_contested"]:
print(change["label"], change["user_value"], "->", change["mode_value"])
# Option A: keep every value you set
client.process("data.csv", "label", advanced_params=params, confirm_mode_changes=True)
# Option B: decide per knob
params["mode_decisions"] = {"dtc.OHElimit": "user"}
client.process("data.csv", "label", advanced_params=params)
Image and sound files
When your table has columns holding image or audio file names, upload the files themselves with image_files=[...] / sound_files=[...]. Pair them with advanced_params={'image_mode': 1-3, 'image_columns': [...]} or {'audio_mode': 1-3, 'audio_columns': [...]}. Without the files, the server has nothing to match the names against, and every media feature is encoded as zeros: the run succeeds on empty data.
| Kind | Accepted extensions |
|---|---|
Images (image_files) | .jpg .jpeg .png .gif .bmp .webp .tiff .svg |
Sound (sound_files) | .mp3 .wav .ogg .flac .aac .m4a .wma |
Data file (file_path) | .csv .xlsx .xls .parquet .json .jsonl .ndjson .feather .orc. Compressed uploads (.zip, .gz) are not accepted. |
client.process(
"listings.csv", # has a column 'photo' with file names
target_columns="price",
advanced_params={"image_mode": 2, "image_columns": ["photo"]},
image_files=["img/1001.jpg", "img/1002.jpg"],
)
# At inference, send only files the parent session has not seen.
# http(s) URLs held in the column are fetched by the server.
client.infer(session_id, "new_listings.csv", image_files=["img/2001.jpg"])
Every path is checked (exists, allowed extension) before anything is sent. Trial accounts are limited to 1 MB of media per request. Media can only travel with a file upload: infer(source=...) / retrain(source=...) with media raises ValueError.
Downloads and DataFrames
Output keys (file stems) you can ask for: anormaly_fixed_output, compleated_dataset, dtc_output, mdh_output, cds_output, dsm_output, dsg_output, multimodal_output, model_ready, model_ready_X, model_ready_y, model_ready_train, model_ready_test, model_ready_X_train, model_ready_X_test, model_ready_y_train, model_ready_y_test, model_ready_excluded_columns, and pipeline_report (PDF). Which files exist depends on your output preferences (get_preferences('outputs')) or the run's output_config, and on the stages that ran. Files are named <input>_<timestamp>_<step>.<ext> by default (for example customers_20261005_142233_dsg_output.csv), or <step>.<ext> with output_name_pattern: 'plain'. dsg_output is mostly synthetic rows; evaluate a model on dsm_output (real rows only).
compressed=True (the default everywhere) downloads one ZIP from /download-archive/<session_id> and extracts it. compressed=False downloads each file from /download/<session_id>/<name>. Empty files, truncated transfers and invalid archives raise AutoDataError. Names you pass can be a bare stem (dsg_output), which matches whatever format and name pattern the run used, or a full file name.
Run modes for automations
process, process_from_connector, create_trigger, create_scheduled_run, update_scheduled_run, create_listener, update_listener, create_sftp_credential and update_sftp_pipeline take mode and base_session_id:
full_pipeline(default): train a new session from the incoming data. Needs target columns.inference: score the incoming data with a completed session's fitted transforms. Needsbase_session_id. Targets and stage settings come from that session.retraining: refresh a completed session with the incoming data (replayed transforms, outlier-row removal, synthetic top-up). Needsbase_session_id.
The client checks the combination before sending: an unknown mode, or a replay mode without base_session_id, raises ValueError. Inference costs about a fifth, and retraining about two fifths, of preparing the same data.
Enterprise options (Enterprise)
Enterprise accounts can pass preparation-policy fields in the advanced_params of process() and process_from_connector(). For other accounts the server removes them. None of them is on unless you send it.
| Key | Shape | What it does |
|---|---|---|
model_families | list[str], up to 4 | The model families the preparation targets. enterprise_families() lists the valid names. |
imputation_overrides | {column: method} | Pins the imputation method for a column. Valid methods come from enterprise_families(). |
scaling_overrides | {column: scaler} | Pins the scaler for a column. |
pivot_config | {column, leakage_action: 'warn'|'relative'|'drop', relative_deltas: bool, auto_suggest: bool} | The as-of column (when a row came into existence). Other dates are measured from it, and columns whose values fall after it are flagged, reduced to a time difference or removed. pivot_suggest() proposes one. |
validation_policy | {split_strategy: 'random'|'stratified'|'time_ordered'|'grouped', eval_protocol: 'holdout'|'kfold'|'walk_forward', test_size, group_column, order_column, eval_folds (3), eval_max_rows, eval_time_budget_s, walk_forward_splits (3), random_state} | How validation data is split and evaluated. |
selection_config | {methods: [...], top_k, known_leaky, leakage_action} | Target-aware column selection in DSM. Methods: leakage_guard, variance_threshold, correlation_pruning, mutual_information, permutation_importance, rfe, shap_importance. They run in order, each on what the previous one left. |
probe_config | {penalty: 'l2'|'l1'|'elasticnet'|'none', C, l1_ratio} | Experimental tuning of the evaluation AutoData uses to compare methods. It does not change your model. |
dataset_metadata | schema v1: per column role, type, semantic_type, min, max, impute, scaler, keep | A declared data dictionary. Declared values win at every stage and are replayed at inference. An invalid document is rejected with 400. |
Method reference
All 146 public methods of AutoDataClient, grouped by area. Every method raises AutoDataError (with status_code) when the server answers with an error, unless its entry says otherwise. Endpoints are relative to base_url.
Processing
Start a pipeline from a local file or a connector, check a file first, and preview speed-mode changes.
process
process(file_path, target_columns=None, output_rows=10000, tools=None, advanced_params=None, wait=True, poll_interval=2, download_path=None, auto_download=True, output_preferences=None, compressed=True, return_dataframes=False, confirm_mode_changes=False, image_files=None, sound_files=None, lm_readiness=None, history_file=None, *, task_types=None, session_name=None, anomaly_params=None, completion_params=None, output_config=None, webhook_ids=None, mode='full_pipeline', base_session_id=None, idempotency_key=None)
Uploads a data file and runs the full pipeline, or replays an existing session over the file (mode='inference' / 'retraining'). The arguments after history_file are keyword-only. With wait=True it polls until the job ends, downloads the outputs and returns the final result. With wait=False it returns as soon as the job is queued.
Endpoint: POST /api/v1/process (multipart: file, config JSON, repeated imageFiles / soundFiles; header Idempotency-Key), then /status, /result and /download-archive
| Parameter | Type | Default | Meaning |
|---|---|---|---|
file_path | str | - | Local data file: .csv .xlsx .xls .parquet .json .jsonl .ndjson .feather .orc. |
target_columns | str or list[str] | None | Target (y) column or columns. Required for a full_pipeline run (ValueError otherwise). May be omitted for a replay, which takes its targets from the base session. |
output_rows | int | 10000 | Rows in the synthetic output (DSG target). To skip synthesis use tools={'dsg': False}. |
tools | dict[str, bool] | None (server defaults) | Stage switches. See the tools table. |
advanced_params | dict | None | Stage options, speed mode, Enterprise policy. See the advanced_params table. Any pipeline setting may be passed by the name the dashboard's configuration uses. |
wait | bool | True | Block until the job completes, fails or is cancelled. |
poll_interval | int (s) | 2 | Seconds between status polls while waiting. |
download_path | str | ./auto_data_outputs/<session_id>/ | Where outputs are extracted. |
auto_download | bool | True | Download outputs after a successful wait. |
output_preferences | list[str] | None (all files) | Subset of outputs to download and load, by stem ('dsg_output') or full name. |
compressed | bool | True | Download one ZIP archive instead of file by file. |
return_dataframes | bool | False | Also load the tabular outputs into result['dataframes'] (needs pandas). |
confirm_mode_changes | bool | False | If the speed mode would override a value you set, re-submit keeping your values instead of raising ModeConfirmationRequired. The file is uploaded a second time. |
image_files | list[str] | None | Image files named by your image column(s). |
sound_files | list[str] | None | Audio files named by your audio column(s). |
lm_readiness | dict | None | Switches the run to LM Readiness preparation (see prepare_lm_readiness) instead of the normal pipeline. Sent inside config. |
history_file | str | None | Prior-events file for an LM Readiness replay (multipart part history). Only valid with lm_readiness={'action': 'replay', ...}. |
task_types | str or list[str] | None ('c' for every target) | 'c' (classification) or 'r' (regression), one per target, or one value for all targets. 'auto' lets the pipeline detect the type of each target. |
session_name | str | None | Name shown for the run in the dashboard. |
anomaly_params | dict | None (saved cleaning defaults for API runs) | Per-run data-cleaning switches: the keys get_preferences('api_anomaly') returns. Cleaning itself is turned on with tools={'anomaly': True}. |
completion_params | dict | None | Settings for the Data Completion & Verification stage. |
output_config | dict | None (saved output preferences) | Which outputs this run writes and which stages it runs: the object get_preferences('outputs') returns, for example {'model_ready_output': True, 'model_ready_split': 'train_test', 'output_format': 'parquet'}. Not the same as output_preferences, which only filters what is downloaded afterwards. |
webhook_ids | list[str] | None | Webhooks notified when the run ends. |
mode | str | 'full_pipeline' | full_pipeline trains a new session. inference scores the file against an existing session and retraining refreshes one. Both need base_session_id, take their columns and settings from it, and are queued and polled like any other run. infer() and retrain() do the same synchronously. |
base_session_id | str | None | The completed session a replay runs against. Required for the two replay modes. |
idempotency_key | str | None (generated per call) | Identifies this submission. A retry carrying the same key gets the first run's session back instead of starting a second run. |
Returns: wait=False: {success, session_id, status: 'queued', message, config, mode, mode_adjustments, policy_signature}; for a replay mode {success, session_id, mode, base_session_id, rows_read}. A submission whose idempotency key was already used returns the first run's session_id with idempotent_replay: True. wait=True: the get_result() dict, plus dataframes ({step: DataFrame}) when return_dataframes=True.
Raises: FileNotFoundError, ValueError (extension, missing target_columns for a full pipeline run, unknown mode, replay mode without base_session_id), ModeConfirmationRequired, AutoDataError (server error, job failed or cancelled, stall).
Without
task_types, every target is treated as classification ('c'). Withoutanomaly_params, the cleaning checks come from your saved cleaning defaults for API runs (get_preferences('api_anomaly')).
result = client.process(
"sales.parquet",
target_columns=["revenue"],
task_types="r",
session_name="Sales forecast",
output_rows=50_000,
tools={"anomaly": True, "feature_selection": True},
advanced_params={
"text_mode": 3,
"mdh_mode": 0,
"similarity_p": 0.95,
"output_format": "parquet",
"feature_selection_config": {"method": "auto", "top_k": 30},
},
output_preferences=["dsg_output", "dsm_output", "pipeline_report"],
)
print(result["session_id"], result["row_count"], result.get("amount_charged_usd"))
# Queue an inference run against that session and poll it yourself
queued = client.process("new_sales.parquet", mode="inference",
base_session_id=result["session_id"], wait=False)
scored = client.wait_for_completion(queued["session_id"])
process_from_connector
process_from_connector(connector_type, table, target_columns=None, secrets=None, credential_id=None, custom_query=None, output_rows=10000, tools=None, advanced_params=None, wait=True, poll_interval=2, download_path=None, auto_download=True, output_preferences=None, compressed=True, incremental=False, type_casts=None, return_dataframes=False, mode='full_pipeline', base_session_id=None, lm_readiness=None, *, watermark_column=None, task_types=None, session_name=None, anomaly_params=None, completion_params=None, output_config=None, webhook_ids=None)
Runs the pipeline on rows the server reads directly from a connector, or replays an existing session over them (mode='inference' / 'retraining').
Endpoint: POST /api/v1/process-connector (JSON), then /status, /result and /download-archive
| Parameter | Type | Default | Meaning |
|---|---|---|---|
connector_type | str | - | One of the 38 connector types (see test_connector). |
table | str | - | Table, collection, object key, topic or sheet to read. |
target_columns | str or list[str] | None | Required for full_pipeline (ValueError otherwise). Not needed for replays, which take their targets from the base session. |
secrets | dict | None | Inline connection secrets (sent as inline_secrets). With credential_id too, they overlay the saved values. |
credential_id | str | None | A saved credential (save_credential). |
custom_query | str | None | Read-only SQL (SELECT/WITH) used instead of table on SQL-style sources. |
output_rows | int | 10000 | Synthetic output size. |
tools | dict | None | Stage switches, as in process. Every key you set is forwarded. |
advanced_params | dict | None | As in process. Every key you pass is forwarded. |
wait / poll_interval / download_path / auto_download / output_preferences / compressed / return_dataframes | - | as in process | Same behaviour as in process. With wait=True and a replay, the wait follows the new replay session. |
incremental | bool | False | Read only rows newer than the last sync watermark. Needs watermark_column (ValueError otherwise). |
watermark_column | str | None | Keyword-only. The column that orders the source (a timestamp or an increasing id). The highest value read is remembered once the run completes, and the next incremental run starts after it. |
task_types / session_name / anomaly_params / completion_params / output_config / webhook_ids | - | None | Keyword-only. As in process. |
type_casts | dict[str, str] | None | Column casts applied before the pipeline: int, float, str, bool, datetime, category (and their aliases). Columns that fail to cast are left unchanged. |
mode | str | 'full_pipeline' | full_pipeline, inference or retraining. |
base_session_id | str | None | Required for the two replay modes. |
lm_readiness | dict | None | Run LM Readiness preparation on the connector rows instead of the pipeline (same options as prepare_lm_readiness's config). A target is required, because LM Readiness runs the classic pipeline first. mode='inference' or 'retraining' with an LM Readiness base_session_id replays or re-prepares that session as LM Readiness. |
Returns: Without waiting: {success, run: {session_id}, session_id, message}; for replays {success, session_id, mode, base_session_id, rows_read}; session_id is None when an incremental read found no new rows. With wait=True: the get_result() dict (+ dataframes).
Raises: ValueError (mode, missing targets, incremental=True without watermark_column), AutoDataError (400 bad source or no rows, 404 credential not found, 403 stage not allowed for the key, 429 key limit).
This method sends no defaults of its own: a setting you leave out takes the server default, the same as a file run (
text_mode2,similarity_p0.99). Versions before 0.15.0 senttext_mode=0andsimilarity_p=95when you set neither. To keep text features off as before, passadvanced_params={"text_mode": 0}.
result = client.process_from_connector(
"snowflake", "ORDERS",
target_columns="churned",
credential_id="c0ffee00-...",
task_types="c",
advanced_params={"datetime_mode": "full"},
type_casts={"order_date": "datetime"},
)
# Read only the rows added since the last run
client.process_from_connector("snowflake", "ORDERS", target_columns="churned",
credential_id="c0ffee00-...",
incremental=True, watermark_column="updated_at")
# Score new rows against an existing session
client.process_from_connector("snowflake", "NEW_ORDERS", credential_id="c0ffee00-...",
mode="inference", base_session_id=result["session_id"])
prepare_lm_readiness
prepare_lm_readiness(file_path, config=None, target_columns=None, **options)
Prepares a dataset for transformer (language-model style) training: train-fitted encodings, a train/validation/test split, and a saved preparation contract you can replay on new data. It runs instead of the normal pipeline (no anomaly detection, column pruning or synthetic rows). A thin wrapper around process(file_path, target_columns, lm_readiness=config, **options).
Endpoint: POST /api/v1/process (multipart, config.lm_readiness set)
| Parameter | Type | Default | Meaning |
|---|---|---|---|
file_path | str | - | Local data file (same formats as process). Limits: 200,000 rows and 100 MB by default (1 MB for trial accounts). |
config | dict | {'profile': 'tabular_transformer'} | LM Readiness options: profile (tabular_transformer or event_sequence), preset (general or transaction_fraud), representation (typed or feature_tokens), level (E1, E2 or auto), sequence_length (1 to 128, default 32), roles ({entity, time, amount, category, counterparty} mapped to column names), excluded_columns, text_columns, test_size and validation_size (fractions, default 0.15 each, sum below 1). event_sequence, transaction_fraud and E2 need the entity and time roles. |
target_columns | str or list[str] | None | Targets. Required: the client raises ValueError before any request when it is missing. |
**options | - | - | Any other process() keyword: wait, download_path, return_dataframes, advanced_params (only validation_policy, dataset_metadata and task_types are read), ... |
Returns: As process(). Without waiting: {success, session_id, status: 'processing', module: 'lm_readiness', dataset_info: {rows, columns}, heartbeat_interval_seconds}. The outputs are a bundle ZIP plus its files under lm_readiness/ (manifest.json, the contract, the split data and an example load_bundle.py).
Raises: ValueError (no target_columns); AutoDataError 400 for an invalid config, unknown mapped columns or too few eligible rows.
New in the current code and not yet described in the package README.
res = client.prepare_lm_readiness(
"transactions.csv",
config={"profile": "event_sequence", "preset": "transaction_fraud",
"roles": {"entity": "card_id", "time": "event_ts", "amount": "amount"}},
target_columns="is_fraud",
)
base = res["session_id"]
replay_lm_readiness
replay_lm_readiness(file_path, base_session_id, history_file=None, **options)
Applies a completed LM Readiness session's saved contract to new data, so the new rows are encoded exactly like the originals.
Endpoint: POST /api/v1/process (multipart, config.lm_readiness = {action: 'replay', base_session_id}, optional history part)
| Parameter | Type | Default | Meaning |
|---|---|---|---|
file_path | str | - | New rows. |
base_session_id | str | - | A completed LM Readiness session you own. |
history_file | str | None | Earlier events for the same entities, so history-based features of the new rows can see them. |
**options | - | - | Other process() keywords. |
Returns: As prepare_lm_readiness.
Raises: FileNotFoundError (history file); AutoDataError 400 when the base session has no completed contract or is not yours.
client.replay_lm_readiness("new_transactions.csv", base, history_file="last_30_days.csv")
preview_mode
preview_mode(mode='balanced', advanced_params=None)
Shows what a speed mode would change about your parameters, without uploading or running anything.
Endpoint: POST /api/v1/mode-preview
| Parameter | Type | Default | Meaning |
|---|---|---|---|
mode | str | 'balanced' | rapid, balanced or accurate. |
advanced_params | dict | None | The same dict you would pass to process. |
Returns: {success, mode, mode_groups, mode_adjustments, mode_contested, policy_signature, knob_specs}. Each adjustment has path, label, user_value, mode_value. knob_specs gives the type and range of every knob.
preview = client.preview_mode("rapid", {"dtc_OHElimit": 250})
print(preview["policy_signature"], len(preview["mode_contested"]))
validate
validate(file_path)
Dry run: uploads a file and returns how the server reads it. No job is started and nothing is charged.
Endpoint: POST /api/v1/validate (multipart file)
| Parameter | Type | Default | Meaning |
|---|---|---|---|
file_path | str | - | Local data file. |
Returns: {columns, shape, missing_values, file_size_mb, validation_level ('full' or 'sampled'), nested_flattening, column_scan, dataset_info: {rows, columns}, has_audio_column, has_image_column, audio_file_count, image_file_count, pricing_profile_id, pricing_status}. Large files are sampled for the preview and carry a note; for the price, every row is measured. pricing_status is ready (measured during the call) or measuring (a larger file; see wait_for_price). Pass pricing_profile_id to estimate() for the exact price.
Raises: AutoDataError (unreadable file; trial accounts are limited to 1 MB files).
wait_for_price
wait_for_price(pricing_profile_id, timeout=600, poll_interval=2.0)
Waits until an uploaded file has been measured for its exact price. Needed only when validate() returned pricing_status: 'measuring'.
Endpoint: GET /api/v1/pricing/profile/<pricing_profile_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
pricing_profile_id | str | - | From validate(). |
timeout | float | 600 | Seconds to wait in total. |
poll_interval | float | 2.0 | Seconds between polls. |
Returns: {status: 'ready', progress_pct, exact, rows, columns}. Then call estimate(pricing_profile_id=...).
Raises: RuntimeError if the file could not be measured (upload it again); TimeoutError after timeout.
validate_media_files
validate_media_files(file_path, media_type, filenames, columns=None)
Checks the media file names a dataset refers to against the files you hold, before a run. Only names are sent; no media is uploaded and no job is started.
Endpoint: POST /api/v1/media/validate (multipart: csv, mediaType, repeated uploadedFilenames, imageColumns / audioColumns)
| Parameter | Type | Default | Meaning |
|---|---|---|---|
file_path | str | - | The dataset whose cells name the media files. |
media_type | str | - | 'image' or 'audio' (ValueError otherwise). |
filenames | list[str] | - | The files you intend to upload. Paths are reduced to their file names. |
columns | list[str] | None (the default media column) | The media columns to check. |
Returns: {valid, column_found, uploaded_files_without_match, total_uploaded, total_matched, message}.
check = client.validate_media_files("listings.csv", "image",
["img/1001.jpg", "img/1002.jpg"], columns=["photo"])
print(check["valid"], check["uploaded_files_without_match"])
Status, results and downloads
Follow a job, fetch its results and load them.
get_status
get_status(session_id)
Returns the current progress of a session once.
Endpoint: GET /api/v1/status/<session_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: {session_id, status, message, current_step, total_steps, progress_percent, duration_seconds} plus, when present, error_info, substep and run_warnings. status is queued, processing, completed, error/failed, cancelled or unknown.
Raises: AutoDataError 403 when the session is not yours.
wait_for_completion
wait_for_completion(session_id, poll_interval=2, timeout=None, stall_timeout=3600.0)
Polls get_status until the job ends, printing a progress line to stdout, then returns get_result().
Endpoint: GET /api/v1/status/<id>, then GET /api/v1/result/<id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
poll_interval | int (s) | 2 | Seconds between polls. |
timeout | float (s) | None (no limit) | Ceiling for the whole wait. |
stall_timeout | float (s) | 3600.0 | Raise if status, percent and message have not changed for this long. The method's docstring says 10 minutes; the real default is one hour. |
Returns: The get_result() dict.
Raises: AutoDataError when the job fails, is cancelled, stalls or the timeout passes. The wait does not cancel the job.
get_result
get_result(session_id)
Returns the final result of a session. Fetching it for a completed session also settles billing (idempotent).
Endpoint: GET /api/v1/result/<session_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: Completed: {session_id, status: 'completed', files: [{name, url, size, description}], row_count, duration_seconds}, plus when relevant warning/message (for example DSG skipped), dsg_provenance (real vs synthetic rows), audio_decode_report / image_encode_report (only when files failed), dsm_column_drops, warnings, run_warnings, media_staging, amount_charged_usd, balance_usd, balance_warning (text in dollars, set when the balance is negative), billing ({amount_charged_usd, quote_low_usd, quote_high_usd, list_usd, capped}: the amount charged), llm (data-residency audit). The credit-named fields (credits_charged, credits_remaining, credit_warning, and *_credits in billing) are deprecated: still returned, equal to the dollar amount × 1,000,000, and will be removed in a future version. Not complete: {session_id, status, message} (+ error_info).
Raises: AutoDataError 403 when the session is not yours.
cancel
cancel(session_id)
Asks the server to cancel a queued or running session.
Endpoint: POST /api/v1/cancel/<session_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: True if the server acknowledged the cancel, otherwise False.
Raises: Never raises for API errors: it prints a warning and returns False.
cancel_session
cancel_session(session_id)
Deprecated alias of cancel().
Endpoint: POST /api/v1/cancel/<session_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: Same as cancel().
retry_session
retry_session(session_id)
Re-submits a failed session from its last checkpoint, with the configuration it was first run with (targets, task types, outputs, cleaning and speed mode).
Endpoint: POST /api/v1/sessions/<session_id>/retry
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: {success, session_id, message}. Follow it with wait_for_completion(session_id).
Raises: AutoDataError 404 when the session's saved request or input file no longer exists.
download_results
download_results(session_id, download_path=None, output_preferences=None, compressed=True)
Downloads a completed session's outputs to a folder.
Endpoint: POST /api/v1/download-archive/<session_id> (body {files: [...]}), or GET /api/v1/download/<session_id>/<name> per file
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
download_path | str | ./auto_data_outputs/<session_id>/ | Folder, created if needed. |
output_preferences | list[str] | None (all) | Stems or file names. 'pipeline_report' means the PDF report. |
compressed | bool | True | One ZIP, extracted locally. With False, files are downloaded one by one. Either way output_preferences entries match by step name ('dsg_output' or 'dsg_output.csv'), whatever format the run exported. |
Returns: Absolute path of the download folder (str).
Raises: AutoDataError on an empty or invalid archive (outputs expired or not generated).
download_file
download_file(url, output_path, timeout=None, retries=2)
Downloads one file. Checks Content-Length, rejects empty files and retries partial transfers.
Endpoint: GET <url>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
url | str | - | Absolute URL, or a path starting with / (for example a url from get_result()['files']). |
output_path | str | - | Local file path; parent folders are created. |
timeout | float (s) | max(client timeout, 900) | Per attempt. |
retries | int | 2 | Extra attempts after a network error or a truncated download. A 4xx answer (other than 429) is raised at once, with no further attempt. |
Returns: None.
Raises: AutoDataError (empty or partial file), network errors after the last retry.
get_dataframe
get_dataframe(session_id, output_name, *, timeout=None, **read_csv_kwargs)
Downloads one output straight into a pandas DataFrame, whatever format the run used.
Endpoint: GET /api/v1/download/<session_id>/<file name>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
output_name | str | - | Stem ('dsg_output', resolved against the session's file list) or full file name. |
timeout | float (s) | max(client timeout, 900) | Request timeout. |
**read_csv_kwargs | - | - | Passed to pandas.read_csv (also to read_json / read_excel; not used for Parquet, Feather and ORC). |
Returns: pandas.DataFrame.
Raises: AutoDataError when pandas is missing or the file is missing or empty.
df = client.get_dataframe(sid, "dsm_output", dtype={"zip": str})
get_dataframes
get_dataframes(session_id, names=None, *, timeout=None)
Loads several outputs at once.
Endpoint: GET /api/v1/result/<id>, then GET /api/v1/download/<id>/<name> per file
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
names | list[str] | None (every tabular file; PDF, JSON and Markdown skipped) | Stems or file names. |
timeout | float (s) | as get_dataframe | Per file. |
Returns: {step: DataFrame}, keyed by the logical step (dsg_output, model_ready_X_train, ...) whatever the file was named.
Raises: AutoDataError.
Inference and retraining
Apply a completed session's fitted transforms to new data. Both calls are synchronous: they return when the server has finished.
infer
infer(original_session_id, file_path=None, enable_anomaly_detection=None, download_path=None, return_dataframes=False, timeout=None, *, source=None, image_files=None, sound_files=None)
Scores new rows with the same encoding, imputation and scaling as the training run (anomaly, DTC, MDH, CDS replayed). Column names must match the training data.
Endpoint: POST /api/v1/inference (multipart file, original_session_id, enableAnomalyDetection, imageFiles, soundFiles; or JSON with source)
| Parameter | Type | Default | Meaning |
|---|---|---|---|
original_session_id | str | - | A completed session you own. |
file_path | str | None | Local data file (same formats as process). Give this or source, not both. |
enable_anomaly_detection | bool | None (inherit) | Override the parent's anomaly setting. |
download_path | str | None | Download the outputs here. When None and return_dataframes=False, nothing is downloaded. |
return_dataframes | bool | False | Load the CSV outputs into result['dataframes'] (downloads to a temp folder if no path is given). |
timeout | float (s) | max(client timeout, 900) | Request timeout. |
source | dict | None | Read rows from a connector instead: {source_type, credential_id or inline_secrets, table or custom_query}. |
image_files / sound_files | list[str] | None | Media named by the data. Optional: the parent session's own media and http(s) URLs in the column are used too. Only with file_path. |
Returns: {status: 'success', new_session_id, output_files: [{name, ...}], ...} plus amount_charged_usd, balance_usd and, when the balance is negative, balance_warning (credits_charged, credits_remaining and credit_warning are deprecated), and files ({name: local path}) / dataframes ({'features': df, 'y_columns': df}) when requested.
Raises: ValueError (both or neither of file_path/source, media with source, bad extension); AutoDataError (400 when the job fails, 403/404 parent not yours or not found).
out = client.infer(sid, "new_rows.csv", return_dataframes=True)
X = out["dataframes"]["features"]
# From a connector
client.infer(sid, source={"source_type": "sql",
"credential_id": "c0ffee00-...", "table": "new_orders"})
retrain
retrain(original_session_id, file_path=None, run_dsm=True, run_dsg=True, output_rows=None, enable_anomaly_detection=None, download_path=None, return_dataframes=False, timeout=None, run_dor=True, dor_eps=None, dor_min_samples=None, *, source=None, image_files=None, sound_files=None)
Refreshes a training set: the same replay as infer, then outlier-row removal (DOR) and synthetic top-up (DSG). Outputs have the same shape as infer (features.csv, y_columns.csv).
Endpoint: POST /api/v1/retraining (multipart or JSON, as infer)
| Parameter | Type | Default | Meaning |
|---|---|---|---|
original_session_id | str | - | A completed session you own. |
file_path / source / image_files / sound_files | - | None | As in infer. |
run_dsm | bool | True | Accepted and ignored: column pruning is no longer part of retraining, so the output schema stays the parent's. |
run_dsg | bool | True | Add synthetic rows. |
output_rows | int | None (server default) | DSG target row count (sent as output_sample_size). |
enable_anomaly_detection | bool | None (inherit) | Override the parent's setting. |
run_dor | bool | True | Remove outlier rows. |
dor_eps | float | None (parent's setting, else 1.0) | Outlier-removal distance, in standardized units. Smaller values remove more rows. |
dor_min_samples | int | None (parent's setting, else 5) | Minimum neighbourhood size. Larger values remove more rows. |
download_path / return_dataframes / timeout | - | as infer | As in infer. |
Returns: Same shape as infer.
Raises: As infer.
client.retrain(sid, "march.csv", output_rows=200_000, dor_eps=0.8, download_path="out/march")
Connectors and credentials
38 connector types: sql, snowflake, bigquery, mongodb, s3, gcs, databricks, delta, fabric, synapse, kafka, mqtt, opcua, influxdb, pi_web_api, kinesis, azure_blob, adls_gen2, oracle, sap_hana, clickhouse, redshift, timescaledb, cassandra, dynamodb, elasticsearch, salesforce, hubspot, stripe, google_sheets, ga4, netsuite, sharepoint, http, google_drive, dropbox, onedrive, box. The secret field names per type are listed in the Connectors section. Every call takes either inline secrets or a saved credential_id (both: the inline values overlay the saved ones).
test_connector
test_connector(connector_type, secrets=None, credential_id=None)
Checks that the server can connect. Testing a saved credential records the result on it.
Endpoint: POST /api/v1/connectors/<type>/test
| Parameter | Type | Default | Meaning |
|---|---|---|---|
connector_type | str | - | Connector type. |
secrets | dict | None | Inline secrets. |
credential_id | str | None | Saved credential. |
Returns: {success: bool, message}.
ok = client.test_connector("s3", secrets={"bucket": "my-data", "region": "eu-west-1",
"access_key_id": "AKIA...", "secret_access_key": "..."})
discover
discover(connector_type, secrets=None, credential_id=None)
Lists what can be read: tables and views (warehouses), collections or indexes (NoSQL), object keys under the prefix, capped at 500 (object stores), topics or streams (Kafka, Kinesis), the URL itself (HTTP) or the table path (Delta).
Endpoint: POST /api/v1/connectors/<type>/discover
| Parameter | Type | Default | Meaning |
|---|---|---|---|
connector_type | str | - | Connector type. |
secrets | dict | None | Inline secrets. |
credential_id | str | None | Saved credential. |
Returns: list[str] (the response's tables).
preview
preview(connector_type, table, secrets=None, credential_id=None, custom_query=None)
Returns the columns and row count of one table, file or index without reading it all.
Endpoint: POST /api/v1/connectors/<type>/preview
| Parameter | Type | Default | Meaning |
|---|---|---|---|
connector_type | str | - | Connector type. |
table | str | - | Name from discover(). |
secrets / credential_id | - | None | Authentication. |
custom_query | str | None | Read-only SELECT/WITH on SQL sources; writes are rejected. |
Returns: {columns: list[str], row_count: int or None}. Cassandra always returns row_count: None.
list_credentials
list_credentials()
Lists saved connector credentials. Secret values are never returned.
Endpoint: GET /api/v1/credentials
Returns: list of credential dicts: id, name, connector_type, the last test result, hints, stored_fields (the names of the fields that are set, never their values) and write (what the connection can do as a destination: modes, target_kind, formats, options, guarded, max_rows, enabled).
save_credential
save_credential(name, connector_type, secrets)
Saves connection secrets on the server (encrypted) for reuse by every method and automation. The server checks the required fields for the type when saving.
Endpoint: POST /api/v1/credentials
| Parameter | Type | Default | Meaning |
|---|---|---|---|
name | str | - | Label. |
connector_type | str | - | Connector type. |
secrets | dict | - | Field names as in the dashboard's connection form. |
Returns: The saved credential (the response's credential), including id.
Raises: AutoDataError 400 when a required field is missing.
cred = client.save_credential("warehouse", "snowflake", {
"account": "xy12345.eu-west-1", "user": "etl", "password": "...",
"database": "SALES", "schema": "PUBLIC", "warehouse": "COMPUTE_WH"})
cred_id = cred["id"]
update_credential
update_credential(credential_id, name=None, secrets=None, merge_secrets=False, clear_secrets=None)
Renames a credential and/or changes its secrets. Empty values are not sent.
Endpoint: PUT /api/v1/credentials/<credential_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
credential_id | str | - | Credential. |
name | str | None | New label. |
secrets | dict | None | New secrets. They replace the stored set unless merge_secrets is true. |
merge_secrets | bool | False | Lay secrets over the stored ones, so only the fields you send change. |
clear_secrets | list[str] | None | Names of stored fields to remove, for example a SAS token after moving the connection to an account key. |
Returns: The full response body, {success, credential}; credential has the same fields as a list_credentials() entry.
# change one field and keep the others
client.update_credential(cred_id, secrets={"password": "..."}, merge_secrets=True)
# move a Snowflake connection from a password to a key pair
client.update_credential(cred_id, secrets={"private_key_pem": pem_text},
merge_secrets=True, clear_secrets=["password"])
# let a Salesforce connection be written to
client.update_credential(cred_id, secrets={"allow_write": True}, merge_secrets=True)
delete_credential
delete_credential(credential_id)
Deletes a saved credential.
Endpoint: DELETE /api/v1/credentials/<credential_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
credential_id | str | - | Credential. |
Returns: True when the server reports success.
write_output
write_output(session_id, connector_type, table_name, secrets=None, credential_id=None, output_stage='dsg', if_exists=None, merge_keys=None, file_format=None, partition_by=None, partition_format=None, options=None)
Writes a completed session's output to a connector destination: a table, a file, a worksheet, an index, a topic or stream, an endpoint, or records in a business or plant system. Every connector type except ga4 can be one. A saved credential's write entry in list_credentials() says which modes, file formats and options it takes.
Endpoint: POST /api/v1/sessions/<session_id>/write-output
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
connector_type | str | - | Target connector type. |
table_name | str | - | Destination name: a table (schema.table), a file name or path, a worksheet, an index, a topic, an object type, depending on the connector. |
secrets / credential_id | - | None | Authentication for the target. Salesforce, HubSpot, NetSuite, SharePoint lists, Stripe, PI Web API and OPC UA take credential_id only, and that credential must have allow_write set. |
output_stage | str | 'dsg' | Stage to export (dsg, dsm, cds, mdh, dtc, ...). If it is not available, the server falls back to dsg, then cds, mdh, dtc. |
if_exists | str | None (the destination's own default) | replace, append, fail or merge (upsert). Left out, the destination's default applies: replace where it is offered, otherwise append. |
merge_keys | list[str] | None | Key columns; required when if_exists='merge' (400 otherwise). |
file_format | str | None (csv) | For file destinations (S3, GCS, Azure Blob, ADLS Gen2, Google Drive, Dropbox, OneDrive, Box): csv, parquet, xlsx, json, jsonl, ndjson, feather or orc. With append each write is a new timestamped file; replace overwrites one fixed file. |
partition_by | str | None | Write one target per distinct value of this column, up to 1,000. |
partition_format | str | None (%Y-%m-%d) | Date pattern for a date partition column. |
options | dict | None | Destination-specific settings, for example {'partition_key_column': 'id'} for Kafka and Kinesis or {'time_column': 'ts', 'tag_columns': ['site']} for InfluxDB. |
Returns: {success, rows_written, table, ...}, with partitions when partition_by was used.
Raises: AutoDataError 400 (bad options, a mode the destination does not offer, or an output over its row limit), 403 (not your session, or a destination whose credential does not allow writing), 404 (no output, or credential not found), 409 (if_exists='fail' and the target exists).
client.write_output(sid, "sql", "ml.training_set", credential_id=cred_id,
if_exists="merge", merge_keys=["customer_id"])
# one Parquet file per day, next to the earlier ones
client.write_output(sid, "s3", "scored/orders", credential_id=s3_cred,
if_exists="append", file_format="parquet", partition_by="order_date")
Column mapping
Match your column names to a target schema.
suggest_mapping
suggest_mapping(source_columns, target_columns, threshold=0.6)
Proposes a source-to-target column pairing from name similarity.
Endpoint: POST /api/v1/mapping/suggest
| Parameter | Type | Default | Meaning |
|---|---|---|---|
source_columns | list[str] | - | Your column names. |
target_columns | list[str] | - | The schema's column names. |
threshold | float 0 to 1 | 0.6 (the server's own default is 0.4) | Minimum similarity for a suggestion. |
Returns: {success, mappings: [{source, target, confidence, auto}], unmapped_sources, unmapped_targets}. The method's docstring mentions mapping / scores; those keys do not exist.
apply_mapping
apply_mapping(session_id, mapping, output_stage='dsg')
Renames a session's output columns and returns a preview.
Endpoint: POST /api/v1/mapping/apply
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
mapping | dict[str, str] or list[dict] | - | {old_name: new_name}, or the list of {'source', 'target'} pairs that suggest_mapping returns. |
output_stage | str | 'dsg' | Stage to read. |
Returns: {success, columns, row_count, preview (5 rows), applied_mappings, dropped_columns}. Nothing is saved on the server; the result is a preview.
Raises: AutoDataError 400 for an empty mapping.
preview = client.apply_mapping(sid, {"cust_id": "customer_id"})
print(preview["columns"])
# Or pass the pairs suggest_mapping returned
pairs = client.suggest_mapping(["cust_id", "amt"], ["customer_id", "amount"])["mappings"]
client.apply_mapping(sid, pairs)
Feature selection
Rank a finished session's columns without re-running it. To select features inside a run, use tools={'feature_selection': True} and advanced_params['feature_selection_config'].
recommend_features
recommend_features(session_id, target_columns, methods=None, top_k=20)
Ranks the columns of a completed session against the targets.
Endpoint: POST /api/v1/feature-selection/recommend
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
target_columns | list[str] | - | Targets to rank against. |
methods | list[str] | ['mutual_information'] | Any of variance_threshold, correlation_pruning, pca, mutual_information, rfe, shap_importance, permutation_importance, leakage_guard. |
top_k | int | 20 | How many columns to return. |
Returns: {ranking (combined), per_method, warnings, methods, top_k, row_count, ...}.
rec = client.recommend_features(sid, ["churn"], methods=["mutual_information", "leakage_guard"], top_k=15)
for row in rec["ranking"]:
print(row)
feature_columns
feature_columns(session_id)
Lists the columns of a completed session that recommend_features can rank.
Endpoint: GET /api/v1/feature-selection/columns?session_id=...
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: {success, session_id, columns, row_count}.
Raises: AutoDataError 404 when the session is not found or has no output to rank.
Templates and guided configuration
Build a configuration once and reuse it: in automations, as a saved template, or proposed for a dataset.
pipeline_config
AutoDataClient.pipeline_config(target_columns=None, output_rows=None, tools=None, advanced_params=None, *, task_types=None, anomaly_params=None, completion_params=None, output_config=None, webhook_ids=None)
Static method. Builds the flat pipeline_config object that scheduled runs, folder listeners, SFTP inboxes and triggers store, from the same arguments process() takes. No request is sent.
| Parameter | Type | Default | Meaning |
|---|---|---|---|
target_columns | str or list[str] | None | Targets. |
output_rows | int | None | Synthetic output size. |
tools | dict[str, bool] | None | Stage switches. See the tools table. |
advanced_params | dict | None | Stage options and speed mode. See the advanced_params table. |
task_types / anomaly_params / completion_params / output_config / webhook_ids | - | None | Keyword-only. As in process. |
Returns: dict, ready to pass as pipeline_config to an automation or as config to create_template.
config = AutoDataClient.pipeline_config(
"churned", output_rows=5000, tools={"anomaly": True},
advanced_params={"text_mode": 2, "mode": "rapid"},
output_config={"model_ready_output": True})
client.create_scheduled_run("nightly", "snowflake", "CUSTOMERS", "churned",
credential_id=cred_id, pipeline_config=config)
list_templates
list_templates()
Lists your saved configurations.
Endpoint: GET /api/v1/pipeline-templates
Returns: list (the response's templates).
create_template
create_template(name, config, description=None)
Saves a configuration under a name. An account can hold up to 50 templates.
Endpoint: POST /api/v1/pipeline-templates
| Parameter | Type | Default | Meaning |
|---|---|---|---|
name | str | - | Label. |
config | dict | - | A flat settings object: what pipeline_config() returns, or the config that get_session_config() gives back for a past run. |
description | str | None | Free text. |
Returns: The saved template dict, including id.
Raises: AutoDataError 400 (invalid name or config, or the template limit is reached).
tpl = client.create_template("churn-default", config, description="Rapid, model-ready output")
update_template
update_template(template_id, **fields)
Changes only the fields you pass: name, description, config.
Endpoint: PUT /api/v1/pipeline-templates/<template_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
template_id | str | - | Template. |
**fields | - | - | Fields to change. |
Returns: The updated template dict.
Raises: AutoDataError 404 (template not found).
delete_template
delete_template(template_id)
Deletes a template.
Endpoint: DELETE /api/v1/pipeline-templates/<template_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
template_id | str | - | Template. |
Returns: {ok: True}.
decide_config
decide_config(columns, y_columns=None, excluded_columns=None, rows=None, profile_columns=None, dataset_metadata=None)
Proposes a configuration for a dataset: the dashboard's "Decide everything". Give it what validate() returned for the file and it proposes the target, the task type and the settings. Nothing is started; pass what you accept on to process().
Endpoint: POST /api/v1/pipeline-config/decide
| Parameter | Type | Default | Meaning |
|---|---|---|---|
columns | list[str] | - | Column names, as validate() returned them. |
y_columns | list[str] | None | Targets you have already chosen. They are kept. |
excluded_columns | list[str] | None | Columns you have already left out. |
rows | int | None | Row count of the dataset. |
profile_columns | list[dict] | None | The per-column profile from validate(), when it returned one. |
dataset_metadata | dict | None | A declared data dictionary (Enterprise). |
Returns: {success, tier, config, decided, needs}: config holds the proposed settings, decided names what was chosen for you and needs what is still yours to choose. Enterprise accounts also get reasons and warnings.
Raises: AutoDataError 400 when columns is empty.
info = client.validate("customers.csv")
decision = client.decide_config(
info["columns"], rows=info["dataset_info"]["rows"],
profile_columns=(info.get("profile") or {}).get("columns"))
print(decision["config"], decision["needs"])
Triggers
A trigger starts a run when something happens: an inbound webhook call, a watermark advancing, another pipeline completing, or a quality alert. After 5 consecutive failures a trigger disables itself.
list_triggers
list_triggers()
Lists your triggers. Webhook secrets are never included.
Endpoint: GET /api/v1/triggers
Returns: list of trigger dicts.
create_trigger
create_trigger(name, trigger_type, config=None, target_pipeline_config=None, enabled=True, mode='full_pipeline', base_session_id=None)
Creates a trigger.
Endpoint: POST /api/v1/triggers
| Parameter | Type | Default | Meaning |
|---|---|---|---|
name | str | - | Label. |
trigger_type | str | - | inbound_webhook, watermark_advance, pipeline_completion or quality_alert. |
config | dict | {} | Per type. inbound_webhook: optional hmac_key (string, 16+ chars). watermark_advance: credential_id, table, watermark_column, min_advance (positive int), all required; optional source_key. pipeline_completion: source_session_name or source_template_id. quality_alert: nothing required. |
target_pipeline_config | dict | {} | What to run: source_type (or connector_type), credential_id, table / custom_query, target_columns (or y_columns), output_rows, and any pipeline option in snake_case (text_mode, enable_dsg, ...). For an inbound webhook, a CSV or JSON body posted to the trigger is used as the data instead of the connector. |
enabled | bool | True | Active on creation. |
mode | str | 'full_pipeline' | Run mode; stored inside target_pipeline_config. |
base_session_id | str | None | Parent for replay modes. |
Returns: The trigger dict (id, name, trigger_type, config, target_pipeline_config, enabled, fire_count, consecutive_failures, last_error, ...). For inbound_webhook also webhook_secret and inbound_url (/trigger/inbound/<secret>). This is the only time the secret is shown.
Raises: ValueError (mode); AutoDataError 400 (invalid type or config).
Inbound calls are limited to 10 per minute and 64 KB per body. With
hmac_keyset, each call must sendX-Signature: t=<unix time>,v1=<hex HMAC-SHA256 of '<t>.' + body>, within 300 s oft.
tr = client.create_trigger(
"drop-zone", "inbound_webhook", config={"hmac_key": "a-long-shared-secret-123"},
target_pipeline_config={"target_columns": ["label"], "output_rows": 20000})
print(client.base_url + tr["inbound_url"]) # POST CSV or a JSON array of rows here
get_trigger
get_trigger(trigger_id)
Fetches one trigger (without its secret).
Endpoint: GET /api/v1/triggers/<trigger_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
trigger_id | str | - | Trigger. |
Returns: Trigger dict.
update_trigger
update_trigger(trigger_id, **fields)
Updates only the fields you pass: name, enabled, config, target_pipeline_config. The webhook secret cannot change.
Endpoint: PUT /api/v1/triggers/<trigger_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
trigger_id | str | - | Trigger. |
**fields | - | - | Fields to change. |
Returns: Updated trigger dict.
client.update_trigger(tr["id"], enabled=False)
delete_trigger
delete_trigger(trigger_id)
Deletes a trigger.
Endpoint: DELETE /api/v1/triggers/<trigger_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
trigger_id | str | - | Trigger. |
Returns: True.
test_trigger
test_trigger(trigger_id)
Dry run: validates the trigger's config and target without firing it.
Endpoint: POST /api/v1/triggers/<trigger_id>/test
| Parameter | Type | Default | Meaning |
|---|---|---|---|
trigger_id | str | - | Trigger. |
Returns: {would_fire, checks: {...}}.
Scheduled runs
Run a connector source on an interval or a cron schedule.
list_scheduled_runs
list_scheduled_runs()
Lists your schedules.
Endpoint: GET /api/v1/scheduled-runs
Returns: list (the response's scheduled_runs).
create_scheduled_run
create_scheduled_run(name, connector_type, source_table, y_columns=None, schedule_type='interval', schedule_value='1h', credential_id=None, output_rows=10000, sql_query=None, pipeline_config=None, description=None, mode='full_pipeline', base_session_id=None)
Creates a schedule. It starts active and first fires one interval (or the next cron time) from now.
Endpoint: POST /api/v1/scheduled-runs
| Parameter | Type | Default | Meaning |
|---|---|---|---|
name | str | - | Label. |
connector_type | str | - | Source connector type. |
source_table | str | - | Table or object key. |
y_columns | str or list[str] | None | Targets. Required for a full_pipeline schedule unless pipeline_config names its own y_columns (ValueError otherwise); a replay takes its targets from the base session. Added to pipeline_config unless the config names its own. |
schedule_type | str | 'interval' | interval or cron. |
schedule_value | str | '1h' | Interval: a number with m, h or d ('30m', '6h', '2d'). Cron: five fields ('0 2 * * 1'), validated by the server. |
credential_id | str | None | Saved credential for the source. |
output_rows | int | 10000 | Synthetic output size. Added to pipeline_config unless the config names its own row count. |
sql_query | str | None | Custom SQL used instead of source_table. |
pipeline_config | dict | {'yColumns': y_columns, 'outputNumber': output_rows} | Pipeline options for each run, in snake_case or camelCase (y_columns, output_rows, text_mode, mdh_mode, dsg_mode, similarity_p, excluded_columns, enable_dsg, mode = speed preset, ...). AutoDataClient.pipeline_config() builds one from the arguments process() takes. |
description | str | None | Free text. |
mode | str | 'full_pipeline' | Run mode. |
base_session_id | str | None | Parent for replay modes. |
Returns: {success, scheduled_run: {...}}.
Raises: ValueError (mode, no targets for a full_pipeline schedule); AutoDataError 400 (invalid cron), 404 (credential not yours).
The
y_columnsandoutput_rowsarguments apply whether or not you passpipeline_config. A config that names its own targets or row count takes precedence.
client.create_scheduled_run(
name="Weekly refresh", connector_type="snowflake", source_table="ORDERS",
y_columns="churned", schedule_type="cron", schedule_value="0 2 * * 1",
credential_id=cred_id, mode="retraining", base_session_id=sid)
update_scheduled_run
update_scheduled_run(run_id, schedule_type=None, schedule_value=None, mode=None, base_session_id=None, pipeline_config=None)
Changes when a schedule fires and what it runs. Only the arguments you pass are sent. pipeline_config replaces the stored one.
Endpoint: PUT /api/v1/scheduled-runs/<run_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
run_id | str | - | Schedule. |
schedule_type / schedule_value | str | None | As in create. |
mode | str | None | New run mode; a replay mode needs base_session_id. |
base_session_id | str | None | Parent. |
pipeline_config | dict | None | Full replacement, including media options (image_processing_mode, image_columns, ...). |
Returns: Server response with the updated schedule.
delete_scheduled_run
delete_scheduled_run(run_id)
Deletes a schedule.
Endpoint: DELETE /api/v1/scheduled-runs/<run_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
run_id | str | - | Schedule. |
Returns: Server response.
toggle_scheduled_run
toggle_scheduled_run(run_id)
Switches a schedule between active and paused. It takes only the id and flips the current state.
Endpoint: POST /api/v1/scheduled-runs/<run_id>/toggle
| Parameter | Type | Default | Meaning |
|---|---|---|---|
run_id | str | - | Schedule. |
Returns: {success, is_active, ...}.
run_scheduled_now
run_scheduled_now(run_id)
Runs a schedule immediately, outside its timetable.
Endpoint: POST /api/v1/scheduled-runs/<run_id>/run-now
| Parameter | Type | Default | Meaning |
|---|---|---|---|
run_id | str | - | Schedule. |
Returns: The run result: the new session_id, or status: 'skipped' when the source returned no rows.
Folder listeners
Watch a folder or bucket and run the pipeline on each new file.
list_listeners
list_listeners()
Lists your folder listeners.
Endpoint: GET /api/v1/listeners/folder
Returns: list (the response's listeners).
create_listener
create_listener(name, source_type, folder_path=None, y_columns=None, *, bucket_name=None, prefix=None, credential_id=None, pipeline_config=None, output_rows=10000, poll_interval_seconds=60, allowed_extensions='.csv,.parquet,.json', max_file_size_mb=500, cooldown_seconds=60, max_files_per_hour=10, enabled=True, auto_sink_config=None, mode='full_pipeline', base_session_id=None, watch_path=None)
Creates a listener that runs the pipeline on each file that appears.
Endpoint: POST /api/v1/listeners/folder
| Parameter | Type | Default | Meaning |
|---|---|---|---|
name | str | - | Label. |
source_type | str | - | local, s3, gcs, azure_blob, adls_gen2 or sftp. |
folder_path | str | None | Required for local, and must be under the server's uploads/, results/ or listener_data/ folders. Derived from the bucket for cloud sources and set to your inbox for sftp. |
y_columns | str or list[str] | None | Required for full_pipeline (ValueError otherwise). |
bucket_name | str | None | Bucket, container or filesystem. Required for cloud sources. |
prefix | str | None | Only watch keys under this prefix. |
credential_id | str | None | Required for cloud sources; the server checks it fits the source type. |
pipeline_config | dict | {} | Pipeline options for each run (text_mode, mdh_mode, excluded_columns, task_types, similarity_p, dsg_mode, enable_* stage flags, webhook_ids, ...). |
output_rows | int | 10000 | Stored as pipeline_config['output_number'] unless already set there. |
poll_interval_seconds | int | 60 | How often the source is checked. |
allowed_extensions | str | '.csv,.parquet,.json' | Comma-separated extensions to pick up. |
max_file_size_mb | int | 500 | Larger files are skipped. |
cooldown_seconds | int | 60 | Minimum time between two runs. |
max_files_per_hour | int | 10 | Cap on runs per hour. |
enabled | bool | True | Start watching immediately. |
auto_sink_config | dict | None | {enabled, connector_type, credential_id, table_name, if_exists}: write each output to a sink. |
mode / base_session_id | str | 'full_pipeline' / None | Run mode and parent. |
watch_path | str | None | Deprecated alias of folder_path. |
Returns: Server response with the listener under listener.
Raises: ValueError; AutoDataError 400 (missing bucket or credential, folder not allowed), 404 (credential not found).
client.create_listener("landing", "s3", y_columns="label", bucket_name="acme-landing",
prefix="daily/", credential_id=cred_id, allowed_extensions=".csv,.parquet")
update_listener
update_listener(listener_id, *, name=None, folder_path=None, y_columns=None, bucket_name=None, prefix=None, credential_id=None, pipeline_config=None, poll_interval_seconds=None, allowed_extensions=None, max_file_size_mb=None, cooldown_seconds=None, max_files_per_hour=None, enabled=None, auto_sink_config=None, watch_path=None, mode=None, base_session_id=None)
Changes only the fields you pass. Setting enabled starts or stops the watcher.
Endpoint: PUT /api/v1/listeners/folder/<listener_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
listener_id | str | - | Listener. |
other keywords | - | None (unchanged) | As in create_listener. |
Returns: Server response.
delete_listener
delete_listener(listener_id)
Deletes a listener.
Endpoint: DELETE /api/v1/listeners/folder/<listener_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
listener_id | str | - | Listener. |
Returns: Server response.
SFTP inbox
Each SFTP credential is an inbox: every file uploaded to it starts a run.
sftp_info
sftp_info()
Returns how to connect to the SFTP server.
Endpoint: GET /api/v1/sftp/info
Returns: {enabled, host, port, instructions}.
list_sftp_credentials
list_sftp_credentials()
Lists your SFTP inboxes (no passwords).
Endpoint: GET /api/v1/sftp/credentials
Returns: list (the response's credentials).
create_sftp_credential
create_sftp_credential(name, y_columns=None, pipeline_config=None, mode='full_pipeline', base_session_id=None)
Creates an inbox with a generated username and password.
Endpoint: POST /api/v1/sftp/credentials
| Parameter | Type | Default | Meaning |
|---|---|---|---|
name | str | - | Label. |
y_columns | str or list[str] | None | Targets for full_pipeline. Without them the inbox stays inactive until it is configured. |
pipeline_config | dict | None | Pipeline options for each file. |
mode / base_session_id | str | 'full_pipeline' / None | Run mode and parent. |
Returns: {success, credential: {id, sftp_username, sftp_password, connection_string, inbox_path, is_active, listener_id, listener_enabled, y_columns, pipeline_config, mode, base_session_id, ...}}. sftp_password is shown only once.
inbox = client.create_sftp_credential("partner-feed", y_columns=["label"])["credential"]
print(inbox["connection_string"]) # sftp -P <port> <user>@<host>
password = inbox["sftp_password"] # store it now
update_sftp_pipeline
update_sftp_pipeline(credential_id, y_columns=None, pipeline_config=None, mode=None, base_session_id=None)
Changes what an inbox runs on each file. Only the arguments you pass are sent.
Endpoint: PUT /api/v1/sftp/credentials/<credential_id>/pipeline
| Parameter | Type | Default | Meaning |
|---|---|---|---|
credential_id | str | - | SFTP credential. |
y_columns | str or list[str] | None | Targets. |
pipeline_config | dict | None | Options. |
mode / base_session_id | str | None | Run mode and parent. |
Returns: The updated inbox (mode, base_session_id, listener_enabled, ...).
delete_sftp_credential
delete_sftp_credential(credential_id)
Deletes an inbox and its login.
Endpoint: DELETE /api/v1/sftp/credentials/<credential_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
credential_id | str | - | SFTP credential. |
Returns: Server response.
Runs held for approval
An automation with a price limit (max_price_dollars in its pipeline config) holds a run priced above the limit instead of starting it. Review, approve or discard held runs here.
list_pending_approvals
list_pending_approvals()
Lists your held runs.
Endpoint: GET /api/v1/pending-approvals
Returns: list (the response's approvals); each entry carries id, session_id, source, name, price_usd, cap_usd, status, created_at and decided_at.
approve_run
approve_run(approval_id)
Starts a held run at the price it was held at. Exactly that price is charged.
Endpoint: POST /api/v1/pending-approvals/<approval_id>/approve
| Parameter | Type | Default | Meaning |
|---|---|---|---|
approval_id | str | - | The held run's id. |
Returns: {success, session_id, price_usd, approval}. Follow it with wait_for_completion(session_id).
Raises: AutoDataError 404 (unknown or not yours), 409 (already approved or discarded), 410 (the held data is no longer available).
discard_run
discard_run(approval_id)
Drops a held run without starting it. Nothing is charged.
Endpoint: POST /api/v1/pending-approvals/<approval_id>/discard
| Parameter | Type | Default | Meaning |
|---|---|---|---|
approval_id | str | - | The held run's id. |
Returns: Server response.
Raises: As approve_run.
for held in client.list_pending_approvals():
if held["price_usd"] <= 5:
client.approve_run(held["id"])
else:
client.discard_run(held["id"])
Streaming inference (Enterprise)
Continuous inference: rows arriving on a Kafka, Kinesis or MQTT source are prepared with a completed session's fitted transforms and written to a sink. These calls need an Enterprise account and answer 403 otherwise. Streams are billed hourly for the rows scored, at the inference rate. The Streaming page describes every field.
list_streams (Enterprise)
list_streams()
Lists your streams.
Endpoint: GET /api/v1/streams
Returns: list (the response's streams).
create_stream (Enterprise)
create_stream(name, base_session_id, source_type, source_credential_id, source_config, sink_type, sink_credential_id, sink_config, dead_letter_config=None, start=True, **options)
Creates a stream that scores a live source against a completed session, continuously.
Endpoint: POST /api/v1/streams
| Parameter | Type | Default | Meaning |
|---|---|---|---|
name | str | - | Label. |
base_session_id | str | - | The completed session whose transforms are applied. |
source_type | str | - | kafka, kinesis or mqtt. |
source_credential_id | str | - | Saved credential for the source. |
source_config | dict | - | Where to read, for example {'topic': 'events'} or {'stream_name': 'events'}. |
sink_type / sink_credential_id / sink_config | - | - | The same, for where prepared rows are written. sink_config['output_format'] is json, values or sparse. |
dead_letter_config | dict | None | Optional sink for rows that could not be prepared: {sink_type, credential_id, topic or stream_name}. |
start | bool | True | Start consuming immediately. |
**options | - | - | max_batch_rows, max_batch_ms, parallelism, enable_anomaly. |
Returns: dict with the created stream (HTTP 201).
Raises: AutoDataError 400 (invalid configuration), 403 (not an Enterprise account).
created = client.create_stream(
"orders-live", base_session_id=sid,
source_type="kafka", source_credential_id=kafka_cred, source_config={"topic": "orders"},
sink_type="kafka", sink_credential_id=kafka_cred,
sink_config={"topic": "orders-prepared", "output_format": "json"})
stream_id = created["stream"]["id"]
print(client.stream_status(stream_id))
get_stream (Enterprise)
get_stream(stream_id)
One stream's configuration and state.
Endpoint: GET /api/v1/streams/<stream_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
stream_id | str | - | Stream. |
Returns: Stream dict.
update_stream (Enterprise)
update_stream(stream_id, **fields)
Changes only the fields you pass (the fields of create_stream). dead_letter_config=None removes the dead-letter sink.
Endpoint: PUT /api/v1/streams/<stream_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
stream_id | str | - | Stream. |
**fields | - | - | Fields to change. |
Returns: Server response.
delete_stream (Enterprise)
delete_stream(stream_id)
Stops and deletes a stream.
Endpoint: DELETE /api/v1/streams/<stream_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
stream_id | str | - | Stream. |
Returns: Server response.
start_stream (Enterprise)
start_stream(stream_id)
Starts consuming.
Endpoint: POST /api/v1/streams/<stream_id>/start
| Parameter | Type | Default | Meaning |
|---|---|---|---|
stream_id | str | - | Stream. |
Returns: Stream dict.
stop_stream (Enterprise)
stop_stream(stream_id)
Stops consuming after the batch in hand is delivered.
Endpoint: POST /api/v1/streams/<stream_id>/stop
| Parameter | Type | Default | Meaning |
|---|---|---|---|
stream_id | str | - | Stream. |
Returns: Stream dict.
restart_stream (Enterprise)
restart_stream(stream_id)
Restarts a stream, picking up a changed configuration.
Endpoint: POST /api/v1/streams/<stream_id>/restart
| Parameter | Type | Default | Meaning |
|---|---|---|---|
stream_id | str | - | Stream. |
Returns: Stream dict.
stream_status (Enterprise)
stream_status(stream_id)
Live state of a stream: workers, throughput, lag and the last error.
Endpoint: GET /api/v1/streams/<stream_id>/status
| Parameter | Type | Default | Meaning |
|---|---|---|---|
stream_id | str | - | Stream. |
Returns: Status dict.
stream_schema (Enterprise)
stream_schema(stream_id)
The columns a stream writes, in order: what the values and sparse output formats index into.
Endpoint: GET /api/v1/streams/<stream_id>/schema
| Parameter | Type | Default | Meaning |
|---|---|---|---|
stream_id | str | - | Stream. |
Returns: Schema dict.
test_stream_connection (Enterprise)
test_stream_connection(connector_type, credential_id, config=None, role='source')
Checks that a saved credential can reach a topic or stream before you build a stream on it.
Endpoint: POST /api/v1/streams/test-connection
| Parameter | Type | Default | Meaning |
|---|---|---|---|
connector_type | str | - | kafka, kinesis or mqtt. |
credential_id | str | - | Saved credential. |
config | dict | {} | The source or sink config to check. |
role | str | 'source' | source or sink. |
Returns: The result of the check.
Sync watermarks
Watermarks record how far an incremental read has got.
list_watermarks
list_watermarks()
Lists your sync watermarks.
Endpoint: GET /api/v1/sync/watermarks/list
Returns: list of watermark dicts (id, connector_type, source_key, watermark_column, last value, row count, ...).
reset_watermark
reset_watermark(watermark_id)
Resets a watermark so the next incremental read is a full read.
Endpoint: POST /api/v1/sync/watermarks/<watermark_id>/reset
| Parameter | Type | Default | Meaning |
|---|---|---|---|
watermark_id | str | - | Watermark. |
Returns: {success, ...} with the updated watermark.
get_watermark
get_watermark(connector_type, source_key)
Reads the stored watermark of one source.
Endpoint: GET /api/v1/sync/watermarks?connector_type=...&source_key=...
| Parameter | Type | Default | Meaning |
|---|---|---|---|
connector_type | str | - | Connector type. |
source_key | str | - | The saved credential's id when one was used, otherwise the connector type. |
Returns: The watermark dict, or None when none is stored.
set_watermark
set_watermark(connector_type, source_key, **fields)
Creates or moves a watermark.
Endpoint: PUT /api/v1/sync/watermarks
| Parameter | Type | Default | Meaning |
|---|---|---|---|
connector_type | str | - | Connector type. |
source_key | str | - | The saved credential's id when one was used, otherwise the connector type. |
**fields | - | - | table_name, watermark_column, watermark_value, watermark_type (timestamp, integer or string), row_count, last_sync_at (ISO 8601). |
Returns: The watermark dict.
create_watermark
create_watermark(connector_type, source_key, watermark_column, watermark_value, table_name=None, watermark_type='timestamp')
Sets the starting point of an incremental read before its first run.
Endpoint: POST /api/v1/sync/watermarks
| Parameter | Type | Default | Meaning |
|---|---|---|---|
connector_type | str | - | Connector type. |
source_key | str | - | The saved credential's id when one was used, otherwise the connector type. |
watermark_column | str | - | The column that orders the source. |
watermark_value | any | - | The value the first incremental read starts after. |
table_name | str | None | Table the watermark belongs to. |
watermark_type | str | 'timestamp' | timestamp, integer or string. |
Returns: The watermark dict.
Raises: AutoDataError 409 when a watermark already exists for the source; use set_watermark to move one.
client.create_watermark("snowflake", cred_id, "updated_at", "2026-10-01T00:00:00")
print(client.get_watermark("snowflake", cred_id))
delete_watermark
delete_watermark(connector_type, source_key)
Removes a watermark; the next incremental read is a full read.
Endpoint: DELETE /api/v1/sync/watermarks?connector_type=...&source_key=...
| Parameter | Type | Default | Meaning |
|---|---|---|---|
connector_type | str | - | Connector type. |
source_key | str | - | The saved credential's id when one was used, otherwise the connector type. |
Returns: {success, message}.
Raises: AutoDataError 404 when no watermark is stored for the source.
Quality alerts
Rules checked against each run's metrics.
list_quality_alerts
list_quality_alerts()
Lists your alert rules.
Endpoint: GET /api/v1/quality-alerts
Returns: list (the response's rules).
create_quality_alert
create_quality_alert(name, metric, operator, threshold, severity='warning', stage=None)
Creates a rule.
Endpoint: POST /api/v1/quality-alerts
| Parameter | Type | Default | Meaning |
|---|---|---|---|
name | str | - | Label. |
metric | str | - | row_loss_pct, null_pct, column_drop_count, duration_seconds or drift_score. |
operator | str | - | >, <, >=, <= or ==. |
threshold | float | - | Value to compare with. |
severity | str | 'warning' | warning or critical. |
stage | str | None (every stage) | anomaly, dtc, mdh, cds, dsm or dsg. |
Returns: {success, rule: {...}}.
Raises: AutoDataError 400 (invalid metric or operator).
The SDK cannot attach
webhook_idsor setenabled. Rules with webhooks are stored, butquality.breachnotifications are not delivered yet.
client.create_quality_alert("too many rows lost", "row_loss_pct", ">", 20, severity="critical", stage="mdh")
update_quality_alert
update_quality_alert(rule_id, name=None, metric=None, operator=None, threshold=None, severity=None, stage=None)
Changes only the fields you pass.
Endpoint: PUT /api/v1/quality-alerts/<rule_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
rule_id | str | - | Rule. |
other arguments | - | None (unchanged) | As in create. |
Returns: {success, rule}.
delete_quality_alert
delete_quality_alert(rule_id)
Deletes a rule.
Endpoint: DELETE /api/v1/quality-alerts/<rule_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
rule_id | str | - | Rule. |
Returns: {success, message}.
get_alert_events
get_alert_events(session_id=None)
Lists the most recent alerts that fired (up to 100).
Endpoint: GET /api/v1/quality-alerts/events
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | None | Only events of this session. |
Returns: list (the response's events).
Webhooks
Get a POST when something happens instead of polling. Each delivery body is {event, data, timestamp, webhook_id}. If the webhook has a secret, the request carries X-Webhook-Signature: sha256=<hex>: the HMAC-SHA256 of the body, re-serialized with sorted keys (Python json.dumps(body, sort_keys=True)), keyed with the secret.
list_webhooks
list_webhooks()
Lists your webhooks.
Endpoint: GET /api/v1/webhooks
Returns: list of webhook dicts.
create_webhook
create_webhook(url, events=None, name=None, secret=None)
Registers an endpoint. The URL must resolve to a public address (checked again at every delivery; redirects are not followed).
Endpoint: POST /api/v1/webhooks
| Parameter | Type | Default | Meaning |
|---|---|---|---|
url | str | - | HTTPS endpoint. |
events | list[str] | None, which subscribes to ['error'] only | Any of job.completed, job.failed, quality.breach, error, warning, info, success. |
name | str | 'Unnamed Webhook' | Label. |
secret | str | None | Signing secret. Set it, and verify the signature. |
Returns: {webhook_id, message, webhook: {...}} (HTTP 201).
Raises: AutoDataError 400 (URL not allowed, unknown event).
Pass
eventsexplicitly. Leaving it out subscribes only toerror, so you will not hear about completed jobs.
import hashlib, hmac, json
client.create_webhook("https://example.com/hooks/autodata",
events=["job.completed", "job.failed"], secret="s3cret")
# In your receiver:
def verify(raw_body: bytes, header: str, secret: str) -> bool:
canonical = json.dumps(json.loads(raw_body), sort_keys=True)
mac = hmac.new(secret.encode(), canonical.encode(), hashlib.sha256).hexdigest()
return hmac.compare_digest(header, f"sha256={mac}")
get_webhook
get_webhook(webhook_id)
Fetches one webhook.
Endpoint: GET /api/v1/webhooks/<webhook_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
webhook_id | str | - | Webhook. |
Returns: Webhook dict.
update_webhook
update_webhook(webhook_id, **fields)
Changes only the fields you pass: url, events, name, secret, active (is_active is accepted as another spelling of active).
Endpoint: PUT /api/v1/webhooks/<webhook_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
webhook_id | str | - | Webhook. |
**fields | - | - | Fields to change. |
Returns: {message, webhook}.
Pass
active=to enable or disable the webhook.
client.update_webhook(wh_id, active=False) # pause deliveries
delete_webhook
delete_webhook(webhook_id)
Deletes a webhook.
Endpoint: DELETE /api/v1/webhooks/<webhook_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
webhook_id | str | - | Webhook. |
Returns: {message}.
test_webhook
test_webhook(webhook_id)
Sends a signed test delivery now.
Endpoint: POST /api/v1/webhooks/<webhook_id>/test
| Parameter | Type | Default | Meaning |
|---|---|---|---|
webhook_id | str | - | Webhook. |
Returns: {..., response_status}, or {warning, response_status} when your endpoint answered with an error status.
webhook_deliveries
webhook_deliveries(webhook_id)
Lists recent delivery attempts.
Endpoint: GET /api/v1/webhooks/<webhook_id>/deliveries
| Parameter | Type | Default | Meaning |
|---|---|---|---|
webhook_id | str | - | Webhook. |
Returns: list of deliveries (HTTP status, response, duration, error).
Notifications
E-mail when a run completes or ends in an error.
get_notification_preferences
get_notification_preferences()
Whether the account is e-mailed when a run completes or ends in an error.
Endpoint: GET /api/v1/notification-preferences
Returns: {email_on_complete, email_on_fail, configured}; configured tells whether the server can send e-mail.
set_notification_preferences
set_notification_preferences(email_on_complete=None, email_on_fail=None)
Turns the two e-mails on or off. Only the settings you pass are changed.
Endpoint: PUT /api/v1/notification-preferences
| Parameter | Type | Default | Meaning |
|---|---|---|---|
email_on_complete | bool | None (unchanged) | E-mail when a run completes. |
email_on_fail | bool | None (unchanged) | E-mail when a run ends in an error or is held for approval. |
Returns: The same dict as get_notification_preferences().
client.set_notification_preferences(email_on_fail=True)
Sessions
Your run history.
list_sessions
list_sessions()
Lists your sessions, newest first.
Endpoint: GET /api/v1/sessions
Returns: list of session dicts (session_id, session_name, status, total_rows, processing_date, ...).
get_session
get_session(session_id)
Full detail of one session, including the parameters it ran with.
Endpoint: GET /api/v1/sessions/<session_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: Session detail dict.
rename_session
rename_session(session_id, name)
Sets a session's display name.
Endpoint: PUT /api/v1/sessions/<session_id>/name
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
name | str | - | New name. |
Returns: Server response.
set_session_retraining
set_session_retraining(session_id, enabled)
Marks or unmarks a session as a retraining base.
Endpoint: PUT /api/v1/sessions/<session_id>/retraining
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
enabled | bool | - | Flag. |
Returns: Server response.
list_session_outputs
list_session_outputs(session_id)
Lists a session's outputs kept in memory rather than written to disk.
Endpoint: GET /api/v1/sessions/<session_id>/outputs
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: Server response listing memory outputs.
get_session_report
get_session_report(session_id)
What a finished run did to the data, stage by stage: the report the dashboard shows.
Endpoint: GET /api/v1/sessions/<session_id>/report
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: {session_id, tier, status, summary, warnings}, plus details (per-column detail) for Enterprise accounts.
get_session_config
get_session_config(session_id)
The configuration a session ran with, ready to reuse for a new run or to save as a template.
Endpoint: GET /api/v1/sessions/<session_id>/config
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: {session_id, config, run_mode}.
set_session_meta
set_session_meta(session_id, tags=None, note=None)
Tags a session and/or attaches a note to it. Only what you pass is changed; tags=[] clears the tags and note='' clears the note.
Endpoint: PUT /api/v1/sessions/<session_id>/meta
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
tags | list[str] | None (unchanged) | Up to 10 tags of up to 32 characters each. |
note | str | None (unchanged) | Up to 500 characters. |
Returns: {tags, note}.
Raises: AutoDataError 400 (too many or too long), 404 (session not found).
client.set_session_meta(sid, tags=["churn", "q4"], note="Baseline for the Q4 model")
list_session_files
list_session_files(session_id)
Every downloadable file of a session you own, or the shared files of a session shared with you.
Endpoint: GET /api/v1/sessions/<session_id>/files
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session you own or that is shared with you. |
Returns: list (the response's files); each entry carries name, size_mb and url.
download_all
download_all(session_id, download_path=None, extract=True)
Downloads every output of a session as one archive.
Endpoint: GET /api/v1/sessions/<session_id>/download-all
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
download_path | str | ./auto_data_outputs/<session_id>/ | Folder, created if needed. |
extract | bool | True | Unpack the archive. With False the .zip itself is saved. |
Returns: Absolute path of the folder, or of the .zip when extract=False.
Raises: AutoDataError on an empty or invalid archive.
download_session_output
download_session_output(session_id, output_name, output_path=None)
Downloads one output that is held in memory rather than on disk.
Endpoint: GET /api/v1/sessions/<session_id>/outputs/<output_name>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
output_name | str | - | A name from list_session_outputs(). |
output_path | str | ./auto_data_outputs/<session_id>/<output_name>.csv | Local file path. |
Returns: Absolute path of the file written.
clear_session_outputs
clear_session_outputs(session_id)
Releases a session's in-memory outputs. Files on disk are not affected.
Endpoint: DELETE /api/v1/sessions/<session_id>/outputs
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: Server response.
Sharing
Share a session with another account or through a link.
list_shared_sessions
list_shared_sessions()
Shares you created.
Endpoint: GET /api/v1/shared-sessions
Returns: list of shares (with session_name, shared_with_name, permission, expires_at, share_token for links, ...).
list_received_sessions
list_received_sessions()
Sessions other accounts shared with you.
Endpoint: GET /api/v1/shared-sessions/received
Returns: list of shares (with owner_name, session_name, permission, ...).
share_session
share_session(session_id, shared_with_id=None, permission='read', expires_at=None, note=None, shared_files=None)
Shares one of your sessions.
Endpoint: POST /api/v1/shared-sessions
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
shared_with_id | str | None | Recipient's user id. Omit to create a link token instead. |
permission | str | 'read' | read, comment or edit. Unknown values become read. |
expires_at | str (ISO 8601) | None (never) | Expiry. Strongly recommended for links. |
note | str | None | Shown to the recipient. |
shared_files | list[str] | None | File names the recipient may download (names from list_session_files()). A share without any shows the run and its report and offers no file, so pass this whenever the data itself is what you are sharing. |
Returns: {success, shared_session: {...}} (with share_token for a link).
Raises: AutoDataError 404 (not your session).
unshare_session
unshare_session(share_id)
Revokes a share; a link stops working at once.
Endpoint: DELETE /api/v1/shared-sessions/<share_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
share_id | str | - | The share's id. |
Returns: Server response.
update_share
update_share(share_id, **fields)
Changes a share you created. A change to a link applies to everyone who joined it.
Endpoint: PUT /api/v1/shared-sessions/<share_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
share_id | str | - | The share's id. |
**fields | - | - | permission (read, comment or edit), expires_at (ISO 8601, or None to remove the expiry), note, shared_files. |
Returns: Server response.
get_shared_session
get_shared_session(token)
Opens a share link: the session's summary and the files it offers.
Endpoint: GET /api/v1/shared-sessions/access/<token>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
token | str | - | The share link's token. |
Returns: dict with the shared session and its files.
Raises: AutoDataError 404 (unknown link), 410 (expired link).
join_shared_session
join_shared_session(token)
Adds a share link to your own account, so it appears in list_received_sessions() and its session can be the base of an inference or retraining run.
Endpoint: POST /api/v1/shared-sessions/access/<token>/join
| Parameter | Type | Default | Meaning |
|---|---|---|---|
token | str | - | The share link's token. |
Returns: Server response.
download_shared_file
download_shared_file(token, filename, output_path=None)
Downloads one file through a share link.
Endpoint: GET /api/v1/shared-sessions/access/<token>/download/<filename>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
token | str | - | The share link's token. |
filename | str | - | One of the shared file names. |
output_path | str | ./<filename> | Local file path. |
Returns: Absolute path of the file written.
list_share_comments
list_share_comments(session_id)
Comments on a session you own or that is shared with you.
Endpoint: GET /api/v1/shared-sessions/by-session/<session_id>/comments
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session you own or that is shared with you. |
Returns: list (the response's comments).
add_share_comment
add_share_comment(session_id, body)
Comments on a session. On a session that is not your own you need comment or edit permission.
Endpoint: POST /api/v1/shared-sessions/by-session/<session_id>/comments
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session you own or that is shared with you. |
body | str | - | The comment, up to 2000 characters. |
Returns: Server response with the new comment.
delete_share_comment
delete_share_comment(session_id, comment_id)
Deletes a comment: your own, or any comment on a session you own.
Endpoint: DELETE /api/v1/shared-sessions/by-session/<session_id>/comments/<comment_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | The session. |
comment_id | str | - | The comment. |
Returns: Server response.
search_users
search_users(query)
Finds the account to share with, by exact e-mail address. It is a lookup, not a directory: it answers only for a complete address, and is limited to 20 searches per minute.
Endpoint: GET /api/v1/users/search?q=...
| Parameter | Type | Default | Meaning |
|---|---|---|---|
query | str | - | The recipient's complete e-mail address. |
Returns: list (the response's users); each entry carries id, name and email_hint. Pass id as shared_with_id.
user = client.search_users("ana@example.com")[0]
files = [f["name"] for f in client.list_session_files(sid)]
client.share_session(sid, shared_with_id=user["id"], permission="comment", shared_files=files)
Preferences
Saved preference profiles. AutoDataClient.PREFERENCE_KINDS is ('outputs', 'advanced', 'anomaly', 'completion', 'api_anomaly'); any other kind raises ValueError.
get_preferences
get_preferences(kind)
Reads one profile.
Endpoint: GET /api/v1/preferences/<kind> (api_anomaly is /api/v1/preferences/api-anomaly)
| Parameter | Type | Default | Meaning |
|---|---|---|---|
kind | str | - | outputs (which files are kept, output format and name pattern), advanced (dashboard pipeline defaults), anomaly (data-cleaning checks of runs started in the dashboard), completion (Data Completion & Verification settings), api_anomaly (data-cleaning defaults of runs started through the API or this client: what process() uses when anomaly_params is not passed). |
Returns: {success, preferences} for outputs or {success, parameters} for the others, plus the available names.
set_preferences
set_preferences(kind, values)
Replaces one profile.
Endpoint: POST /api/v1/preferences/<kind>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
kind | str | - | As above. |
values | dict | - | The profile itself: the inner preferences / parameters dict from get_preferences, not the whole response. |
Returns: {success, message, preferences|parameters}.
Two profiles apply to runs started with
process():outputsdecides which files are written unless the call passesoutput_config, andapi_anomalysupplies the cleaning checks unless the call passesanomaly_params. Theadvanced,anomalyandcompletionprofiles are dashboard defaults.
prefs = client.get_preferences("outputs")["preferences"]
prefs["output_format"] = "parquet"
client.set_preferences("outputs", prefs)
# Cleaning defaults for runs started through the API or the SDK
cleaning = client.get_preferences("api_anomaly")["parameters"]
reset_preferences
reset_preferences(kind)
Restores one profile to the defaults.
Endpoint: POST /api/v1/preferences/<kind>/reset
| Parameter | Type | Default | Meaning |
|---|---|---|---|
kind | str | - | As above. |
Returns: Server response.
Account, usage and cost
list_keys
list_keys()
Lists your active API keys (no secrets).
Endpoint: GET /api/v1/keys
Returns: list (the response's api_keys); each key carries daily_limit_usd, lifetime_limit_usd, daily_spent_usd and lifetime_spent_usd in US dollars (None = no limit).
get_usage
get_usage()
Spend, in US dollars, of the key making the call.
Endpoint: GET /api/v1/usage
Returns: {api_key_id, name, key_prefix, today_usd, daily_limit_usd, daily_remaining_usd, has_daily_limit, lifetime_spent_usd, lifetime_limit_usd, lifetime_remaining_usd, has_lifetime_limit, avg_daily_usd_7d, avg_daily_usd_30d, chart_data: {labels, usd}, daily_request_count, last_used_at}. A limit of None means no limit. {} with passcode auth. Deprecated, still returned, equal to the dollar amount × 1,000,000, and will be removed in a future version: daily_credits_used, daily_credit_limit, daily_remaining, lifetime_credits_used, lifetime_credit_limit, lifetime_remaining.
account_profile
account_profile()
Balance in US dollars, role and API limits of the whole account.
Endpoint: GET /api/v1/account/profile
Returns: Profile dict with balance_usd, spent_usd and total_usd (balance + spent), the role and the API limits. api_credits_remaining, api_credits_used and api_credits_total are deprecated.
account_usage
account_usage()
Account-wide usage across every key and the dashboard.
Endpoint: GET /api/v1/account/usage
Returns: {summary, by_tool}: summary.total_spent_usd and a spent_usd per stage (credits_used and total_credits_used are deprecated).
account_contracts
account_contracts()
Your pricing contracts, with usage against each allowance.
Endpoint: GET /api/v1/account/contracts
Returns: list (the response's contracts); empty when the account has none.
trial_status
trial_status()
Whether the account is on a trial, and how many free runs remain.
Endpoint: GET /api/v1/account/trial-status
Returns: {success, is_trial, trial_uploads_remaining, role, role_name, total_runs, completed_runs, failed_runs, total_data_mb, limits}.
estimate
estimate(input_rows=None, columns=None, column_types=None, missing_rate=None, tools=None, output_rows=None, text_mode=None, session_type='api', pricing_profile_id=None, media_bytes=None, completion=None)
Prices a run before you start it. Runs and charges nothing. With pricing_profile_id (from validate()) it returns the exact price of that file with these settings; start the run with the returned quote_id and it is charged exactly that. With a description of the data instead, it returns an estimated range, which binds nothing.
Endpoint: POST /api/v1/estimate
| Parameter | Type | Default | Meaning |
|---|---|---|---|
input_rows | int | None | Row count. Required unless pricing_profile_id is given. |
columns | int | None | Column count. Required unless column_types or pricing_profile_id is given. |
column_types | dict | None | Optional counts by kind: numeric, boolean, categorical_low, categorical_high, datetime, text, identifier. Columns you don't describe are priced across the whole plausible range, so describing them narrows it. |
missing_rate | float | None | Optional share of missing cells (0–1, or a percentage). |
tools | dict | normal run (anomaly off) | anomaly, dtc, mdh, cds, dsm, dsg. |
output_rows | int | None | Target row count; synthetic rows are charged per doubling of the input rows. |
text_mode | int | None | 0 drop text, 1 neural, 2 TF-IDF. In a range, automatic is counted at the neural price. |
session_type | str | 'api' | api (full pipeline), inference (about a fifth) or retraining (about two fifths). |
pricing_profile_id | str | None | The profile validate() returned for a file, once ready (see wait_for_price). Gives the exact price. |
media_bytes | dict | None | {'image': bytes, 'audio': bytes} of the media the run will upload. |
completion | dict | None | Dataset completion and validation settings; included in the exact price. |
Returns: with pricing_profile_id, the exact quote {quote_id, exact: True, price_usd, breakdown: [{item, usd}], rate_card_version, ...}, or {pricing_status: 'measuring', progress_pct} while the file is still being measured. With a description, a range {quote_id, exact: False, low_usd, high_usd, point_usd, breakdown, rate_card_version}. All prices are US dollars. The *_credits fields (price_credits, low_credits, high_credits, point_credits) are deprecated: still returned, equal to the dollar amount × 1,000,000, and will be removed in a future version.
Price with the same stages,
output_rows,text_mode, media and completion settings you will run with. An exact quote holds for the same file (checked by content) and settings that price the same, for 7 days; otherwiseprocess()raisesAutoDataErrorwithstatus_code409 and nothing starts. Without aquote_id, a run is charged the exact price of its data, measured when it starts.
info = client.validate("data.csv")
if info.get("pricing_status") == "measuring":
client.wait_for_price(info["pricing_profile_id"], timeout=600)
q = client.estimate(pricing_profile_id=info["pricing_profile_id"], output_rows=100_000, text_mode=2)
print(q["price_usd"])
result = client.process("data.csv", target_columns="label", output_rows=100_000,
advanced_params={"text_mode": 2, "quote_id": q["quote_id"]})
worker_status
worker_status()
State of the processing fleet.
Endpoint: GET /api/v1/workers/status
Returns: Worker status dict with cpu_worker queue statistics.
Enterprise: reading what a run decided (Enterprise)
These calls need an Enterprise account and answer 403 otherwise. To use Enterprise preparation, pass the fields listed under Enterprise options in process()'s advanced_params.
pivot_suggest (Enterprise)
pivot_suggest(session_id, target_columns=None)
Ranks the columns that could be the as-of time of a session's rows, with the reasons for each.
Endpoint: POST /api/v1/pivot/suggest
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
target_columns | list[str] | [] | Targets, excluded from the candidates. |
Returns: {chosen, candidates, temporal_positions}.
enterprise_families (Enterprise)
enterprise_families()
Model families a run can target, and the per-column imputation and scaling methods you can pin.
Endpoint: GET /api/v1/enterprise/families
Returns: Families and method lists.
enterprise_columns (Enterprise)
enterprise_columns(session_id, family=None)
Per-column record: what each column received and what else was considered.
Endpoint: GET /api/v1/enterprise/runs/<session_id>/columns
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
family | str | None (recommended family) | Model family to read. |
Returns: {session_id, family, ...ledger}.
Raises: AutoDataError 404 when the session or family has no record.
enterprise_leaderboard (Enterprise)
enterprise_leaderboard(session_id)
How the model families compared, and which is recommended.
Endpoint: GET /api/v1/enterprise/runs/<session_id>/leaderboard
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: {session_id, ...leaderboard}.
enterprise_lineage (Enterprise)
enterprise_lineage(session_id)
Everything recorded about a run: stages, families and decisions.
Endpoint: GET /api/v1/enterprise/lineage/<session_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: Lineage dict.
enterprise_diff (Enterprise)
enterprise_diff(session_id, against, family=None)
What changed between this run's decisions and another run's.
Endpoint: GET /api/v1/enterprise/lineage/<session_id>/diff?against=...
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
against | str | - | The other session. |
family | str | None | Restrict to one model family. |
Returns: Diff dict.
Raises: AutoDataError 404 (comparison session not found).
enterprise_audit_report (Enterprise)
enterprise_audit_report(session_id, fmt='json')
Downloadable audit trail of a run.
Endpoint: GET /api/v1/enterprise/audit-report/<session_id>?format=...
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
fmt | str | 'json' | json (returns a dict) or html (returns the page as a string). |
Returns: dict or str.
validate_dataset_metadata (Enterprise)
validate_dataset_metadata(metadata, columns=None, y_columns=None)
Checks a dataset-metadata declaration before you run with it. Nothing is stored.
Endpoint: POST /api/v1/dataset-metadata/validate
| Parameter | Type | Default | Meaning |
|---|---|---|---|
metadata | dict | - | The declaration (see dataset_metadata under Enterprise options). |
columns | list[str] | None | The dataset's column names, so declared columns can be checked against them. |
y_columns | list[str] | None | The targets, so declared roles can be checked against them. |
Returns: {metadata, warnings, summary}: the declaration as a run will read it, plus warnings.
enterprise_dataset_metadata (Enterprise)
enterprise_dataset_metadata(session_id)
The dataset metadata a finished run used, and how each declaration was applied.
Endpoint: GET /api/v1/enterprise/runs/<session_id>/dataset-metadata
| Parameter | Type | Default | Meaning |
|---|---|---|---|
session_id | str | - | A session owned by the caller. |
Returns: Metadata report dict.
catalog_test (Enterprise)
catalog_test(catalog, credential_id=None, secrets=None)
Checks that a data catalog can be reached with a connection.
Endpoint: POST /api/v1/dataset-metadata/catalog/test
| Parameter | Type | Default | Meaning |
|---|---|---|---|
catalog | str | - | Catalog kind, for example snowflake, databricks, bigquery or glue. |
credential_id | str | None | Saved credential for the catalog. |
secrets | dict | None | Inline secrets (sent as inline_secrets). |
Returns: The result of the check.
catalog_tables (Enterprise)
catalog_tables(catalog, credential_id=None, secrets=None, query=None)
Lists the tables a data catalog describes.
Endpoint: POST /api/v1/dataset-metadata/catalog/tables
| Parameter | Type | Default | Meaning |
|---|---|---|---|
catalog | str | - | Catalog kind, for example snowflake, databricks, bigquery or glue. |
credential_id | str | None | Saved credential for the catalog. |
secrets | dict | None | Inline secrets (sent as inline_secrets). |
query | str | None | Optional filter on the table names. |
Returns: dict with the tables found.
catalog_import (Enterprise)
catalog_import(catalog, table, credential_id=None, secrets=None, columns=None, y_columns=None, existing=None)
Reads a table's description from a data catalog as dataset metadata, ready to pass as advanced_params={'dataset_metadata': ...}.
Endpoint: POST /api/v1/dataset-metadata/catalog/import
| Parameter | Type | Default | Meaning |
|---|---|---|---|
catalog | str | - | Catalog kind, for example snowflake, databricks, bigquery or glue. |
table | str | - | The table to import, as the catalog names it. |
credential_id | str | None | Saved credential for the catalog. |
secrets | dict | None | Inline secrets (sent as inline_secrets). |
columns / y_columns | list[str] | None | The dataset's columns and targets, as in validate_dataset_metadata. |
existing | dict | None | A declaration to merge into. What you already declared is kept. |
Returns: dict with the imported declaration.
enterprise_docs (Enterprise)
enterprise_docs()
Index of the "How It Works" notes on the pipeline's modules.
Endpoint: GET /api/v1/enterprise/docs
Returns: Index dict.
enterprise_doc (Enterprise)
enterprise_doc(module_id)
One module's "How It Works" note.
Endpoint: GET /api/v1/enterprise/docs/<module_id>
| Parameter | Type | Default | Meaning |
|---|---|---|---|
module_id | str | - | A module id from enterprise_docs(). |
Returns: The note as a dict.
Spark (large files on the server)
Only available when the server has Spark enabled; otherwise spark_read and spark_transform answer 503. Paths are server-side and must be inside the server's upload or results folders (403 otherwise).
spark_status
spark_status()
Whether Spark is available.
Endpoint: GET /api/v1/spark/status
Returns: {available: bool, ...} (reason when not).
spark_read
spark_read(file_path, sample_rows=20)
Reads a large server-side file and returns its statistics.
Endpoint: POST /api/v1/spark/read
| Parameter | Type | Default | Meaning |
|---|---|---|---|
file_path | str | - | Server-side path. |
sample_rows | int | 20 | Rows to return as a sample. |
Returns: {row_count, columns, dtypes, sample}.
spark_transform
spark_transform(file_path, output_path, operations, output_format='csv')
Applies a list of operations to a large server-side file.
Endpoint: POST /api/v1/spark/transform
| Parameter | Type | Default | Meaning |
|---|---|---|---|
file_path | str | - | Server-side input path. |
output_path | str | - | Server-side output path. |
operations | list[dict] | - | In order, e.g. {'op': 'cast', 'type_casts': {...}}, {'op': 'drop_nulls', 'subset': [...]}, {'op': 'fill_nulls', 'strategy': 'mean'}, {'op': 'normalize', 'method': 'minmax'}, {'op': 'sample', 'n': 50000}. |
output_format | str | 'csv' | csv or parquet. |
Returns: Output statistics (row_count, columns, ...).
Client lifecycle
close
close()
Closes the HTTP session. Called automatically when a with block ends.
Returns: None.
Not available in the SDK
API keys are created, rotated, edited and revoked in the dashboard (API Account page), including team keys and per-key limits; profile changes and a new passcode are made there too. The SDK lists your keys (list_keys()) and reads their usage (get_usage()). The API Settings and API Config tabs of the API Account page are reserved and are not applied to runs yet; set options on each call.
These options have REST fields but no SDK argument yet. Send them with any HTTP client using the same Authorization: Bearer dtpk_... header:
- Quality-alert
webhook_idsandenabled. - Folder-listener
watermark_columnfor incremental reads.
Usage notes for 0.15.0
process_from_connector()sends no defaults of its own: an unsettext_modeorsimilarity_ptakes the server default (2 and 0.99), the same as a file run. Connector runs that relied on text features being off should passadvanced_params={"text_mode": 0}.process()treats every target as classification unless you passtask_types, and takes its cleaning checks from your saved defaults for API runs unless you passanomaly_params.output_configselects what a run writes;output_preferencesonly filters what is downloaded.update_webhook(): passactive=(oris_active=) to enable or disable a webhook. Passeventsexplicitly when creating one; the default iserroronly.wait_for_completion'sstall_timeoutis 3600 s.