Browse documentation
2026-09-28AutoDataOpen in dashboard

Python SDK

Reference for the datatoolpack Python library: install, authentication, retries and timeouts, pipeline options, speed modes, media uploads, and every AutoDataClient method with its parameters, endpoint, return value and errors.

The official datatoolpack Python library wraps the AutoData REST API. Use it to run the pipeline on files or connectors, follow jobs, and download results as files or DataFrames. It can also replay a session for inference or retraining, and manage connectors, automations, webhooks, sharing and account settings.

Install

pip install datatoolpack

Package datatoolpack, version 0.15.0. Requires Python 3.8 or later. The only required dependency is requests. The DataFrame helpers (get_dataframe, get_dataframes, return_dataframes=True) also need pandas; install it yourself (pip install pandas, plus pyarrow for Parquet, Feather and ORC outputs). The package exports AutoDataClient, AutoDataError, ModeConfirmationRequired and __version__.

Authentication and client setup

Every call is authenticated. With an API key the client sends Authorization: Bearer dtpk_...; with a passcode it sends X-Passcode: <passcode>. Create API keys on the API Account page. A key can carry its own spending limits (in US dollars) and per-stage permissions, so a call can fail with 429 (limit reached) or 403 (stage not allowed for this key).

from datatoolpack import AutoDataClient

# 1. API key (recommended)
client = AutoDataClient(api_key="dtpk_your_key")

# 2. Passcode
client = AutoDataClient(passcode="123456789012")

# 3. Environment variables: AUTODATA_API_KEY or AUTODATA_PASSCODE,
#    optional AUTODATA_BASE_URL
client = AutoDataClient()

# Context manager closes the HTTP session for you
with AutoDataClient(api_key="dtpk_your_key") as client:
    print(client.worker_status())
AutoDataClient(api_key=None, passcode=None, base_url=None, timeout=300, max_retries=3)
ParameterTypeDefaultMeaning
api_keystr$AUTODATA_API_KEYAPI key. Must start with dtpk_, otherwise ValueError. Wins over passcode when both are given.
passcodestr$AUTODATA_PASSCODEAccount passcode, used only when no API key is given. Passcode calls have no per-key limits. get_usage() needs an API key and returns {} with a passcode.
base_urlstr$AUTODATA_BASE_URL, else https://autodata.datatoolpack.comServer address. A trailing slash is removed. Every endpoint is called under <base_url>/api/v1.
timeoutint (seconds)300Timeout for each HTTP request. Downloads, ZIP archives and infer/retrain use at least 900 s.
max_retriesint3How many times a request is retried after a transient failure. See Errors, retries and timeouts.

ValueError is raised when neither an API key nor a passcode is available, or when the key does not start with dtpk_. The client keeps one requests.Session; call close() or use with to release it. repr(client) masks the credentials.

Errors, retries and timeouts

ExceptionWhenAttributes
AutoDataErrorThe server answered with a non-2xx status, a job failed, was cancelled, stalled or timed out while waiting, or a download was empty, truncated or not a valid ZIP. The message is the server's error field when there is one.status_code (int or None for errors that are not HTTP errors)
ModeConfirmationRequired (subclass of AutoDataError)process() only: the speed mode would change a parameter you set yourself, and you did not say how to resolve it. HTTP 409.mode, adjustments (every change the mode makes), contested (only the ones that override your values), policy_signature, status_code (409). str(e) lists each contested parameter with your value and the mode's value.
ValueErrorRaised by the client before any request: unsupported file extension, bad mode/base_session_id combination, missing required argument, unknown preference kind, both file_path and source.-
FileNotFoundErrorprocess() input file or a media file does not exist.-
requests.ConnectionError / requests.TimeoutThe network still fails after all retries.-
from datatoolpack import AutoDataClient, AutoDataError, ModeConfirmationRequired

client = AutoDataClient(api_key="dtpk_your_key")
try:
    result = client.process("data.csv", target_columns="label")
except ModeConfirmationRequired as e:
    print(e)                 # lists every contested parameter
    print(e.policy_signature)
except AutoDataError as e:
    if e.status_code == 429:
        print("Spending or rate limit reached:", e)
    elif e.status_code == 403:
        print("This key may not run one of the requested stages:", e)
    else:
        raise

Retries. Every request is retried automatically on HTTP 429, 502, 503 and 504 and on connection errors and timeouts, up to max_retries times. Before each retry the client waits max(Retry-After, 2**attempt) seconds (1, 2, 4, ...). Retry-After is honoured as a number of seconds or as an HTTP date, and a single wait is capped at 300 s. When the last attempt still returns a retryable status, it is raised as AutoDataError.

Retries also apply to job submission. process() sends an Idempotency-Key header with every submission (one is generated per call unless you pass idempotency_key), so a retry returns the first run instead of starting a second one, and a retried upload re-sends the file from its first byte. For process_from_connector, infer and retrain, a submit that times out or gets a 502/504 may still have created the job on the server. Check list_sessions() before you submit again.

Timeouts. timeout covers one HTTP request, not a whole job. wait_for_completion() has its own timeout (whole wait, off by default) and stall_timeout (default 3600 s with no change in status, percent or message). infer() and retrain() are synchronous: the server holds the request open for up to about 25 minutes while the job runs, so the client waits at least 900 s per request. Raise timeout= on those calls for large files.

Quick start

from datatoolpack import AutoDataClient

with AutoDataClient(api_key="dtpk_your_key") as client:
    # Check how the server reads the file (no job, no charge)
    info = client.validate("customers.csv")
    print(info["columns"], info["dataset_info"])

    # Get the exact price of this file with these settings
    if info.get("pricing_status") == "measuring":
        client.wait_for_price(info["pricing_profile_id"], timeout=600)
    quote = client.estimate(pricing_profile_id=info["pricing_profile_id"],
                            tools={"anomaly": True}, output_rows=100_000, text_mode=2)
    print("price: $", quote["price_usd"])

    # Run the pipeline, wait, download, and load the outputs as DataFrames
    result = client.process(
        "customers.csv",
        target_columns="churn",
        output_rows=100_000,
        tools={"anomaly": True},
        advanced_params={"excluded_columns": ["customer_id"], "datetime_mode": "basic",
                         "text_mode": 2, "quote_id": quote["quote_id"]},  # charged exactly that price
        return_dataframes=True,
    )
    print(result["files"])                  # [{name, url, size, description}, ...]
    train = result["dataframes"]["dsg_output"]

    # Later: score new rows with the same fitted transforms
    scored = client.infer(result["session_id"], "new_customers.csv", return_dataframes=True)
    features = scored["dataframes"]["features"]

By default process() blocks until the job ends, prints progress to stdout, and extracts every output into ./auto_data_outputs/<session_id>/. Pressing Ctrl+C while it waits cancels the job on the server and re-raises KeyboardInterrupt.

Pipeline options: tools and advanced_params

process() and process_from_connector() take two option dicts. tools switches stages on and off. advanced_params tunes them. Both methods accept the same keys and forward every key you pass.

tools

KeyDefaultEffectprocess / process_from_connector
anomalyFalseAnomaly detection and format repair. Which checks run comes from anomaly_params when you pass it; a file run without it uses your saved cleaning defaults for API runs (get_preferences('api_anomaly')).both
dtcTrueData Type Conversion: encoding, text, date and media features.both
mdhTrueMissing Data Handler. See mdh_mode.both
dorFalseAccepted, but outlier-row removal does not run in a normal pipeline. It runs only in retraining (retrain(run_dor=True)).both
cdsTrueColumn scaling.both
dsmTrueColumn pruning (drops near-duplicate and low-information columns; rows are never split).both
dsgTrueSynthetic row generation up to output_rows.both
feature_selectionFalseOptional feature-selection stage between DTC and MDH, configured by advanced_params['feature_selection_config'].both
preliminary_cleaningTrueRemoves duplicates, normalises missing-value tokens, drops empty, constant and ID-like columns.both
nested_flatteningTrueExpands JSON objects held in cells into their own columns (details becomes details.guest.email). Lists stay as JSON text; one input row is always one output row.both
result_evaluatorFalseAdds a multi-model benchmark report.both

Both methods send only the tools keys you set. A key you leave out takes the default above.

advanced_params

KeyType / valuesDefaultEffectprocess / process_from_connector
excluded_columnslist[str][]Columns removed before processing.both
text_mode0 off, 1 neural tokenization, 2 TF-IDF, 3 auto2How free-text columns are turned into features.both
text_cleaningboolTrueReserved; not available yet.both
mdh_mode0 Imputation, 1 2D-Removal, 2 Imputation/dropping00 fills missing values and keeps every row. 1 removes the rows and columns with too many missing values instead. 2 (Enterprise) removes only what is too empty to fill and fills the rest.both
zscore_limitfloat3.0Outlier limit for generated synthetic values (in standard deviations).both
datetime_modenone, minimal, basic, fullbasicHow many feature columns each date column becomes.both
dsg_modecopula, gancopulaSynthetic-data generator.both
similarity_pfloat, 0 to 10.99Correlation above which DSM drops one of two near-duplicate columns. A value of 1 or more turns this pruning off.both
task_typeslist of 'c'/'r', one per target; one value for all targets; or 'auto''c' for every targetClassification or regression per target. 'auto' lets the pipeline detect the type of each target. Also a keyword argument of both methods.both
anomaly_paramsdict of per-check togglessaved cleaning defaults for API runsWhich data-cleaning checks run in this run. Also a keyword argument of both methods.both
feature_selection_config{method, top_k, threshold}top_k 20Configures the feature-selection stage when tools['feature_selection'] is on. method is a name, a list, or 'auto'.both
llm_enabledboolTrueFalse guarantees no stage sends data to a language model. The result carries an llm audit block.both
strict_llmboolFalseTrue fails the run instead of silently falling back when a model is unreachable.both
moderapid, balanced, accuratebalancedSpeed preset. See Speed modes.both
speed-mode knobs (for example dtc_OHElimit)see Speed modesset by the modeYour own value for a knob the mode would otherwise choose.both
mode_decisions{'section.knob': 'user' or 'mode'}{}Per-knob answer when the mode disagrees with your value.both
touched_fieldslist[str][]Names of the knobs you set on purpose.both
mode_ackstr-The policy_signature from a 409, meaning keep all your values. confirm_mode_changes=True sends it for you.both
cds_scalerauto, standard, minmax, robust, maxabs, yeojohnson, quantile_normal, quantile_uniform, log1p, noneautoForces one scaler for every column.both
output_formatcsv, parquet, xlsx, json, jsonl, ndjson, feather, orcaccount preference, else csvFile format of the outputs.both
output_name_patternprefixed, plainaccount preference, else prefixedprefixed names files <input>_<timestamp>_<step>.<ext>. plain names them <step>.<ext>.both
flatten_columnslist[str] or NoneNone (detect)Columns to expand with nested flattening. [] expands none.both
dsg_allow_downsampleboolFalseWhen output_rows is below the processed row count, take a class-preserving subsample instead of skipping synthesis.both
dsg_y_oversample_thresholdfloat, 0 < v < 1offRaises every target class to at least this share of the generated rows.both
dsg_y_num_binsint, 2 to 10010Bins used to apply that rule to a continuous target.both
dsg_rare_oversample_thresholdfloat, 0 to 0.5offOversamples rare categories in generated rows.both
image_mode / audio_mode0 none, 1 basic features, 2 advanced processing, 3 neural embeddings (image) or Mel-spectrogram tokenization (audio)0How media columns are encoded. See Image and sound files.both
image_columns / audio_columnslist[str][]The columns that hold media file names.both
model_families, imputation_overrides, scaling_overrides, pivot_config, validation_policy, selection_config, probe_config, dataset_metadatasee Enterprise optionsoffEnterprise preparation policy. Removed by the server for other accounts.both

advanced_params accepts any pipeline setting by the name the dashboard's configuration uses, and both methods forward every key you pass. Inside advanced_params, mode is the speed preset. The run mode (full_pipeline, inference, retraining) is the separate mode argument of process() and process_from_connector().

Speed modes

advanced_params['mode'] picks a preset: rapid, balanced (default) or accurate. A preset sets the knobs below. To keep a value yourself, pass the knob's field name in advanced_params. If the preset disagrees with a value you set, process() raises ModeConfirmationRequired (HTTP 409) instead of running with changed parameters. There are three ways to answer: pass confirm_mode_changes=True to keep all your values, pass advanced_params['mode_decisions'] to decide per knob ('user' keeps yours, 'mode' takes the preset's), or pass advanced_params['mode_ack'] with the policy_signature from the error. Call preview_mode() to see the changes without running anything. Replays (infer, retrain, replay modes) reuse the parent session's settings, so they never ask.

Decision key (mode_decisions)Field name (advanced_params)Type / rangeWhat it controls
dtc.OHElimitdtc_OHElimitint 2 to 100000One-hot / free-text cardinality limit: which object columns are treated as free text.
dtc.sample_rows_detect_textdtc_sample_rows_detect_textint 10 to 100000Rows read when detecting free-text columns.
dtc.sample_rows_text_modedtc_sample_rows_text_modeint 10 to 100000Rows read when text mode 3 (auto) chooses a method.
dtc.seq_max_featuresdtc_seq_max_featuresint 100 to 1000000Vocabulary size for neural text tokenization (text mode 1).
dtc.tfidf_max_featuresdtc_tfidf_max_featuresint 100 to 1000000Number of TF-IDF features (text mode 2).
dtc.datetime_modedtc_datetime_modenone, minimal, basic, fullDatetime expansion (same as datetime_mode).
anomaly.force_no_llmanomaly_force_no_llmboolTurns the language-model anomaly checks off; local checks still run.
anomaly.sample_scaleanomaly_sample_scalefloat 0.1 to 10Scales how many rows each anomaly check inspects.
mdh.nooftrialsmdh_nooftrialsint 1 to 100Missing-data threshold trials (2D-Removal mode).
imputer.skip_expensiveimputer_skip_expensiveboolLeaves out the slowest imputation strategies.
imputer.force_allimputer_force_allboolEvaluates every imputation strategy, even on large data.
cds.eval_top_ncds_eval_top_nint 1 to 9How many candidate scalers are tried.
dsg.iess_budget_multiplierdsg_iess_budget_multiplierfloat 0.1 to 10Scales the fitting budget for synthetic generation.
multimodal.text_synth_methodtext_synth_methodstatistical, vaeGenerator for synthetic text.
multimodal.image_synth_methodimage_synth_methodstatistical, vaeGenerator for synthetic image features.
multimodal.audio_synth_methodaudio_synth_methodstatistical, vaeGenerator for synthetic audio features.
from datatoolpack import ModeConfirmationRequired

params = {"mode": "rapid", "dtc_OHElimit": 250}

preview = client.preview_mode("rapid", params)
for change in preview["mode_contested"]:
    print(change["label"], change["user_value"], "->", change["mode_value"])

# Option A: keep every value you set
client.process("data.csv", "label", advanced_params=params, confirm_mode_changes=True)

# Option B: decide per knob
params["mode_decisions"] = {"dtc.OHElimit": "user"}
client.process("data.csv", "label", advanced_params=params)

Image and sound files

When your table has columns holding image or audio file names, upload the files themselves with image_files=[...] / sound_files=[...]. Pair them with advanced_params={'image_mode': 1-3, 'image_columns': [...]} or {'audio_mode': 1-3, 'audio_columns': [...]}. Without the files, the server has nothing to match the names against, and every media feature is encoded as zeros: the run succeeds on empty data.

KindAccepted extensions
Images (image_files).jpg .jpeg .png .gif .bmp .webp .tiff .svg
Sound (sound_files).mp3 .wav .ogg .flac .aac .m4a .wma
Data file (file_path).csv .xlsx .xls .parquet .json .jsonl .ndjson .feather .orc. Compressed uploads (.zip, .gz) are not accepted.
client.process(
    "listings.csv",                        # has a column 'photo' with file names
    target_columns="price",
    advanced_params={"image_mode": 2, "image_columns": ["photo"]},
    image_files=["img/1001.jpg", "img/1002.jpg"],
)

# At inference, send only files the parent session has not seen.
# http(s) URLs held in the column are fetched by the server.
client.infer(session_id, "new_listings.csv", image_files=["img/2001.jpg"])

Every path is checked (exists, allowed extension) before anything is sent. Trial accounts are limited to 1 MB of media per request. Media can only travel with a file upload: infer(source=...) / retrain(source=...) with media raises ValueError.

Downloads and DataFrames

Output keys (file stems) you can ask for: anormaly_fixed_output, compleated_dataset, dtc_output, mdh_output, cds_output, dsm_output, dsg_output, multimodal_output, model_ready, model_ready_X, model_ready_y, model_ready_train, model_ready_test, model_ready_X_train, model_ready_X_test, model_ready_y_train, model_ready_y_test, model_ready_excluded_columns, and pipeline_report (PDF). Which files exist depends on your output preferences (get_preferences('outputs')) or the run's output_config, and on the stages that ran. Files are named <input>_<timestamp>_<step>.<ext> by default (for example customers_20261005_142233_dsg_output.csv), or <step>.<ext> with output_name_pattern: 'plain'. dsg_output is mostly synthetic rows; evaluate a model on dsm_output (real rows only).

compressed=True (the default everywhere) downloads one ZIP from /download-archive/<session_id> and extracts it. compressed=False downloads each file from /download/<session_id>/<name>. Empty files, truncated transfers and invalid archives raise AutoDataError. Names you pass can be a bare stem (dsg_output), which matches whatever format and name pattern the run used, or a full file name.

Run modes for automations

process, process_from_connector, create_trigger, create_scheduled_run, update_scheduled_run, create_listener, update_listener, create_sftp_credential and update_sftp_pipeline take mode and base_session_id:

  • full_pipeline (default): train a new session from the incoming data. Needs target columns.
  • inference: score the incoming data with a completed session's fitted transforms. Needs base_session_id. Targets and stage settings come from that session.
  • retraining: refresh a completed session with the incoming data (replayed transforms, outlier-row removal, synthetic top-up). Needs base_session_id.

The client checks the combination before sending: an unknown mode, or a replay mode without base_session_id, raises ValueError. Inference costs about a fifth, and retraining about two fifths, of preparing the same data.

Enterprise options (Enterprise)

Enterprise accounts can pass preparation-policy fields in the advanced_params of process() and process_from_connector(). For other accounts the server removes them. None of them is on unless you send it.

KeyShapeWhat it does
model_familieslist[str], up to 4The model families the preparation targets. enterprise_families() lists the valid names.
imputation_overrides{column: method}Pins the imputation method for a column. Valid methods come from enterprise_families().
scaling_overrides{column: scaler}Pins the scaler for a column.
pivot_config{column, leakage_action: 'warn'|'relative'|'drop', relative_deltas: bool, auto_suggest: bool}The as-of column (when a row came into existence). Other dates are measured from it, and columns whose values fall after it are flagged, reduced to a time difference or removed. pivot_suggest() proposes one.
validation_policy{split_strategy: 'random'|'stratified'|'time_ordered'|'grouped', eval_protocol: 'holdout'|'kfold'|'walk_forward', test_size, group_column, order_column, eval_folds (3), eval_max_rows, eval_time_budget_s, walk_forward_splits (3), random_state}How validation data is split and evaluated.
selection_config{methods: [...], top_k, known_leaky, leakage_action}Target-aware column selection in DSM. Methods: leakage_guard, variance_threshold, correlation_pruning, mutual_information, permutation_importance, rfe, shap_importance. They run in order, each on what the previous one left.
probe_config{penalty: 'l2'|'l1'|'elasticnet'|'none', C, l1_ratio}Experimental tuning of the evaluation AutoData uses to compare methods. It does not change your model.
dataset_metadataschema v1: per column role, type, semantic_type, min, max, impute, scaler, keepA declared data dictionary. Declared values win at every stage and are replayed at inference. An invalid document is rejected with 400.

Method reference

All 146 public methods of AutoDataClient, grouped by area. Every method raises AutoDataError (with status_code) when the server answers with an error, unless its entry says otherwise. Endpoints are relative to base_url.

Processing

Start a pipeline from a local file or a connector, check a file first, and preview speed-mode changes.

process

process(file_path, target_columns=None, output_rows=10000, tools=None, advanced_params=None, wait=True, poll_interval=2, download_path=None, auto_download=True, output_preferences=None, compressed=True, return_dataframes=False, confirm_mode_changes=False, image_files=None, sound_files=None, lm_readiness=None, history_file=None, *, task_types=None, session_name=None, anomaly_params=None, completion_params=None, output_config=None, webhook_ids=None, mode='full_pipeline', base_session_id=None, idempotency_key=None)

Uploads a data file and runs the full pipeline, or replays an existing session over the file (mode='inference' / 'retraining'). The arguments after history_file are keyword-only. With wait=True it polls until the job ends, downloads the outputs and returns the final result. With wait=False it returns as soon as the job is queued.

Endpoint: POST /api/v1/process (multipart: file, config JSON, repeated imageFiles / soundFiles; header Idempotency-Key), then /status, /result and /download-archive

ParameterTypeDefaultMeaning
file_pathstr-Local data file: .csv .xlsx .xls .parquet .json .jsonl .ndjson .feather .orc.
target_columnsstr or list[str]NoneTarget (y) column or columns. Required for a full_pipeline run (ValueError otherwise). May be omitted for a replay, which takes its targets from the base session.
output_rowsint10000Rows in the synthetic output (DSG target). To skip synthesis use tools={'dsg': False}.
toolsdict[str, bool]None (server defaults)Stage switches. See the tools table.
advanced_paramsdictNoneStage options, speed mode, Enterprise policy. See the advanced_params table. Any pipeline setting may be passed by the name the dashboard's configuration uses.
waitboolTrueBlock until the job completes, fails or is cancelled.
poll_intervalint (s)2Seconds between status polls while waiting.
download_pathstr./auto_data_outputs/<session_id>/Where outputs are extracted.
auto_downloadboolTrueDownload outputs after a successful wait.
output_preferenceslist[str]None (all files)Subset of outputs to download and load, by stem ('dsg_output') or full name.
compressedboolTrueDownload one ZIP archive instead of file by file.
return_dataframesboolFalseAlso load the tabular outputs into result['dataframes'] (needs pandas).
confirm_mode_changesboolFalseIf the speed mode would override a value you set, re-submit keeping your values instead of raising ModeConfirmationRequired. The file is uploaded a second time.
image_fileslist[str]NoneImage files named by your image column(s).
sound_fileslist[str]NoneAudio files named by your audio column(s).
lm_readinessdictNoneSwitches the run to LM Readiness preparation (see prepare_lm_readiness) instead of the normal pipeline. Sent inside config.
history_filestrNonePrior-events file for an LM Readiness replay (multipart part history). Only valid with lm_readiness={'action': 'replay', ...}.
task_typesstr or list[str]None ('c' for every target)'c' (classification) or 'r' (regression), one per target, or one value for all targets. 'auto' lets the pipeline detect the type of each target.
session_namestrNoneName shown for the run in the dashboard.
anomaly_paramsdictNone (saved cleaning defaults for API runs)Per-run data-cleaning switches: the keys get_preferences('api_anomaly') returns. Cleaning itself is turned on with tools={'anomaly': True}.
completion_paramsdictNoneSettings for the Data Completion & Verification stage.
output_configdictNone (saved output preferences)Which outputs this run writes and which stages it runs: the object get_preferences('outputs') returns, for example {'model_ready_output': True, 'model_ready_split': 'train_test', 'output_format': 'parquet'}. Not the same as output_preferences, which only filters what is downloaded afterwards.
webhook_idslist[str]NoneWebhooks notified when the run ends.
modestr'full_pipeline'full_pipeline trains a new session. inference scores the file against an existing session and retraining refreshes one. Both need base_session_id, take their columns and settings from it, and are queued and polled like any other run. infer() and retrain() do the same synchronously.
base_session_idstrNoneThe completed session a replay runs against. Required for the two replay modes.
idempotency_keystrNone (generated per call)Identifies this submission. A retry carrying the same key gets the first run's session back instead of starting a second run.

Returns: wait=False: {success, session_id, status: 'queued', message, config, mode, mode_adjustments, policy_signature}; for a replay mode {success, session_id, mode, base_session_id, rows_read}. A submission whose idempotency key was already used returns the first run's session_id with idempotent_replay: True. wait=True: the get_result() dict, plus dataframes ({step: DataFrame}) when return_dataframes=True.

Raises: FileNotFoundError, ValueError (extension, missing target_columns for a full pipeline run, unknown mode, replay mode without base_session_id), ModeConfirmationRequired, AutoDataError (server error, job failed or cancelled, stall).

Without task_types, every target is treated as classification ('c'). Without anomaly_params, the cleaning checks come from your saved cleaning defaults for API runs (get_preferences('api_anomaly')).

result = client.process(
    "sales.parquet",
    target_columns=["revenue"],
    task_types="r",
    session_name="Sales forecast",
    output_rows=50_000,
    tools={"anomaly": True, "feature_selection": True},
    advanced_params={
        "text_mode": 3,
        "mdh_mode": 0,
        "similarity_p": 0.95,
        "output_format": "parquet",
        "feature_selection_config": {"method": "auto", "top_k": 30},
    },
    output_preferences=["dsg_output", "dsm_output", "pipeline_report"],
)
print(result["session_id"], result["row_count"], result.get("amount_charged_usd"))

# Queue an inference run against that session and poll it yourself
queued = client.process("new_sales.parquet", mode="inference",
                        base_session_id=result["session_id"], wait=False)
scored = client.wait_for_completion(queued["session_id"])

process_from_connector

process_from_connector(connector_type, table, target_columns=None, secrets=None, credential_id=None, custom_query=None, output_rows=10000, tools=None, advanced_params=None, wait=True, poll_interval=2, download_path=None, auto_download=True, output_preferences=None, compressed=True, incremental=False, type_casts=None, return_dataframes=False, mode='full_pipeline', base_session_id=None, lm_readiness=None, *, watermark_column=None, task_types=None, session_name=None, anomaly_params=None, completion_params=None, output_config=None, webhook_ids=None)

Runs the pipeline on rows the server reads directly from a connector, or replays an existing session over them (mode='inference' / 'retraining').

Endpoint: POST /api/v1/process-connector (JSON), then /status, /result and /download-archive

ParameterTypeDefaultMeaning
connector_typestr-One of the 38 connector types (see test_connector).
tablestr-Table, collection, object key, topic or sheet to read.
target_columnsstr or list[str]NoneRequired for full_pipeline (ValueError otherwise). Not needed for replays, which take their targets from the base session.
secretsdictNoneInline connection secrets (sent as inline_secrets). With credential_id too, they overlay the saved values.
credential_idstrNoneA saved credential (save_credential).
custom_querystrNoneRead-only SQL (SELECT/WITH) used instead of table on SQL-style sources.
output_rowsint10000Synthetic output size.
toolsdictNoneStage switches, as in process. Every key you set is forwarded.
advanced_paramsdictNoneAs in process. Every key you pass is forwarded.
wait / poll_interval / download_path / auto_download / output_preferences / compressed / return_dataframes-as in processSame behaviour as in process. With wait=True and a replay, the wait follows the new replay session.
incrementalboolFalseRead only rows newer than the last sync watermark. Needs watermark_column (ValueError otherwise).
watermark_columnstrNoneKeyword-only. The column that orders the source (a timestamp or an increasing id). The highest value read is remembered once the run completes, and the next incremental run starts after it.
task_types / session_name / anomaly_params / completion_params / output_config / webhook_ids-NoneKeyword-only. As in process.
type_castsdict[str, str]NoneColumn casts applied before the pipeline: int, float, str, bool, datetime, category (and their aliases). Columns that fail to cast are left unchanged.
modestr'full_pipeline'full_pipeline, inference or retraining.
base_session_idstrNoneRequired for the two replay modes.
lm_readinessdictNoneRun LM Readiness preparation on the connector rows instead of the pipeline (same options as prepare_lm_readiness's config). A target is required, because LM Readiness runs the classic pipeline first. mode='inference' or 'retraining' with an LM Readiness base_session_id replays or re-prepares that session as LM Readiness.

Returns: Without waiting: {success, run: {session_id}, session_id, message}; for replays {success, session_id, mode, base_session_id, rows_read}; session_id is None when an incremental read found no new rows. With wait=True: the get_result() dict (+ dataframes).

Raises: ValueError (mode, missing targets, incremental=True without watermark_column), AutoDataError (400 bad source or no rows, 404 credential not found, 403 stage not allowed for the key, 429 key limit).

This method sends no defaults of its own: a setting you leave out takes the server default, the same as a file run (text_mode 2, similarity_p 0.99). Versions before 0.15.0 sent text_mode=0 and similarity_p=95 when you set neither. To keep text features off as before, pass advanced_params={"text_mode": 0}.

result = client.process_from_connector(
    "snowflake", "ORDERS",
    target_columns="churned",
    credential_id="c0ffee00-...",
    task_types="c",
    advanced_params={"datetime_mode": "full"},
    type_casts={"order_date": "datetime"},
)

# Read only the rows added since the last run
client.process_from_connector("snowflake", "ORDERS", target_columns="churned",
                              credential_id="c0ffee00-...",
                              incremental=True, watermark_column="updated_at")

# Score new rows against an existing session
client.process_from_connector("snowflake", "NEW_ORDERS", credential_id="c0ffee00-...",
                              mode="inference", base_session_id=result["session_id"])

prepare_lm_readiness

prepare_lm_readiness(file_path, config=None, target_columns=None, **options)

Prepares a dataset for transformer (language-model style) training: train-fitted encodings, a train/validation/test split, and a saved preparation contract you can replay on new data. It runs instead of the normal pipeline (no anomaly detection, column pruning or synthetic rows). A thin wrapper around process(file_path, target_columns, lm_readiness=config, **options).

Endpoint: POST /api/v1/process (multipart, config.lm_readiness set)

ParameterTypeDefaultMeaning
file_pathstr-Local data file (same formats as process). Limits: 200,000 rows and 100 MB by default (1 MB for trial accounts).
configdict{'profile': 'tabular_transformer'}LM Readiness options: profile (tabular_transformer or event_sequence), preset (general or transaction_fraud), representation (typed or feature_tokens), level (E1, E2 or auto), sequence_length (1 to 128, default 32), roles ({entity, time, amount, category, counterparty} mapped to column names), excluded_columns, text_columns, test_size and validation_size (fractions, default 0.15 each, sum below 1). event_sequence, transaction_fraud and E2 need the entity and time roles.
target_columnsstr or list[str]NoneTargets. Required: the client raises ValueError before any request when it is missing.
**options--Any other process() keyword: wait, download_path, return_dataframes, advanced_params (only validation_policy, dataset_metadata and task_types are read), ...

Returns: As process(). Without waiting: {success, session_id, status: 'processing', module: 'lm_readiness', dataset_info: {rows, columns}, heartbeat_interval_seconds}. The outputs are a bundle ZIP plus its files under lm_readiness/ (manifest.json, the contract, the split data and an example load_bundle.py).

Raises: ValueError (no target_columns); AutoDataError 400 for an invalid config, unknown mapped columns or too few eligible rows.

New in the current code and not yet described in the package README.

res = client.prepare_lm_readiness(
    "transactions.csv",
    config={"profile": "event_sequence", "preset": "transaction_fraud",
            "roles": {"entity": "card_id", "time": "event_ts", "amount": "amount"}},
    target_columns="is_fraud",
)
base = res["session_id"]

replay_lm_readiness

replay_lm_readiness(file_path, base_session_id, history_file=None, **options)

Applies a completed LM Readiness session's saved contract to new data, so the new rows are encoded exactly like the originals.

Endpoint: POST /api/v1/process (multipart, config.lm_readiness = {action: 'replay', base_session_id}, optional history part)

ParameterTypeDefaultMeaning
file_pathstr-New rows.
base_session_idstr-A completed LM Readiness session you own.
history_filestrNoneEarlier events for the same entities, so history-based features of the new rows can see them.
**options--Other process() keywords.

Returns: As prepare_lm_readiness.

Raises: FileNotFoundError (history file); AutoDataError 400 when the base session has no completed contract or is not yours.

client.replay_lm_readiness("new_transactions.csv", base, history_file="last_30_days.csv")

preview_mode

preview_mode(mode='balanced', advanced_params=None)

Shows what a speed mode would change about your parameters, without uploading or running anything.

Endpoint: POST /api/v1/mode-preview

ParameterTypeDefaultMeaning
modestr'balanced'rapid, balanced or accurate.
advanced_paramsdictNoneThe same dict you would pass to process.

Returns: {success, mode, mode_groups, mode_adjustments, mode_contested, policy_signature, knob_specs}. Each adjustment has path, label, user_value, mode_value. knob_specs gives the type and range of every knob.

preview = client.preview_mode("rapid", {"dtc_OHElimit": 250})
print(preview["policy_signature"], len(preview["mode_contested"]))

validate

validate(file_path)

Dry run: uploads a file and returns how the server reads it. No job is started and nothing is charged.

Endpoint: POST /api/v1/validate (multipart file)

ParameterTypeDefaultMeaning
file_pathstr-Local data file.

Returns: {columns, shape, missing_values, file_size_mb, validation_level ('full' or 'sampled'), nested_flattening, column_scan, dataset_info: {rows, columns}, has_audio_column, has_image_column, audio_file_count, image_file_count, pricing_profile_id, pricing_status}. Large files are sampled for the preview and carry a note; for the price, every row is measured. pricing_status is ready (measured during the call) or measuring (a larger file; see wait_for_price). Pass pricing_profile_id to estimate() for the exact price.

Raises: AutoDataError (unreadable file; trial accounts are limited to 1 MB files).

wait_for_price

wait_for_price(pricing_profile_id, timeout=600, poll_interval=2.0)

Waits until an uploaded file has been measured for its exact price. Needed only when validate() returned pricing_status: 'measuring'.

Endpoint: GET /api/v1/pricing/profile/<pricing_profile_id>

ParameterTypeDefaultMeaning
pricing_profile_idstr-From validate().
timeoutfloat600Seconds to wait in total.
poll_intervalfloat2.0Seconds between polls.

Returns: {status: 'ready', progress_pct, exact, rows, columns}. Then call estimate(pricing_profile_id=...).

Raises: RuntimeError if the file could not be measured (upload it again); TimeoutError after timeout.

validate_media_files

validate_media_files(file_path, media_type, filenames, columns=None)

Checks the media file names a dataset refers to against the files you hold, before a run. Only names are sent; no media is uploaded and no job is started.

Endpoint: POST /api/v1/media/validate (multipart: csv, mediaType, repeated uploadedFilenames, imageColumns / audioColumns)

ParameterTypeDefaultMeaning
file_pathstr-The dataset whose cells name the media files.
media_typestr-'image' or 'audio' (ValueError otherwise).
filenameslist[str]-The files you intend to upload. Paths are reduced to their file names.
columnslist[str]None (the default media column)The media columns to check.

Returns: {valid, column_found, uploaded_files_without_match, total_uploaded, total_matched, message}.

check = client.validate_media_files("listings.csv", "image",
                                    ["img/1001.jpg", "img/1002.jpg"], columns=["photo"])
print(check["valid"], check["uploaded_files_without_match"])

Status, results and downloads

Follow a job, fetch its results and load them.

get_status

get_status(session_id)

Returns the current progress of a session once.

Endpoint: GET /api/v1/status/<session_id>

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: {session_id, status, message, current_step, total_steps, progress_percent, duration_seconds} plus, when present, error_info, substep and run_warnings. status is queued, processing, completed, error/failed, cancelled or unknown.

Raises: AutoDataError 403 when the session is not yours.

wait_for_completion

wait_for_completion(session_id, poll_interval=2, timeout=None, stall_timeout=3600.0)

Polls get_status until the job ends, printing a progress line to stdout, then returns get_result().

Endpoint: GET /api/v1/status/<id>, then GET /api/v1/result/<id>

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
poll_intervalint (s)2Seconds between polls.
timeoutfloat (s)None (no limit)Ceiling for the whole wait.
stall_timeoutfloat (s)3600.0Raise if status, percent and message have not changed for this long. The method's docstring says 10 minutes; the real default is one hour.

Returns: The get_result() dict.

Raises: AutoDataError when the job fails, is cancelled, stalls or the timeout passes. The wait does not cancel the job.

get_result

get_result(session_id)

Returns the final result of a session. Fetching it for a completed session also settles billing (idempotent).

Endpoint: GET /api/v1/result/<session_id>

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: Completed: {session_id, status: 'completed', files: [{name, url, size, description}], row_count, duration_seconds}, plus when relevant warning/message (for example DSG skipped), dsg_provenance (real vs synthetic rows), audio_decode_report / image_encode_report (only when files failed), dsm_column_drops, warnings, run_warnings, media_staging, amount_charged_usd, balance_usd, balance_warning (text in dollars, set when the balance is negative), billing ({amount_charged_usd, quote_low_usd, quote_high_usd, list_usd, capped}: the amount charged), llm (data-residency audit). The credit-named fields (credits_charged, credits_remaining, credit_warning, and *_credits in billing) are deprecated: still returned, equal to the dollar amount × 1,000,000, and will be removed in a future version. Not complete: {session_id, status, message} (+ error_info).

Raises: AutoDataError 403 when the session is not yours.

cancel

cancel(session_id)

Asks the server to cancel a queued or running session.

Endpoint: POST /api/v1/cancel/<session_id>

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: True if the server acknowledged the cancel, otherwise False.

Raises: Never raises for API errors: it prints a warning and returns False.

cancel_session

cancel_session(session_id)

Deprecated alias of cancel().

Endpoint: POST /api/v1/cancel/<session_id>

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: Same as cancel().

retry_session

retry_session(session_id)

Re-submits a failed session from its last checkpoint, with the configuration it was first run with (targets, task types, outputs, cleaning and speed mode).

Endpoint: POST /api/v1/sessions/<session_id>/retry

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: {success, session_id, message}. Follow it with wait_for_completion(session_id).

Raises: AutoDataError 404 when the session's saved request or input file no longer exists.

download_results

download_results(session_id, download_path=None, output_preferences=None, compressed=True)

Downloads a completed session's outputs to a folder.

Endpoint: POST /api/v1/download-archive/<session_id> (body {files: [...]}), or GET /api/v1/download/<session_id>/<name> per file

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
download_pathstr./auto_data_outputs/<session_id>/Folder, created if needed.
output_preferenceslist[str]None (all)Stems or file names. 'pipeline_report' means the PDF report.
compressedboolTrueOne ZIP, extracted locally. With False, files are downloaded one by one. Either way output_preferences entries match by step name ('dsg_output' or 'dsg_output.csv'), whatever format the run exported.

Returns: Absolute path of the download folder (str).

Raises: AutoDataError on an empty or invalid archive (outputs expired or not generated).

download_file

download_file(url, output_path, timeout=None, retries=2)

Downloads one file. Checks Content-Length, rejects empty files and retries partial transfers.

Endpoint: GET <url>

ParameterTypeDefaultMeaning
urlstr-Absolute URL, or a path starting with / (for example a url from get_result()['files']).
output_pathstr-Local file path; parent folders are created.
timeoutfloat (s)max(client timeout, 900)Per attempt.
retriesint2Extra attempts after a network error or a truncated download. A 4xx answer (other than 429) is raised at once, with no further attempt.

Returns: None.

Raises: AutoDataError (empty or partial file), network errors after the last retry.

get_dataframe

get_dataframe(session_id, output_name, *, timeout=None, **read_csv_kwargs)

Downloads one output straight into a pandas DataFrame, whatever format the run used.

Endpoint: GET /api/v1/download/<session_id>/<file name>

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
output_namestr-Stem ('dsg_output', resolved against the session's file list) or full file name.
timeoutfloat (s)max(client timeout, 900)Request timeout.
**read_csv_kwargs--Passed to pandas.read_csv (also to read_json / read_excel; not used for Parquet, Feather and ORC).

Returns: pandas.DataFrame.

Raises: AutoDataError when pandas is missing or the file is missing or empty.

df = client.get_dataframe(sid, "dsm_output", dtype={"zip": str})

get_dataframes

get_dataframes(session_id, names=None, *, timeout=None)

Loads several outputs at once.

Endpoint: GET /api/v1/result/<id>, then GET /api/v1/download/<id>/<name> per file

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
nameslist[str]None (every tabular file; PDF, JSON and Markdown skipped)Stems or file names.
timeoutfloat (s)as get_dataframePer file.

Returns: {step: DataFrame}, keyed by the logical step (dsg_output, model_ready_X_train, ...) whatever the file was named.

Raises: AutoDataError.

Inference and retraining

Apply a completed session's fitted transforms to new data. Both calls are synchronous: they return when the server has finished.

infer

infer(original_session_id, file_path=None, enable_anomaly_detection=None, download_path=None, return_dataframes=False, timeout=None, *, source=None, image_files=None, sound_files=None)

Scores new rows with the same encoding, imputation and scaling as the training run (anomaly, DTC, MDH, CDS replayed). Column names must match the training data.

Endpoint: POST /api/v1/inference (multipart file, original_session_id, enableAnomalyDetection, imageFiles, soundFiles; or JSON with source)

ParameterTypeDefaultMeaning
original_session_idstr-A completed session you own.
file_pathstrNoneLocal data file (same formats as process). Give this or source, not both.
enable_anomaly_detectionboolNone (inherit)Override the parent's anomaly setting.
download_pathstrNoneDownload the outputs here. When None and return_dataframes=False, nothing is downloaded.
return_dataframesboolFalseLoad the CSV outputs into result['dataframes'] (downloads to a temp folder if no path is given).
timeoutfloat (s)max(client timeout, 900)Request timeout.
sourcedictNoneRead rows from a connector instead: {source_type, credential_id or inline_secrets, table or custom_query}.
image_files / sound_fileslist[str]NoneMedia named by the data. Optional: the parent session's own media and http(s) URLs in the column are used too. Only with file_path.

Returns: {status: 'success', new_session_id, output_files: [{name, ...}], ...} plus amount_charged_usd, balance_usd and, when the balance is negative, balance_warning (credits_charged, credits_remaining and credit_warning are deprecated), and files ({name: local path}) / dataframes ({'features': df, 'y_columns': df}) when requested.

Raises: ValueError (both or neither of file_path/source, media with source, bad extension); AutoDataError (400 when the job fails, 403/404 parent not yours or not found).

out = client.infer(sid, "new_rows.csv", return_dataframes=True)
X = out["dataframes"]["features"]

# From a connector
client.infer(sid, source={"source_type": "sql",
                          "credential_id": "c0ffee00-...", "table": "new_orders"})

retrain

retrain(original_session_id, file_path=None, run_dsm=True, run_dsg=True, output_rows=None, enable_anomaly_detection=None, download_path=None, return_dataframes=False, timeout=None, run_dor=True, dor_eps=None, dor_min_samples=None, *, source=None, image_files=None, sound_files=None)

Refreshes a training set: the same replay as infer, then outlier-row removal (DOR) and synthetic top-up (DSG). Outputs have the same shape as infer (features.csv, y_columns.csv).

Endpoint: POST /api/v1/retraining (multipart or JSON, as infer)

ParameterTypeDefaultMeaning
original_session_idstr-A completed session you own.
file_path / source / image_files / sound_files-NoneAs in infer.
run_dsmboolTrueAccepted and ignored: column pruning is no longer part of retraining, so the output schema stays the parent's.
run_dsgboolTrueAdd synthetic rows.
output_rowsintNone (server default)DSG target row count (sent as output_sample_size).
enable_anomaly_detectionboolNone (inherit)Override the parent's setting.
run_dorboolTrueRemove outlier rows.
dor_epsfloatNone (parent's setting, else 1.0)Outlier-removal distance, in standardized units. Smaller values remove more rows.
dor_min_samplesintNone (parent's setting, else 5)Minimum neighbourhood size. Larger values remove more rows.
download_path / return_dataframes / timeout-as inferAs in infer.

Returns: Same shape as infer.

Raises: As infer.

client.retrain(sid, "march.csv", output_rows=200_000, dor_eps=0.8, download_path="out/march")

Connectors and credentials

38 connector types: sql, snowflake, bigquery, mongodb, s3, gcs, databricks, delta, fabric, synapse, kafka, mqtt, opcua, influxdb, pi_web_api, kinesis, azure_blob, adls_gen2, oracle, sap_hana, clickhouse, redshift, timescaledb, cassandra, dynamodb, elasticsearch, salesforce, hubspot, stripe, google_sheets, ga4, netsuite, sharepoint, http, google_drive, dropbox, onedrive, box. The secret field names per type are listed in the Connectors section. Every call takes either inline secrets or a saved credential_id (both: the inline values overlay the saved ones).

test_connector

test_connector(connector_type, secrets=None, credential_id=None)

Checks that the server can connect. Testing a saved credential records the result on it.

Endpoint: POST /api/v1/connectors/<type>/test

ParameterTypeDefaultMeaning
connector_typestr-Connector type.
secretsdictNoneInline secrets.
credential_idstrNoneSaved credential.

Returns: {success: bool, message}.

ok = client.test_connector("s3", secrets={"bucket": "my-data", "region": "eu-west-1",
    "access_key_id": "AKIA...", "secret_access_key": "..."})

discover

discover(connector_type, secrets=None, credential_id=None)

Lists what can be read: tables and views (warehouses), collections or indexes (NoSQL), object keys under the prefix, capped at 500 (object stores), topics or streams (Kafka, Kinesis), the URL itself (HTTP) or the table path (Delta).

Endpoint: POST /api/v1/connectors/<type>/discover

ParameterTypeDefaultMeaning
connector_typestr-Connector type.
secretsdictNoneInline secrets.
credential_idstrNoneSaved credential.

Returns: list[str] (the response's tables).

preview

preview(connector_type, table, secrets=None, credential_id=None, custom_query=None)

Returns the columns and row count of one table, file or index without reading it all.

Endpoint: POST /api/v1/connectors/<type>/preview

ParameterTypeDefaultMeaning
connector_typestr-Connector type.
tablestr-Name from discover().
secrets / credential_id-NoneAuthentication.
custom_querystrNoneRead-only SELECT/WITH on SQL sources; writes are rejected.

Returns: {columns: list[str], row_count: int or None}. Cassandra always returns row_count: None.

list_credentials

list_credentials()

Lists saved connector credentials. Secret values are never returned.

Endpoint: GET /api/v1/credentials

Returns: list of credential dicts: id, name, connector_type, the last test result, hints, stored_fields (the names of the fields that are set, never their values) and write (what the connection can do as a destination: modes, target_kind, formats, options, guarded, max_rows, enabled).

save_credential

save_credential(name, connector_type, secrets)

Saves connection secrets on the server (encrypted) for reuse by every method and automation. The server checks the required fields for the type when saving.

Endpoint: POST /api/v1/credentials

ParameterTypeDefaultMeaning
namestr-Label.
connector_typestr-Connector type.
secretsdict-Field names as in the dashboard's connection form.

Returns: The saved credential (the response's credential), including id.

Raises: AutoDataError 400 when a required field is missing.

cred = client.save_credential("warehouse", "snowflake", {
    "account": "xy12345.eu-west-1", "user": "etl", "password": "...",
    "database": "SALES", "schema": "PUBLIC", "warehouse": "COMPUTE_WH"})
cred_id = cred["id"]

update_credential

update_credential(credential_id, name=None, secrets=None, merge_secrets=False, clear_secrets=None)

Renames a credential and/or changes its secrets. Empty values are not sent.

Endpoint: PUT /api/v1/credentials/<credential_id>

ParameterTypeDefaultMeaning
credential_idstr-Credential.
namestrNoneNew label.
secretsdictNoneNew secrets. They replace the stored set unless merge_secrets is true.
merge_secretsboolFalseLay secrets over the stored ones, so only the fields you send change.
clear_secretslist[str]NoneNames of stored fields to remove, for example a SAS token after moving the connection to an account key.

Returns: The full response body, {success, credential}; credential has the same fields as a list_credentials() entry.

# change one field and keep the others
client.update_credential(cred_id, secrets={"password": "..."}, merge_secrets=True)

# move a Snowflake connection from a password to a key pair
client.update_credential(cred_id, secrets={"private_key_pem": pem_text},
                         merge_secrets=True, clear_secrets=["password"])

# let a Salesforce connection be written to
client.update_credential(cred_id, secrets={"allow_write": True}, merge_secrets=True)

delete_credential

delete_credential(credential_id)

Deletes a saved credential.

Endpoint: DELETE /api/v1/credentials/<credential_id>

ParameterTypeDefaultMeaning
credential_idstr-Credential.

Returns: True when the server reports success.

write_output

write_output(session_id, connector_type, table_name, secrets=None, credential_id=None, output_stage='dsg', if_exists=None, merge_keys=None, file_format=None, partition_by=None, partition_format=None, options=None)

Writes a completed session's output to a connector destination: a table, a file, a worksheet, an index, a topic or stream, an endpoint, or records in a business or plant system. Every connector type except ga4 can be one. A saved credential's write entry in list_credentials() says which modes, file formats and options it takes.

Endpoint: POST /api/v1/sessions/<session_id>/write-output

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
connector_typestr-Target connector type.
table_namestr-Destination name: a table (schema.table), a file name or path, a worksheet, an index, a topic, an object type, depending on the connector.
secrets / credential_id-NoneAuthentication for the target. Salesforce, HubSpot, NetSuite, SharePoint lists, Stripe, PI Web API and OPC UA take credential_id only, and that credential must have allow_write set.
output_stagestr'dsg'Stage to export (dsg, dsm, cds, mdh, dtc, ...). If it is not available, the server falls back to dsg, then cds, mdh, dtc.
if_existsstrNone (the destination's own default)replace, append, fail or merge (upsert). Left out, the destination's default applies: replace where it is offered, otherwise append.
merge_keyslist[str]NoneKey columns; required when if_exists='merge' (400 otherwise).
file_formatstrNone (csv)For file destinations (S3, GCS, Azure Blob, ADLS Gen2, Google Drive, Dropbox, OneDrive, Box): csv, parquet, xlsx, json, jsonl, ndjson, feather or orc. With append each write is a new timestamped file; replace overwrites one fixed file.
partition_bystrNoneWrite one target per distinct value of this column, up to 1,000.
partition_formatstrNone (%Y-%m-%d)Date pattern for a date partition column.
optionsdictNoneDestination-specific settings, for example {'partition_key_column': 'id'} for Kafka and Kinesis or {'time_column': 'ts', 'tag_columns': ['site']} for InfluxDB.

Returns: {success, rows_written, table, ...}, with partitions when partition_by was used.

Raises: AutoDataError 400 (bad options, a mode the destination does not offer, or an output over its row limit), 403 (not your session, or a destination whose credential does not allow writing), 404 (no output, or credential not found), 409 (if_exists='fail' and the target exists).

client.write_output(sid, "sql", "ml.training_set", credential_id=cred_id,
                    if_exists="merge", merge_keys=["customer_id"])

# one Parquet file per day, next to the earlier ones
client.write_output(sid, "s3", "scored/orders", credential_id=s3_cred,
                    if_exists="append", file_format="parquet", partition_by="order_date")

Column mapping

Match your column names to a target schema.

suggest_mapping

suggest_mapping(source_columns, target_columns, threshold=0.6)

Proposes a source-to-target column pairing from name similarity.

Endpoint: POST /api/v1/mapping/suggest

ParameterTypeDefaultMeaning
source_columnslist[str]-Your column names.
target_columnslist[str]-The schema's column names.
thresholdfloat 0 to 10.6 (the server's own default is 0.4)Minimum similarity for a suggestion.

Returns: {success, mappings: [{source, target, confidence, auto}], unmapped_sources, unmapped_targets}. The method's docstring mentions mapping / scores; those keys do not exist.

apply_mapping

apply_mapping(session_id, mapping, output_stage='dsg')

Renames a session's output columns and returns a preview.

Endpoint: POST /api/v1/mapping/apply

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
mappingdict[str, str] or list[dict]-{old_name: new_name}, or the list of {'source', 'target'} pairs that suggest_mapping returns.
output_stagestr'dsg'Stage to read.

Returns: {success, columns, row_count, preview (5 rows), applied_mappings, dropped_columns}. Nothing is saved on the server; the result is a preview.

Raises: AutoDataError 400 for an empty mapping.

preview = client.apply_mapping(sid, {"cust_id": "customer_id"})
print(preview["columns"])

# Or pass the pairs suggest_mapping returned
pairs = client.suggest_mapping(["cust_id", "amt"], ["customer_id", "amount"])["mappings"]
client.apply_mapping(sid, pairs)

Feature selection

Rank a finished session's columns without re-running it. To select features inside a run, use tools={'feature_selection': True} and advanced_params['feature_selection_config'].

recommend_features

recommend_features(session_id, target_columns, methods=None, top_k=20)

Ranks the columns of a completed session against the targets.

Endpoint: POST /api/v1/feature-selection/recommend

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
target_columnslist[str]-Targets to rank against.
methodslist[str]['mutual_information']Any of variance_threshold, correlation_pruning, pca, mutual_information, rfe, shap_importance, permutation_importance, leakage_guard.
top_kint20How many columns to return.

Returns: {ranking (combined), per_method, warnings, methods, top_k, row_count, ...}.

rec = client.recommend_features(sid, ["churn"], methods=["mutual_information", "leakage_guard"], top_k=15)
for row in rec["ranking"]:
    print(row)

feature_columns

feature_columns(session_id)

Lists the columns of a completed session that recommend_features can rank.

Endpoint: GET /api/v1/feature-selection/columns?session_id=...

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: {success, session_id, columns, row_count}.

Raises: AutoDataError 404 when the session is not found or has no output to rank.

Templates and guided configuration

Build a configuration once and reuse it: in automations, as a saved template, or proposed for a dataset.

pipeline_config

AutoDataClient.pipeline_config(target_columns=None, output_rows=None, tools=None, advanced_params=None, *, task_types=None, anomaly_params=None, completion_params=None, output_config=None, webhook_ids=None)

Static method. Builds the flat pipeline_config object that scheduled runs, folder listeners, SFTP inboxes and triggers store, from the same arguments process() takes. No request is sent.

ParameterTypeDefaultMeaning
target_columnsstr or list[str]NoneTargets.
output_rowsintNoneSynthetic output size.
toolsdict[str, bool]NoneStage switches. See the tools table.
advanced_paramsdictNoneStage options and speed mode. See the advanced_params table.
task_types / anomaly_params / completion_params / output_config / webhook_ids-NoneKeyword-only. As in process.

Returns: dict, ready to pass as pipeline_config to an automation or as config to create_template.

config = AutoDataClient.pipeline_config(
    "churned", output_rows=5000, tools={"anomaly": True},
    advanced_params={"text_mode": 2, "mode": "rapid"},
    output_config={"model_ready_output": True})
client.create_scheduled_run("nightly", "snowflake", "CUSTOMERS", "churned",
                            credential_id=cred_id, pipeline_config=config)

list_templates

list_templates()

Lists your saved configurations.

Endpoint: GET /api/v1/pipeline-templates

Returns: list (the response's templates).

create_template

create_template(name, config, description=None)

Saves a configuration under a name. An account can hold up to 50 templates.

Endpoint: POST /api/v1/pipeline-templates

ParameterTypeDefaultMeaning
namestr-Label.
configdict-A flat settings object: what pipeline_config() returns, or the config that get_session_config() gives back for a past run.
descriptionstrNoneFree text.

Returns: The saved template dict, including id.

Raises: AutoDataError 400 (invalid name or config, or the template limit is reached).

tpl = client.create_template("churn-default", config, description="Rapid, model-ready output")

update_template

update_template(template_id, **fields)

Changes only the fields you pass: name, description, config.

Endpoint: PUT /api/v1/pipeline-templates/<template_id>

ParameterTypeDefaultMeaning
template_idstr-Template.
**fields--Fields to change.

Returns: The updated template dict.

Raises: AutoDataError 404 (template not found).

delete_template

delete_template(template_id)

Deletes a template.

Endpoint: DELETE /api/v1/pipeline-templates/<template_id>

ParameterTypeDefaultMeaning
template_idstr-Template.

Returns: {ok: True}.

decide_config

decide_config(columns, y_columns=None, excluded_columns=None, rows=None, profile_columns=None, dataset_metadata=None)

Proposes a configuration for a dataset: the dashboard's "Decide everything". Give it what validate() returned for the file and it proposes the target, the task type and the settings. Nothing is started; pass what you accept on to process().

Endpoint: POST /api/v1/pipeline-config/decide

ParameterTypeDefaultMeaning
columnslist[str]-Column names, as validate() returned them.
y_columnslist[str]NoneTargets you have already chosen. They are kept.
excluded_columnslist[str]NoneColumns you have already left out.
rowsintNoneRow count of the dataset.
profile_columnslist[dict]NoneThe per-column profile from validate(), when it returned one.
dataset_metadatadictNoneA declared data dictionary (Enterprise).

Returns: {success, tier, config, decided, needs}: config holds the proposed settings, decided names what was chosen for you and needs what is still yours to choose. Enterprise accounts also get reasons and warnings.

Raises: AutoDataError 400 when columns is empty.

info = client.validate("customers.csv")
decision = client.decide_config(
    info["columns"], rows=info["dataset_info"]["rows"],
    profile_columns=(info.get("profile") or {}).get("columns"))
print(decision["config"], decision["needs"])

Triggers

A trigger starts a run when something happens: an inbound webhook call, a watermark advancing, another pipeline completing, or a quality alert. After 5 consecutive failures a trigger disables itself.

list_triggers

list_triggers()

Lists your triggers. Webhook secrets are never included.

Endpoint: GET /api/v1/triggers

Returns: list of trigger dicts.

create_trigger

create_trigger(name, trigger_type, config=None, target_pipeline_config=None, enabled=True, mode='full_pipeline', base_session_id=None)

Creates a trigger.

Endpoint: POST /api/v1/triggers

ParameterTypeDefaultMeaning
namestr-Label.
trigger_typestr-inbound_webhook, watermark_advance, pipeline_completion or quality_alert.
configdict{}Per type. inbound_webhook: optional hmac_key (string, 16+ chars). watermark_advance: credential_id, table, watermark_column, min_advance (positive int), all required; optional source_key. pipeline_completion: source_session_name or source_template_id. quality_alert: nothing required.
target_pipeline_configdict{}What to run: source_type (or connector_type), credential_id, table / custom_query, target_columns (or y_columns), output_rows, and any pipeline option in snake_case (text_mode, enable_dsg, ...). For an inbound webhook, a CSV or JSON body posted to the trigger is used as the data instead of the connector.
enabledboolTrueActive on creation.
modestr'full_pipeline'Run mode; stored inside target_pipeline_config.
base_session_idstrNoneParent for replay modes.

Returns: The trigger dict (id, name, trigger_type, config, target_pipeline_config, enabled, fire_count, consecutive_failures, last_error, ...). For inbound_webhook also webhook_secret and inbound_url (/trigger/inbound/<secret>). This is the only time the secret is shown.

Raises: ValueError (mode); AutoDataError 400 (invalid type or config).

Inbound calls are limited to 10 per minute and 64 KB per body. With hmac_key set, each call must send X-Signature: t=<unix time>,v1=<hex HMAC-SHA256 of '<t>.' + body>, within 300 s of t.

tr = client.create_trigger(
    "drop-zone", "inbound_webhook", config={"hmac_key": "a-long-shared-secret-123"},
    target_pipeline_config={"target_columns": ["label"], "output_rows": 20000})
print(client.base_url + tr["inbound_url"])   # POST CSV or a JSON array of rows here

get_trigger

get_trigger(trigger_id)

Fetches one trigger (without its secret).

Endpoint: GET /api/v1/triggers/<trigger_id>

ParameterTypeDefaultMeaning
trigger_idstr-Trigger.

Returns: Trigger dict.

update_trigger

update_trigger(trigger_id, **fields)

Updates only the fields you pass: name, enabled, config, target_pipeline_config. The webhook secret cannot change.

Endpoint: PUT /api/v1/triggers/<trigger_id>

ParameterTypeDefaultMeaning
trigger_idstr-Trigger.
**fields--Fields to change.

Returns: Updated trigger dict.

client.update_trigger(tr["id"], enabled=False)

delete_trigger

delete_trigger(trigger_id)

Deletes a trigger.

Endpoint: DELETE /api/v1/triggers/<trigger_id>

ParameterTypeDefaultMeaning
trigger_idstr-Trigger.

Returns: True.

test_trigger

test_trigger(trigger_id)

Dry run: validates the trigger's config and target without firing it.

Endpoint: POST /api/v1/triggers/<trigger_id>/test

ParameterTypeDefaultMeaning
trigger_idstr-Trigger.

Returns: {would_fire, checks: {...}}.

Scheduled runs

Run a connector source on an interval or a cron schedule.

list_scheduled_runs

list_scheduled_runs()

Lists your schedules.

Endpoint: GET /api/v1/scheduled-runs

Returns: list (the response's scheduled_runs).

create_scheduled_run

create_scheduled_run(name, connector_type, source_table, y_columns=None, schedule_type='interval', schedule_value='1h', credential_id=None, output_rows=10000, sql_query=None, pipeline_config=None, description=None, mode='full_pipeline', base_session_id=None)

Creates a schedule. It starts active and first fires one interval (or the next cron time) from now.

Endpoint: POST /api/v1/scheduled-runs

ParameterTypeDefaultMeaning
namestr-Label.
connector_typestr-Source connector type.
source_tablestr-Table or object key.
y_columnsstr or list[str]NoneTargets. Required for a full_pipeline schedule unless pipeline_config names its own y_columns (ValueError otherwise); a replay takes its targets from the base session. Added to pipeline_config unless the config names its own.
schedule_typestr'interval'interval or cron.
schedule_valuestr'1h'Interval: a number with m, h or d ('30m', '6h', '2d'). Cron: five fields ('0 2 * * 1'), validated by the server.
credential_idstrNoneSaved credential for the source.
output_rowsint10000Synthetic output size. Added to pipeline_config unless the config names its own row count.
sql_querystrNoneCustom SQL used instead of source_table.
pipeline_configdict{'yColumns': y_columns, 'outputNumber': output_rows}Pipeline options for each run, in snake_case or camelCase (y_columns, output_rows, text_mode, mdh_mode, dsg_mode, similarity_p, excluded_columns, enable_dsg, mode = speed preset, ...). AutoDataClient.pipeline_config() builds one from the arguments process() takes.
descriptionstrNoneFree text.
modestr'full_pipeline'Run mode.
base_session_idstrNoneParent for replay modes.

Returns: {success, scheduled_run: {...}}.

Raises: ValueError (mode, no targets for a full_pipeline schedule); AutoDataError 400 (invalid cron), 404 (credential not yours).

The y_columns and output_rows arguments apply whether or not you pass pipeline_config. A config that names its own targets or row count takes precedence.

client.create_scheduled_run(
    name="Weekly refresh", connector_type="snowflake", source_table="ORDERS",
    y_columns="churned", schedule_type="cron", schedule_value="0 2 * * 1",
    credential_id=cred_id, mode="retraining", base_session_id=sid)

update_scheduled_run

update_scheduled_run(run_id, schedule_type=None, schedule_value=None, mode=None, base_session_id=None, pipeline_config=None)

Changes when a schedule fires and what it runs. Only the arguments you pass are sent. pipeline_config replaces the stored one.

Endpoint: PUT /api/v1/scheduled-runs/<run_id>

ParameterTypeDefaultMeaning
run_idstr-Schedule.
schedule_type / schedule_valuestrNoneAs in create.
modestrNoneNew run mode; a replay mode needs base_session_id.
base_session_idstrNoneParent.
pipeline_configdictNoneFull replacement, including media options (image_processing_mode, image_columns, ...).

Returns: Server response with the updated schedule.

delete_scheduled_run

delete_scheduled_run(run_id)

Deletes a schedule.

Endpoint: DELETE /api/v1/scheduled-runs/<run_id>

ParameterTypeDefaultMeaning
run_idstr-Schedule.

Returns: Server response.

toggle_scheduled_run

toggle_scheduled_run(run_id)

Switches a schedule between active and paused. It takes only the id and flips the current state.

Endpoint: POST /api/v1/scheduled-runs/<run_id>/toggle

ParameterTypeDefaultMeaning
run_idstr-Schedule.

Returns: {success, is_active, ...}.

run_scheduled_now

run_scheduled_now(run_id)

Runs a schedule immediately, outside its timetable.

Endpoint: POST /api/v1/scheduled-runs/<run_id>/run-now

ParameterTypeDefaultMeaning
run_idstr-Schedule.

Returns: The run result: the new session_id, or status: 'skipped' when the source returned no rows.

Folder listeners

Watch a folder or bucket and run the pipeline on each new file.

list_listeners

list_listeners()

Lists your folder listeners.

Endpoint: GET /api/v1/listeners/folder

Returns: list (the response's listeners).

create_listener

create_listener(name, source_type, folder_path=None, y_columns=None, *, bucket_name=None, prefix=None, credential_id=None, pipeline_config=None, output_rows=10000, poll_interval_seconds=60, allowed_extensions='.csv,.parquet,.json', max_file_size_mb=500, cooldown_seconds=60, max_files_per_hour=10, enabled=True, auto_sink_config=None, mode='full_pipeline', base_session_id=None, watch_path=None)

Creates a listener that runs the pipeline on each file that appears.

Endpoint: POST /api/v1/listeners/folder

ParameterTypeDefaultMeaning
namestr-Label.
source_typestr-local, s3, gcs, azure_blob, adls_gen2 or sftp.
folder_pathstrNoneRequired for local, and must be under the server's uploads/, results/ or listener_data/ folders. Derived from the bucket for cloud sources and set to your inbox for sftp.
y_columnsstr or list[str]NoneRequired for full_pipeline (ValueError otherwise).
bucket_namestrNoneBucket, container or filesystem. Required for cloud sources.
prefixstrNoneOnly watch keys under this prefix.
credential_idstrNoneRequired for cloud sources; the server checks it fits the source type.
pipeline_configdict{}Pipeline options for each run (text_mode, mdh_mode, excluded_columns, task_types, similarity_p, dsg_mode, enable_* stage flags, webhook_ids, ...).
output_rowsint10000Stored as pipeline_config['output_number'] unless already set there.
poll_interval_secondsint60How often the source is checked.
allowed_extensionsstr'.csv,.parquet,.json'Comma-separated extensions to pick up.
max_file_size_mbint500Larger files are skipped.
cooldown_secondsint60Minimum time between two runs.
max_files_per_hourint10Cap on runs per hour.
enabledboolTrueStart watching immediately.
auto_sink_configdictNone{enabled, connector_type, credential_id, table_name, if_exists}: write each output to a sink.
mode / base_session_idstr'full_pipeline' / NoneRun mode and parent.
watch_pathstrNoneDeprecated alias of folder_path.

Returns: Server response with the listener under listener.

Raises: ValueError; AutoDataError 400 (missing bucket or credential, folder not allowed), 404 (credential not found).

client.create_listener("landing", "s3", y_columns="label", bucket_name="acme-landing",
                       prefix="daily/", credential_id=cred_id, allowed_extensions=".csv,.parquet")

update_listener

update_listener(listener_id, *, name=None, folder_path=None, y_columns=None, bucket_name=None, prefix=None, credential_id=None, pipeline_config=None, poll_interval_seconds=None, allowed_extensions=None, max_file_size_mb=None, cooldown_seconds=None, max_files_per_hour=None, enabled=None, auto_sink_config=None, watch_path=None, mode=None, base_session_id=None)

Changes only the fields you pass. Setting enabled starts or stops the watcher.

Endpoint: PUT /api/v1/listeners/folder/<listener_id>

ParameterTypeDefaultMeaning
listener_idstr-Listener.
other keywords-None (unchanged)As in create_listener.

Returns: Server response.

delete_listener

delete_listener(listener_id)

Deletes a listener.

Endpoint: DELETE /api/v1/listeners/folder/<listener_id>

ParameterTypeDefaultMeaning
listener_idstr-Listener.

Returns: Server response.

SFTP inbox

Each SFTP credential is an inbox: every file uploaded to it starts a run.

sftp_info

sftp_info()

Returns how to connect to the SFTP server.

Endpoint: GET /api/v1/sftp/info

Returns: {enabled, host, port, instructions}.

list_sftp_credentials

list_sftp_credentials()

Lists your SFTP inboxes (no passwords).

Endpoint: GET /api/v1/sftp/credentials

Returns: list (the response's credentials).

create_sftp_credential

create_sftp_credential(name, y_columns=None, pipeline_config=None, mode='full_pipeline', base_session_id=None)

Creates an inbox with a generated username and password.

Endpoint: POST /api/v1/sftp/credentials

ParameterTypeDefaultMeaning
namestr-Label.
y_columnsstr or list[str]NoneTargets for full_pipeline. Without them the inbox stays inactive until it is configured.
pipeline_configdictNonePipeline options for each file.
mode / base_session_idstr'full_pipeline' / NoneRun mode and parent.

Returns: {success, credential: {id, sftp_username, sftp_password, connection_string, inbox_path, is_active, listener_id, listener_enabled, y_columns, pipeline_config, mode, base_session_id, ...}}. sftp_password is shown only once.

inbox = client.create_sftp_credential("partner-feed", y_columns=["label"])["credential"]
print(inbox["connection_string"])   # sftp -P <port> <user>@<host>
password = inbox["sftp_password"]     # store it now

update_sftp_pipeline

update_sftp_pipeline(credential_id, y_columns=None, pipeline_config=None, mode=None, base_session_id=None)

Changes what an inbox runs on each file. Only the arguments you pass are sent.

Endpoint: PUT /api/v1/sftp/credentials/<credential_id>/pipeline

ParameterTypeDefaultMeaning
credential_idstr-SFTP credential.
y_columnsstr or list[str]NoneTargets.
pipeline_configdictNoneOptions.
mode / base_session_idstrNoneRun mode and parent.

Returns: The updated inbox (mode, base_session_id, listener_enabled, ...).

delete_sftp_credential

delete_sftp_credential(credential_id)

Deletes an inbox and its login.

Endpoint: DELETE /api/v1/sftp/credentials/<credential_id>

ParameterTypeDefaultMeaning
credential_idstr-SFTP credential.

Returns: Server response.

Runs held for approval

An automation with a price limit (max_price_dollars in its pipeline config) holds a run priced above the limit instead of starting it. Review, approve or discard held runs here.

list_pending_approvals

list_pending_approvals()

Lists your held runs.

Endpoint: GET /api/v1/pending-approvals

Returns: list (the response's approvals); each entry carries id, session_id, source, name, price_usd, cap_usd, status, created_at and decided_at.

approve_run

approve_run(approval_id)

Starts a held run at the price it was held at. Exactly that price is charged.

Endpoint: POST /api/v1/pending-approvals/<approval_id>/approve

ParameterTypeDefaultMeaning
approval_idstr-The held run's id.

Returns: {success, session_id, price_usd, approval}. Follow it with wait_for_completion(session_id).

Raises: AutoDataError 404 (unknown or not yours), 409 (already approved or discarded), 410 (the held data is no longer available).

discard_run

discard_run(approval_id)

Drops a held run without starting it. Nothing is charged.

Endpoint: POST /api/v1/pending-approvals/<approval_id>/discard

ParameterTypeDefaultMeaning
approval_idstr-The held run's id.

Returns: Server response.

Raises: As approve_run.

for held in client.list_pending_approvals():
    if held["price_usd"] <= 5:
        client.approve_run(held["id"])
    else:
        client.discard_run(held["id"])

Streaming inference (Enterprise)

Continuous inference: rows arriving on a Kafka, Kinesis or MQTT source are prepared with a completed session's fitted transforms and written to a sink. These calls need an Enterprise account and answer 403 otherwise. Streams are billed hourly for the rows scored, at the inference rate. The Streaming page describes every field.

list_streams (Enterprise)

list_streams()

Lists your streams.

Endpoint: GET /api/v1/streams

Returns: list (the response's streams).

create_stream (Enterprise)

create_stream(name, base_session_id, source_type, source_credential_id, source_config, sink_type, sink_credential_id, sink_config, dead_letter_config=None, start=True, **options)

Creates a stream that scores a live source against a completed session, continuously.

Endpoint: POST /api/v1/streams

ParameterTypeDefaultMeaning
namestr-Label.
base_session_idstr-The completed session whose transforms are applied.
source_typestr-kafka, kinesis or mqtt.
source_credential_idstr-Saved credential for the source.
source_configdict-Where to read, for example {'topic': 'events'} or {'stream_name': 'events'}.
sink_type / sink_credential_id / sink_config--The same, for where prepared rows are written. sink_config['output_format'] is json, values or sparse.
dead_letter_configdictNoneOptional sink for rows that could not be prepared: {sink_type, credential_id, topic or stream_name}.
startboolTrueStart consuming immediately.
**options--max_batch_rows, max_batch_ms, parallelism, enable_anomaly.

Returns: dict with the created stream (HTTP 201).

Raises: AutoDataError 400 (invalid configuration), 403 (not an Enterprise account).

created = client.create_stream(
    "orders-live", base_session_id=sid,
    source_type="kafka", source_credential_id=kafka_cred, source_config={"topic": "orders"},
    sink_type="kafka", sink_credential_id=kafka_cred,
    sink_config={"topic": "orders-prepared", "output_format": "json"})
stream_id = created["stream"]["id"]
print(client.stream_status(stream_id))

get_stream (Enterprise)

get_stream(stream_id)

One stream's configuration and state.

Endpoint: GET /api/v1/streams/<stream_id>

ParameterTypeDefaultMeaning
stream_idstr-Stream.

Returns: Stream dict.

update_stream (Enterprise)

update_stream(stream_id, **fields)

Changes only the fields you pass (the fields of create_stream). dead_letter_config=None removes the dead-letter sink.

Endpoint: PUT /api/v1/streams/<stream_id>

ParameterTypeDefaultMeaning
stream_idstr-Stream.
**fields--Fields to change.

Returns: Server response.

delete_stream (Enterprise)

delete_stream(stream_id)

Stops and deletes a stream.

Endpoint: DELETE /api/v1/streams/<stream_id>

ParameterTypeDefaultMeaning
stream_idstr-Stream.

Returns: Server response.

start_stream (Enterprise)

start_stream(stream_id)

Starts consuming.

Endpoint: POST /api/v1/streams/<stream_id>/start

ParameterTypeDefaultMeaning
stream_idstr-Stream.

Returns: Stream dict.

stop_stream (Enterprise)

stop_stream(stream_id)

Stops consuming after the batch in hand is delivered.

Endpoint: POST /api/v1/streams/<stream_id>/stop

ParameterTypeDefaultMeaning
stream_idstr-Stream.

Returns: Stream dict.

restart_stream (Enterprise)

restart_stream(stream_id)

Restarts a stream, picking up a changed configuration.

Endpoint: POST /api/v1/streams/<stream_id>/restart

ParameterTypeDefaultMeaning
stream_idstr-Stream.

Returns: Stream dict.

stream_status (Enterprise)

stream_status(stream_id)

Live state of a stream: workers, throughput, lag and the last error.

Endpoint: GET /api/v1/streams/<stream_id>/status

ParameterTypeDefaultMeaning
stream_idstr-Stream.

Returns: Status dict.

stream_schema (Enterprise)

stream_schema(stream_id)

The columns a stream writes, in order: what the values and sparse output formats index into.

Endpoint: GET /api/v1/streams/<stream_id>/schema

ParameterTypeDefaultMeaning
stream_idstr-Stream.

Returns: Schema dict.

test_stream_connection (Enterprise)

test_stream_connection(connector_type, credential_id, config=None, role='source')

Checks that a saved credential can reach a topic or stream before you build a stream on it.

Endpoint: POST /api/v1/streams/test-connection

ParameterTypeDefaultMeaning
connector_typestr-kafka, kinesis or mqtt.
credential_idstr-Saved credential.
configdict{}The source or sink config to check.
rolestr'source'source or sink.

Returns: The result of the check.

Sync watermarks

Watermarks record how far an incremental read has got.

list_watermarks

list_watermarks()

Lists your sync watermarks.

Endpoint: GET /api/v1/sync/watermarks/list

Returns: list of watermark dicts (id, connector_type, source_key, watermark_column, last value, row count, ...).

reset_watermark

reset_watermark(watermark_id)

Resets a watermark so the next incremental read is a full read.

Endpoint: POST /api/v1/sync/watermarks/<watermark_id>/reset

ParameterTypeDefaultMeaning
watermark_idstr-Watermark.

Returns: {success, ...} with the updated watermark.

get_watermark

get_watermark(connector_type, source_key)

Reads the stored watermark of one source.

Endpoint: GET /api/v1/sync/watermarks?connector_type=...&source_key=...

ParameterTypeDefaultMeaning
connector_typestr-Connector type.
source_keystr-The saved credential's id when one was used, otherwise the connector type.

Returns: The watermark dict, or None when none is stored.

set_watermark

set_watermark(connector_type, source_key, **fields)

Creates or moves a watermark.

Endpoint: PUT /api/v1/sync/watermarks

ParameterTypeDefaultMeaning
connector_typestr-Connector type.
source_keystr-The saved credential's id when one was used, otherwise the connector type.
**fields--table_name, watermark_column, watermark_value, watermark_type (timestamp, integer or string), row_count, last_sync_at (ISO 8601).

Returns: The watermark dict.

create_watermark

create_watermark(connector_type, source_key, watermark_column, watermark_value, table_name=None, watermark_type='timestamp')

Sets the starting point of an incremental read before its first run.

Endpoint: POST /api/v1/sync/watermarks

ParameterTypeDefaultMeaning
connector_typestr-Connector type.
source_keystr-The saved credential's id when one was used, otherwise the connector type.
watermark_columnstr-The column that orders the source.
watermark_valueany-The value the first incremental read starts after.
table_namestrNoneTable the watermark belongs to.
watermark_typestr'timestamp'timestamp, integer or string.

Returns: The watermark dict.

Raises: AutoDataError 409 when a watermark already exists for the source; use set_watermark to move one.

client.create_watermark("snowflake", cred_id, "updated_at", "2026-10-01T00:00:00")
print(client.get_watermark("snowflake", cred_id))

delete_watermark

delete_watermark(connector_type, source_key)

Removes a watermark; the next incremental read is a full read.

Endpoint: DELETE /api/v1/sync/watermarks?connector_type=...&source_key=...

ParameterTypeDefaultMeaning
connector_typestr-Connector type.
source_keystr-The saved credential's id when one was used, otherwise the connector type.

Returns: {success, message}.

Raises: AutoDataError 404 when no watermark is stored for the source.

Quality alerts

Rules checked against each run's metrics.

list_quality_alerts

list_quality_alerts()

Lists your alert rules.

Endpoint: GET /api/v1/quality-alerts

Returns: list (the response's rules).

create_quality_alert

create_quality_alert(name, metric, operator, threshold, severity='warning', stage=None)

Creates a rule.

Endpoint: POST /api/v1/quality-alerts

ParameterTypeDefaultMeaning
namestr-Label.
metricstr-row_loss_pct, null_pct, column_drop_count, duration_seconds or drift_score.
operatorstr->, <, >=, <= or ==.
thresholdfloat-Value to compare with.
severitystr'warning'warning or critical.
stagestrNone (every stage)anomaly, dtc, mdh, cds, dsm or dsg.

Returns: {success, rule: {...}}.

Raises: AutoDataError 400 (invalid metric or operator).

The SDK cannot attach webhook_ids or set enabled. Rules with webhooks are stored, but quality.breach notifications are not delivered yet.

client.create_quality_alert("too many rows lost", "row_loss_pct", ">", 20, severity="critical", stage="mdh")

update_quality_alert

update_quality_alert(rule_id, name=None, metric=None, operator=None, threshold=None, severity=None, stage=None)

Changes only the fields you pass.

Endpoint: PUT /api/v1/quality-alerts/<rule_id>

ParameterTypeDefaultMeaning
rule_idstr-Rule.
other arguments-None (unchanged)As in create.

Returns: {success, rule}.

delete_quality_alert

delete_quality_alert(rule_id)

Deletes a rule.

Endpoint: DELETE /api/v1/quality-alerts/<rule_id>

ParameterTypeDefaultMeaning
rule_idstr-Rule.

Returns: {success, message}.

get_alert_events

get_alert_events(session_id=None)

Lists the most recent alerts that fired (up to 100).

Endpoint: GET /api/v1/quality-alerts/events

ParameterTypeDefaultMeaning
session_idstrNoneOnly events of this session.

Returns: list (the response's events).

Webhooks

Get a POST when something happens instead of polling. Each delivery body is {event, data, timestamp, webhook_id}. If the webhook has a secret, the request carries X-Webhook-Signature: sha256=<hex>: the HMAC-SHA256 of the body, re-serialized with sorted keys (Python json.dumps(body, sort_keys=True)), keyed with the secret.

list_webhooks

list_webhooks()

Lists your webhooks.

Endpoint: GET /api/v1/webhooks

Returns: list of webhook dicts.

create_webhook

create_webhook(url, events=None, name=None, secret=None)

Registers an endpoint. The URL must resolve to a public address (checked again at every delivery; redirects are not followed).

Endpoint: POST /api/v1/webhooks

ParameterTypeDefaultMeaning
urlstr-HTTPS endpoint.
eventslist[str]None, which subscribes to ['error'] onlyAny of job.completed, job.failed, quality.breach, error, warning, info, success.
namestr'Unnamed Webhook'Label.
secretstrNoneSigning secret. Set it, and verify the signature.

Returns: {webhook_id, message, webhook: {...}} (HTTP 201).

Raises: AutoDataError 400 (URL not allowed, unknown event).

Pass events explicitly. Leaving it out subscribes only to error, so you will not hear about completed jobs.

import hashlib, hmac, json

client.create_webhook("https://example.com/hooks/autodata",
                      events=["job.completed", "job.failed"], secret="s3cret")

# In your receiver:
def verify(raw_body: bytes, header: str, secret: str) -> bool:
    canonical = json.dumps(json.loads(raw_body), sort_keys=True)
    mac = hmac.new(secret.encode(), canonical.encode(), hashlib.sha256).hexdigest()
    return hmac.compare_digest(header, f"sha256={mac}")

get_webhook

get_webhook(webhook_id)

Fetches one webhook.

Endpoint: GET /api/v1/webhooks/<webhook_id>

ParameterTypeDefaultMeaning
webhook_idstr-Webhook.

Returns: Webhook dict.

update_webhook

update_webhook(webhook_id, **fields)

Changes only the fields you pass: url, events, name, secret, active (is_active is accepted as another spelling of active).

Endpoint: PUT /api/v1/webhooks/<webhook_id>

ParameterTypeDefaultMeaning
webhook_idstr-Webhook.
**fields--Fields to change.

Returns: {message, webhook}.

Pass active= to enable or disable the webhook.

client.update_webhook(wh_id, active=False)   # pause deliveries

delete_webhook

delete_webhook(webhook_id)

Deletes a webhook.

Endpoint: DELETE /api/v1/webhooks/<webhook_id>

ParameterTypeDefaultMeaning
webhook_idstr-Webhook.

Returns: {message}.

test_webhook

test_webhook(webhook_id)

Sends a signed test delivery now.

Endpoint: POST /api/v1/webhooks/<webhook_id>/test

ParameterTypeDefaultMeaning
webhook_idstr-Webhook.

Returns: {..., response_status}, or {warning, response_status} when your endpoint answered with an error status.

webhook_deliveries

webhook_deliveries(webhook_id)

Lists recent delivery attempts.

Endpoint: GET /api/v1/webhooks/<webhook_id>/deliveries

ParameterTypeDefaultMeaning
webhook_idstr-Webhook.

Returns: list of deliveries (HTTP status, response, duration, error).

Notifications

E-mail when a run completes or ends in an error.

get_notification_preferences

get_notification_preferences()

Whether the account is e-mailed when a run completes or ends in an error.

Endpoint: GET /api/v1/notification-preferences

Returns: {email_on_complete, email_on_fail, configured}; configured tells whether the server can send e-mail.

set_notification_preferences

set_notification_preferences(email_on_complete=None, email_on_fail=None)

Turns the two e-mails on or off. Only the settings you pass are changed.

Endpoint: PUT /api/v1/notification-preferences

ParameterTypeDefaultMeaning
email_on_completeboolNone (unchanged)E-mail when a run completes.
email_on_failboolNone (unchanged)E-mail when a run ends in an error or is held for approval.

Returns: The same dict as get_notification_preferences().

client.set_notification_preferences(email_on_fail=True)

Sessions

Your run history.

list_sessions

list_sessions()

Lists your sessions, newest first.

Endpoint: GET /api/v1/sessions

Returns: list of session dicts (session_id, session_name, status, total_rows, processing_date, ...).

get_session

get_session(session_id)

Full detail of one session, including the parameters it ran with.

Endpoint: GET /api/v1/sessions/<session_id>

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: Session detail dict.

rename_session

rename_session(session_id, name)

Sets a session's display name.

Endpoint: PUT /api/v1/sessions/<session_id>/name

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
namestr-New name.

Returns: Server response.

set_session_retraining

set_session_retraining(session_id, enabled)

Marks or unmarks a session as a retraining base.

Endpoint: PUT /api/v1/sessions/<session_id>/retraining

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
enabledbool-Flag.

Returns: Server response.

list_session_outputs

list_session_outputs(session_id)

Lists a session's outputs kept in memory rather than written to disk.

Endpoint: GET /api/v1/sessions/<session_id>/outputs

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: Server response listing memory outputs.

get_session_report

get_session_report(session_id)

What a finished run did to the data, stage by stage: the report the dashboard shows.

Endpoint: GET /api/v1/sessions/<session_id>/report

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: {session_id, tier, status, summary, warnings}, plus details (per-column detail) for Enterprise accounts.

get_session_config

get_session_config(session_id)

The configuration a session ran with, ready to reuse for a new run or to save as a template.

Endpoint: GET /api/v1/sessions/<session_id>/config

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: {session_id, config, run_mode}.

set_session_meta

set_session_meta(session_id, tags=None, note=None)

Tags a session and/or attaches a note to it. Only what you pass is changed; tags=[] clears the tags and note='' clears the note.

Endpoint: PUT /api/v1/sessions/<session_id>/meta

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
tagslist[str]None (unchanged)Up to 10 tags of up to 32 characters each.
notestrNone (unchanged)Up to 500 characters.

Returns: {tags, note}.

Raises: AutoDataError 400 (too many or too long), 404 (session not found).

client.set_session_meta(sid, tags=["churn", "q4"], note="Baseline for the Q4 model")

list_session_files

list_session_files(session_id)

Every downloadable file of a session you own, or the shared files of a session shared with you.

Endpoint: GET /api/v1/sessions/<session_id>/files

ParameterTypeDefaultMeaning
session_idstr-A session you own or that is shared with you.

Returns: list (the response's files); each entry carries name, size_mb and url.

download_all

download_all(session_id, download_path=None, extract=True)

Downloads every output of a session as one archive.

Endpoint: GET /api/v1/sessions/<session_id>/download-all

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
download_pathstr./auto_data_outputs/<session_id>/Folder, created if needed.
extractboolTrueUnpack the archive. With False the .zip itself is saved.

Returns: Absolute path of the folder, or of the .zip when extract=False.

Raises: AutoDataError on an empty or invalid archive.

download_session_output

download_session_output(session_id, output_name, output_path=None)

Downloads one output that is held in memory rather than on disk.

Endpoint: GET /api/v1/sessions/<session_id>/outputs/<output_name>

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
output_namestr-A name from list_session_outputs().
output_pathstr./auto_data_outputs/<session_id>/<output_name>.csvLocal file path.

Returns: Absolute path of the file written.

clear_session_outputs

clear_session_outputs(session_id)

Releases a session's in-memory outputs. Files on disk are not affected.

Endpoint: DELETE /api/v1/sessions/<session_id>/outputs

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: Server response.

Sharing

Share a session with another account or through a link.

list_shared_sessions

list_shared_sessions()

Shares you created.

Endpoint: GET /api/v1/shared-sessions

Returns: list of shares (with session_name, shared_with_name, permission, expires_at, share_token for links, ...).

list_received_sessions

list_received_sessions()

Sessions other accounts shared with you.

Endpoint: GET /api/v1/shared-sessions/received

Returns: list of shares (with owner_name, session_name, permission, ...).

share_session

share_session(session_id, shared_with_id=None, permission='read', expires_at=None, note=None, shared_files=None)

Shares one of your sessions.

Endpoint: POST /api/v1/shared-sessions

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
shared_with_idstrNoneRecipient's user id. Omit to create a link token instead.
permissionstr'read'read, comment or edit. Unknown values become read.
expires_atstr (ISO 8601)None (never)Expiry. Strongly recommended for links.
notestrNoneShown to the recipient.
shared_fileslist[str]NoneFile names the recipient may download (names from list_session_files()). A share without any shows the run and its report and offers no file, so pass this whenever the data itself is what you are sharing.

Returns: {success, shared_session: {...}} (with share_token for a link).

Raises: AutoDataError 404 (not your session).

unshare_session

unshare_session(share_id)

Revokes a share; a link stops working at once.

Endpoint: DELETE /api/v1/shared-sessions/<share_id>

ParameterTypeDefaultMeaning
share_idstr-The share's id.

Returns: Server response.

update_share

update_share(share_id, **fields)

Changes a share you created. A change to a link applies to everyone who joined it.

Endpoint: PUT /api/v1/shared-sessions/<share_id>

ParameterTypeDefaultMeaning
share_idstr-The share's id.
**fields--permission (read, comment or edit), expires_at (ISO 8601, or None to remove the expiry), note, shared_files.

Returns: Server response.

get_shared_session

get_shared_session(token)

Opens a share link: the session's summary and the files it offers.

Endpoint: GET /api/v1/shared-sessions/access/<token>

ParameterTypeDefaultMeaning
tokenstr-The share link's token.

Returns: dict with the shared session and its files.

Raises: AutoDataError 404 (unknown link), 410 (expired link).

join_shared_session

join_shared_session(token)

Adds a share link to your own account, so it appears in list_received_sessions() and its session can be the base of an inference or retraining run.

Endpoint: POST /api/v1/shared-sessions/access/<token>/join

ParameterTypeDefaultMeaning
tokenstr-The share link's token.

Returns: Server response.

download_shared_file

download_shared_file(token, filename, output_path=None)

Downloads one file through a share link.

Endpoint: GET /api/v1/shared-sessions/access/<token>/download/<filename>

ParameterTypeDefaultMeaning
tokenstr-The share link's token.
filenamestr-One of the shared file names.
output_pathstr./<filename>Local file path.

Returns: Absolute path of the file written.

list_share_comments

list_share_comments(session_id)

Comments on a session you own or that is shared with you.

Endpoint: GET /api/v1/shared-sessions/by-session/<session_id>/comments

ParameterTypeDefaultMeaning
session_idstr-A session you own or that is shared with you.

Returns: list (the response's comments).

add_share_comment

add_share_comment(session_id, body)

Comments on a session. On a session that is not your own you need comment or edit permission.

Endpoint: POST /api/v1/shared-sessions/by-session/<session_id>/comments

ParameterTypeDefaultMeaning
session_idstr-A session you own or that is shared with you.
bodystr-The comment, up to 2000 characters.

Returns: Server response with the new comment.

delete_share_comment

delete_share_comment(session_id, comment_id)

Deletes a comment: your own, or any comment on a session you own.

Endpoint: DELETE /api/v1/shared-sessions/by-session/<session_id>/comments/<comment_id>

ParameterTypeDefaultMeaning
session_idstr-The session.
comment_idstr-The comment.

Returns: Server response.

search_users

search_users(query)

Finds the account to share with, by exact e-mail address. It is a lookup, not a directory: it answers only for a complete address, and is limited to 20 searches per minute.

Endpoint: GET /api/v1/users/search?q=...

ParameterTypeDefaultMeaning
querystr-The recipient's complete e-mail address.

Returns: list (the response's users); each entry carries id, name and email_hint. Pass id as shared_with_id.

user = client.search_users("ana@example.com")[0]
files = [f["name"] for f in client.list_session_files(sid)]
client.share_session(sid, shared_with_id=user["id"], permission="comment", shared_files=files)

Preferences

Saved preference profiles. AutoDataClient.PREFERENCE_KINDS is ('outputs', 'advanced', 'anomaly', 'completion', 'api_anomaly'); any other kind raises ValueError.

get_preferences

get_preferences(kind)

Reads one profile.

Endpoint: GET /api/v1/preferences/<kind> (api_anomaly is /api/v1/preferences/api-anomaly)

ParameterTypeDefaultMeaning
kindstr-outputs (which files are kept, output format and name pattern), advanced (dashboard pipeline defaults), anomaly (data-cleaning checks of runs started in the dashboard), completion (Data Completion & Verification settings), api_anomaly (data-cleaning defaults of runs started through the API or this client: what process() uses when anomaly_params is not passed).

Returns: {success, preferences} for outputs or {success, parameters} for the others, plus the available names.

set_preferences

set_preferences(kind, values)

Replaces one profile.

Endpoint: POST /api/v1/preferences/<kind>

ParameterTypeDefaultMeaning
kindstr-As above.
valuesdict-The profile itself: the inner preferences / parameters dict from get_preferences, not the whole response.

Returns: {success, message, preferences|parameters}.

Two profiles apply to runs started with process(): outputs decides which files are written unless the call passes output_config, and api_anomaly supplies the cleaning checks unless the call passes anomaly_params. The advanced, anomaly and completion profiles are dashboard defaults.

prefs = client.get_preferences("outputs")["preferences"]
prefs["output_format"] = "parquet"
client.set_preferences("outputs", prefs)

# Cleaning defaults for runs started through the API or the SDK
cleaning = client.get_preferences("api_anomaly")["parameters"]

reset_preferences

reset_preferences(kind)

Restores one profile to the defaults.

Endpoint: POST /api/v1/preferences/<kind>/reset

ParameterTypeDefaultMeaning
kindstr-As above.

Returns: Server response.

Account, usage and cost

list_keys

list_keys()

Lists your active API keys (no secrets).

Endpoint: GET /api/v1/keys

Returns: list (the response's api_keys); each key carries daily_limit_usd, lifetime_limit_usd, daily_spent_usd and lifetime_spent_usd in US dollars (None = no limit).

get_usage

get_usage()

Spend, in US dollars, of the key making the call.

Endpoint: GET /api/v1/usage

Returns: {api_key_id, name, key_prefix, today_usd, daily_limit_usd, daily_remaining_usd, has_daily_limit, lifetime_spent_usd, lifetime_limit_usd, lifetime_remaining_usd, has_lifetime_limit, avg_daily_usd_7d, avg_daily_usd_30d, chart_data: {labels, usd}, daily_request_count, last_used_at}. A limit of None means no limit. {} with passcode auth. Deprecated, still returned, equal to the dollar amount × 1,000,000, and will be removed in a future version: daily_credits_used, daily_credit_limit, daily_remaining, lifetime_credits_used, lifetime_credit_limit, lifetime_remaining.

account_profile

account_profile()

Balance in US dollars, role and API limits of the whole account.

Endpoint: GET /api/v1/account/profile

Returns: Profile dict with balance_usd, spent_usd and total_usd (balance + spent), the role and the API limits. api_credits_remaining, api_credits_used and api_credits_total are deprecated.

account_usage

account_usage()

Account-wide usage across every key and the dashboard.

Endpoint: GET /api/v1/account/usage

Returns: {summary, by_tool}: summary.total_spent_usd and a spent_usd per stage (credits_used and total_credits_used are deprecated).

account_contracts

account_contracts()

Your pricing contracts, with usage against each allowance.

Endpoint: GET /api/v1/account/contracts

Returns: list (the response's contracts); empty when the account has none.

trial_status

trial_status()

Whether the account is on a trial, and how many free runs remain.

Endpoint: GET /api/v1/account/trial-status

Returns: {success, is_trial, trial_uploads_remaining, role, role_name, total_runs, completed_runs, failed_runs, total_data_mb, limits}.

estimate

estimate(input_rows=None, columns=None, column_types=None, missing_rate=None, tools=None, output_rows=None, text_mode=None, session_type='api', pricing_profile_id=None, media_bytes=None, completion=None)

Prices a run before you start it. Runs and charges nothing. With pricing_profile_id (from validate()) it returns the exact price of that file with these settings; start the run with the returned quote_id and it is charged exactly that. With a description of the data instead, it returns an estimated range, which binds nothing.

Endpoint: POST /api/v1/estimate

ParameterTypeDefaultMeaning
input_rowsintNoneRow count. Required unless pricing_profile_id is given.
columnsintNoneColumn count. Required unless column_types or pricing_profile_id is given.
column_typesdictNoneOptional counts by kind: numeric, boolean, categorical_low, categorical_high, datetime, text, identifier. Columns you don't describe are priced across the whole plausible range, so describing them narrows it.
missing_ratefloatNoneOptional share of missing cells (0–1, or a percentage).
toolsdictnormal run (anomaly off)anomaly, dtc, mdh, cds, dsm, dsg.
output_rowsintNoneTarget row count; synthetic rows are charged per doubling of the input rows.
text_modeintNone0 drop text, 1 neural, 2 TF-IDF. In a range, automatic is counted at the neural price.
session_typestr'api'api (full pipeline), inference (about a fifth) or retraining (about two fifths).
pricing_profile_idstrNoneThe profile validate() returned for a file, once ready (see wait_for_price). Gives the exact price.
media_bytesdictNone{'image': bytes, 'audio': bytes} of the media the run will upload.
completiondictNoneDataset completion and validation settings; included in the exact price.

Returns: with pricing_profile_id, the exact quote {quote_id, exact: True, price_usd, breakdown: [{item, usd}], rate_card_version, ...}, or {pricing_status: 'measuring', progress_pct} while the file is still being measured. With a description, a range {quote_id, exact: False, low_usd, high_usd, point_usd, breakdown, rate_card_version}. All prices are US dollars. The *_credits fields (price_credits, low_credits, high_credits, point_credits) are deprecated: still returned, equal to the dollar amount × 1,000,000, and will be removed in a future version.

Price with the same stages, output_rows, text_mode, media and completion settings you will run with. An exact quote holds for the same file (checked by content) and settings that price the same, for 7 days; otherwise process() raises AutoDataError with status_code 409 and nothing starts. Without a quote_id, a run is charged the exact price of its data, measured when it starts.

info = client.validate("data.csv")
if info.get("pricing_status") == "measuring":
    client.wait_for_price(info["pricing_profile_id"], timeout=600)
q = client.estimate(pricing_profile_id=info["pricing_profile_id"], output_rows=100_000, text_mode=2)
print(q["price_usd"])
result = client.process("data.csv", target_columns="label", output_rows=100_000,
                        advanced_params={"text_mode": 2, "quote_id": q["quote_id"]})

worker_status

worker_status()

State of the processing fleet.

Endpoint: GET /api/v1/workers/status

Returns: Worker status dict with cpu_worker queue statistics.

Enterprise: reading what a run decided (Enterprise)

These calls need an Enterprise account and answer 403 otherwise. To use Enterprise preparation, pass the fields listed under Enterprise options in process()'s advanced_params.

pivot_suggest (Enterprise)

pivot_suggest(session_id, target_columns=None)

Ranks the columns that could be the as-of time of a session's rows, with the reasons for each.

Endpoint: POST /api/v1/pivot/suggest

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
target_columnslist[str][]Targets, excluded from the candidates.

Returns: {chosen, candidates, temporal_positions}.

enterprise_families (Enterprise)

enterprise_families()

Model families a run can target, and the per-column imputation and scaling methods you can pin.

Endpoint: GET /api/v1/enterprise/families

Returns: Families and method lists.

enterprise_columns (Enterprise)

enterprise_columns(session_id, family=None)

Per-column record: what each column received and what else was considered.

Endpoint: GET /api/v1/enterprise/runs/<session_id>/columns

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
familystrNone (recommended family)Model family to read.

Returns: {session_id, family, ...ledger}.

Raises: AutoDataError 404 when the session or family has no record.

enterprise_leaderboard (Enterprise)

enterprise_leaderboard(session_id)

How the model families compared, and which is recommended.

Endpoint: GET /api/v1/enterprise/runs/<session_id>/leaderboard

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: {session_id, ...leaderboard}.

enterprise_lineage (Enterprise)

enterprise_lineage(session_id)

Everything recorded about a run: stages, families and decisions.

Endpoint: GET /api/v1/enterprise/lineage/<session_id>

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: Lineage dict.

enterprise_diff (Enterprise)

enterprise_diff(session_id, against, family=None)

What changed between this run's decisions and another run's.

Endpoint: GET /api/v1/enterprise/lineage/<session_id>/diff?against=...

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
againststr-The other session.
familystrNoneRestrict to one model family.

Returns: Diff dict.

Raises: AutoDataError 404 (comparison session not found).

enterprise_audit_report (Enterprise)

enterprise_audit_report(session_id, fmt='json')

Downloadable audit trail of a run.

Endpoint: GET /api/v1/enterprise/audit-report/<session_id>?format=...

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.
fmtstr'json'json (returns a dict) or html (returns the page as a string).

Returns: dict or str.

validate_dataset_metadata (Enterprise)

validate_dataset_metadata(metadata, columns=None, y_columns=None)

Checks a dataset-metadata declaration before you run with it. Nothing is stored.

Endpoint: POST /api/v1/dataset-metadata/validate

ParameterTypeDefaultMeaning
metadatadict-The declaration (see dataset_metadata under Enterprise options).
columnslist[str]NoneThe dataset's column names, so declared columns can be checked against them.
y_columnslist[str]NoneThe targets, so declared roles can be checked against them.

Returns: {metadata, warnings, summary}: the declaration as a run will read it, plus warnings.

enterprise_dataset_metadata (Enterprise)

enterprise_dataset_metadata(session_id)

The dataset metadata a finished run used, and how each declaration was applied.

Endpoint: GET /api/v1/enterprise/runs/<session_id>/dataset-metadata

ParameterTypeDefaultMeaning
session_idstr-A session owned by the caller.

Returns: Metadata report dict.

catalog_test (Enterprise)

catalog_test(catalog, credential_id=None, secrets=None)

Checks that a data catalog can be reached with a connection.

Endpoint: POST /api/v1/dataset-metadata/catalog/test

ParameterTypeDefaultMeaning
catalogstr-Catalog kind, for example snowflake, databricks, bigquery or glue.
credential_idstrNoneSaved credential for the catalog.
secretsdictNoneInline secrets (sent as inline_secrets).

Returns: The result of the check.

catalog_tables (Enterprise)

catalog_tables(catalog, credential_id=None, secrets=None, query=None)

Lists the tables a data catalog describes.

Endpoint: POST /api/v1/dataset-metadata/catalog/tables

ParameterTypeDefaultMeaning
catalogstr-Catalog kind, for example snowflake, databricks, bigquery or glue.
credential_idstrNoneSaved credential for the catalog.
secretsdictNoneInline secrets (sent as inline_secrets).
querystrNoneOptional filter on the table names.

Returns: dict with the tables found.

catalog_import (Enterprise)

catalog_import(catalog, table, credential_id=None, secrets=None, columns=None, y_columns=None, existing=None)

Reads a table's description from a data catalog as dataset metadata, ready to pass as advanced_params={'dataset_metadata': ...}.

Endpoint: POST /api/v1/dataset-metadata/catalog/import

ParameterTypeDefaultMeaning
catalogstr-Catalog kind, for example snowflake, databricks, bigquery or glue.
tablestr-The table to import, as the catalog names it.
credential_idstrNoneSaved credential for the catalog.
secretsdictNoneInline secrets (sent as inline_secrets).
columns / y_columnslist[str]NoneThe dataset's columns and targets, as in validate_dataset_metadata.
existingdictNoneA declaration to merge into. What you already declared is kept.

Returns: dict with the imported declaration.

enterprise_docs (Enterprise)

enterprise_docs()

Index of the "How It Works" notes on the pipeline's modules.

Endpoint: GET /api/v1/enterprise/docs

Returns: Index dict.

enterprise_doc (Enterprise)

enterprise_doc(module_id)

One module's "How It Works" note.

Endpoint: GET /api/v1/enterprise/docs/<module_id>

ParameterTypeDefaultMeaning
module_idstr-A module id from enterprise_docs().

Returns: The note as a dict.

Spark (large files on the server)

Only available when the server has Spark enabled; otherwise spark_read and spark_transform answer 503. Paths are server-side and must be inside the server's upload or results folders (403 otherwise).

spark_status

spark_status()

Whether Spark is available.

Endpoint: GET /api/v1/spark/status

Returns: {available: bool, ...} (reason when not).

spark_read

spark_read(file_path, sample_rows=20)

Reads a large server-side file and returns its statistics.

Endpoint: POST /api/v1/spark/read

ParameterTypeDefaultMeaning
file_pathstr-Server-side path.
sample_rowsint20Rows to return as a sample.

Returns: {row_count, columns, dtypes, sample}.

spark_transform

spark_transform(file_path, output_path, operations, output_format='csv')

Applies a list of operations to a large server-side file.

Endpoint: POST /api/v1/spark/transform

ParameterTypeDefaultMeaning
file_pathstr-Server-side input path.
output_pathstr-Server-side output path.
operationslist[dict]-In order, e.g. {'op': 'cast', 'type_casts': {...}}, {'op': 'drop_nulls', 'subset': [...]}, {'op': 'fill_nulls', 'strategy': 'mean'}, {'op': 'normalize', 'method': 'minmax'}, {'op': 'sample', 'n': 50000}.
output_formatstr'csv'csv or parquet.

Returns: Output statistics (row_count, columns, ...).

Client lifecycle

close

close()

Closes the HTTP session. Called automatically when a with block ends.

Returns: None.

Not available in the SDK

API keys are created, rotated, edited and revoked in the dashboard (API Account page), including team keys and per-key limits; profile changes and a new passcode are made there too. The SDK lists your keys (list_keys()) and reads their usage (get_usage()). The API Settings and API Config tabs of the API Account page are reserved and are not applied to runs yet; set options on each call.

These options have REST fields but no SDK argument yet. Send them with any HTTP client using the same Authorization: Bearer dtpk_... header:

  • Quality-alert webhook_ids and enabled.
  • Folder-listener watermark_column for incremental reads.

Usage notes for 0.15.0

  • process_from_connector() sends no defaults of its own: an unset text_mode or similarity_p takes the server default (2 and 0.99), the same as a file run. Connector runs that relied on text features being off should pass advanced_params={"text_mode": 0}.
  • process() treats every target as classification unless you pass task_types, and takes its cleaning checks from your saved defaults for API runs unless you pass anomaly_params.
  • output_config selects what a run writes; output_preferences only filters what is downloaded.
  • update_webhook(): pass active= (or is_active=) to enable or disable a webhook. Pass events explicitly when creating one; the default is error only.
  • wait_for_completion's stall_timeout is 3600 s.