Browse documentation
2026-09-29AutoDataOpen in dashboard

Folder & Cloud Listeners

Watch an S3, GCS, Azure Blob or ADLS Gen2 location, or a server folder, and start a run for every new file.

What a listener does

A listener watches a storage location and starts a run for every new file that appears. It can watch an Amazon S3 bucket, a Google Cloud Storage bucket, an Azure Blob container, an Azure Data Lake Storage Gen2 filesystem, or (on self-hosted installs) a folder on the AutoData server. Each file becomes one run in the Runs tab.

Before you start

  • Save a connection of the matching type (s3, gcs, azure_blob, adls_gen2) on the Connections tab. The listener checks its secrets when you create it.
  • A server folder must be inside the server's uploads/, results/ or listener_data/ directory.
  • For inference or retraining, a completed session whose stored transforms are still available.

Creating a listener

  1. Open Listeners → Folder Listeners and click Add Folder Listener.
  2. Pick the source type: Amazon S3, Google Cloud Storage, Azure Blob, Azure Data Lake (Gen2) or Server Folder.
  3. Enter the bucket, container or filesystem, an optional prefix and the saved credential; or the folder path.
  4. Set allowed extensions and the poll interval.
  5. Choose what each file runs, and for a new model open Configure to set target columns and other settings.
  6. Click Create. Polling starts within a few seconds.

Run modes

OptionmodeEach runNeeds
Prepare a new modelfull_pipelineTrains a new session with your pipeline configuration.Target columns
Run inferenceinferenceScores the new data with the stored transforms of a completed session. Output columns match that session.base_session_id
Retrain a sessionretrainingRefreshes a completed session with the new data, keeping its output schema. Options: remove outlier rows (run_dor), add synthetic rows (run_dsg), target rows (output_number, default 10000).base_session_id

If the base session's stored transforms have been removed, the listener records an error and skips the file rather than retrying it.

How files are picked up

  • The location is listed every poll_interval_seconds. A server folder is read without sub-folders; a cloud listener sees everything under its prefix.
  • Files with a disallowed extension or larger than max_file_size_mb are skipped.
  • The last 500 file names are remembered, so each file is processed once.
  • A new run starts only when the cooldown since the last one has passed and the hourly cap is not reached. Held-back files are picked up by a later poll.
  • A new-model listener skips a file that lacks a target column and records why.
  • Images or audio named by the file are collected from the same folder or prefix; upload them first. URL cells are fetched directly.

Accepted file types

The API default is every readable format: .csv, .feather, .json, .jsonl, .ndjson, .orc, .parquet, .xls, .xlsx. The dashboard form and the Python SDK pre-fill .csv,.parquet,.json; widen it if you drop other formats. .tsv, .tab and .txt are read when listed.

Price per run

Each run is charged the exact price of the data it actually processes, measured when the run starts. Nobody confirms the price first, so you can set an optional Maximum price per run ($) (max_price_dollars). When a listener's run is priced above it, the run is held instead of running and listed under Runs → Waiting for approval with its price and the limit. Approve starts it at that price, and it is charged exactly that; Discard drops it. You are notified by a job.held webhook event and, if you receive failure emails, by email. Leave the field empty for no limit. The limit applies to runs that prepare a new model. See Account and billing.

Status

StatusMeaning
idleCreated or paused.
watchingPolling normally.
errorThe last poll or file failed or was skipped; the error message says why. Polling continues.

After 5 consecutive failed polls the listener switches itself off and its error begins with [auto-disabled after N consecutive failures]. If a bucket is not found, the error lists the buckets the credential can see.

Fields

KeyType / defaultMeaning
namestringDefault: the folder or bucket name.
source_typelocal | s3 | gcs | azure_blob | adls_gen2; default localWhere to watch. local is the server folder.
folder_pathstringRequired for local. Set automatically for cloud sources. watch_path is an old alias.
bucket_namestringRequired for cloud sources.
prefixstringOptional object prefix.
credential_idstringRequired for cloud sources.
y_columnsstring[]Required for full_pipeline.
modedefault full_pipelineSee run modes.
base_session_idstringRequired for inference and retraining.
pipeline_configobjectPipeline settings; defaults output_number 10000 and classification task type. Also webhook_ids and max_price_dollars (Maximum price per run).
allowed_extensionscomma-separatedSee above.
poll_interval_secondsinteger, 60Seconds between polls.
max_file_size_mbinteger, 500Larger files are skipped.
cooldown_secondsinteger, 60Minimum gap between runs.
max_files_per_hourinteger, 100 means no cap.
enabledboolean, trueFalse creates it paused.
auto_sink_configobjectWrite each output to a connection on completion; needs "enabled": true. See Write output.

Pipeline settings use the shared keys described in the Configuration reference (for example text_mode 0 drop / 1 neural tokenization / 2 TF-IDF / 3 auto, mdh_mode 0 imputation / 1 2D-removal / 2 imputation/dropping, dsg_mode copula or gan, output_number). Every key the pipeline accepts is passed through. Inference and retraining take all of their settings from the base session.

API

MethodPathResult
GET/api/v1/listeners/folder{ success, listeners }
POST/api/v1/listeners/folder201 { success, listener }
PUT/api/v1/listeners/folder/{id}Only the keys you send change.
DELETE/api/v1/listeners/folder/{id}{ success, message }

Python SDK: list_listeners(), create_listener(name, source_type, folder_path=None, y_columns=None, bucket_name=, prefix=, credential_id=, pipeline_config=, ...), update_listener(listener_id, ...), delete_listener(listener_id).

Troubleshooting

SymptomFix
"folder_path must be under an allowed directory"Use a path inside uploads/, results/ or listener_data/.
"The selected credential is invalid"Re-save the connection with all its fields.
Target column missingA file with another schema arrived; use one listener per schema.
Files not picked upCheck extensions, size limit, cooldown and hourly cap.
Listener switched offFive failed polls in a row; fix the error and resume.

LM Readiness runs

A new-model run can be LM Readiness instead of the classic pipeline: choose it under Job at the top of the configuration pop-up. It uses the same settings as the LM Readiness tab and is stored as lm_readiness in pipeline_config.

For inference and retraining, the base session decides: an LM Readiness session is replayed as LM Readiness (inference), or prepared again with its own LM Readiness settings (retraining). Outlier removal and synthetic rows do not apply to it. Each run accepts up to 200,000 rows.