Folder & Cloud Listeners
Watch an S3, GCS, Azure Blob or ADLS Gen2 location, or a server folder, and start a run for every new file.
What a listener does
A listener watches a storage location and starts a run for every new file that appears. It can watch an Amazon S3 bucket, a Google Cloud Storage bucket, an Azure Blob container, an Azure Data Lake Storage Gen2 filesystem, or (on self-hosted installs) a folder on the AutoData server. Each file becomes one run in the Runs tab.
Before you start
- Save a connection of the matching type (
s3,gcs,azure_blob,adls_gen2) on the Connections tab. The listener checks its secrets when you create it. - A server folder must be inside the server's
uploads/,results/orlistener_data/directory. - For inference or retraining, a completed session whose stored transforms are still available.
Creating a listener
- Open Listeners → Folder Listeners and click Add Folder Listener.
- Pick the source type: Amazon S3, Google Cloud Storage, Azure Blob, Azure Data Lake (Gen2) or Server Folder.
- Enter the bucket, container or filesystem, an optional prefix and the saved credential; or the folder path.
- Set allowed extensions and the poll interval.
- Choose what each file runs, and for a new model open Configure to set target columns and other settings.
- Click Create. Polling starts within a few seconds.
Run modes
| Option | mode | Each run | Needs |
|---|---|---|---|
| Prepare a new model | full_pipeline | Trains a new session with your pipeline configuration. | Target columns |
| Run inference | inference | Scores the new data with the stored transforms of a completed session. Output columns match that session. | base_session_id |
| Retrain a session | retraining | Refreshes a completed session with the new data, keeping its output schema. Options: remove outlier rows (run_dor), add synthetic rows (run_dsg), target rows (output_number, default 10000). | base_session_id |
If the base session's stored transforms have been removed, the listener records an error and skips the file rather than retrying it.
How files are picked up
- The location is listed every
poll_interval_seconds. A server folder is read without sub-folders; a cloud listener sees everything under its prefix. - Files with a disallowed extension or larger than
max_file_size_mbare skipped. - The last 500 file names are remembered, so each file is processed once.
- A new run starts only when the cooldown since the last one has passed and the hourly cap is not reached. Held-back files are picked up by a later poll.
- A new-model listener skips a file that lacks a target column and records why.
- Images or audio named by the file are collected from the same folder or prefix; upload them first. URL cells are fetched directly.
Accepted file types
The API default is every readable format: .csv, .feather, .json, .jsonl, .ndjson, .orc, .parquet, .xls, .xlsx. The dashboard form and the Python SDK pre-fill .csv,.parquet,.json; widen it if you drop other formats. .tsv, .tab and .txt are read when listed.
Price per run
Each run is charged the exact price of the data it actually processes, measured when the run starts. Nobody confirms the price first, so you can set an optional Maximum price per run ($) (max_price_dollars). When a listener's run is priced above it, the run is held instead of running and listed under Runs → Waiting for approval with its price and the limit. Approve starts it at that price, and it is charged exactly that; Discard drops it. You are notified by a job.held webhook event and, if you receive failure emails, by email. Leave the field empty for no limit. The limit applies to runs that prepare a new model. See Account and billing.
Status
| Status | Meaning |
|---|---|
idle | Created or paused. |
watching | Polling normally. |
error | The last poll or file failed or was skipped; the error message says why. Polling continues. |
After 5 consecutive failed polls the listener switches itself off and its error begins with [auto-disabled after N consecutive failures]. If a bucket is not found, the error lists the buckets the credential can see.
Fields
| Key | Type / default | Meaning |
|---|---|---|
name | string | Default: the folder or bucket name. |
source_type | local | s3 | gcs | azure_blob | adls_gen2; default local | Where to watch. local is the server folder. |
folder_path | string | Required for local. Set automatically for cloud sources. watch_path is an old alias. |
bucket_name | string | Required for cloud sources. |
prefix | string | Optional object prefix. |
credential_id | string | Required for cloud sources. |
y_columns | string[] | Required for full_pipeline. |
mode | default full_pipeline | See run modes. |
base_session_id | string | Required for inference and retraining. |
pipeline_config | object | Pipeline settings; defaults output_number 10000 and classification task type. Also webhook_ids and max_price_dollars (Maximum price per run). |
allowed_extensions | comma-separated | See above. |
poll_interval_seconds | integer, 60 | Seconds between polls. |
max_file_size_mb | integer, 500 | Larger files are skipped. |
cooldown_seconds | integer, 60 | Minimum gap between runs. |
max_files_per_hour | integer, 10 | 0 means no cap. |
enabled | boolean, true | False creates it paused. |
auto_sink_config | object | Write each output to a connection on completion; needs "enabled": true. See Write output. |
Pipeline settings use the shared keys described in the Configuration reference (for example text_mode 0 drop / 1 neural tokenization / 2 TF-IDF / 3 auto, mdh_mode 0 imputation / 1 2D-removal / 2 imputation/dropping, dsg_mode copula or gan, output_number). Every key the pipeline accepts is passed through. Inference and retraining take all of their settings from the base session.
API
| Method | Path | Result |
|---|---|---|
| GET | /api/v1/listeners/folder | { success, listeners } |
| POST | /api/v1/listeners/folder | 201 { success, listener } |
| PUT | /api/v1/listeners/folder/{id} | Only the keys you send change. |
| DELETE | /api/v1/listeners/folder/{id} | { success, message } |
Python SDK: list_listeners(), create_listener(name, source_type, folder_path=None, y_columns=None, bucket_name=, prefix=, credential_id=, pipeline_config=, ...), update_listener(listener_id, ...), delete_listener(listener_id).
Troubleshooting
| Symptom | Fix |
|---|---|
| "folder_path must be under an allowed directory" | Use a path inside uploads/, results/ or listener_data/. |
| "The selected credential is invalid" | Re-save the connection with all its fields. |
| Target column missing | A file with another schema arrived; use one listener per schema. |
| Files not picked up | Check extensions, size limit, cooldown and hourly cap. |
| Listener switched off | Five failed polls in a row; fix the error and resume. |
LM Readiness runs
A new-model run can be LM Readiness instead of the classic pipeline: choose it under Job at the top of the configuration pop-up. It uses the same settings as the LM Readiness tab and is stored as lm_readiness in pipeline_config.
For inference and retraining, the base session decides: an LM Readiness session is replayed as LM Readiness (inference), or prepared again with its own LM Readiness settings (retraining). Outlier removal and synthetic rows do not apply to it. Each run accepts up to 200,000 rows.