SFTP Inbox
Upload files over SFTP and start a run for each one: set-up, accepted file types, run modes and API.
What the SFTP inbox is
A private upload folder for your account, reachable with any SFTP client or ETL tool. Each data file you upload starts a run when the upload finishes. The service must be enabled by your server administrator; the Listeners → SFTP Inbox tab shows the host, port (2222 by default) and whether it is enabled.
Set-up
- Click Generate SFTP Credentials. The username, password and an
sftp -P <port> <user>@<host>command are shown once. Save the password; it cannot be shown again. - Click the gear icon, choose what each upload runs and set the pipeline configuration, then Save & Enable Pipeline. The card shows Pipeline Active.
- Connect and upload.
sftp -P 2222 ad_1a2b3c4d_9f8e7d@your-server.com
sftp> put orders.csv
Which uploads start a run
Only .csv, .parquet, .json, .xlsx and .tsv. Other files are stored but start nothing, including .xls, .jsonl, .ndjson, .feather and .orc; use a folder or cloud listener for those.
- Uploads in sub-folders also start runs.
- Each upload starts its own run; listener cooldowns and caps do not apply.
- A new-model inbox skips files without the target columns; the reason appears on the SFTP Inbox row under Folder Listeners.
- Upload images or audio before the data file that names them, into the same folder.
Run modes
| Option | mode | Each run | Needs |
|---|---|---|---|
| Prepare a new model | full_pipeline | Trains a new session with your pipeline configuration. | Target columns |
| Run inference | inference | Scores the new data with the stored transforms of a completed session. Output columns match that session. | base_session_id |
| Retrain a session | retraining | Refreshes a completed session with the new data, keeping its output schema. Options: remove outlier rows (run_dor), add synthetic rows (run_dsg), target rows (output_number, default 10000). | base_session_id |
The inbox is enabled once it has target columns or a base session. All credentials on your account share one inbox and one configuration; revoking the last one removes the configuration.
Security
- Password login, confined to your inbox.
- 5 failed logins from an IP block it for 10 minutes (server defaults).
- A limited number of simultaneous connections (20 by default).
Pipeline settings use the shared keys described in the Configuration reference (for example text_mode 0 drop / 1 neural tokenization / 2 TF-IDF / 3 auto, mdh_mode 0 imputation / 1 2D-removal / 2 imputation/dropping, dsg_mode copula or gan, output_number). Every key the pipeline accepts is passed through. Inference and retraining take all of their settings from the base session.
API
| Method | Path | Body / result |
|---|---|---|
| GET | /api/v1/sftp/info | { enabled, host, port, instructions } |
| GET | /api/v1/sftp/credentials | Credentials with their inbox configuration. |
| POST | /api/v1/sftp/credentials | Optional y_columns, pipeline_config, mode, base_session_id. Returns sftp_password and connection_string once. |
| PUT | /api/v1/sftp/credentials/{id}/pipeline | y_columns, pipeline_config, mode, base_session_id. Replaces the whole configuration. |
| DELETE | /api/v1/sftp/credentials/{id} | Revokes the login. |
Python SDK: sftp_info(), list_sftp_credentials(), create_sftp_credential(name, y_columns=None, pipeline_config=None, mode="full_pipeline", base_session_id=None), update_sftp_pipeline(credential_id, ...), delete_sftp_credential(credential_id).
The "files processed" count on the card is not currently increased by SFTP uploads; the Runs tab lists what ran.
LM Readiness runs
A new-model run can be LM Readiness instead of the classic pipeline: choose it under Job at the top of the configuration pop-up. It uses the same settings as the LM Readiness tab and is stored as lm_readiness in pipeline_config.
For inference and retraining, the base session decides: an LM Readiness session is replayed as LM Readiness (inference), or prepared again with its own LM Readiness settings (retraining). Outlier removal and synthetic rows do not apply to it. Each run accepts up to 200,000 rows.