Scheduled Runs
Run the pipeline on a connector table or query on an interval or cron schedule: set-up, statuses, run modes and API.
What scheduled runs do
A schedule reads a table or a read-only SQL query from a saved connection on an interval or cron schedule and runs the pipeline on the rows. Each execution appears in the Runs tab as Scheduled: <name>.
Creating a schedule
- On the Scheduled tab click New Schedule.
- Name it and pick Interval or Cron.
- Pick a saved connection (required) and a source table, or tick Use custom SQL query instead of table.
- Choose the run type and, for a new model, the pipeline configuration. Empty targets use the source's last column.
- Click Create. It starts active.
Schedules
| Type | Value | Notes |
|---|---|---|
interval | <n>m, <n>h or <n>d; default 1h. Presets: 15m, 30m, 1h, 6h, 12h, 1d, 7d. | Other formats are saved but never fire. |
cron | 5-field expression, e.g. 0 2 * * * | Validated on save. Times are UTC. |
Due schedules are checked every 30 seconds.
Run modes
| Option | mode | Each run | Needs |
|---|---|---|---|
| Prepare a new model | full_pipeline | Trains a new session with your pipeline configuration. | Target columns |
| Run inference | inference | Scores the new data with the stored transforms of a completed session. Output columns match that session. | base_session_id |
| Retrain a session | retraining | Refreshes a completed session with the new data, keeping its output schema. Options: remove outlier rows (run_dor), add synthetic rows (run_dsg), target rows (output_number, default 10000). | base_session_id |
Each execution
- Reads the whole table or query (up to the server's connector row limit, 1,000,000 by default). Watermarks are not applied to schedules.
- No rows: status
skipped, nothing runs. - While a run is in progress the schedule does not fire again; that tick is skipped, not queued.
- A run still going after an hour is marked failed with "Timed out (>1h)".
last_run_status | Meaning |
|---|---|
running | In progress. |
completed | Finished. |
failed | Read or run failed, or timed out; see last_error. |
skipped | No rows returned. |
held | The price was above the Maximum price per run; the run is waiting under Runs → Waiting for approval. |
Actions: view last run results, pause or resume, Run Now (works while paused), edit and delete.
Fields
| Key | Default | Meaning |
|---|---|---|
name | required | |
description | ||
schedule_type | interval | |
schedule_value | 1h | |
credential_id | required in practice | Without it the run fails. |
connector_type | the credential's type | |
source_table | ||
sql_query | Single read-only statement; used instead of the table. | |
mode | full_pipeline | |
base_session_id | Required for replay modes. | |
pipeline_config | {} | Pipeline settings; output_number 10000; a mode key here is the speed preset. Also webhook_ids, auto_sink (written when each run completes, in every run mode; see Write Output & Auto-Sink) and max_price_dollars (Maximum price per run). |
is_active | true |
Pipeline settings use the shared keys described in the Configuration reference (for example text_mode 0 drop / 1 neural tokenization / 2 TF-IDF / 3 auto, mdh_mode 0 imputation / 1 2D-removal / 2 imputation/dropping, dsg_mode copula or gan, output_number). Every key the pipeline accepts is passed through. Inference and retraining take all of their settings from the base session.
Price per run
Each run is charged the exact price of the data it actually processes, measured when the run starts. Nobody confirms the price first, so you can set an optional Maximum price per run ($) (max_price_dollars). When a scheduled run is priced above it, the run is held instead of running and listed under Runs → Waiting for approval with its price and the limit. Approve starts it at that price, and it is charged exactly that; Discard drops it. You are notified by a job.held webhook event and, if you receive failure emails, by email. Leave the field empty for no limit. The limit applies to runs that prepare a new model. See Account and billing.
API
| Method | Path | Result |
|---|---|---|
| GET | /api/v1/scheduled-runs | { scheduled_runs } |
| POST | /api/v1/scheduled-runs | 201 { success, scheduled_run } |
| PUT | /api/v1/scheduled-runs/{id} | Update any field. |
| DELETE | /api/v1/scheduled-runs/{id} | Delete. |
| POST | /api/v1/scheduled-runs/{id}/toggle | { is_active, next_run_at } |
| POST | /api/v1/scheduled-runs/{id}/run-now | { success, session_id, status } |
Python SDK: list_scheduled_runs(), create_scheduled_run(name, connector_type, source_table, y_columns, schedule_type="interval", schedule_value="1h", credential_id=None, output_rows=10000, sql_query=None, pipeline_config=None, description=None, mode="full_pipeline", base_session_id=None), update_scheduled_run(run_id, schedule_type=, schedule_value=, mode=, base_session_id=, pipeline_config=), toggle_scheduled_run(run_id), run_scheduled_now(run_id), delete_scheduled_run(run_id).
LM Readiness runs
A new-model run can be LM Readiness instead of the classic pipeline: choose it under Job at the top of the configuration pop-up. It uses the same settings as the LM Readiness tab and is stored as lm_readiness in pipeline_config.
For inference and retraining, the base session decides: an LM Readiness session is replayed as LM Readiness (inference), or prepared again with its own LM Readiness settings (retraining). Outlier removal and synthetic rows do not apply to it. Each run accepts up to 200,000 rows.