Browse documentation
2026-09-28AutoDataOpen in dashboard

Connector Data Sources

Read data directly from 38 databases, warehouses, storage services, streams and SaaS apps — connect, test, discover, preview and run.

Connectors let AutoData read a dataset directly from where it lives — a database, a data warehouse, cloud storage, a stream, a SaaS application or a REST endpoint — instead of you exporting and uploading a file. The rows are then prepared by exactly the same pipeline as an upload.

Supported sources (38)

CategoryConnectors
DatabasesSQL Database (PostgreSQL, MySQL/MariaDB, SQL Server, SQLite), MongoDB, Oracle, Apache Cassandra, Amazon DynamoDB, Elasticsearch / OpenSearch
Data warehousesSnowflake, Google BigQuery, Amazon Redshift, ClickHouse, Databricks, Microsoft Fabric, Azure Synapse, SAP HANA
Storage and filesAmazon S3, Google Cloud Storage, Azure Blob Storage, Azure Data Lake Storage Gen2, Delta Lake, Google Sheets, Google Drive, Dropbox, OneDrive / SharePoint Files, Box
Streaming, IoT and historiansApache Kafka, Amazon Kinesis, MQTT, OPC-UA, InfluxDB, TimescaleDB, PI Web API
SaaS and appsSalesforce, HubSpot, Stripe, Google Analytics 4, NetSuite, SharePoint Lists, REST / HTTP API

The fields each one needs are listed in the Connector Reference.

Connecting in five steps

  1. Choose the connector. On the dashboard’s Pipeline tab set Source to Connector and pick a type (grouped by category, with search).
  2. Authenticate. Pick a saved connection of that type, or use Quick Connect to type the fields for this run. Quick Connect fields can be saved as a connection.
  3. Test. A lightweight check (for example SELECT 1) confirms the credentials and network path without reading data.
  4. Discover. Lists what the connection can read — tables, collections, files, topics, indices, worksheets, objects. For the SQL Database connector you can write a query instead.
  5. Pick and preview. Selecting an item returns its column names and, where cheap, a row count. Choose your target columns and settings, then start the run.

The dashboard needs at least 5 rows in the preview. Streaming and IoT connectors report partitions, shards or sampled messages in that count, so small topics or streams may be refused in the dashboard; start those runs through the API.

Saved connections

The Connections tab stores credentials once and reuses them in pipeline runs, inference and retraining, scheduled runs, triggers and incremental sync, folder listeners (S3, GCS, Azure Blob, ADLS Gen2), write-back and streaming. You can add, test, edit (blank fields keep the stored value, and a stored field can be cleared) and delete connections there.

  • Secrets are encrypted at rest and are never returned to the browser or API after saving; listings show only non-secret hints such as host or bucket.
  • Connections belong to one user account and are not visible to teammates or administrators.
  • Passwords, tokens and keys are masked in every error message.
  • Use a read-only database user or least-privilege key: connectors only read, except when you write results back.

Sign-in methods

Besides a user and password or an access key, connections accept the sign-in methods platform administrators usually hand out. Where a platform has several, the connection form has a Sign-in method choice. Certificates and private keys are pasted as PEM text.

PlatformMethods
SnowflakePassword, key pair, or token; an optional role
Amazon S3, DynamoDB, Kinesis, Delta Lake on S3Access keys, temporary keys with a session token, or an assumed role with an optional external id; a custom endpoint for S3-compatible stores
Azure Blob Storage, ADLS Gen2Connection string, SAS token, account key, or service principal
DatabricksPersonal access token or service principal
Kafka, MQTTA custom CA certificate and a client certificate, for batch runs and for streams
SQL Database (PostgreSQL), TimescaleDB, RedshiftSSL server and client certificates
REST / HTTP APIAPI key in a header or a query parameter, bearer token, OAuth2 client credentials, or basic authentication
Elasticsearch / OpenSearchURL or Elastic Cloud ID; API key, bearer token, or user and password; a custom CA certificate
OPC-UAAnonymous, user and password, and a client certificate

The keys for each method are listed in the Connector Reference.

Custom queries

SQL-family connectors accept a read-only query instead of a table: SQL Database (in the dashboard and API), and through the API Snowflake, BigQuery, Databricks, Microsoft Fabric, Azure Synapse, Oracle, Redshift, SAP HANA, ClickHouse and TimescaleDB. Salesforce accepts SOQL, NetSuite SuiteQL and S3 an S3 Select expression on a CSV object. A query must be a single SELECT or WITH statement; statements with write or DDL keywords (including the word REPLACE) or a second statement are rejected.

Incremental sync

Database and warehouse connectors (SQL Database, Snowflake, BigQuery, Databricks, Fabric, Synapse, Oracle, Redshift, SAP HANA, ClickHouse, TimescaleDB, Cassandra, DynamoDB, Elasticsearch) can read only rows whose watermark column — for example updated_at or an increasing id — is greater than the value seen last time. Send incremental: true and watermark_column with the run. The new watermark is stored only when the run completes, so a failed run re-reads the same rows. Watermarks are kept per connection (or per source_key for Quick Connect sources) and are visible on the Sync tab; a watermark_advance trigger can start a run whenever new rows arrive. Incremental reads do not apply to custom queries.

Starting a run through the API

POST /api/v1/process-connector
Authorization: Bearer <API key>
Content-Type: application/json

{
  "source_type": "snowflake",
  "credential_id": "<saved connection id>",
  "table": "SALES.ORDERS",
  "y_columns": ["CHURNED"],
  "incremental": true,
  "watermark_column": "UPDATED_AT",
  "type_casts": {"CUSTOMER_AGE": "int"}
}
  • source_type — connector type id; credential_id or inline_secrets; table or custom_query; y_columns (not needed for inference/retraining).
  • run_mode — full_pipeline (default), inference or retraining with base_session_id.
  • type_casts — column → int, float, string, bool, datetime or category, applied before the pipeline; unconvertible values become empty.
  • source_key — names the watermark of a Quick Connect source.
  • All pipeline settings of /api/v1/process are accepted as well.

Test, discover and preview are available as POST /api/v1/connectors/<type>/test, /discover and /preview; saved connections are managed with /api/v1/credentials. The Python SDK wraps all of these.

Writing results back

After a run, outputs can be written to any connector type except Google Analytics 4: databases and warehouses, file stores and cloud drives (as CSV, Parquet, Excel, JSON, JSON Lines, Feather or ORC), Google Sheets, search indexes, topics, streams, HTTP endpoints, and business and plant systems. Destinations offer replace, append, fail and merge modes as they apply, manually or automatically when a run completes. Salesforce, HubSpot, NetSuite, SharePoint lists, Stripe, PI Web API and OPC UA create or change live records, so each of their connections has an Allow writing switch that must be on before it can be a destination. See Write Output & Auto-Sink.

Streaming (Enterprise)

Batch connectors read a bounded batch per run. For continuous, per-message scoring, Kafka, Kinesis and MQTT connections can be the source and sink of a streaming inference job.

Limits

  • Up to 1,000,000 rows per read.
  • File connectors read one file per run in CSV, TSV/TXT, Parquet, JSON, JSON Lines, Excel, Feather or ORC.
  • Nested JSON (MongoDB, Elasticsearch, REST, JSON inside file cells) is expanded into dotted columns when the run starts, so the column list can grow after the preview.
  • The source must be reachable from AutoData’s servers; allow AutoData in your firewall or IP access list.

Known issues (2026-09-28)

  • MongoDB: runs that read from MongoDB currently fail; test, discover, preview and writing to MongoDB work.
  • S3, Azure Blob and ADLS Gen2: with a prefix / path set on the connection, a file picked from Discover fails at run time. Leave the prefix empty on connections you read from.

Connection errors

When a test, discover, preview or write cannot proceed and the cause is recognised, the answer carries a category — auth, permission, not_found, tls, timeout or network — and a plain-language hint beside the source's own message.

Unattended runs watch for rejected credentials. When a scheduled run, listener, trigger or automatic write is rejected by the source, the saved connection is marked as failing with the reason, and owners who have opted into failure e-mails are notified. A schedule, listener, trigger or stream that keeps being rejected is switched off, with the reason shown on it.

Troubleshooting

MessageWhat to do
Credential validation failed — … is requiredA required field for that connector is missing; the message names it.
Credential not foundThe connection id is wrong or belongs to another account.
Connection refused / timed outHost, port or firewall. Allow AutoData’s servers to reach the source.
Authentication failedCheck the user, password, key or token, and that it has read permission.
A connection shows Failed with “Rejected during …”An unattended run had its credentials rejected by the source. Update the connection and test it; switch the schedule, listener, trigger or stream back on if it was switched off.
custom_query contains a disallowed keywordRewrite the query as a plain SELECT.
Connector returned no dataThe table is empty, or nothing is newer than the watermark.
… not installed on serverSelf-hosted installs only: install the connector’s driver package.