Browse documentation
2026-09-28AutoDataOpen in dashboard

Enterprise

What the Enterprise plan adds: data prepared for your target model family, per-column decisions you can review and pin, validation and as-of policies, declared dataset metadata, per-column lineage with an audit report, and the Enterprise API.

What Enterprise is (Enterprise)

Enterprise is a plan assigned to your account by an administrator. It is not a separate area of the product: with it, extra controls appear wherever you already configure a run (the Pipeline tab, Advanced Parameters, scheduled runs, folder and cloud listeners, triggers and SFTP inboxes), and what each run decided appears wherever you already inspect a run. The same settings are accepted by the REST API and the Python SDK. A plan change takes effect immediately, without signing out.

Enterprise sessions are billed at 2x the standard price. A run is priced on what is in its data, as on every plan, and the 2x applies to the whole price: the estimate, the exact price, the cost breakdown and the charge all include it. Comparing several model families in one run costs nothing extra.

For accounts without Enterprise, the Enterprise settings are removed from a run before it starts (the run itself is never refused), and the Enterprise read endpoints answer 403.

What Enterprise adds (Enterprise)

CapabilityWhat you get
Target model familyData prepared for the kind of model you will train. Choose up to four families to have the data prepared for each, scored, and one recommended.
Column decisionsBefore a run, the proposed encoding, imputation and scaling for every column, with the reason. Pin imputation or scaling per column, or drop a column instead of filling it; a mostly-empty column is proposed for dropping.
Pivot (as-of) columnDates measured from the moment each row came into existence, and handling for columns recorded after it.
Validation methodRandom, stratified, time-ordered or grouped train/test splits, and holdout, cross-validation or walk-forward evaluation.
Column selectionAn optional chain of selection methods on the fully prepared columns, including a leakage guard.
Dataset metadataA data dictionary you declare, or import from a data catalog: Snowflake Horizon, Databricks Unity Catalog, BigQuery, AWS Glue, OpenMetadata, DataHub, Atlan, Alation or Collibra. Declared facts take precedence over automatic detection at every stage and are replayed on inference and retraining.
Per-column lineage and audit reportFor every column, the method used, where it came from, the reason and the alternatives considered, plus a downloadable report for reviewers.
Detailed run report and previewPer-column detail in the run report and in the upload preview.
StreamingContinuous preparation of Kafka, Kinesis and MQTT records with a trained session.
How It WorksModule-by-module technical guides, readable only by Enterprise accounts.

Target model families (Enterprise)

Without Enterprise, the automatic imputation and scaling choices are made for a linear model. Choosing a target family makes them be made for the model you name.

FamilyAPI nameChoose it for
Linear / GLMlinearLinear and logistic regression, GLMs
Tree ensembles / GBDTtree_gbdtRandom forests and gradient-boosted trees
Feed-forward neural (MLP)mlpDense neural networks on tabular features
Sequence / text neuralsequence_neuralModels that read text as token sequences
  • One family: the data is prepared once, for that family.
  • Two to four families: the data is prepared once per family, each result is scored on the same targets, and one family is recommended with its reason and its margin over the runner-up. The recommended family's files become the run's outputs, so downloads, inference and retraining use it automatically. If one family fails, the others are still delivered.
  • No families sent: the run compares Linear and Tree ensembles. The Pipeline tab pre-selects these two; the automation forms pre-select one.
  • More than four families are cut to the first four; unknown names are ignored and listed in the comparison.

The family does not change settings you configure elsewhere, such as the scaling method, text tokenization, feature selection or the one-hot limit.

Reviewing column decisions (Enterprise)

When you validate a file, Enterprise accounts also receive a proposed plan: for every column, how it will be encoded, imputed and scaled, and why. Open Advanced Parameters, then Experimental, then Column decisions to review it and pin a different imputation or scaling method for any column. Nothing is charged until you start the run.

A pin that cannot be honoured is refused and recorded with the reason, never silently replaced. Examples: log1p scaling on a column with negative values, a whole-frame time-series imputation on a single column, or a method name that does not exist. The column then gets the automatic choice, and the refusal appears in the lineage and the audit report.

Imputation methods you can pin: Mean, Median, Mode, KNN, Iterative (MICE), Random Forest (MissForest), XGBoost and LightGBM Imputation. Scalers: standard, minmax, robust, maxabs, yeojohnson, quantile_normal, quantile_uniform, log1p and none. The exact names are returned by GET /api/v1/enterprise/families.

Preparation policy (Enterprise)

Four optional settings. Each is sent only when you configure it; a run that sets none of them prepares the data like a standard run.

Pivot (as-of) column

The column that records when each row came into existence, such as an application date. Other dates are then measured from it, and columns whose values fall after it (information a model could not have had at prediction time) are flagged, reduced to a time difference, or removed.

Key (pivot_config)ValuesDefault
columnA column namenone
leakage_actionwarn (keep and warn), relative (keep only the time difference), drop (remove)warn
relative_deltastrue or false: add a days-from-pivot column for each other dateon for a date pivot
auto_suggesttrue or false: when no column is given, let AutoData pick one and say whyfalse

The pivot is replayed on inference, retraining and streaming. POST /api/v1/pivot/suggest ranks the candidate columns with reasons before you choose.

Validation method

Key (validation_policy)ValuesDefault
split_strategyrandom, stratified, time_ordered (newest rows are the test set), grouped (a group stays on one side)random
eval_protocolholdout, kfold, walk_forwardholdout
order_column, group_columnColumn names for the time-ordered and grouped strategiesthe pivot column for ordering, if set
test_size0.05 to 0.50.25
eval_folds, walk_forward_splits2 to 5, 2 to 103, 3
eval_max_rows, eval_time_budget_s, random_stateBounds on evaluation cost, and the seed25000, 120, 42

A strategy the data cannot support falls back to a random split, with a warning and a recorded refusal.

Column selection

An ordered chain of methods applied to the fully prepared columns, each on what the previous one kept. Methods (selection_config.methods): leakage_guard, variance_threshold, correlation_pruning, mutual_information, permutation_importance, rfe. Through the API, auto runs a sensible default chain. Optional top_k keeps the best N columns per target, and leakage_action (warn or drop) decides what the leakage guard does with a suspected leak. Targets, excluded columns, text and media columns, and columns you declared as kept are never removed. With no methods, only the standard constant and near-duplicate column pruning happens.

Probe regularization

probe_config (penalty: l2, l1, elasticnet or none, with C and l1_ratio) is reserved and not applied to runs yet.

Dataset metadata (Enterprise)

Declare what your columns are: role (feature, target, identifier, timestamp, group key, metadata, free text, excluded, drop), type, semantic type, allowed range, ordinal order, allowed values, missing-value codes, unit, description, and optional imputation or scaling choices. Add dataset-level context such as a description and domain. Declared facts win over automatic detection for that column; undeclared columns are handled as before.

  • Rows are never dropped: a value that breaks a declared constraint becomes missing and is imputed. Only a column declared drop is removed.
  • Targets are never modified, except by a value map you declare.
  • Recode values: give a column's values your own codes with value_map, for example {"map": {"A": 0}, "other": 1} to turn a four-class target into A versus the rest. Values you do not list get the other code, or keep their own value when there is none. A target recoded into whole numbers or names is prepared for classification, and predictions come back in your codes. In the app, use Recode values in the Dataset metadata section: after an upload it lists each column's values with their counts. In a CSV dictionary, write A=0|*=1.
  • Provide it as JSON or a CSV data dictionary (up to 2 MB), in the Advanced tier of Advanced Parameters or as the dataset_metadata field.
  • Import from a data catalog: in the Dataset metadata section, Import from catalog reads a table's column descriptions, data types and key columns from Snowflake Horizon, Databricks Unity Catalog, BigQuery, AWS Glue, OpenMetadata, DataHub, Atlan, Alation or Collibra, through one of your saved connections. The result fills the editor for you to review before a run, and anything you already declared is kept. The import is a snapshot: import again to pick up later changes in the catalog. From the API, use POST /api/v1/dataset-metadata/catalog/import.
  • POST /api/v1/dataset-metadata/validate checks a dictionary against your columns before a run; GET /api/v1/enterprise/runs/<session_id>/dataset-metadata shows what a run was given and what it did.

Lineage, comparison and audit report (Enterprise)

Open the Lineage tab and expand a session to see the per-column decisions (method, source, reason and alternatives considered), the family comparison and links to the audit report as HTML or JSON. Sources read automatic, default, your choice or refused. The diff endpoint shows which decisions changed between two runs.

Currently the lineage panel, the audit report and the diff cover runs that compared two or more families. A single-family run records its decisions, but they are not yet shown there; its run report still shows per-column detail.

Enterprise API (Enterprise)

Start an Enterprise run with the ordinary endpoints (/process-data, /api/v1/process, the SDK's process() and every automation surface) by adding model_families, imputation_overrides, scaling_overrides, pivot_config, validation_policy, selection_config, probe_config or dataset_metadata. The read endpoints accept a session cookie or an API key, answer 401 when you are not signed in, 403 without Enterprise, and 404 for a session that is not yours.

EndpointReturns
GET /api/v1/enterprise/familiesThe families, the default pair, the per-run cap and the method names you can pin
GET /api/v1/enterprise/runs/<session_id>/columnsPer-column decisions for the recommended family, or ?family=
GET /api/v1/enterprise/runs/<session_id>/leaderboardHow the families compared and which is recommended
GET /api/v1/enterprise/lineage/<session_id>Everything recorded about a run, per family
GET /api/v1/enterprise/lineage/<session_id>/diff?against=<session_id>Decisions changed, added and removed between two runs
GET /api/v1/enterprise/audit-report/<session_id>The audit report, ?format=html (default) or json
GET /api/v1/enterprise/runs/<session_id>/dataset-metadataThe metadata a run was given and its report
POST /api/v1/dataset-metadata/validateA dictionary normalized, with warnings
POST /api/v1/pivot/suggestRanked as-of column candidates with reasons
GET /api/v1/enterprise/docs and /api/v1/enterprise/docs/<module>The How It Works guides

SDK: enterprise_families(), enterprise_columns(), enterprise_leaderboard(), enterprise_lineage(), enterprise_diff(), enterprise_audit_report() and pivot_suggest().

How It Works (Enterprise)

Enterprise accounts can read technical guides for every pipeline module in the dashboard: open Docs and choose How It Works, or follow a module's How it works (Enterprise) link. The same guides are available from GET /api/v1/enterprise/docs.