Index
dbt-bouncer runs checks against artifacts from dbt. Every check also has a unique rule code (e.g. MO021) that can be used in place of its name. These checks fall into four categories:
- Catalog checks:
- Information Schema checks (dbt 2.0+):
- Manifest checks:
- Run Results checks:
Information Schema checks#
dbt 2.0 and later can write the dbt Information Schema: a set of Parquet tables that describe the project, including column-level lineage, inferred column types and the grain of each model. The JSON artifacts do not contain this data. info_schema_checks read it from info_schema/v1/ in the dbt target directory, so dbt_artifacts_dir must point at a target directory that contains it.
To generate the Information Schema, add --generate-info-schema to a dbt command. Add --static-analysis strict to include column types and column-level lineage:
dbt-bouncer reads info_schema/v1/ and not the Parquet files under target/private/:
- dbt documents
info_schema/v1/as a contracted interface with a versioned directory. dbt-bouncer refuses an Information Schema whoseschema_versionit does not support. target/private/is undocumented. It also holds only the data of the last dbt command, so a later command such asdbt source freshnessreplaces the lineage thatdbt buildwrote.
If info_schema_checks are configured and the directory does not exist, for example with dbt 1.x artifacts, dbt-bouncer exits with an artifact error.
Known limitations, from dbt:
- Column-level lineage needs
--static-analysis strict. dbt skips a model whose SQL it cannot analyse, and records the lineage of an ephemeral model on the models that select from it. dbt.node_columns.metais empty, so checks that use columnmetaread it frommanifest.json.- Inferred column types use Apache Arrow names, for example
Int64(dbt-labs/dbt#16515).