Skip to content

Index

dbt-bouncer runs checks against artifacts from dbt. Every check also has a unique rule code (e.g. MO021) that can be used in place of its name. These checks fall into four categories:

Information Schema checks#

dbt 2.0 and later can write the dbt Information Schema: a set of Parquet tables that describe the project, including column-level lineage, inferred column types and the grain of each model. The JSON artifacts do not contain this data. info_schema_checks read it from info_schema/v1/ in the dbt target directory, so dbt_artifacts_dir must point at a target directory that contains it.

To generate the Information Schema, add --generate-info-schema to a dbt command. Add --static-analysis strict to include column types and column-level lineage:

dbt build --static-analysis strict --generate-info-schema

dbt-bouncer reads info_schema/v1/ and not the Parquet files under target/private/:

  • dbt documents info_schema/v1/ as a contracted interface with a versioned directory. dbt-bouncer refuses an Information Schema whose schema_version it does not support.
  • target/private/ is undocumented. It also holds only the data of the last dbt command, so a later command such as dbt source freshness replaces the lineage that dbt build wrote.

If info_schema_checks are configured and the directory does not exist, for example with dbt 1.x artifacts, dbt-bouncer exits with an artifact error.

Known limitations, from dbt:

  • Column-level lineage needs --static-analysis strict. dbt skips a model whose SQL it cannot analyse, and records the lineage of an ephemeral model on the models that select from it.
  • dbt.node_columns.meta is empty, so checks that use column meta read it from manifest.json.
  • Inferred column types use Apache Arrow names, for example Int64 (dbt-labs/dbt#16515).