Information Schema Checks: Columns#
Note
The below checks require manifest.json and the dbt Information Schema (info_schema/v1/ in the dbt target directory) to be present. dbt 2.0 and later write the Information Schema when a command runs with --generate-info-schema. Add --static-analysis strict to include column types and column-level lineage. See Information Schema checks for details.
Column checks that use the dbt Information Schema (column types and column-level lineage).
Functions:
| Name | Description |
|---|---|
check_model_column_descriptions_propagated |
Columns that copy an upstream column must have a description when the upstream column has one. |
check_model_column_meta_propagated |
Columns derived from an upstream column that sets a |
check_model_column_types_match_inferred |
The declared |
check_model_columns_have_lineage |
Each column of a model with upstream dependencies must have column-level lineage. |
check_model_public_columns_not_derived_from_meta |
Public models must not expose columns derived from a column that sets a |
check_model_column_descriptions_propagated
#
Columns that copy an upstream column must have a description when the upstream column has one.
Rationale
A column that is passed through unchanged from an upstream model or source means the same thing downstream. When the upstream column is documented but the downstream column is not, the description is lost one step later in the DAG and users of the downstream model see an undocumented column. This check uses dbt's column-level lineage to find these columns, so you can copy the description or reference a shared doc block.
Note
This check requires the dbt Information Schema (dbt 2.0+, --generate-info-schema). Only copy lineage edges are followed: a column that transforms its input (mod) can mean something different and is not checked.
Receives at execution time:
| Name | Type | Description |
|---|---|---|
model |
ModelNode
|
The ModelNode object to check. |
Other Parameters (passed via config file):
| Name | Type | Description |
|---|---|---|
description |
str | None
|
Description of what the check does and why it is implemented. |
exclude |
str | list[str] | None
|
Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked. |
include |
str | list[str] | None
|
Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked. |
materialization |
Literal[ephemeral, incremental, table, view] | None
|
Limit check to models with the specified materialization. |
severity |
Literal[error, warn] | None
|
Severity level of the check. Default: |
Example(s):
Source code in src/dbt_bouncer/checks/info_schema/columns.py
check_model_column_meta_propagated
#
Columns derived from an upstream column that sets a meta key must set the same key.
Rationale
Column meta often classifies data, for example pii: true or contains_financial_data: true. When a classified column flows into a downstream model, the downstream column holds the same data, but nothing in dbt copies the classification. Masking policies, access reviews and data catalogs that read the meta key then miss the downstream column. This check uses dbt's column-level lineage to make sure the classification follows the data.
Note
This check requires the dbt Information Schema (dbt 2.0+, --generate-info-schema). It follows copy and mod lineage edges (the column value flows downstream), not scan edges (the column is only read, e.g. in a join). Column meta is read from manifest.json.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
meta_key
|
str
|
The |
required |
Receives at execution time:
| Name | Type | Description |
|---|---|---|
model |
ModelNode
|
The ModelNode object to check. |
Other Parameters (passed via config file):
| Name | Type | Description |
|---|---|---|
description |
str | None
|
Description of what the check does and why it is implemented. |
exclude |
str | list[str] | None
|
Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked. |
include |
str | list[str] | None
|
Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked. |
materialization |
Literal[ephemeral, incremental, table, view] | None
|
Limit check to models with the specified materialization. |
severity |
Literal[error, warn] | None
|
Severity level of the check. Default: |
Example(s):
Source code in src/dbt_bouncer/checks/info_schema/columns.py
check_model_column_types_match_inferred
#
The declared data_type of a column must match the type that the model SQL produces.
Rationale
The data_type declared in YAML is what consumers rely on, and with an enforced contract dbt casts the column to that type. When the SQL produces a different type, for example an integer where a double is declared, the declaration hides a modelling mistake or a silent cast. dbt's static analysis infers the type that the SQL produces, so this check can compare the two without running the model.
Note
This check requires the dbt Information Schema (dbt 2.0+, --generate-info-schema). Types are compared by family (boolean, date, decimal, float, integer, string, timestamp), so bigint and Int64 match while double and Int64 do not. Columns without both a declared and an inferred type, or with a type outside these families, are not checked.
Receives at execution time:
| Name | Type | Description |
|---|---|---|
model |
ModelNode
|
The ModelNode object to check. |
Other Parameters (passed via config file):
| Name | Type | Description |
|---|---|---|
description |
str | None
|
Description of what the check does and why it is implemented. |
exclude |
str | list[str] | None
|
Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked. |
include |
str | list[str] | None
|
Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked. |
materialization |
Literal[ephemeral, incremental, table, view] | None
|
Limit check to models with the specified materialization. |
severity |
Literal[error, warn] | None
|
Severity level of the check. Default: |
Example(s):
Source code in src/dbt_bouncer/checks/info_schema/columns.py
check_model_columns_have_lineage
#
Each column of a model with upstream dependencies must have column-level lineage.
Rationale
Column-level lineage powers impact analysis and the propagation checks in this category. A model that dbt's static analysis cannot analyse has no lineage at all, and a column declared in YAML that the SQL does not produce has no lineage either. This check reports both, so gaps in lineage are visible instead of silently weakening every check that depends on it.
Note
This check requires the dbt Information Schema (dbt 2.0+, --generate-info-schema). Models without upstream dependencies are not checked. Ephemeral models are not checked either: dbt inlines them, so their lineage is recorded on the models that select from them. Python models are not checked, because dbt's static analysis reads only SQL and cannot produce lineage for them. Columns built only from literals (e.g. 'web' as channel) have no upstream column: exclude them with exclude_column_name_pattern.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
exclude_column_name_pattern
|
str | None
|
Regex pattern to match column names that do not need lineage. |
None
|
Receives at execution time:
| Name | Type | Description |
|---|---|---|
model |
ModelNode
|
The ModelNode object to check. |
Other Parameters (passed via config file):
| Name | Type | Description |
|---|---|---|
description |
str | None
|
Description of what the check does and why it is implemented. |
exclude |
str | list[str] | None
|
Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked. |
include |
str | list[str] | None
|
Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked. |
materialization |
Literal[ephemeral, incremental, table, view] | None
|
Limit check to models with the specified materialization. |
severity |
Literal[error, warn] | None
|
Severity level of the check. Default: |
Example(s):
info_schema_checks:
- name: check_model_columns_have_lineage
exclude_column_name_pattern: ^_loaded_at$
Source code in src/dbt_bouncer/checks/info_schema/columns.py
check_model_public_columns_not_derived_from_meta
#
Public models must not expose columns derived from a column that sets a meta key.
Rationale
Public models are the interface that other teams and projects build on, so anything they expose spreads beyond the owning team. When a column classified with a meta key, for example pii: true, flows into a public model, possibly through several intermediate models, sensitive data leaves the team's control. This check walks dbt's column-level lineage upstream from every column of a public model and fails when any ancestor column sets the key.
Note
This check requires the dbt Information Schema (dbt 2.0+, --generate-info-schema). It follows copy and mod lineage edges through any number of models, not scan edges. Column meta is read from manifest.json. Models that are not public are not checked.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
meta_key
|
str
|
The |
required |
Receives at execution time:
| Name | Type | Description |
|---|---|---|
model |
ModelNode
|
The ModelNode object to check. |
Other Parameters (passed via config file):
| Name | Type | Description |
|---|---|---|
description |
str | None
|
Description of what the check does and why it is implemented. |
exclude |
str | list[str] | None
|
Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked. |
include |
str | list[str] | None
|
Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked. |
materialization |
Literal[ephemeral, incremental, table, view] | None
|
Limit check to models with the specified materialization. |
severity |
Literal[error, warn] | None
|
Severity level of the check. Default: |
Example(s):