Skip to content

Information Schema Checks: Columns#

Note

The below checks require manifest.json and the dbt Information Schema (info_schema/v1/ in the dbt target directory) to be present. dbt 2.0 and later write the Information Schema when a command runs with --generate-info-schema. Add --static-analysis strict to include column types and column-level lineage. See Information Schema checks for details.

Column checks that use the dbt Information Schema (column types and column-level lineage).

Functions:

Name Description
check_model_column_descriptions_propagated

Columns that copy an upstream column must have a description when the upstream column has one.

check_model_column_meta_propagated

Columns derived from an upstream column that sets a meta key must set the same key.

check_model_column_types_match_inferred

The declared data_type of a column must match the type that the model SQL produces.

check_model_columns_have_lineage

Each column of a model with upstream dependencies must have column-level lineage.

check_model_public_columns_not_derived_from_meta

Public models must not expose columns derived from a column that sets a meta key.

check_model_column_descriptions_propagated #

Columns that copy an upstream column must have a description when the upstream column has one.

Rationale

A column that is passed through unchanged from an upstream model or source means the same thing downstream. When the upstream column is documented but the downstream column is not, the description is lost one step later in the DAG and users of the downstream model see an undocumented column. This check uses dbt's column-level lineage to find these columns, so you can copy the description or reference a shared doc block.

Note

This check requires the dbt Information Schema (dbt 2.0+, --generate-info-schema). Only copy lineage edges are followed: a column that transforms its input (mod) can mean something different and is not checked.

Receives at execution time:

Name Type Description
model ModelNode

The ModelNode object to check.

Other Parameters (passed via config file):

Name Type Description
description str | None

Description of what the check does and why it is implemented.

exclude str | list[str] | None

Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked.

include str | list[str] | None

Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked.

materialization Literal[ephemeral, incremental, table, view] | None

Limit check to models with the specified materialization.

severity Literal[error, warn] | None

Severity level of the check. Default: error.

Example(s):

info_schema_checks:
    - name: check_model_column_descriptions_propagated
      include: ^models/marts

Source code in src/dbt_bouncer/checks/info_schema/columns.py
@check(code="IS002")
def check_model_column_descriptions_propagated(model, ctx):
    """Columns that copy an upstream column must have a description when the upstream column has one.

    !!! info "Rationale"

        A column that is passed through unchanged from an upstream model or source means the same thing downstream. When the upstream column is documented but the downstream column is not, the description is lost one step later in the DAG and users of the downstream model see an undocumented column. This check uses dbt's column-level lineage to find these columns, so you can copy the description or reference a shared `doc` block.

    !!! note

        This check requires the dbt Information Schema (dbt 2.0+, `--generate-info-schema`). Only `copy` lineage edges are followed: a column that transforms its input (`mod`) can mean something different and is not checked.

    Receives:
        model (ModelNode): The ModelNode object to check.

    Other Parameters:
        description (str | None): Description of what the check does and why it is implemented.
        exclude (str | list[str] | None): Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked.
        include (str | list[str] | None): Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked.
        materialization (Literal["ephemeral", "incremental", "table", "view"] | None): Limit check to models with the specified materialization.
        severity (Literal["error", "warn"] | None): Severity level of the check. Default: `error`.

    Example(s):
        ```yaml
        info_schema_checks:
            - name: check_model_column_descriptions_propagated
              include: ^models/marts
        ```

    """
    info_schema: "InfoSchema" = ctx.info_schema
    columns = info_schema.node_columns.get(model.unique_id, {})
    undocumented: list[str] = []
    for key, edges in sorted(
        info_schema.lineage_by_child.get(model.unique_id, {}).items()
    ):
        if (columns.get(key, {}).get("description") or "").strip():
            continue
        for edge in edges:
            if edge.evolution != "copy":
                continue
            parent = info_schema.node_columns.get(edge.parent_node_unique_id, {}).get(
                edge.parent_column_name.casefold(), {}
            )
            if (parent.get("description") or "").strip():
                undocumented.append(
                    f"`{_column_name(info_schema, model.unique_id, key)}` (from `{edge.parent_node_unique_id}.{edge.parent_column_name}`)"
                )
                break
    if undocumented:
        fail(
            f"`{get_clean_model_name(model.unique_id)}` has undocumented columns whose upstream column is documented: {', '.join(undocumented)}."
        )

check_model_column_meta_propagated #

Columns derived from an upstream column that sets a meta key must set the same key.

Rationale

Column meta often classifies data, for example pii: true or contains_financial_data: true. When a classified column flows into a downstream model, the downstream column holds the same data, but nothing in dbt copies the classification. Masking policies, access reviews and data catalogs that read the meta key then miss the downstream column. This check uses dbt's column-level lineage to make sure the classification follows the data.

Note

This check requires the dbt Information Schema (dbt 2.0+, --generate-info-schema). It follows copy and mod lineage edges (the column value flows downstream), not scan edges (the column is only read, e.g. in a join). Column meta is read from manifest.json.

Parameters:

Name Type Description Default
meta_key str

The meta key to propagate, e.g. pii. A column is classified when the key has a truthy value.

required

Receives at execution time:

Name Type Description
model ModelNode

The ModelNode object to check.

Other Parameters (passed via config file):

Name Type Description
description str | None

Description of what the check does and why it is implemented.

exclude str | list[str] | None

Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked.

include str | list[str] | None

Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked.

materialization Literal[ephemeral, incremental, table, view] | None

Limit check to models with the specified materialization.

severity Literal[error, warn] | None

Severity level of the check. Default: error.

Example(s):

info_schema_checks:
    - name: check_model_column_meta_propagated
      meta_key: pii

Source code in src/dbt_bouncer/checks/info_schema/columns.py
@check(code="IS003")
def check_model_column_meta_propagated(model, ctx, *, meta_key: str):
    """Columns derived from an upstream column that sets a `meta` key must set the same key.

    !!! info "Rationale"

        Column `meta` often classifies data, for example `pii: true` or `contains_financial_data: true`. When a classified column flows into a downstream model, the downstream column holds the same data, but nothing in dbt copies the classification. Masking policies, access reviews and data catalogs that read the `meta` key then miss the downstream column. This check uses dbt's column-level lineage to make sure the classification follows the data.

    !!! note

        This check requires the dbt Information Schema (dbt 2.0+, `--generate-info-schema`). It follows `copy` and `mod` lineage edges (the column value flows downstream), not `scan` edges (the column is only read, e.g. in a join). Column `meta` is read from `manifest.json`.

    Parameters:
        meta_key (str): The `meta` key to propagate, e.g. `pii`. A column is classified when the key has a truthy value.

    Receives:
        model (ModelNode): The ModelNode object to check.

    Other Parameters:
        description (str | None): Description of what the check does and why it is implemented.
        exclude (str | list[str] | None): Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked.
        include (str | list[str] | None): Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked.
        materialization (Literal["ephemeral", "incremental", "table", "view"] | None): Limit check to models with the specified materialization.
        severity (Literal["error", "warn"] | None): Severity level of the check. Default: `error`.

    Example(s):
        ```yaml
        info_schema_checks:
            - name: check_model_column_meta_propagated
              meta_key: pii
        ```

    """
    info_schema: "InfoSchema" = ctx.info_schema
    missing: list[str] = []
    for key, edges in sorted(
        info_schema.lineage_by_child.get(model.unique_id, {}).items()
    ):
        column_name = _column_name(info_schema, model.unique_id, key)
        if _has_meta_key(ctx, model.unique_id, column_name, meta_key):
            continue
        for edge in _value_edges(edges):
            if _has_meta_key(
                ctx, edge.parent_node_unique_id, edge.parent_column_name, meta_key
            ):
                missing.append(
                    f"`{column_name}` (from `{edge.parent_node_unique_id}.{edge.parent_column_name}`)"
                )
                break
    if missing:
        fail(
            f"`{get_clean_model_name(model.unique_id)}` has columns derived from a column with `meta.{meta_key}` that do not set it: {', '.join(missing)}."
        )

check_model_column_types_match_inferred #

The declared data_type of a column must match the type that the model SQL produces.

Rationale

The data_type declared in YAML is what consumers rely on, and with an enforced contract dbt casts the column to that type. When the SQL produces a different type, for example an integer where a double is declared, the declaration hides a modelling mistake or a silent cast. dbt's static analysis infers the type that the SQL produces, so this check can compare the two without running the model.

Note

This check requires the dbt Information Schema (dbt 2.0+, --generate-info-schema). Types are compared by family (boolean, date, decimal, float, integer, string, timestamp), so bigint and Int64 match while double and Int64 do not. Columns without both a declared and an inferred type, or with a type outside these families, are not checked.

Receives at execution time:

Name Type Description
model ModelNode

The ModelNode object to check.

Other Parameters (passed via config file):

Name Type Description
description str | None

Description of what the check does and why it is implemented.

exclude str | list[str] | None

Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked.

include str | list[str] | None

Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked.

materialization Literal[ephemeral, incremental, table, view] | None

Limit check to models with the specified materialization.

severity Literal[error, warn] | None

Severity level of the check. Default: error.

Example(s):

info_schema_checks:
    - name: check_model_column_types_match_inferred

Source code in src/dbt_bouncer/checks/info_schema/columns.py
@check(code="IS004")
def check_model_column_types_match_inferred(model, ctx):
    """The declared `data_type` of a column must match the type that the model SQL produces.

    !!! info "Rationale"

        The `data_type` declared in YAML is what consumers rely on, and with an enforced contract dbt casts the column to that type. When the SQL produces a different type, for example an integer where a `double` is declared, the declaration hides a modelling mistake or a silent cast. dbt's static analysis infers the type that the SQL produces, so this check can compare the two without running the model.

    !!! note

        This check requires the dbt Information Schema (dbt 2.0+, `--generate-info-schema`). Types are compared by family (boolean, date, decimal, float, integer, string, timestamp), so `bigint` and `Int64` match while `double` and `Int64` do not. Columns without both a declared and an inferred type, or with a type outside these families, are not checked.

    Receives:
        model (ModelNode): The ModelNode object to check.

    Other Parameters:
        description (str | None): Description of what the check does and why it is implemented.
        exclude (str | list[str] | None): Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked.
        include (str | list[str] | None): Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked.
        materialization (Literal["ephemeral", "incremental", "table", "view"] | None): Limit check to models with the specified materialization.
        severity (Literal["error", "warn"] | None): Severity level of the check. Default: `error`.

    Example(s):
        ```yaml
        info_schema_checks:
            - name: check_model_column_types_match_inferred
        ```

    """
    info_schema: "InfoSchema" = ctx.info_schema
    mismatches: list[str] = []
    for _, row in sorted(info_schema.node_columns.get(model.unique_id, {}).items()):
        declared = row.get("data_type_declared")
        inferred = row.get("data_type_inferred")
        declared_family = _type_family(declared)
        inferred_family = _type_family(inferred)
        # Skip columns without both types, or with a type outside the known families.
        if declared_family and inferred_family and declared_family != inferred_family:
            mismatches.append(
                f"`{row['column_name']}` (declared `{declared}`, inferred `{inferred}`)"
            )
    if mismatches:
        fail(
            f"`{get_clean_model_name(model.unique_id)}` has columns whose declared type does not match the type its SQL produces: {', '.join(mismatches)}."
        )

check_model_columns_have_lineage #

Each column of a model with upstream dependencies must have column-level lineage.

Rationale

Column-level lineage powers impact analysis and the propagation checks in this category. A model that dbt's static analysis cannot analyse has no lineage at all, and a column declared in YAML that the SQL does not produce has no lineage either. This check reports both, so gaps in lineage are visible instead of silently weakening every check that depends on it.

Note

This check requires the dbt Information Schema (dbt 2.0+, --generate-info-schema). Models without upstream dependencies are not checked. Ephemeral models are not checked either: dbt inlines them, so their lineage is recorded on the models that select from them. Python models are not checked, because dbt's static analysis reads only SQL and cannot produce lineage for them. Columns built only from literals (e.g. 'web' as channel) have no upstream column: exclude them with exclude_column_name_pattern.

Parameters:

Name Type Description Default
exclude_column_name_pattern str | None

Regex pattern to match column names that do not need lineage.

None

Receives at execution time:

Name Type Description
model ModelNode

The ModelNode object to check.

Other Parameters (passed via config file):

Name Type Description
description str | None

Description of what the check does and why it is implemented.

exclude str | list[str] | None

Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked.

include str | list[str] | None

Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked.

materialization Literal[ephemeral, incremental, table, view] | None

Limit check to models with the specified materialization.

severity Literal[error, warn] | None

Severity level of the check. Default: error.

Example(s):

info_schema_checks:
    - name: check_model_columns_have_lineage
      exclude_column_name_pattern: ^_loaded_at$

Source code in src/dbt_bouncer/checks/info_schema/columns.py
@check(code="IS005")
def check_model_columns_have_lineage(
    model, ctx, *, exclude_column_name_pattern: RegexPattern | None = None
):
    """Each column of a model with upstream dependencies must have column-level lineage.

    !!! info "Rationale"

        Column-level lineage powers impact analysis and the propagation checks in this category. A model that dbt's static analysis cannot analyse has no lineage at all, and a column declared in YAML that the SQL does not produce has no lineage either. This check reports both, so gaps in lineage are visible instead of silently weakening every check that depends on it.

    !!! note

        This check requires the dbt Information Schema (dbt 2.0+, `--generate-info-schema`). Models without upstream dependencies are not checked. Ephemeral models are not checked either: dbt inlines them, so their lineage is recorded on the models that select from them. Python models are not checked, because dbt's static analysis reads only SQL and cannot produce lineage for them. Columns built only from literals (e.g. `'web' as channel`) have no upstream column: exclude them with `exclude_column_name_pattern`.

    Parameters:
        exclude_column_name_pattern (str | None): Regex pattern to match column names that do not need lineage.

    Receives:
        model (ModelNode): The ModelNode object to check.

    Other Parameters:
        description (str | None): Description of what the check does and why it is implemented.
        exclude (str | list[str] | None): Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked.
        include (str | list[str] | None): Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked.
        materialization (Literal["ephemeral", "incremental", "table", "view"] | None): Limit check to models with the specified materialization.
        severity (Literal["error", "warn"] | None): Severity level of the check. Default: `error`.

    Example(s):
        ```yaml
        info_schema_checks:
            - name: check_model_columns_have_lineage
              exclude_column_name_pattern: ^_loaded_at$
        ```

    """
    if not (model.depends_on and model.depends_on.nodes):
        return
    if model.config and model.config.materialized == "ephemeral":
        return
    # Static analysis reads only SQL, so Python models never have lineage.
    if model.language == "python":
        return
    info_schema: "InfoSchema" = ctx.info_schema
    model_name = get_clean_model_name(model.unique_id)
    lineage = info_schema.lineage_by_child.get(model.unique_id)
    if not lineage:
        fail(
            f"`{model_name}` has no column-level lineage. Run dbt with `--static-analysis strict` and check the dbt logs for static analysis errors on this model."
        )
    exclude = (
        re.compile(exclude_column_name_pattern.strip())
        if exclude_column_name_pattern
        else None
    )
    without_lineage = [
        row["column_name"]
        for key, row in sorted(
            info_schema.node_columns.get(model.unique_id, {}).items()
        )
        if key not in lineage and not (exclude and exclude.match(row["column_name"]))
    ]
    if without_lineage:
        fail(
            f"`{model_name}` has columns without column-level lineage: {', '.join(f'`{c}`' for c in without_lineage)}."
        )

check_model_public_columns_not_derived_from_meta #

Public models must not expose columns derived from a column that sets a meta key.

Rationale

Public models are the interface that other teams and projects build on, so anything they expose spreads beyond the owning team. When a column classified with a meta key, for example pii: true, flows into a public model, possibly through several intermediate models, sensitive data leaves the team's control. This check walks dbt's column-level lineage upstream from every column of a public model and fails when any ancestor column sets the key.

Note

This check requires the dbt Information Schema (dbt 2.0+, --generate-info-schema). It follows copy and mod lineage edges through any number of models, not scan edges. Column meta is read from manifest.json. Models that are not public are not checked.

Parameters:

Name Type Description Default
meta_key str

The meta key that marks a sensitive column, e.g. pii. A column is sensitive when the key has a truthy value.

required

Receives at execution time:

Name Type Description
model ModelNode

The ModelNode object to check.

Other Parameters (passed via config file):

Name Type Description
description str | None

Description of what the check does and why it is implemented.

exclude str | list[str] | None

Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked.

include str | list[str] | None

Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked.

materialization Literal[ephemeral, incremental, table, view] | None

Limit check to models with the specified materialization.

severity Literal[error, warn] | None

Severity level of the check. Default: error.

Example(s):

info_schema_checks:
    - name: check_model_public_columns_not_derived_from_meta
      meta_key: pii

Source code in src/dbt_bouncer/checks/info_schema/columns.py
@check(code="IS008")
def check_model_public_columns_not_derived_from_meta(model, ctx, *, meta_key: str):
    """Public models must not expose columns derived from a column that sets a `meta` key.

    !!! info "Rationale"

        Public models are the interface that other teams and projects build on, so anything they expose spreads beyond the owning team. When a column classified with a `meta` key, for example `pii: true`, flows into a public model, possibly through several intermediate models, sensitive data leaves the team's control. This check walks dbt's column-level lineage upstream from every column of a public model and fails when any ancestor column sets the key.

    !!! note

        This check requires the dbt Information Schema (dbt 2.0+, `--generate-info-schema`). It follows `copy` and `mod` lineage edges through any number of models, not `scan` edges. Column `meta` is read from `manifest.json`. Models that are not public are not checked.

    Parameters:
        meta_key (str): The `meta` key that marks a sensitive column, e.g. `pii`. A column is sensitive when the key has a truthy value.

    Receives:
        model (ModelNode): The ModelNode object to check.

    Other Parameters:
        description (str | None): Description of what the check does and why it is implemented.
        exclude (str | list[str] | None): Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked.
        include (str | list[str] | None): Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked.
        materialization (Literal["ephemeral", "incremental", "table", "view"] | None): Limit check to models with the specified materialization.
        severity (Literal["error", "warn"] | None): Severity level of the check. Default: `error`.

    Example(s):
        ```yaml
        info_schema_checks:
            - name: check_model_public_columns_not_derived_from_meta
              meta_key: pii
        ```

    """
    if model.access != "public":
        return
    info_schema: "InfoSchema" = ctx.info_schema
    exposed: list[str] = []
    for key in sorted(info_schema.node_columns.get(model.unique_id, {})):
        column_name = _column_name(info_schema, model.unique_id, key)
        # Breadth-first walk upstream; `seen` stops cycles and repeated paths.
        queue: deque[tuple[str, str]] = deque([(model.unique_id, key)])
        seen: set[tuple[str, str]] = set()
        while queue:
            unique_id, column_key = queue.popleft()
            if (unique_id, column_key) in seen:
                continue
            seen.add((unique_id, column_key))
            source_column = _column_name(info_schema, unique_id, column_key)
            if _has_meta_key(ctx, unique_id, source_column, meta_key):
                exposed.append(f"`{column_name}` (from `{unique_id}.{source_column}`)")
                break
            queue.extend(
                (e.parent_node_unique_id, e.parent_column_name.casefold())
                for e in _value_edges(
                    info_schema.lineage_by_child.get(unique_id, {}).get(column_key, [])
                )
            )
    if exposed:
        fail(
            f"Public model `{get_clean_model_name(model.unique_id)}` exposes columns derived from a column with `meta.{meta_key}`: {', '.join(exposed)}."
        )