Skip to content

Frequently Asked Questions#

Can other tools perform the same checks as dbt-bouncer?#

dbt-bouncer is designed to be the most complete and flexible way to enforce conventions in a dbt project: it ships 120+ built-in checks, is configured via a YAML or TOML file format that dbt developers are already familiar with, never needs a database connection, can run checks against any of dbt's artifacts (manifest, catalog, and run results), can run both locally and in a CI pipeline, and lets you write custom checks in python.

dbt-bouncer dbt-checkpoint dbt-project-evaluator dbt-score
Built-in checks 120+ Collection of pre-commit hooks Fixed set of SQL/Jinja tests Fixed set of model-only checks
Configuration YAML or TOML config file .pre-commit-config.yaml dbt_project.yml + seed files (CSV) pyproject.toml
Checks written in Python Python SQL + Jinja Python
Requires a database connection ❌ ❌ ✅ ❌
Checks any dbt artifact (manifest, catalog, run results) ✅ Manifest and catalog ❌ (queries the database) Manifest only
Where it runs Locally and in a CI pipeline Only as part of pre-commit/prek Locally and in a CI pipeline Locally, from the command line

Tip

dbt-bouncer can perform all the checks currently included in dbt-checkpoint, dbt-project-evaluator and dbt-score. If you see an existing check that is not possible with dbt-bouncer, open an issue and we'll add it!

How do I onboard an existing dbt project?#

An existing project usually fails many checks on the first run. Do not try to fix every failure first. Adopt the checks with a baseline, then reduce the backlog over time. A baseline records the current failures, so the run reports only failures that are new.

Follow these steps:

  1. Create a starter config. Run dbt-bouncer init, or start from a preset with dbt-bouncer init --preset standard.
  2. Generate the dbt artifacts. Run dbt parse for manifest checks, or dbt build when you also use catalog or run-results checks.
  3. Record the current failures as a baseline:

    dbt-bouncer baseline --output-file .dbt-bouncer-baseline.json
    
  4. Commit the config file and the baseline file.

  5. Run dbt-bouncer with the baseline so only new failures fail the build:

    dbt-bouncer run --baseline .dbt-bouncer-baseline.json
    

Now the check run is green, and any new violation fails the build. To reduce the backlog, fix some failures and regenerate the baseline with the same command. The baseline shrinks as the project improves.

Two more options help during onboarding:

  • To compare against a previous run without a stored file, use --state <dir> in place of --baseline. It points at a directory of dbt artifacts from a previous run, the same way dbt's own --state flag does. See the CLI reference.
  • To warn instead of fail while the team adapts, set severity: warn at the top of the config file. Change it to error once the team is ready.

Does dbt-bouncer work with dbt Cloud?#

Yes! As dbt-bouncer runs on the artifacts generated by dbt, it can be used with dbt Cloud as long as the artifacts generated by the CI job in dbt Cloud are available.

For GitHub this can be achieved using the pgoslatara/dbt-cloud-download-artifacts-action action:

name: CI pipeline

on:
  pull_request:
      branches:
          - main

jobs:
    download-artifacts:
        runs-on: ubuntu-latest
        permissions:
            pull-requests: write
        steps:
          - name: Checkout
            uses: actions/checkout@v6

          - name: Download dbt artifacts
            uses: pgoslatara/dbt-cloud-download-artifacts-action@v1
            with:
              commit-sha: ${{ github.event.pull_request.head.sha }}
              dbt-cloud-api-token: ${{ secrets.DBT_CLOUD_API_TOKEN }}

          - name: Run dbt-bouncer
            uses: godatadriven/dbt-bouncer@vX.X

Warning

dbt Cloud now supports a "versionless" option, which allows dbt projects to be run with the latest version of dbt. One effect of choosing this option is that dbt artifacts may receive non-breaking changes (source), these may or may not be compatible with dbt-bouncer. If you encounter a bug as a result of this, please open an issue and we'll investigate.

Does dbt-bouncer work with dbt 2.0 / Fusion?#

Yes. dbt 2.0 (the Rust-based Fusion engine) emits the same manifest.json, catalog.json and run_results.json artifacts — at manifest schema v12, catalog schema v1 and run results schema v6 — as dbt-core 1.x. Because dbt-bouncer consumes those artifacts directly, every check that works against dbt-core 1.x also works against dbt 2.0. Our CI builds the test project with dbt 2.0 and runs dbt-bouncer against the result on every pull request.

Three details changed at dbt 2.0. Plan for them:

  • dbt 2.0 ships as the dbt and dbt-oss distributions. The dbt-core distribution stays on the 1.x line.
  • --write-catalog moved off dbt build. Use dbt compile --write-catalog or dbt docs generate --write-catalog.
  • The Apache-2.0 dbt-oss distribution writes no catalog.json. Catalog checks need the dbt distribution.

dbt 2.0 also writes Parquet artifacts under target/private/. These back the dbt docs v2 site. dbt-bouncer reads the JSON artifacts and ignores the Parquet directory.

Does dbt-bouncer support Python models?#

Yes! dbt-bouncer fully supports Python models (introduced in dbt 1.3). All checks that work with SQL models also work with Python models, this means dbt-bouncer can enforce conventions on your Python models just as it does for SQL models.

Example#

A Python model can be validated with the same checks as SQL models:

manifest_checks:
  - name: check_model_has_meta_keys
    keys:
      - maturity
  - name: check_model_description_populated
  - name: check_model_code_does_not_contain_regexp_pattern
    regexp_pattern: .*\.to_pandas\(\)  # Discourage memory-intensive operations

The check check_model_code_does_not_contain_regexp_pattern will match against the Python code in raw_code, allowing you to enforce conventions like avoiding certain libraries or patterns in your Python models.

Tip

You can use regex patterns to enforce Python-specific conventions, such as:

  • Preventing certain imports: regexp_pattern: .*import\s+os.*
  • Discouraging memory-intensive operations: regexp_pattern: .*\.to_pandas\(\).*
  • Enforcing use of specific libraries: Use include/exclude patterns to target only Python models

Can AI coding agents use dbt-bouncer?#

Yes! dbt-bouncer ships a Model Context Protocol server. An AI coding agent can read which conventions your project enforces before it generates dbt code, and can run the checks to verify its work. Install the optional dependency with pip install 'dbt-bouncer[mcp]' and start the server with dbt-bouncer mcp. See the CLI documentation for the available tools and an example client configuration.

How to configure dbt-bouncer for use in a CI pipeline?#

dbt-bouncer is designed to be use primarily in a CI pipeline such as GitHub Actions or Azure DevOps. To do this we create a config file such as:

catalog_checks:
  - name: check_column_description_populated
    include: ^models/marts

manifest_checks:
  - name: check_model_directories
    include: ^models
    permitted_sub_directories:
      - intermediate
      - marts
      - staging
      - utilities

run_results_checks:
  - name: check_run_results_max_execution_time
    max_execution_time_seconds: 10

The goal of a CI pipeline is to test the changes in a pull request but also to provide feedback to the developer as quickly as possible without incurring unnecessary costs (time, financial, compute, etc.). To achieve this we can combine several features of dbt and dbt-bouncer:

  1. By running dbt parse, dbt can generate a manifest.json without a database connection. We can then run our manifest checks via:

    dbt-bouncer --only manifest_checks
    
  2. dbt requires models to be materialised before it can generate a catalog.json file. By running dbt run --empty we can materialise every model without processing any data. Once these materializations are performed we can run our catalog checks via:

    dbt-bouncer --only catalog_checks
    
  3. Typically a CI pipeline will run a dbt build command with flags such as --state and/or --defer. After this command has completed we can run our run results checks via:

    dbt-bouncer --only run_results_checks
    

Additionally, you can use the --check flag to run only specific checks by name. This is useful for debugging or validating a single convention:

  ```shell
  dbt-bouncer --check check_model_has_unique_test
  ```

Multiple checks can be specified as a comma-separated list:

  ```shell
  dbt-bouncer --check check_model_has_unique_test,check_model_description_populated
  ```

The --check and --only flags can be combined: --only restricts to the specified categories, then --check further narrows to only the named checks within those categories.

By using this approach, and combining with your own unique constraints and desires, dbt-bouncer can be used efficiently as part of your CI pipeline.

How do I get machine-readable results to parse in a script?#

Do not parse the console results table. That table is for humans. It truncates check names to fit the terminal width.

Set --output-file and select a structured --output-format. The supported formats are csv, json, junit, sarif, and tap.

dbt-bouncer run --output-file results/check-results.json --output-format json

Then read the output file in your script. See the CLI reference for full option details.

How to set up dbt-bouncer in a monorepo?#

A monorepo may consist of one directory with a dbt project and other directories with unrelated code. It may be desired for dbt-bouncer to be configured from the root directory. Sample directory tree:

.
├── dbt-bouncer.yml
├── README.md
├── dbt-project
│   ├── models
│   ├── dbt_project.yml
│   └── profiles.yml
└── package-a
    ├── src
    ├── tests
    └── package.json

To ease configuration you can use exclude or include at the global level (see Config File for more details). For the above example dbt-bouncer.yml could be configured as:

dbt_artifacts_dir: dbt-project/target
include: ^dbt-project

manifest_checks:
    - name: check_exposure_based_on_non_public_models

dbt-bouncer can now be run from the root directory.

How to set up dbt-bouncer in a dbt Mesh?#

A dbt Mesh is a collection of dbt projects in an organization, some of which can read models from other dbt projects. Natively supported by dbt Cloud, a dbt Mesh can also be set up with dbt Core using a plugin such as dbt-loom.

One challenge in a dbt Mesh is the large number of developers working across multiple dbt projects leading to differing conventions being implemented. There are multiple approaches to using dbt-bouncer in a dbt Mesh, two are outlined below.

Approach 1: Individual dbt-bouncer.yml configuration file#

Each dbt project can have its own dbt-bouncer.yml configuration file. This allows each project to adopt and implement its own conventions in addition to any conventions to be shared across all dbt projects. Should a breaking change be required to the config file then each dbt project can be updated independently at a time that makes sense.

This is the recommended approach due to its simplicity and ability to update each dbt project independently.

Approach 2: Centralised dbt-bouncer.yml configuration file shared via git submodule#

Warning

With this approach, a change to the centralised dbt-bouncer.yml file may result in CI pipelines in dbt projects failing despite no changes being made to these projects. As such we recommend implementing this approach only after extensive discussion with all dbt project developers so that all dbt projects can be brought into line before dbt-bouncer is enforced in the CI pipeline.

Should it be necessary for a breaking change to be made to the centralised dbt-bouncer.yml configuration file, we recommend setting the severity of the relevant check to warn so that CI pipelines in dbt projects will not fail and maintainers have sufficient time to make the necessary changes.

Git submodules allow the contents from one repository to be accessible from a different repository. Such a setup for dbt-bouncer can be achieved as follows (this example uses GitHub, similar setups can be achieved with other providers):

  1. Set up a dedicated repository to store a centralised dbt-bouncer.yml configuration file that will be used by all dbt projects. Let's call this repository dbt-bouncer-config.

  2. The contents of the dbt-bouncer.yml file in dbt-bouncer-config should contain the following configuration for dbt_artifacts_dir:

    dbt_artifacts_dir: ../target
    
    manifest_checks:
      - name: check_model_directories
        include: ^models
        permitted_sub_directories:
          - intermediate
          - marts
          - staging
          - utilities
      ...
    
  3. In every repository add a git submodule via:

    git submodule add git@github.com:<YOUR_ORG>/dbt-bouncer-config.git
    
  4. Run dbt-bouncer:

    dbt-bouncer --config-file dbt-bouncer-config/dbt-bouncer.yml
    

Your directory tree should look like this:

.
├── dbt-bouncer-config
│   └── dbt-bouncer.yml
├── dbt_project.yml
├── macros
│   └── ...
├── models
│   └── ...
├── profiles.yml
├── README.md
└── target
    ├── catalog.json
    ├── manifest.json
    └── run_results.json

Note: if you update your central dbt-bouncer.yml file, you will need to run git submodule update --remote in every repository to update the submodule.

How to set up dbt-bouncer with prek/pre-commit?#

You can use the official pre-commit hook, in your .pre-commit-config.yaml file:

repos:
  - repo: https://github.com/godatadriven/dbt-bouncer
    rev: v4.0.0 # Check https://github.com/godatadriven/dbt-bouncer/releases for latest version
    hooks:
      - id: dbt-bouncer
        args: ["--config-file", "<PATH_TO_CONFIG_FILE>"] # Optional

Alternatively, you can use a local hook to run automatically run dbt-bouncer before your commits get added to the git tree.

- repo: local
  hooks:
    - id: dbt-bouncer
      name: dbt-bouncer
      entry: dbt-bouncer # --config-file <PATH_TO_CONFIG_FILE>
      language: system
      pass_filenames: false
      always_run: true

Can I skip specific checks for an exposure/model/source/etc.?#

Yes! Many dbt objects permit adding a meta config field (docs), this can be used to skip checks for the object. For example, for a model:

models:
  - name: my_model
    config:
      meta:
        dbt-bouncer:
          skip_checks:
            - check_model_description_populated
            - check_model_has_meta_keys
          reason: We recommend documenting why these checks are being skipped.

And for a source:

version: 2

sources:
  - name: source_system
    tables:
      - name: source_1
        config:
          meta:
            dbt-bouncer:
              skip_checks:
                - check_source_description_populated
                - check_source_has_meta_keys
                - check_source_has_tags
                - check_source_names
              reason: We recommend documenting why these checks are being skipped.

Similar can be done for other objects that support the meta value.

How to add a custom check to dbt-bouncer?#

In addition to the checks built into dbt-bouncer, you can write custom checks specific to your project's conventions. To add a custom check:

  1. Create an empty directory and add a custom_checks_dir key to your config file. The value should be the path to the directory, relative to the config file.
  2. In this directory create an empty __init__.py file.
  3. Create a subdirectory named catalog, manifest, or run_results depending on the artifact type you want to check.
  4. In that subdirectory create a Python file that defines a check using the @check decorator:

    • The function name must start with check_.
    • The function must be decorated with @check from dbt_bouncer.check_framework.decorator.
    • The first positional parameter determines the resource type to iterate over (e.g. model, source, exposure, seed).
    • Keyword-only arguments (after *) become user-configurable parameters, with types inferred from type hints.
    • Add ctx as a parameter only if the function needs access to the full check context (e.g. all models, all sources).
    • Use fail() from dbt_bouncer.check_framework.decorator to signal a check failure with a clear message.
    • Include a docstring describing what the check does.
  5. Add the check name and any desired arguments to your config file.

  6. Run dbt-bouncer — your custom check will be executed.

Example#

Directory tree:

.
├── dbt-bouncer.yml
├── dbt_project.yml
├── my_custom_checks
|   ├── __init__.py
|   └── manifest
|       └── check_custom_to_me.py
└── target
    └── manifest.json

Contents of check_custom_to_me.py:

import re

from dbt_bouncer.check_framework.decorator import check, fail


@check
def check_model_naming_convention(model, *, model_name_pattern: str = "^(stg|int|fct|dim)_"):
    """Model names must match the supplied regex."""
    if not re.match(model_name_pattern, str(model.name)):
        fail(
            f"`{model.unique_id}` does not match the required pattern "
            f"`{model_name_pattern}`."
        )

Contents of dbt-bouncer.yml:

custom_checks_dir: my_custom_checks

manifest_checks:
    - name: check_model_naming_convention
      include: ^models/staging
      model_name_pattern: ^stg_

All custom checks automatically support the following parameters (no need to declare them): description, exclude, include, and severity.

To contribute a new check back to dbt-bouncer itself, see Contributing.