Frequently Asked Questions#
Can other tools perform the same checks as dbt-bouncer?#
dbt-bouncer is designed to be the most complete and flexible way to enforce conventions in a dbt project: it ships 120+ built-in checks, is configured via a YAML or TOML file format that dbt developers are already familiar with, never needs a database connection, can run checks against any of dbt's artifacts (manifest, catalog, and run results), can run both locally and in a CI pipeline, and lets you write custom checks in python.
dbt-bouncer |
dbt-checkpoint |
dbt-project-evaluator |
dbt-score |
|
|---|---|---|---|---|
| Built-in checks | 120+ | Collection of pre-commit hooks |
Fixed set of SQL/Jinja tests | Fixed set of model-only checks |
| Configuration | YAML or TOML config file |
.pre-commit-config.yaml |
dbt_project.yml + seed files (CSV) |
pyproject.toml |
| Checks written in | Python | Python | SQL + Jinja | Python |
| Requires a database connection | ❌ | ❌ | ✅ | ❌ |
| Checks any dbt artifact (manifest, catalog, run results) | ✅ | Manifest and catalog | ❌ (queries the database) | Manifest only |
| Where it runs | Locally and in a CI pipeline | Only as part of pre-commit/prek |
Locally and in a CI pipeline | Locally, from the command line |
Tip
dbt-bouncer can perform all the checks currently included in dbt-checkpoint, dbt-project-evaluator and dbt-score. If you see an existing check that is not possible with dbt-bouncer, open an issue and we'll add it!
How do I onboard an existing dbt project?#
An existing project usually fails many checks on the first run. Do not try to fix every failure first. Adopt the checks with a baseline, then reduce the backlog over time. A baseline records the current failures, so the run reports only failures that are new.
Follow these steps:
- Create a starter config. Run
dbt-bouncer init, or start from a preset withdbt-bouncer init --preset standard. - Generate the dbt artifacts. Run
dbt parsefor manifest checks, ordbt buildwhen you also use catalog or run-results checks. -
Record the current failures as a baseline:
-
Commit the config file and the baseline file.
-
Run
dbt-bouncerwith the baseline so only new failures fail the build:
Now the check run is green, and any new violation fails the build. To reduce the backlog, fix some failures and regenerate the baseline with the same command. The baseline shrinks as the project improves.
Two more options help during onboarding:
- To compare against a previous run without a stored file, use
--state <dir>in place of--baseline. It points at a directory of dbt artifacts from a previous run, the same way dbt's own--stateflag does. See the CLI reference. - To warn instead of fail while the team adapts, set
severity: warnat the top of the config file. Change it toerroronce the team is ready.
Does dbt-bouncer work with dbt Cloud?#
Yes! As dbt-bouncer runs on the artifacts generated by dbt, it can be used with dbt Cloud as long as the artifacts generated by the CI job in dbt Cloud are available.
For GitHub this can be achieved using the pgoslatara/dbt-cloud-download-artifacts-action action:
name: CI pipeline
on:
pull_request:
branches:
- main
jobs:
download-artifacts:
runs-on: ubuntu-latest
permissions:
pull-requests: write
steps:
- name: Checkout
uses: actions/checkout@v6
- name: Download dbt artifacts
uses: pgoslatara/dbt-cloud-download-artifacts-action@v1
with:
commit-sha: ${{ github.event.pull_request.head.sha }}
dbt-cloud-api-token: ${{ secrets.DBT_CLOUD_API_TOKEN }}
- name: Run dbt-bouncer
uses: godatadriven/dbt-bouncer@vX.X
Warning
dbt Cloud now supports a "versionless" option, which allows dbt projects to be run with the latest version of dbt. One effect of choosing this option is that dbt artifacts may receive non-breaking changes (source), these may or may not be compatible with dbt-bouncer. If you encounter a bug as a result of this, please open an issue and we'll investigate.
Does dbt-bouncer work with dbt 2.0 / Fusion?#
Yes. dbt 2.0 (the Rust-based Fusion engine) emits the same manifest.json, catalog.json and run_results.json artifacts — at manifest schema v12, catalog schema v1 and run results schema v6 — as dbt-core 1.x. Because dbt-bouncer consumes those artifacts directly, every check that works against dbt-core 1.x also works against dbt 2.0. Our CI builds the test project with dbt 2.0 and runs dbt-bouncer against the result on every pull request.
Three details changed at dbt 2.0. Plan for them:
- dbt 2.0 ships as the
dbtanddbt-ossdistributions. Thedbt-coredistribution stays on the 1.x line. --write-catalogmoved offdbt build. Usedbt compile --write-catalogordbt docs generate --write-catalog.- The Apache-2.0
dbt-ossdistribution writes nocatalog.json. Catalog checks need thedbtdistribution.
dbt 2.0 also writes Parquet artifacts under target/private/. These back the dbt docs v2 site. dbt-bouncer reads the JSON artifacts and ignores the Parquet directory.
Does dbt-bouncer support Python models?#
Yes! dbt-bouncer fully supports Python models (introduced in dbt 1.3). All checks that work with SQL models also work with Python models, this means dbt-bouncer can enforce conventions on your Python models just as it does for SQL models.
Example#
A Python model can be validated with the same checks as SQL models:
manifest_checks:
- name: check_model_has_meta_keys
keys:
- maturity
- name: check_model_description_populated
- name: check_model_code_does_not_contain_regexp_pattern
regexp_pattern: .*\.to_pandas\(\) # Discourage memory-intensive operations
The check check_model_code_does_not_contain_regexp_pattern will match against the Python code in raw_code, allowing you to enforce conventions like avoiding certain libraries or patterns in your Python models.
Tip
You can use regex patterns to enforce Python-specific conventions, such as:
- Preventing certain imports:
regexp_pattern: .*import\s+os.* - Discouraging memory-intensive operations:
regexp_pattern: .*\.to_pandas\(\).* - Enforcing use of specific libraries: Use
include/excludepatterns to target only Python models
Can AI coding agents use dbt-bouncer?#
Yes! dbt-bouncer ships a Model Context Protocol server. An AI coding agent can read which conventions your project enforces before it generates dbt code, and can run the checks to verify its work. Install the optional dependency with pip install 'dbt-bouncer[mcp]' and start the server with dbt-bouncer mcp. See the CLI documentation for the available tools and an example client configuration.
How to configure dbt-bouncer for use in a CI pipeline?#
dbt-bouncer is designed to be use primarily in a CI pipeline such as GitHub Actions or Azure DevOps. To do this we create a config file such as:
catalog_checks:
- name: check_column_description_populated
include: ^models/marts
manifest_checks:
- name: check_model_directories
include: ^models
permitted_sub_directories:
- intermediate
- marts
- staging
- utilities
run_results_checks:
- name: check_run_results_max_execution_time
max_execution_time_seconds: 10
The goal of a CI pipeline is to test the changes in a pull request but also to provide feedback to the developer as quickly as possible without incurring unnecessary costs (time, financial, compute, etc.). To achieve this we can combine several features of dbt and dbt-bouncer:
-
By running
dbt parse, dbt can generate amanifest.jsonwithout a database connection. We can then run our manifest checks via: -
dbt requires models to be materialised before it can generate a
catalog.jsonfile. By runningdbt run --emptywe can materialise every model without processing any data. Once these materializations are performed we can run our catalog checks via: -
Typically a CI pipeline will run a
dbt buildcommand with flags such as--stateand/or--defer. After this command has completed we can run our run results checks via:
Additionally, you can use the --check flag to run only specific checks by name. This is useful for debugging or validating a single convention:
```shell
dbt-bouncer --check check_model_has_unique_test
```
Multiple checks can be specified as a comma-separated list:
```shell
dbt-bouncer --check check_model_has_unique_test,check_model_description_populated
```
The --check and --only flags can be combined: --only restricts to the specified categories, then --check further narrows to only the named checks within those categories.
By using this approach, and combining with your own unique constraints and desires, dbt-bouncer can be used efficiently as part of your CI pipeline.
How do I get machine-readable results to parse in a script?#
Do not parse the console results table. That table is for humans. It truncates check names to fit the terminal width.
Set --output-file and select a structured --output-format. The supported formats are csv, json, junit, sarif, and tap.
Then read the output file in your script. See the CLI reference for full option details.
How to set up dbt-bouncer in a monorepo?#
A monorepo may consist of one directory with a dbt project and other directories with unrelated code. It may be desired for dbt-bouncer to be configured from the root directory. Sample directory tree:
.
├── dbt-bouncer.yml
├── README.md
├── dbt-project
│ ├── models
│ ├── dbt_project.yml
│ └── profiles.yml
└── package-a
├── src
├── tests
└── package.json
To ease configuration you can use exclude or include at the global level (see Config File for more details). For the above example dbt-bouncer.yml could be configured as:
dbt_artifacts_dir: dbt-project/target
include: ^dbt-project
manifest_checks:
- name: check_exposure_based_on_non_public_models
dbt-bouncer can now be run from the root directory.
How to set up dbt-bouncer in a dbt Mesh?#
A dbt Mesh is a collection of dbt projects in an organization, some of which can read models from other dbt projects. Natively supported by dbt Cloud, a dbt Mesh can also be set up with dbt Core using a plugin such as dbt-loom.
One challenge in a dbt Mesh is the large number of developers working across multiple dbt projects leading to differing conventions being implemented. There are multiple approaches to using dbt-bouncer in a dbt Mesh, two are outlined below.
Approach 1: Individual dbt-bouncer.yml configuration file#
Each dbt project can have its own dbt-bouncer.yml configuration file. This allows each project to adopt and implement its own conventions in addition to any conventions to be shared across all dbt projects. Should a breaking change be required to the config file then each dbt project can be updated independently at a time that makes sense.
This is the recommended approach due to its simplicity and ability to update each dbt project independently.
Approach 2: Centralised dbt-bouncer.yml configuration file shared via git submodule#
Warning
With this approach, a change to the centralised dbt-bouncer.yml file may result in CI pipelines in dbt projects failing despite no changes being made to these projects. As such we recommend implementing this approach only after extensive discussion with all dbt project developers so that all dbt projects can be brought into line before dbt-bouncer is enforced in the CI pipeline.
Should it be necessary for a breaking change to be made to the centralised dbt-bouncer.yml configuration file, we recommend setting the severity of the relevant check to warn so that CI pipelines in dbt projects will not fail and maintainers have sufficient time to make the necessary changes.
Git submodules allow the contents from one repository to be accessible from a different repository. Such a setup for dbt-bouncer can be achieved as follows (this example uses GitHub, similar setups can be achieved with other providers):
-
Set up a dedicated repository to store a centralised
dbt-bouncer.ymlconfiguration file that will be used by all dbt projects. Let's call this repositorydbt-bouncer-config. -
The contents of the
dbt-bouncer.ymlfile indbt-bouncer-configshould contain the following configuration fordbt_artifacts_dir: -
In every repository add a git submodule via:
-
Run
dbt-bouncer:
Your directory tree should look like this:
.
├── dbt-bouncer-config
│ └── dbt-bouncer.yml
├── dbt_project.yml
├── macros
│ └── ...
├── models
│ └── ...
├── profiles.yml
├── README.md
└── target
├── catalog.json
├── manifest.json
└── run_results.json
Note: if you update your central dbt-bouncer.yml file, you will need to run git submodule update --remote in every repository to update the submodule.
How to set up dbt-bouncer with prek/pre-commit?#
You can use the official pre-commit hook, in your .pre-commit-config.yaml file:
repos:
- repo: https://github.com/godatadriven/dbt-bouncer
rev: v4.0.0 # Check https://github.com/godatadriven/dbt-bouncer/releases for latest version
hooks:
- id: dbt-bouncer
args: ["--config-file", "<PATH_TO_CONFIG_FILE>"] # Optional
Alternatively, you can use a local hook to run automatically run dbt-bouncer before your commits get added to the git tree.
- repo: local
hooks:
- id: dbt-bouncer
name: dbt-bouncer
entry: dbt-bouncer # --config-file <PATH_TO_CONFIG_FILE>
language: system
pass_filenames: false
always_run: true
Can I skip specific checks for an exposure/model/source/etc.?#
Yes! Many dbt objects permit adding a meta config field (docs), this can be used to skip checks for the object. For example, for a model:
models:
- name: my_model
config:
meta:
dbt-bouncer:
skip_checks:
- check_model_description_populated
- check_model_has_meta_keys
reason: We recommend documenting why these checks are being skipped.
And for a source:
version: 2
sources:
- name: source_system
tables:
- name: source_1
config:
meta:
dbt-bouncer:
skip_checks:
- check_source_description_populated
- check_source_has_meta_keys
- check_source_has_tags
- check_source_names
reason: We recommend documenting why these checks are being skipped.
Similar can be done for other objects that support the meta value.
How to add a custom check to dbt-bouncer?#
In addition to the checks built into dbt-bouncer, you can write custom checks specific to your project's conventions. To add a custom check:
- Create an empty directory and add a
custom_checks_dirkey to your config file. The value should be the path to the directory, relative to the config file. - In this directory create an empty
__init__.pyfile. - Create a subdirectory named
catalog,manifest, orrun_resultsdepending on the artifact type you want to check. -
In that subdirectory create a Python file that defines a check using the
@checkdecorator:- The function name must start with
check_. - The function must be decorated with
@checkfromdbt_bouncer.check_framework.decorator. - The first positional parameter determines the resource type to iterate over (e.g.
model,source,exposure,seed). - Keyword-only arguments (after
*) become user-configurable parameters, with types inferred from type hints. - Add
ctxas a parameter only if the function needs access to the full check context (e.g. all models, all sources). - Use
fail()fromdbt_bouncer.check_framework.decoratorto signal a check failure with a clear message. - Include a docstring describing what the check does.
- The function name must start with
-
Add the check name and any desired arguments to your config file.
- Run
dbt-bouncer— your custom check will be executed.
Example#
Directory tree:
.
├── dbt-bouncer.yml
├── dbt_project.yml
├── my_custom_checks
| ├── __init__.py
| └── manifest
| └── check_custom_to_me.py
└── target
└── manifest.json
Contents of check_custom_to_me.py:
import re
from dbt_bouncer.check_framework.decorator import check, fail
@check
def check_model_naming_convention(model, *, model_name_pattern: str = "^(stg|int|fct|dim)_"):
"""Model names must match the supplied regex."""
if not re.match(model_name_pattern, str(model.name)):
fail(
f"`{model.unique_id}` does not match the required pattern "
f"`{model_name_pattern}`."
)
Contents of dbt-bouncer.yml:
custom_checks_dir: my_custom_checks
manifest_checks:
- name: check_model_naming_convention
include: ^models/staging
model_name_pattern: ^stg_
All custom checks automatically support the following parameters (no need to declare them): description, exclude, include, and severity.
To contribute a new check back to dbt-bouncer itself, see Contributing.