Code#
Note
The below checks require manifest.json to be present.
Checks related to model source code content and structure.
Functions:
| Name | Description |
|---|---|
check_model_code_does_not_contain_regexp_pattern |
The raw code for a model must not match the specified regexp pattern. |
check_model_does_not_use_cartesian_join |
Models must not perform Cartesian or CROSS JOINs. |
check_model_does_not_use_select_star |
Models must not use |
check_model_hard_coded_references |
A model must not contain hard-coded table references; use ref() or source() instead. |
check_model_has_semi_colon |
Model may not end with a semi-colon ( |
check_model_incremental_has_unique_key |
Incremental models must declare a |
check_model_materialization_permitted |
Models must use a permitted materialization for their location. |
check_model_max_number_of_lines |
Models may not have more than the specified number of lines. |
check_model_code_does_not_contain_regexp_pattern
#
The raw code for a model must not match the specified regexp pattern.
Rationale
Teams often adopt coding standards that forbid certain SQL patterns — for example, using ifnull instead of coalesce, or using deprecated functions. This check allows those standards to be enforced automatically at CI time, preventing non-compliant code from reaching production.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
regexp_pattern
|
str
|
The regexp pattern that should not be matched by the model code. |
required |
Receives at execution time:
| Name | Type | Description |
|---|---|---|
model |
ModelNode
|
The ModelNode object to check. |
Other Parameters (passed via config file):
| Name | Type | Description |
|---|---|---|
description |
str | None
|
Description of what the check does and why it is implemented. |
exclude |
str | list[str] | None
|
Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked. |
include |
str | list[str] | None
|
Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked. |
materialization |
Literal[ephemeral, incremental, table, view] | None
|
Limit check to models with the specified materialization. |
severity |
Literal[error, warn] | None
|
Severity level of the check. Default: |
Example(s):
manifest_checks:
# Prefer `coalesce` over `ifnull`: https://docs.sqlfluff.com/en/stable/rules.html#sqlfluff.rules.sphinx.Rule_CV02
- name: check_model_code_does_not_contain_regexp_pattern
regexp_pattern: .*[i][f][n][u][l][l].*
Source code in src/dbt_bouncer/checks/manifest/models/code.py
check_model_does_not_use_cartesian_join
#
Models must not perform Cartesian or CROSS JOINs.
Rationale
Cartesian joins (or CROSS JOINs) join every row of one table to every row of another, leading to exponential row explosion, massive warehouse credit consumption, and potential query timeouts. Catching unintended cross joins at lint time prevents costly query execution errors in production.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
allow_explicit_cross_join
|
bool
|
Whether to allow intentional Cartesian joins. When |
False
|
Receives at execution time:
| Name | Type | Description |
|---|---|---|
model |
ModelNode
|
The ModelNode object to check. |
Other Parameters (passed via config file):
| Name | Type | Description |
|---|---|---|
description |
str | None
|
Description of what the check does and why it is implemented. |
exclude |
str | list[str] | None
|
Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked. |
include |
str | list[str] | None
|
Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked. |
materialization |
Literal[ephemeral, incremental, table, view] | None
|
Limit check to models with the specified materialization. |
severity |
Literal[error, warn] | None
|
Severity level of the check. Default: |
Info
Analysis is AST-based (via sqlglot) and flags explicit CROSS JOIN keywords,
missing ON/USING clauses, and constant ON conditions (e.g. ON 1=1,
ON TRUE). A condition is constant only when every operand is a literal,
so a genuine predicate is not flagged whichever side its literal sits on
(ON 1 = b.id and ON b.id = 1 both pass). NATURAL JOIN is not flagged,
as it joins on the columns the two relations share. Non-SQL (e.g. Python)
models are skipped.
Warning
Models that sqlglot cannot parse (e.g. heavy {% ... %} control flow)
fall back to a best-effort regular-expression scan.
Example(s):
Source code in src/dbt_bouncer/checks/manifest/models/code.py
211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 | |
check_model_does_not_use_select_star
#
Models must not use SELECT *.
Rationale
SELECT * makes a model's output schema implicit and brittle to upstream column changes. When a source or upstream model adds, removes, or reorders columns, a SELECT * model silently propagates the change, potentially breaking downstream consumers or introducing unexpected columns into the DAG. Explicit column lists are self-documenting, stable, and make schema changes intentional and reviewable.
Receives at execution time:
| Name | Type | Description |
|---|---|---|
model |
ModelNode
|
The ModelNode object to check. |
Other Parameters (passed via config file):
| Name | Type | Description |
|---|---|---|
description |
str | None
|
Description of what the check does and why it is implemented. |
exclude |
str | list[str] | None
|
Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked. |
include |
str | list[str] | None
|
Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked. |
materialization |
Literal[ephemeral, incremental, table, view] | None
|
Limit check to models with the specified materialization. |
severity |
Literal[error, warn] | None
|
Severity level of the check. Default: |
Info
Analysis is AST-based (via sqlglot) and flags both bare (SELECT *) and
qualified (SELECT t.*) stars, while correctly ignoring stars inside
function calls (count(*)), string literals, and comments. It analyses
the model's raw_code with Jinja neutralized. Non-SQL (e.g. Python)
models are skipped.
Warning
Models that sqlglot cannot parse (e.g. heavy {% ... %} control flow)
fall back to a best-effort regular-expression scan, which retains the
original limitations around string literals and Jinja tags.
Example(s):
Source code in src/dbt_bouncer/checks/manifest/models/code.py
check_model_hard_coded_references
#
A model must not contain hard-coded table references; use ref() or source() instead.
Flags table references qualified with a schema or catalog (e.g.
FROM schema.table or JOIN catalog.schema.table). Hard-coded
references bypass the dbt DAG, break lineage, and are environment-specific.
Rationale
Hard-coded table references bypass dbt's dependency graph, break lineage tracking, and are environment-specific — a reference that works in production will silently read the wrong data in development. Using ref() or source() ensures models run in the correct order, compile to the right environment, and appear correctly in lineage tools.
Info
Analysis is AST-based (via sqlglot) on the model's raw_code with Jinja
neutralized, so that {{ ref(...) }} / {{ source(...) }} are not
mistaken for hard-coded references. Compiled SQL is intentionally not
used: dbt renders every ref()/source() into a physical
schema-qualified relation, which is indistinguishable from a hand-written
hard-coded reference. Non-SQL (e.g. Python) models are skipped.
Warning
Models that sqlglot cannot parse (e.g. heavy {% ... %} control flow)
fall back to a best-effort regular-expression scan, which may miss
references inside complex Jinja logic.
Receives at execution time:
| Name | Type | Description |
|---|---|---|
model |
ModelNode
|
The ModelNode object to check. |
Other Parameters (passed via config file):
| Name | Type | Description |
|---|---|---|
description |
str | None
|
Description of what the check does and why it is implemented. |
exclude |
str | list[str] | None
|
Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked. |
include |
str | list[str] | None
|
Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked. |
materialization |
Literal[ephemeral, incremental, table, view] | None
|
Limit check to models with the specified materialization. |
severity |
Literal[error, warn] | None
|
Severity level of the check. Default: |
Example(s):
Source code in src/dbt_bouncer/checks/manifest/models/code.py
check_model_has_semi_colon
#
Model may not end with a semi-colon (;).
Rationale
dbt automatically wraps model SQL before executing it, so a trailing semi-colon can cause syntax errors in certain warehouse adapters. This check catches the mistake at lint time, preventing obscure build failures that can be hard to diagnose in CI.
Receives at execution time:
| Name | Type | Description |
|---|---|---|
model |
ModelNode
|
The ModelNode object to check. |
Other Parameters (passed via config file):
| Name | Type | Description |
|---|---|---|
description |
str | None
|
Description of what the check does and why it is implemented. |
exclude |
str | list[str] | None
|
Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked. |
include |
str | list[str] | None
|
Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked. |
materialization |
Literal[ephemeral, incremental, table, view] | None
|
Limit check to models with the specified materialization. |
severity |
Literal[error, warn] | None
|
Severity level of the check. Default: |
Example(s):
Source code in src/dbt_bouncer/checks/manifest/models/code.py
check_model_incremental_has_unique_key
#
Incremental models must declare a unique_key.
Rationale
Incremental models without a unique_key perform insert-only loads.
Late-arriving or updated rows are silently duplicated on every run,
leading to over-counting in downstream aggregations and broken
idempotency. Declaring a unique_key enables dbt to merge or
delete-insert matching rows, keeping the table correct across reruns.
Receives at execution time:
| Name | Type | Description |
|---|---|---|
model |
ModelNode
|
The ModelNode object to check. |
Other Parameters (passed via config file):
| Name | Type | Description |
|---|---|---|
description |
str | None
|
Description of what the check does and why it is implemented. |
exclude |
str | list[str] | None
|
Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked. |
include |
str | list[str] | None
|
Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked. |
materialization |
Literal[ephemeral, incremental, table, view] | None
|
Limit check to models with the specified materialization. |
severity |
Literal[error, warn] | None
|
Severity level of the check. Default: |
Example(s):
Source code in src/dbt_bouncer/checks/manifest/models/code.py
check_model_materialization_permitted
#
Models must use a permitted materialization for their location.
Rationale
A project may declare a directory-wide materialization in dbt_project.yml
(e.g. all staging models are views), but that default can be silently
overridden by an in-model config() block or a nested properties file.
Asserting the resolved materialization directly - rather than comparing the
sources of configuration - catches any model whose final materialization was
overridden away from what the directory is supposed to guarantee.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
permitted_materializations
|
list[Materialization]
|
List of materializations that models are permitted to use, e.g. |
required |
Receives at execution time:
| Name | Type | Description |
|---|---|---|
model |
ModelNode
|
The ModelNode object to check. |
Other Parameters (passed via config file):
| Name | Type | Description |
|---|---|---|
description |
str | None
|
Description of what the check does and why it is implemented. |
exclude |
str | list[str] | None
|
Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked. |
include |
str | list[str] | None
|
Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked. |
materialization |
Literal[ephemeral, incremental, table, view] | None
|
Limit check to models with the specified materialization. |
severity |
Literal[error, warn] | None
|
Severity level of the check. Default: |
Example(s):
manifest_checks:
- name: check_model_materialization_permitted
include: ^models/staging
permitted_materializations:
- view
Source code in src/dbt_bouncer/checks/manifest/models/code.py
check_model_max_number_of_lines
#
Models may not have more than the specified number of lines.
Rationale
Very long SQL files are a code smell that often indicates a model is doing too much — mixing staging, joining, and aggregating in a single file. Capping line counts encourages splitting large transformations into smaller, focused models that are easier to test, understand, and reuse.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
max_number_of_lines
|
int
|
The maximum number of permitted lines. |
100
|
Receives at execution time:
| Name | Type | Description |
|---|---|---|
model |
ModelNode
|
The ModelNode object to check. |
Other Parameters (passed via config file):
| Name | Type | Description |
|---|---|---|
description |
str | None
|
Description of what the check does and why it is implemented. |
exclude |
str | list[str] | None
|
Regex pattern(s) to match the model path. Model paths that match any pattern will not be checked. |
include |
str | list[str] | None
|
Regex pattern(s) to match the model path. Only model paths that match any pattern will be checked. |
materialization |
Literal[ephemeral, incremental, table, view] | None
|
Limit check to models with the specified materialization. |
severity |
Literal[error, warn] | None
|
Severity level of the check. Default: |
Example(s):