Contributing: CONTRIBUTING.md · ROADMAP.md · GOVERNANCE.md · CODE_OF_CONDUCT.md · SECURITY.md · NOTICE
- Governed Agent Stack brings together on-prem tools for auditable database agents; sql-sop is a standalone linter that can complement that workflow.
- sql-steward governs SQL generated by agents through a semantic layer and audit trail, complementing sql-sop's checks on SQL written by people.
- sql-explorer-mcp provides read-only database exploration over MCP and uses sql-sop to reject dangerous queries before execution.
The dbt-aware rule pack (DBT001+) extends sql-sop into dbt projects. See the ADR for the broader roadmap.
One bad SQL query can delete production data, expose customer records, or bring down a database. Most teams only find out after the damage is done. sql-sop catches dangerous patterns automatically, before the query ever runs, in 0.08 seconds.
| Rules | 48 (9 errors, 25 warnings, 3 structural, 6 T-SQL, 5 Python-source); 53 with --contract, 55 with --dbt |
| Tests | 445 |
| Scan speed | 0.08s across 200 files |
| PyPI installs | 2,000+ (mirrors excluded) |
| Version | 0.11.0 |
from sql_guard import SqlGuard
result = SqlGuard().enable("E001", "W001").scan("DELETE FROM users")
print(result.passed) # False
print(result.summary()) # "1 error, 0 warnings in 1 statement"48 rules, plus an optional Contracts pack (5 schema-aware rules) under --contract path.yml and a dbt pack (7 rules) under --dbt. Inline disable, project config, git-changed-only mode, and SARIF output for GitHub Code Scanning. Runs as a CLI tool, pre-commit hook, and GitHub Action.
Want the same rules inside your AI assistant? sql-sop-mcp exposes them over MCP, so the model is told a query is unsafe before it suggests it.
If sql-sop catches a real bug for you, a GitHub star is the easiest way to help. It makes the project more discoverable for people with the same problem.
pip install sql-sop
sql-sop check .
# Also scan .py files for SQL hazards in execute()/read_sql() calls:
pip install "sql-sop[python]"
sql-sop check . --include-pythonqueries/create_orders.sql
L3: ERROR [E001] DELETE without WHERE clause -- this will delete all rows
-> Add a WHERE clause to limit affected rows
L7: WARN [W001] SELECT * -- specify columns explicitly
-> Replace with: SELECT col1, col2, col3 FROM ...
Found 2 issues (1 error, 1 warning) in 1 file (0.001s)
Most teams have no SQL review process at all. sql-sop gives you one in two places, both driven by the same rules:
- Pre-commit -- runs in under 0.2s on the files you changed, so a bad
DELETE FROM usersis blocked before it is ever committed. - CI -- runs on every pull request and uploads SARIF, so findings appear inline in the GitHub "Files changed" view through Code Scanning.
Same engine in both, no AI, no API keys, nothing to host. Fast enough to sit in the commit path, strict enough to gate a PR.
# .pre-commit-config.yaml
repos:
- repo: https://github.com/Pawansingh3889/sql-sop
rev: v0.11.0
hooks:
- id: sql-sop
args: [--severity, error] # only block on errors locallypip install pre-commit
pre-commit installNow every git commit with .sql changes runs sql-sop automatically. Errors block the commit. Warnings are shown but don't block.
# .github/workflows/sql-quality.yml
name: SQL Quality
on:
pull_request:
paths: ['**/*.sql']
permissions:
contents: read
pull-requests: write
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: Pawansingh3889/sql-sop@v0.11.0
with:
severity: warningThat's it. Every SQL change gets an instant, rule-based lint on the PR.
sql-sop list-rules prints every registered rule. See Configuration for the flags that change what is reported.
If the sql-sop script is not on PATH (some CI images and Windows setups), python -m sql_guard runs the same CLI under a chosen interpreter: python -m sql_guard check .
| ID | Name | What it catches |
|---|---|---|
| E001 | delete-without-where |
DELETE FROM orders; -- deletes all rows |
| E002 | drop-without-if-exists |
DROP TABLE users; -- fails if table missing |
| E003 | grant-revoke |
GRANT SELECT ON users TO public; -- privilege escalation |
| E004 | string-concat-in-where |
WHERE id = '' + @input -- SQL injection |
| E005 | insert-without-columns |
INSERT INTO t VALUES (...) -- breaks on schema change |
| E006 | update-without-where |
UPDATE orders SET status = 'x'; -- overwrites every row |
| E007 | alter-add-not-null-no-default |
ALTER TABLE t ADD c INT NOT NULL; -- locks table for full rewrite |
| E008 | drop-column |
ALTER TABLE t DROP COLUMN c; -- irreversible, breaks subscribers |
| E009 | update-from-without-join |
UPDATE a SET x = b.x FROM a, b -- comma-separated FROM silently makes a Cartesian product |
| ID | Name | What it catches |
|---|---|---|
| W001 | select-star |
SELECT * FROM users -- pulls unnecessary columns |
| W002 | missing-limit |
Unbounded SELECT -- could return millions of rows |
| W003 | function-on-column |
WHERE YEAR(date) = 2024 -- kills index usage |
| W004 | missing-alias |
JOIN without table aliases -- hard to read |
| W005 | subquery-in-where |
WHERE x IN (SELECT ...) -- often slower than JOIN |
| W006 | orderby-without-limit |
ORDER BY without LIMIT -- sorts entire result |
| W007 | hardcoded-values |
WHERE amount > 10000 -- use parameters |
| W008 | mixed-case-keywords |
select ... FROM -- inconsistent casing |
| W009 | missing-semicolon |
Statement not terminated with ; |
| W010 | commented-out-code |
-- SELECT * FROM old_table -- use version control |
| W011 | union-without-all |
UNION between disjoint sets -- forces a deduplication sort, UNION ALL is faster when uniqueness is guaranteed |
| W012 | group-by-ordinal |
GROUP BY 1, 2 -- fragile to SELECT-list reorders |
| W013 | window-missing-partition |
OVER () -- unpredictable results and unclear intent |
| W014 | case-without-else |
CASE WHEN ... THEN ... END -- unmatched rows return NULL |
| W015 | join-function-on-column |
JOIN customers c ON UPPER(o.email) = UPPER(c.email) -- kills index seek |
| W016 | not-in-with-subquery |
WHERE id NOT IN (SELECT ...) -- silently returns 0 rows on NULL |
| W017 | leading-wildcard-like |
WHERE name LIKE '%smith' -- non-SARGable, full scan |
| W018 | or-across-columns |
WHERE a = 1 OR b = 2 -- defeats single-column indexes |
| W019 | count-distinct-unbounded |
COUNT(DISTINCT col) with no WHERE / GROUP BY / LIMIT -- full sort + distinct over the whole table |
| W020 | truncate-table |
TRUNCATE TABLE staging; -- bypasses triggers, resets identity |
| W021 | having-without-group-by |
HAVING status = 'x' with no GROUP BY -- legal, but usually a misplaced WHERE |
| W022 | cross-join-explicit |
FROM products CROSS JOIN regions -- Cartesian product, confirm intent |
| W023 | scalar-udf-in-where |
WHERE dbo.fn_X(col) = 1 -- row-by-row predicate evaluation |
| W024 | select-distinct-suspicious |
SELECT DISTINCT a, b FROM x JOIN y ON ... -- DISTINCT often masks a missing join condition or GROUP BY |
| W025 | assertion-malformed |
-- @assert: <predicate> comment whose predicate does not match the sql-sop grammar (row_count <op> <int>, unique(<col>), not_null(<col>), <col> <op> <literal>) |
Rules that need an AST view of the statement, parsed via sqlparse. Catch issues that single-line regex matching cannot reliably see.
| ID | Name | What it catches |
|---|---|---|
| S001 | implicit-cross-join |
JOIN customers with no ON / USING -- accidental Cartesian product |
| S002 | deeply-nested-subquery |
Subqueries more than 2 levels deep -- typically a refactor opportunity |
| S003 | unused-cte |
WITH x AS (...) defined but never referenced |
Rules targeting SQL Server anti-patterns common in legacy stored procs and SSRS datasets. Fire on text patterns that do not appear in BigQuery or Postgres code, so they run unconditionally with near-zero false positives on non-T-SQL input.
T002 and T004 have default severity error (they block --severity error
and the documented pre-commit hook). The other T-SQL rules are warnings.
| ID | Name | Severity | What it catches |
|---|---|---|---|
| T001 | with-nolock |
warning | SELECT * FROM t WITH (NOLOCK) -- dirty reads |
| T002 | xp-cmdshell |
error | EXEC xp_cmdshell ... -- shell-exec surface |
| T003 | cursor-declaration |
warning | DECLARE c CURSOR FOR ... -- row-by-row processing |
| T004 | deprecated-outer-join |
error | WHERE a.x *= b.y -- removed in SQL Server 2012+ |
| T005 | create-index-without-online |
warning | CREATE INDEX ix ON t (...) -- locks table; add WITH (ONLINE = ON) |
| T006 | select-into-without-typed-fields |
warning | SELECT * INTO target FROM source -- destination schema is inferred at runtime |
Pass --contract path/to/contract.yml (or set contract: in
.sql-guard.yml) to lint queries against the expected schema. Without a
contract these rules are silent. Format is a thin subset of the open
data-contract space; see tests/fixtures/contract_sample.yml for a
working example.
| ID | Name | What it catches |
|---|---|---|
| C001 | column-not-in-contract |
SELECT o.bogus FROM orders o -- column not declared for that table |
| C002 | table-not-in-contract |
SELECT * FROM ghost_table -- table absent from the contract |
| C003 | not-null-violation |
INSERT INTO orders (id) VALUES (1) -- omits a NOT NULL column |
| C004 | primary-key-missing-on-insert |
INSERT omits a PK column with no default |
| C005 | unmapped-fk |
JOIN ... ON o.id = c.id -- columns have no FK relationship in the contract |
Two helper subcommands round out the workflow:
# Bootstrap a contract from an existing database (requires sql-sop[snapshot]):
sql-sop schema-snapshot \
--dsn "mssql+pyodbc://user:pass@host/db?driver=ODBC+Driver+18+for+SQL+Server" \
--output contract.yml
# Validate a contract YAML structure before running rules (CI-friendly):
sql-sop validate-contract --contract contract.ymlEnable with pip install "sql-sop[python]" and --include-python. Uses
libCST to walk Python source and extract SQL strings from .execute(),
.read_sql(), sqlalchemy.text(...) calls and sql =/query = style
assignments. Then applies every rule above, plus five that only make
sense at the Python level:
| ID | Name | What it catches |
|---|---|---|
| P001 | fstring-in-execute |
cursor.execute(f"... {user_input}") -- SQL injection |
| P002 | concat-in-execute |
cursor.execute("..." + user_input) -- SQL injection |
| P003 | format-in-execute |
.format() or % interpolation into an execute call |
| P004 | bare-variable-in-execute |
cursor.execute(query) where query is an unchecked variable |
| P005 | sqlalchemy-text-fstring |
sqlalchemy.text(f"... {var}") -- SQL injection on the SQLAlchemy text() surface |
Pass --dbt to activate the dbt-aware rule pack. sql-sop walks up from
the first checked path to find dbt_project.yml and reads schema.yml
at lint time; with no project found the pack stays silent, so existing
users see no behaviour change. Rules only fire on .sql files inside
the project's model-paths -- macros, analyses, seeds and snapshots are
skipped.
| ID | Name | What it catches |
|---|---|---|
| DBT001 | model-without-test |
Model absent from schema.yml, or declared with neither tests: nor data_tests: |
| DBT002 | direct-table-ref |
Raw table name in FROM/JOIN instead of ref() / source() -- breaks lineage |
| DBT003 | incremental-without-unique-key |
materialized='incremental' with no unique_key. Error on merge / delete+insert (silent duplicate rows), warning when the strategy is unset, silent on append / insert_overwrite |
| DBT004 | hook-with-ddl |
pre_hook / post_hook containing destructive or structural DDL |
| DBT005 | select-star-in-mart |
SELECT * in a mart model -- downstream contracts drift silently |
| DBT006 | model-without-description |
Model has no description: in schema.yml |
| DBT007 | unquoted-var-interpolation |
{{ var(...) }} interpolated into SQL without surrounding quotes -- breaks the day the var holds whitespace or a quote |
sql-sop check models/ --dbtDBT005 treats a path segment named marts (case-insensitive) as the mart
layer. Projects that name it gold, core, reporting, etc. can override
the segment with a repeatable flag or the dbt_mart_paths config key:
sql-sop check models/ --dbt --dbt-mart-path gold --dbt-mart-path reportingsql-sop check . --disable E002 W008 W010Drop a .sql-guard.yml (or .sql-guard.yaml) at the repo root. The loader walks up from the current directory; CLI flags merge with and override these settings.
disable:
- W005
- T001
ignore:
- migrations/legacy/
- vendor/
include_python: true
severity: warning
dbt_mart_paths: # path segments DBT005 treats as mart layers (default: [marts])
- marts
- goldSilence a known false positive on a single line, no project-wide override needed:
SELECT * FROM lookups; -- sql-guard: disable=W001
SELECT * FROM users -- sql-guard: disable=W001,W002
WHERE name LIKE '%smith';
-- sql-guard: disable-next-line=W017
SELECT * FROM events WHERE name LIKE '%checkout';A bare -- sql-guard: disable (no equals sign) silences every rule on the line. The same directives work in Python with # instead of --.
For pre-commit and CI on big repos:
sql-sop check . --changed-only # working tree
sql-sop check . --changed-only --changed-base main # vs a branch refFalls back to a full scan with a warning when not in a git repo.
sql-sop check . --severity error # only show errors
sql-sop check . --severity warning # show everything (default)sql-sop check . --fail-fast # stop after first error foundRender findings inline on PRs in the GitHub Files Changed view:
sql-sop check . --format sarif --output results.sarifIn a GitHub Actions workflow:
- run: sql-sop check . --format sarif --output sql-guard.sarif
- uses: github/codeql-action/upload-sarif@v3
with:
sarif_file: sql-guard.sarifsql-sop is designed to be fast:
- Compiled regex -- patterns compiled once at startup, reused per file
- Two-pass scanning -- single-line rules run first (16 of 43 SQL rules), multi-line parsing only when needed
- Line-by-line streaming -- files read line by line, not loaded entirely into memory
- Early exit --
--fail-faststops on first error
Benchmark: 200 SQL files, 20 SQL rules
sql-guard: 0.08 seconds
sqlfluff: 45 seconds (560x slower)
In a regulated data environment, sql-sop runs as a pre-commit hook on all SQL that touches the database. Combined with read-only database users and container isolation, it forms part of a layered safety setup that prevents accidental writes to production.
| sql-sop | sqlfluff | sql-lint | |
|---|---|---|---|
| Rules | 48 (focused) | 800+ (comprehensive) | ~20 |
| Speed | <0.1s for 200 files | 45s for 200 files | ~2s |
| Config needed | Zero | Extensive | Minimal |
| Language | Python | Python | JavaScript |
| Pre-commit | Yes | Yes | No |
| GitHub Action | Yes | Community | No |
| AI integration | MCP server (sql-sop-mcp) | No | No |
sql-sop is not a replacement for sqlfluff. It's a fast first pass that catches 80% of real issues with zero setup. If you need dialect-specific formatting and 800 rules, use sqlfluff. If you want instant feedback on dangerous SQL, use sql-sop.
git clone https://github.com/Pawansingh3889/sql-sop.git
cd sql-sop
pip install -e ".[dev]"
pytest- Create a class in
sql_guard/rules/errors.pyorwarnings.py - Inherit from
Rule, setid,name,severity,description - Override
check_line()for single-line rules orcheck_statement()for multi-line - Add to
ALL_RULESinsql_guard/rules/__init__.py - Add a test in
tests/test_rules.py - Add a trigger case in
tests/fixtures/
class MyNewRule(Rule):
id = "W011"
name = "my-rule"
severity = "warning"
description = "What this rule catches"
multiline = False
_pattern = Rule._compile(r"your regex here")
def check_line(self, line, line_number, file):
if self._pattern.search(line):
return Finding(
rule_id=self.id,
severity=self.severity,
file=file,
line=line_number,
message="What went wrong",
suggestion="How to fix it",
)
return NonePRs welcome. Keep rules simple, keep patterns fast.
Thank you to the people who have shipped rules and code to sql-sop.
| Contributor | Contribution |
|---|---|
| @tmchow | W011 union-without-all. Flags UNION where UNION ALL would be safe and faster. |
| @tmchow | P005 sqlalchemy-text-fstring. Catches sqlalchemy.text(f"...{var}") patterns that defeat parameter binding. |
| @mvanhorn | W019 count-distinct-unbounded. Flags COUNT(DISTINCT col) without WHERE, GROUP BY, or LIMIT. |
| @mvanhorn | W015 join-function-on-column. JOIN-side companion to W003. Flags function calls wrapping columns inside JOIN ... ON predicates. |
| @mvanhorn | W023 scalar-udf-in-where. Flags schema-qualified scalar UDF calls inside WHERE, HAVING, and ON predicates. |
| @Prabhu-1409 | W013 window-without-partition. Flags OVER () without PARTITION BY, dialect-aware messaging for Postgres and Redshift. |
| @hellozzm | W014 case-without-else. Walks CASE/END token-by-token; catches outer CASE without ELSE even when an inner CASE does have one. |
| @vibeyclaw | W022 cross-join-explicit. Flags explicit CROSS JOIN. Strips trailing line comments before matching to avoid false positives on commentary. |
See the full contributors graph on GitHub.
Want to add your name here? Pick a good first issue, follow CONTRIBUTING.md, and check the roadmap for the next batch of rules.
MIT