Builtin Syntax Rules
Design and guidance for built-in syntax linting rules shipped with normalize.
Rule Loading Order
Rules are loaded in this order (later overrides earlier by id):
- Embedded builtins - compiled into the normalize binary
- User global -
~/.config/normalize/rules/*.scm - Project -
.normalize/rules/*.scm
To disable a builtin, add to .normalize/config.toml:
[analyze.rules."rust/println-debug"]
enabled = falseOr create a rule file with same id and enabled = false.
Project Type Guidance
Not all rules are appropriate for all project types:
| Rule | Library | CLI Tool | Web App | Notes |
|---|---|---|---|---|
rust/println-debug | ✓ Use | ❌ Disable | ✓ Use | CLI tools use println for output |
rust/dbg-macro | ✓ Use | ✓ Use | ✓ Use | Never commit dbg! |
rust/unwrap-in-impl | ✓ Use | ⚠️ Noisy | ⚠️ Noisy | Many in test code |
js/console-log | ✓ Use | N/A | ✓ Use | Production should use logging |
python/print-debug | ✓ Use | ❌ Disable | ✓ Use | CLI tools use print for output |
python/breakpoint | ✓ Use | ✓ Use | ✓ Use | Never commit breakpoint() |
go/fmt-print | ✓ Use | ❌ Disable | ✓ Use | CLI tools use fmt.Print |
ruby/binding-pry | ✓ Use | ✓ Use | ✓ Use | Never commit binding.pry |
Library code should avoid stdout/stderr side effects - use logging crates instead.
CLI tools legitimately use println! for output - disable rust/println-debug:
[analyze.rules."rust/println-debug"]
enabled = falseSeverity Recommendations
| Severity | Rules | When to use |
|---|---|---|
error | hardcoded-secret | Must fix before commit |
warning | rust/dbg-macro, rust/todo-macro, no-fixme-comment | Fix before merge |
info | rust/unnecessary-let, rust/println-debug, rust/unwrap-in-impl | Consider fixing |
Builtin Rules Reference
Rust Rules
rust/println-debug
Severity: info | Languages: rust
Flags println!, print!, eprintln!, eprint! macros.
Default allow: **/tests/**, **/examples/**, **/bin/**, **/main.rs
When to use: Library code where stdout/stderr side effects are bugs. When to disable: CLI tools, where println is the correct output mechanism.
rust/dbg-macro
Severity: warning | Languages: rust
Flags dbg!() macro calls - these should never be committed.
Default allow: **/tests/**
Always use. The dbg! macro is for temporary debugging only.
rust/todo-macro
Severity: warning | Languages: rust
Flags todo!() macro calls - unfinished code paths.
Always use. Helps track incomplete implementations.
rust/unwrap-in-impl
Severity: info | Languages: rust
Flags .unwrap() calls - suggests using ? or .expect() with context.
Default allow: **/tests/**, **/test_*.rs, **/*_test.rs, **/*_tests.rs, **/examples/**, **/benches/**
Legitimate unwrap uses:
- Lock poisoning:
mutex.lock().unwrap()- panic is correct if lock poisoned - Known-safe conversions: after validation that guarantees success
- Test code: tests should panic on unexpected failures
// normalize-allow: rust/unwrap-in-impl - panic correct if lock poisoned
let guard = CACHE.lock().unwrap();rust/expect-empty
Severity: warning | Languages: rust
Flags .expect("") with empty string - provide meaningful context.
Always use. Empty expect messages waste the opportunity to explain failures.
rust/unnecessary-let
Severity: info | Languages: rust
Flags let x = y; where both are simple identifiers - the binding adds no value.
Exclusions: let mut, underscore-prefixed names, None value.
rust/unnecessary-type-alias
Severity: info | Languages: rust
Flags type Foo = Bar; where Bar is a simple type - adds indirection without value.
Legitimate uses: Generic bounds, documentation, API stability.
rust/chained-if-let
Severity: info | Languages: rust | Requires: rust.edition >= 2024
Flags nested if-let patterns that can be chained with && in Rust 2024+.
// Before (flagged)
if let Some(x) = foo() {
if let Some(y) = bar(x) {
use_both(x, y);
}
}
// After (Rust 2024+)
if let Some(x) = foo() && let Some(y) = bar(x) {
use_both(x, y);
}Only runs on: Rust 2024+ edition (detected from Cargo.toml).
JavaScript/TypeScript Rules
js/console-log
Severity: info | Languages: javascript, typescript, tsx, jsx
Flags console.log, console.debug, console.info calls.
Default allow: **/tests/**, **/*.test.*, **/*.spec.*
js/unnecessary-const
Severity: info | Languages: javascript, typescript, tsx, jsx
Flags const x = y; where both are simple identifiers.
Exclusions: undefined, Infinity, NaN (global constants).
Python Rules
python/print-debug
Severity: info | Languages: python
Flags print() calls.
Default allow: **/tests/**, **/test_*.py, **/*_test.py, **/examples/**, **/__main__.py
When to use: Library code where stdout side effects are unexpected. When to disable: CLI tools, scripts where print is the output mechanism.
python/breakpoint
Severity: warning | Languages: python
Flags breakpoint() calls - these should never be committed.
Default allow: **/tests/**
Always use. The breakpoint() function is for interactive debugging only.
Go Rules
go/fmt-print
Severity: info | Languages: go
Flags fmt.Print, fmt.Println, fmt.Printf calls.
Default allow: **/tests/**, **/*_test.go, **/examples/**, **/cmd/**
When to use: Library code where structured logging is preferred. When to disable: CLI tools in cmd/ directories.
Ruby Rules
ruby/binding-pry
Severity: warning | Languages: ruby
Flags binding.pry and binding.irb calls - debugging breakpoints.
Default allow: **/tests/**, **/test/**, **/spec/**
Always use. Debug breakpoints should never be committed.
Cross-Language Rules
hardcoded-secret
Severity: error | Languages: all
Flags potential hardcoded secrets (API keys, passwords, tokens).
Default allow: **/tests/**, **/*.test.*
no-todo-comment
Severity: info | Languages: all with line comments
Flags // TODO comments in code.
no-fixme-comment
Severity: warning | Languages: all with line comments
Flags // FIXME comments - known bugs that need fixing.
Configuration Examples
Library Project
# Default rules are appropriate
# Optionally upgrade severity
[analyze.rules."rust/unwrap-in-impl"]
severity = "warning"CLI Tool Project
# Disable println-debug
[analyze.rules."rust/println-debug"]
enabled = falsePer-Directory Exclusions
[analyze.rules."rust/unwrap-in-impl"]
allow = ["**/generated/**", "**/proto/**"]Rule Conditionals (requires)
Rules can specify conditions that must be met before they run:
# ---
# id = "rust/chained-if-let"
# requires = { "rust.edition" = ">=2024" }
# ---Available Sources
| Source | Keys | Description |
|---|---|---|
env.* | Any env var | Environment variables (e.g., env.CI) |
path.* | rel, abs, ext, filename | File path components |
git.* | branch, staged, dirty | Repository state |
rust.* | edition, resolver, name, version | Cargo.toml fields |
typescript.* | target, module, strict, moduleResolution, name, version, node_version | tsconfig.json + package.json |
python.* | requires_python, name, version | pyproject.toml fields |
go.* | version, module | go.mod fields |
Operators
| Operator | Example | Description |
|---|---|---|
| (none) | "2024" | Exact match |
>= | ">=2024" | Greater or equal |
<= | "<=2021" | Less or equal |
! | "!2018" | Not equal |
Examples
# Only on Rust 2024+
requires = { "rust.edition" = ">=2024" }
# Only in CI
requires = { "env.CI" = "true" }
# Not on main branch
requires = { "git.branch" = "!main" }
# Combine multiple conditions (all must match)
requires = { "rust.edition" = ">=2024", "env.CI" = "true" }Pluggable Sources
The source system is pluggable via the RuleSource trait. Built-in sources cover common project types. Additional sources can be added by implementing RuleSource.
Known Limitations
In-File Test Detection (#[cfg(test)])
Rust commonly places tests in #[cfg(test)] modules within source files:
// src/lib.rs
pub fn add(a: i32, b: i32) -> i32 { a + b }
#[cfg(test)]
mod tests {
#[test]
fn test_add() {
assert_eq!(add(1, 2).unwrap(), 3); // Flagged but legitimate
}
}Current state: Glob patterns cannot detect #[cfg(test)] structure. These are flagged.
Workarounds:
- Use inline
// normalize-allow: rule-idcomments - Move tests to separate
tests/directory or*_test.rsfiles - Accept some false positives for rules like
rust/unwrap-in-impl
Future work: See "Design: #[cfg(test)] Detection" below.
Query Predicate Limitations
Tree-sitter predicates have limitations:
- No access to type information (can't distinguish
Result::unwrapfromOption::unwrap) - No cross-file analysis
- Limited to syntactic patterns
Design: #[cfg(test)] Detection
Problem: Many false positives come from test code in #[cfg(test)] modules within source files.
Options considered:
Global config flag:
skip_test_code = true- Pros: Simple
- Cons: All-or-nothing, affects all rules
Per-rule option:
skip_cfg_test = truein rule frontmatter- Pros: Fine-grained control
- Cons: Config complexity, must add to each rule
Query-level predicate:
(#not-in-cfg-test?)implemented in evaluator- Pros: Opt-in per query, reusable
- Cons: Rust-specific, implementation complexity
Automatic by category: Rules tagged "code-quality" skip test code by default
- Pros: Sensible defaults
- Cons: Hidden behavior, may surprise users
Proposed design: Option 3 with category defaults from option 4.
Add a (#not-in-test?) predicate that:
- For Rust: walks up tree looking for
#[cfg(test)]attribute on ancestor module/item - For JS/TS: checks if inside
describe(),it(),test()calls - Returns true if NOT in test context
Rules like rust/unwrap-in-impl would use:
((call_expression
function: (field_expression field: (field_identifier) @_method)
(#eq? @_method "unwrap")
(#not-in-test?)) @match)Status: Not implemented. Use inline allows or accept false positives for now.
Implementation Notes
Combined Query Optimization
All rules for a grammar are combined into a single tree-sitter query for single-traversal matching. This gives ~5-6x speedup over running each rule separately.
Key insight: Predicates scope per-pattern. When combining queries like:
; Pattern 0
((call_expression ...) (#eq? @_method "unwrap")) @match)
; Pattern 1
((call_expression ...) (#eq? @_method "expect")) @match)The @_method capture and #eq? predicate are evaluated independently for each pattern. Pattern 0 only matches "unwrap", Pattern 1 only matches "expect", even though they share the capture name. The pattern_index on each match identifies which rule matched.
This means we can concatenate arbitrary rule queries without conflict, as long as each rule uses @match for the primary capture. Cross-language rules (like no-todo-comment) that use node types not present in all grammars are filtered out during query compilation.
Embedding Rules
Rules are in crates/normalize/src/commands/analyze/builtin_rules/:
pub const BUILTIN_RULES: &[BuiltinRule] = &[
BuiltinRule {
id: "rust/todo-macro",
content: include_str!("rust_todo_macro.scm"),
},
// ...
];Testing Rules
To test a rule against the normalize codebase:
normalize analyze rules --rule "rust/unnecessary-let"To get SARIF output for IDE integration:
normalize analyze rules --sarif > results.sarifDebug Output
Enable timing information with --debug:
normalize analyze rules --debug timing
# [timing] file collection: 6ms
# [timing] query compilation: 90ms (4 grammars)
# [timing] file processing: 500ms (554 findings)
# [timing] total: 596msAvailable debug categories: timing, all