Syntax-Based Linting Design
Custom rules using tree-sitter queries, like ESLint's no-restricted-syntax.
Problem
Some patterns should be restricted to specific locations:
GrammarLoader::newshould only appear ingrammar_loader()singletonQuery::newshould only appear in cached query getters- Direct
unwrap()on user input should be flagged
Currently no way to express these rules without external tools (semgrep, etc).
Solution
Tree-sitter query-based rules with location allowlists.
Rule Definition
Rules in .normalize/rules/ as .scm files with TOML frontmatter:
# .normalize/rules/no-grammar-loader-new.scm
# ---
# id = "no-grammar-loader-new"
# severity = "error"
# message = "Use grammar_loader() singleton instead of GrammarLoader::new"
# allow = ["**/grammar_loader.rs"]
# ---
(call_expression
function: (scoped_identifier
path: (identifier) @_type
name: (identifier) @_method)
(#eq? @_type "GrammarLoader")
(#eq? @_method "new")) @matchCLI Interface
# Run all rules
normalize analyze rules
# Run specific rule
normalize analyze rules --rule no-grammar-loader-new
# List available rules
normalize analyze rules --list
# Authoring helpers
normalize analyze ast <file> # Show AST for a file
normalize analyze ast <file> --at 42 # Show AST node at line 42
normalize analyze query <file> <query> # Test query against fileAuthoring Tools
To help write queries, expose AST inspection:
# Dump full AST with node types
$ normalize analyze ast src/main.rs
(source_file
(function_item
name: (identifier) "main"
body: (block
(expression_statement
(call_expression ...)))))
# Show node at cursor/line
$ normalize analyze ast src/main.rs --at 42
Line 42 is inside:
call_expression (L42:5-42:30)
function: scoped_identifier (L42:5-42:22)
path: identifier "GrammarLoader" (L42:5-42:18)
name: identifier "new" (L42:20-42:22)
arguments: arguments (L42:23-42:30)
# Test a query interactively
$ normalize analyze query src/main.rs '(call_expression function: (scoped_identifier) @fn)'
3 matches:
src/main.rs:42 - GrammarLoader::new()
src/main.rs:55 - Config::load()
src/main.rs:78 - Parser::parse()Rule Storage
- Project rules:
.normalize/rules/*.scm - Global rules:
~/.config/normalize/rules/*.scm - Builtin rules: Compiled into normalize (optional, curated set)
Configuration
In .normalize/config.toml:
[rules]
# Enable/disable specific rules
enabled = ["no-grammar-loader-new", "no-unwrap-on-input"]
disabled = ["some-noisy-rule"]
# Severity overrides
[rules.severity]
no-grammar-loader-new = "warn" # downgrade from errorImplementation
- Rule loader: Parse
.scmfiles with TOML frontmatter - Query runner: Execute tree-sitter queries against files
- Match filter: Apply allowlist patterns to matches
- Reporter: Format findings (text, JSON, SARIF)
Core types:
pub struct Rule {
id: String,
query: tree_sitter::Query,
severity: Severity,
message: String,
allow: Vec<glob::Pattern>,
}
pub struct Finding {
rule_id: String,
file: PathBuf,
range: Range,
message: String,
severity: Severity,
}Query Language
Use tree-sitter's native S-expression query syntax:
- Pattern matching:
(function_item name: (identifier) @name) - Predicates:
#eq?,#match?,#any-of? - Captures:
@name,@match(special: marks the finding location)
The @match capture is required and marks where the finding is reported.
Phase 1: MVP
normalize analyze ast <file>- dump ASTnormalize analyze query <file> <query>- test queries- Rule files with basic format
normalize analyze rules- run all rules
Phase 2: Polish
--at <line>for focused AST inspection- Allow patterns (glob-based)
- Severity configuration
- SARIF output for IDE integration
Phase 3: Ecosystem
- Builtin rule library
- Rule sharing/import mechanism
- Auto-fix support (where possible)
Non-Goals
- Semantic analysis (type information, control flow)
- Cross-file analysis (imports, call graphs)
- Auto-fix for complex patterns
These would require deeper integration with the indexer.
Alternatives Considered
Semgrep/ruff integration
Pro: Mature, feature-rich Con: External dependency, different query syntax, overkill for simple patterns
Custom DSL
Pro: Tailored to our needs Con: Yet another language to learn, maintenance burden
Tree-sitter queries (chosen)
Pro: Standard syntax, reusable knowledge, extensive documentation Con: Learning curve for users unfamiliar with S-expressions
Success Criteria
- Can express "X only allowed in Y" rules
- Authoring workflow is discoverable (
analyze ast,analyze query) - Performance: <100ms for typical project
- Rules are portable (share in repos, copy between projects)