DWIM-Driven Agentic Loop
Design for an indefinite agent loop where the LLM outputs terse intents and DWIM handles tool routing.
Philosophy
- LLM makes decisions, not tool selections
- DWIM interprets intent and routes to appropriate tools
- No tool schemas in prompts - saves tokens, reduces coupling
- Terse agent output - "view foo.py" not "please show me the structure"
- Natural language for humans - same DWIM handles both
Architecture
┌──────────────────────────────────────────────────────┐
│ AGENT LOOP │
├──────────────────────────────────────────────────────┤
│ │
│ USER/LLM │
│ │ │
│ ▼ │
│ "view foo.py" ─or─ "show me the structure" │
│ │ │
│ ▼ │
│ DWIM PARSER │
│ ├─ Extract verb: view, edit, analyze │
│ ├─ Extract target: file path, symbol, error desc │
│ └─ Route to tool with confidence score │
│ │ │
│ ▼ │
│ TOOL EXECUTOR │
│ ├─ Core Primitives (view, edit, analyze) │
│ ├─ MCP servers (external capabilities) │
│ └─ LLM (when tool needs generated content) │
│ │ │
│ ▼ │
│ RESULT → fed back to LLM │
│ │ │
│ ▼ │
│ LLM → next intent ─or─ "done" │
│ │
└──────────────────────────────────────────────────────┘Agent Output Format
Terse, token-efficient. Verb + target(s):
view src/normalize/agent_loop.py
view src/normalize/agent_loop.py/Patch
analyze --complexity
edit -f patches.py "add type check for anchor"
doneNo prose, no "I will now...", just action.
DWIM Responsibilities
- Verb extraction - identify action: view, edit, analyze, done
- Target extraction - parse file paths, symbol names, options
- Tool routing - map intent to one of 3 core primitives
- Confidence scoring - know when to ask for clarification
- Parameter construction - build tool call from extracted parts
Integration Points
Existing Infrastructure
normalize.dwim- hasresolve_core_primitive(), simple alias matchingnormalize.session- tracks tool calls, file changes, LLM usagenormalize.agent_loop- hasAgentLoopRunner, executors, metricslitellm- unified LLM access
What Needs Building
- Intent parser - extract verb + targets from terse commands
- Main loop - iterate: LLM → DWIM → execute → result → LLM
- Completion detection - recognize "done" signal
- Context management - what to feed back to LLM (truncation, summarization)
Example Flow
User: "Fix the type error in Patch.apply"
LLM: view src/normalize/patches.py
→ Routes to view command (Rust CLI)
→ Returns: class Patch, def apply(...) skeleton
LLM: view src/normalize/patches.py/Patch/apply
→ Routes to view with symbol path
→ Returns: full function source
LLM: edit -f patches.py "add type check for anchor parameter"
→ Routes to edit command
→ Applies fix via structural editing
LLM: analyze --security
→ Routes to analyze command
→ Returns: no issues found
LLM: done
→ Loop terminatesToken Efficiency
Compared to tool-schema approach:
| Approach | Tokens/turn | Notes |
|---|---|---|
| OpenAI function calling | ~500-2000 | Full schemas every request |
| Claude Code XML | ~200-500 | Tool blocks + formatting |
| DWIM terse | ~10-50 | Just "view foo.py" |
90%+ token reduction for tool selection.
Context Model: Hierarchical Path
The agent does NOT accumulate conversation history. Context is structured as a path from root task to current leaf, with optional attachments.
Core Structure
Task: Fix auth bug
→ Find failure point ✓ (token expires during refresh)
→ Implement fix
→ [now] Patching refresh_token()
[note: refresh_token() is called from 3 places | expires: on_done]Design Principles
- Context-excluded by default: Start lean, pull what's needed
- Path, not history: Chain of refinements, not transcript
- Levels emerge, not predefined: Arbitrary depth, task dictates structure
- Recursive breakdown is fundamental: Agent decomposes until leaf is actionable
Path Components
Each node in the path:
goal: What this step aims to dostatus: pending | active | done | blockedsummary: One-line result (when done)description: Expandable detail (on demand)children: Subtasks (if decomposed)
Attachments
Standalone notes that travel with context:
content: The note itselfcondition: When to expire (on_done, after:N_turns, until:pattern_found, manual)scope: Which subtree it applies to
note("refresh_token calls: auth.py:45, session.py:120, api.py:89", expires="on_done")
note("avoid changing public API", expires="manual")Prompt Structure
Each turn:
[system: terse agent role]
[path: Task → Subtask → Current step]
[notes: active attachments]
[last_result: preview + id:0042]
[action?]~300 tokens typical, scales with path depth not turn count.
State is External
- Path nodes: TaskTree structure
- Full outputs: EphemeralCache (by ID)
- Notes: Attachment store with TTL
- Findings: Working memory (compact)
Open Questions
- Ambiguity handling - when DWIM confidence is low, ask LLM to clarify or just pick best?
- Error recovery - tool fails, how does LLM know? Structured error format?
- Multi-step intents - "fix and validate" in one line?
- Decomposition trigger - when does a task become subtasks? LLM decides? Heuristic?
Related Docs
docs/dwim-architecture.md- DWIM internalsdocs/hybrid-loops.md- CompositeToolExecutor for multi-source toolsdocs/philosophy.md- minimize LLM usage, structure over text