Skip to content

Agent Adaptation Framework ​

Research notes on "Adaptation of Agentic AI" (Stanford, Harvard, Berkeley, Caltech).

Core Thesis ​

Demos use static systems. Production requires continuous adaptation.

Adaptation is "the central mechanism for improving performance, reliability, and generalization" in agentic AI.

Why Agents Fail in Production ​

  1. Tool misalignment: Agents select inappropriate tools when tools lack task-specific optimization
  2. Action distribution mismatch: Agent preferences diverge from tool capabilities
  3. Feedback sparsity: Output-only signals provide insufficient learning information
  4. Generalization gaps: Tools trained on narrow distributions fail on diverse agent queries

The 4 Adaptation Paradigms ​

Two dimensions: what adapts (agent vs tool) × signal source (tool execution vs final output).

ParadigmWhat AdaptsSignal SourceBest For
A1AgentTool execution feedbackTool-intensive tasks (rich intermediate feedback)
A2AgentFinal output onlyComplex reasoning (simpler implementation)
T1ToolsIndependent of agentGeneral-purpose tools (reusable across agents)
T2ToolsAgent-supervisedSpecialized tasks (agent-specific optimization)

A1: Tool Execution Signaled Agent Adaptation ​

Agent improves based on tool execution outcomes.

Examples:

  • Orion: GRPO for information retrieval agents
  • DeepSeek-Prover-V2: GRPO + SFT for theorem proving with Lean compiler
  • Tool-R1: GRPO for multimodal QA with code execution feedback

A2: Agent Output Signaled Agent Adaptation ​

Agent adapts based on final output quality, not intermediate tool feedback.

Examples:

  • VerlTool: Multi-domain agent (math, SQL, web search) via GRPO
  • Self-RAG: Self-reflective retrieval-augmented generation
  • DeepSeek-R1-Zero: Pure reasoning improvement without tools
  • TextGrad: Gradient-based prompt optimization

T1: Agent-Agnostic Tool Adaptation ​

Tools improve independently, reusable across agents.

Examples:

  • CLIP: Vision-language alignment
  • DPR: Dense passage retrieval
  • AlphaFold2: Protein structure prediction
  • HuggingGPT: Multi-model tool orchestration

T2: Agent-Supervised Tool Adaptation ​

Tools improve through agent-generated training signals.

Examples:

  • s3: Small retriever adapted via agent feedback using PPO
  • RA-DIT: Knowledge-intensive tools supervised by LLM agents
  • Proxy-Tuning: Tool adaptation using larger model guidance

Practical Recommendation ​

Combine:

  • Rare A1/A2 updates on strong base models
  • Frequent T1/T2 adaptation of retrievers, search policies, simulators, memory

This balances stability (don't break the base model) with responsiveness (tools stay current).

  • GRPO (Group Relative Policy Optimization): Dominant RL approach across 40+ papers
  • SFT (Supervised Fine-tuning): Foundation for most agent adaptation
  • PPO/DPO: Preference-based learning for agent-tool alignment
  • Test-time methods: Gradient-based and contrastive refinement at inference

Relevance to Normalize ​

Current state:

  • Normalize tools (view, analyze, grep) are T1 - agent-agnostic, reusable across any LLM
  • Index refresh is T1 adaptation - tools improve independently via file watching
  • No agent adaptation (A1/A2) - normalize doesn't fine-tune LLMs

Implications:

  • T1 is the right choice for normalize: general-purpose tools work with any agent
  • If normalize ever needed specialization, T2 (agent-supervised tool adaptation) would be the path
  • A1/A2 require fine-tuning LLMs, which is outside normalize's scope (use the best available models)

Design validation:

  • "Tool misalignment" is exactly what normalize's structural tools address - give agents better tools, not more context
  • "Generalization gaps" supports normalize's approach of broad language support (98 languages)
  • Framework confirms normalize's implicit strategy: invest in T1 (tool quality), outsource A1/A2 to model providers

Friction Signals (Usage Telemetry) ​

The paper's "T2" (agent-supervised tool adaptation) maps to a practical question: how do we know when tools aren't working?

Implicit signals from agent behavior:

SignalWhat It MeansExample
Correction patternsAgent struggled, user fixed"You're right", "Should have", "Fair point" after tool call
Long tool chainsAgent can't get what it needs5+ normalize view calls without acting
Tool avoidanceAgent works around the toolUses grep instead of normalize view, spawns Explore agent
Follow-up patternsOutput was incompleteview --types-only → immediately view <symbol>
Repeated queriesFirst result wasn't usefulSame file/symbol viewed multiple times

What adaptation looks like (not LLM fine-tuning):

  • Adjust defaults: if 80% of --types-only is followed by symbol lookup, include first lines
  • Change output format: if agents parse JSON more successfully, prefer structured output
  • Add context: if agents always need imports after viewing a function, include them

Data collection (not yet implemented):

  • Log tool invocations with session ID
  • Capture next action (another tool call? user message? agent reasoning?)
  • Detect correction keywords in surrounding context
  • Track success/failure outcomes where measurable (edit applied? tests passed?)

This is observational, not intrusive - agents don't change behavior, tools learn from patterns.