跳转至

Architecture Deep Dive

This document provides a detailed explanation of the active tool selection system architecture, inspired by MCP-Zero.

Table of Contents

  1. System Overview
  2. Core Components
  3. Active Discovery Flow
  4. Semantic Routing Algorithm
  5. Comparison: Active vs Passive
  6. Performance Optimization
  7. Design Decisions

System Overview

The active tool selection system consists of four major components working together:

┌─────────────────────────────────────────────────────────┐
│                    User Task                            │
└──────────────────────┬──────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│              Active Tool Agent                          │
│  • Task analysis                                        │
│  • Capability gap identification                        │
│  • Structured tool request generation                   │
│  • Tool usage and task execution                        │
└──────────────────────┬──────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│          Hierarchical Semantic Router                   │
│  Stage 1: Server-level routing (platform matching)     │
│  Stage 2: Tool-level routing (operation matching)      │
└──────────────────────┬──────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│            Tool Knowledge Base                          │
│  8 Servers × 40+ Tools                                 │
│  Organized by domain/platform                           │
└─────────────────────────────────────────────────────────┘

Core Components

1. Active Tool Agent (agent.py)

The agent is responsible for:

Task Analysis

def execute_task(self, task: str):
    # 1. Initialize with empty toolset
    self.available_tools = []

    # 2. Analyze task to identify capability needs
    # 3. Generate structured tool requests
    # 4. Iteratively discover and load tools
    # 5. Execute task with discovered tools

Tool Request Generation

Agent generates structured requests in this format:

<tool_request>
server: [platform/domain description]
tool: [operation description]
</tool_request>

Example:

<tool_request>
server: GitHub for repository operations
tool: search repositories by keywords and filters
</tool_request>

Iterative Discovery

The agent can make multiple tool requests as understanding evolves:

# Iteration 1: Basic need identified
Request: "GitHub repository access"
 Load: github_search_repos, github_list_issues

# Iteration 2: Additional need identified
Request: "File system operations for local storage"
 Load: fs_read_file, fs_write_file

# Iteration 3: Analysis need identified
Request: "Data visualization and statistics"
 Load: analytics_summarize, analytics_visualize

2. Semantic Router (semantic_router.py)

Implements two-stage hierarchical routing:

Stage 1: Server-Level Routing

Matches tool requests to relevant servers (platforms):

def _route_to_servers(self, request: str, top_k: int):
    # 1. Vectorize request using TF-IDF
    request_vector = self.server_vectorizer.transform([request])

    # 2. Calculate cosine similarity with all servers
    similarities = cosine_similarity(request_vector, self.server_embeddings)

    # 3. Return top-K servers by similarity
    top_indices = np.argsort(similarities)[::-1][:top_k]
    return [(self.servers[idx], similarities[idx]) for idx in top_indices]

Why This Works: - Reduces search space from all tools to tools in relevant servers - Platform/domain matching is coarse-grained and reliable - Example: "GitHub" request → GitHub server (not filesystem server)

Stage 2: Tool-Level Routing

Matches requests to specific tools within selected servers:

def _route_to_tools(self, server: ServerDefinition, request: str, top_k: int):
    # 1. Get server-specific vectorizer and embeddings
    vectorizer = self.tool_vectorizers[server.name]
    tool_embeddings = server._tool_embeddings

    # 2. Vectorize request
    request_vector = vectorizer.transform([request])

    # 3. Calculate similarity with tools in this server
    similarities = cosine_similarity(request_vector, tool_embeddings)

    # 4. Return top-K tools
    top_indices = np.argsort(similarities)[::-1][:top_k]
    return [(server.tools[idx], similarities[idx]) for idx in top_indices]

Why This Works: - Fine-grained matching within relevant domain - Tool descriptions are more specific than server descriptions - Example: "search repositories" → github_search_repos (not github_create_issue)

Score Combination

Final tool scores combine both stages:

combined_score = 0.3 * server_score + 0.7 * tool_score

Rationale: - Server score (30%): Ensures tool is from relevant domain - Tool score (70%): Prioritizes operation-level match - Weighted combination prevents cross-domain false positives

3. Tool Knowledge Base (tool_knowledge_base.py)

Organized hierarchically:

Knowledge Base
├── GitHub Server
│   ├── github_search_repos
│   ├── github_create_pr
│   ├── github_list_issues
│   ├── github_get_file
│   └── github_create_issue
├── Filesystem Server
│   ├── fs_read_file
│   ├── fs_write_file
│   ├── fs_list_directory
│   ├── fs_delete_file
│   └── fs_search_files
├── Database Server
│   ├── db_query
│   ├── db_insert
│   ├── db_update
│   ├── db_delete
│   └── db_schema
└── ... (5 more servers)

Design Principles:

  1. Hierarchical Organization: Tools grouped by platform/domain
  2. Rich Descriptions: Both servers and tools have semantic descriptions
  3. Standard Schema: OpenAI function calling format
  4. Extensible: Easy to add new servers/tools

4. Configuration (config.py)

Centralized configuration for all components:

# LLM Settings
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
OPENAI_BASE_URL = os.getenv("OPENAI_BASE_URL")
OPENAI_MODEL = os.getenv("OPENAI_MODEL", "gpt-5.6-luna")

# Routing Thresholds
SIMILARITY_THRESHOLD = 0.3  # Minimum similarity for match
TOP_K_SERVERS = 3           # Servers to search
TOP_K_TOOLS = 5             # Tools per server

# Agent Limits
MAX_TOOL_REQUESTS = 5       # Max discovery iterations

Active Discovery Flow

Detailed flow of active tool discovery:

┌─────────────────────────────────────────────────────────┐
│ Step 1: Task Submission                                 │
│ User: "Search for Python ML repos on GitHub"           │
└──────────────────────┬──────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ Step 2: Task Analysis (Agent)                          │
│ • Identifies need for repository search capability     │
│ • Current tools: None                                   │
│ • Decision: Request GitHub tools                        │
└──────────────────────┬──────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ Step 3: Tool Request Generation                        │
│ <tool_request>                                          │
│   server: GitHub for repository operations             │
│   tool: search repositories by keywords               │
│ </tool_request>                                        │
└──────────────────────┬──────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ Step 4: Semantic Routing                               │
│ Stage 1: Server routing                                 │
│   • github: 0.89 ✓                                     │
│   • filesystem: 0.12                                    │
│   • web: 0.24                                          │
│                                                         │
│ Stage 2: Tool routing (GitHub server)                  │
│   • github_search_repos: 0.94 ✓                       │
│   • github_list_issues: 0.45                           │
│   • github_get_file: 0.31                              │
└──────────────────────┬──────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ Step 5: Tool Loading                                    │
│ Loaded: [github_search_repos]                          │
│ Available tools count: 1                                │
└──────────────────────┬──────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ Step 6: Task Execution                                  │
│ Agent uses github_search_repos to complete task        │
└──────────────────────┬──────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ Step 7: Response                                        │
│ Results returned to user                                │
│ Metrics: 1 tool loaded, ~2000 tokens used              │
└─────────────────────────────────────────────────────────┘

Multi-Iteration Example

Complex task requiring multiple tool discovery iterations:

Task: "Clone repo, analyze code, visualize metrics, email report"

Iteration 1:
  Analysis: Need GitHub access
  Request: GitHub repository operations
  Loaded: github tools (2 tools)

Iteration 2:
  Analysis: Need file system for code storage
  Request: Filesystem operations
  Loaded: filesystem tools (3 tools total)

Iteration 3:
  Analysis: Need analytics for code analysis
  Request: Data analytics and visualization
  Loaded: analytics tools (5 tools total)

Iteration 4:
  Analysis: Need communication for email
  Request: Email communication
  Loaded: communication tools (6 tools total)

Execution: Use all 6 tools to complete task

Semantic Routing Algorithm

TF-IDF Vectorization

Tools and requests are converted to vectors using TF-IDF:

# Build vocabulary from all tool descriptions
vectorizer = TfidfVectorizer(stop_words='english')

# Server descriptions
server_docs = [f"{s.name} {s.description}" for s in servers]
server_matrix = vectorizer.fit_transform(server_docs)

# Tool descriptions (per server)
tool_docs = [f"{t.name} {t.description}" for t in tools]
tool_matrix = vectorizer.fit_transform(tool_docs)

What is TF-IDF?

  • TF (Term Frequency): How often a word appears in a document
  • IDF (Inverse Document Frequency): How rare a word is across documents
  • TF-IDF: Words that are frequent in a document but rare overall get high scores

Example:

Server: "GitHub repository management and version control"
Tool: "search repositories by keywords"
Request: "find GitHub repositories"

TF-IDF vectors capture semantic overlap:
- "repository" appears in all three → medium weight
- "GitHub" appears in server and request → strong match
- "search" appears in tool and request → strong match

Cosine Similarity

Measures similarity between vectors:

similarity = cosine_similarity(request_vector, tool_vector)
# Returns value between 0 (orthogonal) and 1 (identical)

Geometric Interpretation:

If vectors point in same direction → similar (score near 1)
If vectors are perpendicular → dissimilar (score near 0)

Example Scores:

Request: "search for repositories"
  • github_search_repos: 0.92 (strong match)
  • github_create_pr: 0.31 (weak match)
  • fs_read_file: 0.08 (no match)

Threshold Filtering

Tools below similarity threshold are filtered out:

SIMILARITY_THRESHOLD = 0.3

relevant_tools = [
    tool for tool, score in tool_scores 
    if score >= SIMILARITY_THRESHOLD
]

Why 0.3? - Balance between precision and recall - Captures semantic overlap without false positives - Empirically determined from testing

Comparison: Active vs Passive

Passive Tool Injection (Traditional)

class PassiveToolAgent:
    def __init__(self):
        # Load ALL tools at initialization
        self.all_tools = load_all_40_plus_tools()

    def execute_task(self, task):
        # Inject all tool schemas into prompt
        response = llm.complete(
            messages=[{"role": "user", "content": task}],
            tools=self.all_tools  # 40+ tool schemas
        )

Problems: 1. Massive Context: 30k-50k tokens just for tool schemas 2. Poor Scalability: Adding 10 tools increases every request by 5k tokens 3. Lost Autonomy: Agent selects from pre-defined set 4. Cognitive Overload: LLM must process irrelevant tools

Active Tool Discovery (MCP-Zero Approach)

class ActiveToolAgent:
    def __init__(self):
        # Start with empty toolset
        self.available_tools = []

    def execute_task(self, task):
        # Iteratively discover tools as needed
        while not task_complete:
            # Agent identifies capability gaps
            if need_more_tools:
                request = agent.generate_tool_request()
                new_tools = router.discover_tools(request)
                self.available_tools.extend(new_tools)
            else:
                # Use available tools
                execute_with_tools(self.available_tools)

Benefits: 1. Minimal Context: 2k-5k tokens (only needed tools) 2. Efficient Scaling: Adding 100 tools doesn't affect simple tasks 3. Preserved Autonomy: Agent controls capability acquisition 4. Focused Processing: LLM sees only relevant tools

Performance Comparison Table

Metric Passive Active Improvement
Initial Tools 40 0 N/A
Tools for Simple Task 40 2-3 92-95% reduction
Tokens (Simple Task) 45,000 2,500 94% reduction
Tokens (Complex Task) 50,000 8,000 84% reduction
Scalability O(n) O(k) k << n
Agent Autonomy Low High Qualitative

where: - n = total tools in ecosystem - k = tools needed for specific task

Performance Optimization

1. Embedding Precomputation

Tool embeddings are computed once at initialization:

def __init__(self, servers):
    # Precompute all embeddings
    self._build_server_index()
    self._build_tool_indices()

    # Query time: just cosine similarity
    # No re-vectorization needed

Benefit: O(1) query time instead of O(n) vectorization

Two-stage routing reduces complexity:

# Without hierarchy: Search all 40 tools
# Complexity: O(40) similarity comparisons

# With hierarchy: Search 8 servers, then top-3 servers
# Stage 1: O(8) server comparisons
# Stage 2: O(5) tool comparisons per server = O(15)
# Total: O(8 + 15) = O(23)

# Savings: 40 - 23 = 17 comparisons (42% reduction)

Scales Better: - 100 tools, 10 servers: 100 vs 35 comparisons (65% reduction) - 1000 tools, 20 servers: 1000 vs 120 comparisons (88% reduction)

3. Caching Potential

Future optimization: Cache routing results:

# Cache structure
routing_cache = {
    "search GitHub repos": ["github_search_repos", ...],
    "read local file": ["fs_read_file", ...]
}

# Cache hit: O(1) lookup
# Cache miss: Fall back to semantic routing

Design Decisions

Why TF-IDF Instead of Neural Embeddings?

Chosen: TF-IDF with cosine similarity

Alternatives Considered: - Sentence-BERT embeddings - OpenAI embeddings (text-embedding-ada-002)

Rationale: 1. Educational Clarity: TF-IDF is easier to understand and debug 2. No API Calls: Works offline without additional costs 3. Sufficient Performance: Tool descriptions are technical and keyword-rich 4. Fast: No model inference required

When Neural Embeddings Better: - Natural language queries (less technical) - Semantic nuances important - Large corpus with synonyms

Why Two-Stage Routing?

Alternatives Considered: - Flat search over all tools - Clustering-based search - Retrieval-augmented generation (RAG)

Rationale: 1. Matches Mental Model: Users think "GitHub" → "search repos" 2. Reduces False Positives: "search" alone might match wrong domain 3. Improves Precision: Server context narrows tool search 4. Scalable: Logarithmic complexity vs linear

Why Structured Requests?

Format:

<tool_request>
server: [domain]
tool: [operation]
</tool_request>

Alternatives Considered: - Free-form natural language - JSON format - Function calling

Rationale: 1. Explicit Structure: Server + tool decomposition matches routing stages 2. Easy Parsing: Simple string matching 3. LLM-Friendly: Clear format reduces ambiguity 4. Semantic Alignment: Request format matches knowledge base organization

Why Simulated Tool Execution?

Decision: Tools return simulated results instead of real execution

Rationale: 1. Educational Focus: Demonstrates discovery, not execution 2. Safety: No real API calls or file operations 3. Portability: Works without external dependencies 4. Simplicity: Focus on architecture, not integration

Future Enhancement: Connect to real APIs for production use

Why 3 Servers and 5 Tools?

Configuration:

TOP_K_SERVERS = 3
TOP_K_TOOLS = 5

Rationale: 1. Balance: Captures relevant tools without overwhelming context 2. Empirical: Tested on various tasks, 3×5=15 tools usually sufficient 3. Context Window: 15 tool schemas ≈ 3k-5k tokens (manageable) 4. Fallback: Can request more tools if initial set insufficient

Tuning Guidelines: - Simple tasks: Decrease to 2×3 = 6 tools - Complex tasks: Increase to 5×7 = 35 tools - Large ecosystems: Keep ratio, not absolute numbers

Conclusion

The active tool selection architecture demonstrates that:

  1. Hierarchical routing reduces search complexity while maintaining precision
  2. Active discovery preserves agent autonomy and scales efficiently
  3. Iterative extension allows toolchains to evolve with task understanding
  4. Semantic matching (even with simple TF-IDF) works well for tool discovery

This architecture represents a fundamental shift from passive tool injection to active capability acquisition, enabling agents to operate effectively in ecosystems with hundreds or thousands of available tools.