Today for AI

Designing a Simple Knowledge Base Agent: Architecture and Implementation

Intermediate level · Step-by-step guide · Cost: Low cost

Contents11 sections

Core Workflow of a Knowledge Base Agent

TEXT
           [ User Query ]
                 │
                 ▼
       [ Intent Recognition ] ──── (Non-KB Query) ──> [ Direct Reply ]
                 │
         (Knowledge Query)
                 ▼
       [ Query Expansion ] <────────────────────────────┐
                 │                                      │
        ┌────────┴────────┐                             │
        ▼                 ▼                             │
  [ Vector Search ]   [ Fulltext Search ]               │
        │                 │                             │
        └────────┬────────┘                             │
                 ▼                                      │
        [ LLM Reranking ]                               │
                 │                                      │
                 ▼                                      │
      [ Info Assessment ] ───── (Insufficient) ─────────┘
                 │                                 ( queryLoop )
            (Sufficient)
                 ▼
      [ Result Generation ]

1. Intent Recognition

Intent recognition intercepts non-knowledge queries to reduce system latency and retrieval costs.

  • Keyword and Regex Matching: Establishes a fast-track channel. Queries containing specific product names, error codes, or high-frequency phrases trigger regular expression rules to bypass LLM classification and directly enter the retrieval branch. This approach operates with zero latency.
  • LLM Classification: Handles complex or ambiguous expressions. The system calls a lightweight LLM (highly recommending the extremely cheap and fast DeepSeek V4 Flash) combined with Function Calling to output classification tags, distinguishing technical inquiries from casual conversation.

2. Retrieval Strategy

Relying solely on semantic search struggles with complex queries. The system introduces hybrid search and query expansion mechanisms.

  • Synonym Expansion and Query Rewrite: User inputs are often brief or contain typos. Before searching, an LLM rewrites the original query to include synonyms, industry jargon, and missing context, significantly broadening the recall scope.
  • Hybrid Search Implementation:
    • Vector RAG: Uses Embedding models to capture semantic relevance, handling conceptual questions and colloquial phrasing well.
    • Fulltext Search: Uses BM25-based exact matching. When users query specific API names, log codes, or unique serial numbers, full-text search compensates for vector retrieval's weakness with low-frequency terms.
  • LLM Reranking: After retrieving candidate chunks, basic scoring often fails to reflect logical fit. The system passes the candidate chunks to an LLM for secondary evaluation. The LLM filters out noisy or conflicting snippets, preserving only the most critical Top-K chunks for final synthesis.

3. Agent Loop Mechanism (queryLoop)

For multi-step logic or cross-document integration, a single retrieval pass is often inadequate. The queryLoop provides the agent with autonomous planning capabilities.

During the information assessment stage, the agent evaluates whether the reranked text can fully answer the query. If information is missing, it triggers the loop: extracting new search keywords, altering the search strategy, and re-entering the expansion stage. The loop terminates when a complete evidence chain is assembled or a maximum iteration limit (e.g., 3 loops) is reached.

4. Knowledge Base Maintenance and Component Selection

The structural quality of underlying data and component choices dictate the agent's output ceiling and operational cost.

Text Chunking Strategy

  • Semantic Level Chunking: Prioritizes logical chunking based on Markdown syntax (headers, lists, code blocks) to ensure the integrity of knowledge within a single chunk.
  • Fixed Size and Overlap: Applies hard boundaries to excessively long paragraphs (e.g., 500 tokens) while retaining a 50-100 token overlap area. This prevents critical information loss at chunk boundaries.

Component Selection Comparison

1. Large Language Models (LLM)

Stages like intent classification, query expansion, and reranking generate massive requests, making response speed and token costs critical.

ModelProsConsRecommendation
DeepSeek V4 FlashExtremely cheap, lightning-fast, ultra-high cost-effectivenessMay require fallback for extremely heavy logic reasoningHighly recommended for high-frequency stages (intent, rewrite, reranking) and general generation
GPT-4o-miniStable API, exceptionally rich toolchain ecosystemHigher cost compared to DeepSeekBackup option
Claude 3.5 SonnetExcellent at long-text logic assembly and complex instruction followingExpensive and slower generationOnly for complex final synthesis as a fallback

2. Search Engine

Hybrid search is standard for knowledge bases; selection balances development speed and concurrency.

EngineCore FeaturesProsCons
QdrantNative Dense + Sparse hybrid searchDesigned for high-concurrency vector/hybrid search, ready out of the boxRequires managing an independent cluster
PGVectorRelational DB + VectorReuses PostgreSQL architecture, easily joins business dataScaling ultra-large clusters is relatively difficult
ChromaLightweight vector searchExtremely lightweight, fast startup for local testingStruggles with production-level high concurrency and hybrid architectures

3. Embedding Models

Determines the quality of recall foundation and multilingual performance.

ModelDeploymentProsCons
BGE-m3Open-source localFree, data remains on-premise; excellent native multilingual and long-text supportRequires local GPU resources and maintenance
text-embedding-3-smallCloud APILowest integration cost, zero compute maintenanceClosed-source usage-based billing, potential data compliance concerns

5. Open Access and MCP Integration (SSE MCP)

To allow other systems or agent frameworks to directly invoke this knowledge base node, the system adopts the Model Context Protocol (MCP) standard at the interface layer and exposes it via the Server-Sent Events (SSE) transport layer.

  • SSE MCP Architecture Design: The primary entry point of the knowledge base agent is encapsulated as a standard MCP Tool (e.g., search_knowledge_base). Compared to local execution via standard I/O (stdio), SSE allows external clients to remotely invoke this tool over an HTTP persistent connection, providing the capability to cross network boundaries and support independent distributed deployment.
  • Communication and Invocation Chain: External applications establish a persistent listening stream via GET /sse. When a query is needed, they send JSON-RPC standard invocation requests to the POST /messages endpoint. The server takes over and independently executes intent recognition, RAG, reranking, and the queryLoop internally, ultimately pushing the final answer back to the external application through the SSE event stream.
  • Integration Benefits: Any external client compatible with the MCP protocol (such as Claude Desktop or various open-source agent orchestration engines) can directly mount this knowledge base as an "external brain" just by providing the SSE endpoint URL. This requires writing zero custom integration code, bringing integration costs to near zero.

Next step

Continue with related topics

Continue along the same topic.

Browse latest news