Contents11 sections
Core Workflow of a Knowledge Base Agent
[ User Query ]
│
▼
[ Intent Recognition ] ──── (Non-KB Query) ──> [ Direct Reply ]
│
(Knowledge Query)
▼
[ Query Expansion ] <────────────────────────────┐
│ │
┌────────┴────────┐ │
▼ ▼ │
[ Vector Search ] [ Fulltext Search ] │
│ │ │
└────────┬────────┘ │
▼ │
[ LLM Reranking ] │
│ │
▼ │
[ Info Assessment ] ───── (Insufficient) ─────────┘
│ ( queryLoop )
(Sufficient)
▼
[ Result Generation ]1. Intent Recognition
Intent recognition intercepts non-knowledge queries to reduce system latency and retrieval costs.
- Keyword and Regex Matching: Establishes a fast-track channel. Queries containing specific product names, error codes, or high-frequency phrases trigger regular expression rules to bypass LLM classification and directly enter the retrieval branch. This approach operates with zero latency.
- LLM Classification: Handles complex or ambiguous expressions. The system calls a lightweight LLM (highly recommending the extremely cheap and fast DeepSeek V4 Flash) combined with Function Calling to output classification tags, distinguishing technical inquiries from casual conversation.
2. Retrieval Strategy
Relying solely on semantic search struggles with complex queries. The system introduces hybrid search and query expansion mechanisms.
- Synonym Expansion and Query Rewrite: User inputs are often brief or contain typos. Before searching, an LLM rewrites the original query to include synonyms, industry jargon, and missing context, significantly broadening the recall scope.
- Hybrid Search Implementation:
- Vector RAG: Uses Embedding models to capture semantic relevance, handling conceptual questions and colloquial phrasing well.
- Fulltext Search: Uses BM25-based exact matching. When users query specific API names, log codes, or unique serial numbers, full-text search compensates for vector retrieval's weakness with low-frequency terms.
- LLM Reranking: After retrieving candidate chunks, basic scoring often fails to reflect logical fit. The system passes the candidate chunks to an LLM for secondary evaluation. The LLM filters out noisy or conflicting snippets, preserving only the most critical Top-K chunks for final synthesis.
3. Agent Loop Mechanism (queryLoop)
For multi-step logic or cross-document integration, a single retrieval pass is often inadequate. The queryLoop provides the agent with autonomous planning capabilities.
During the information assessment stage, the agent evaluates whether the reranked text can fully answer the query. If information is missing, it triggers the loop: extracting new search keywords, altering the search strategy, and re-entering the expansion stage. The loop terminates when a complete evidence chain is assembled or a maximum iteration limit (e.g., 3 loops) is reached.
4. Knowledge Base Maintenance and Component Selection
The structural quality of underlying data and component choices dictate the agent's output ceiling and operational cost.
Text Chunking Strategy
- Semantic Level Chunking: Prioritizes logical chunking based on Markdown syntax (headers, lists, code blocks) to ensure the integrity of knowledge within a single chunk.
- Fixed Size and Overlap: Applies hard boundaries to excessively long paragraphs (e.g., 500 tokens) while retaining a 50-100 token overlap area. This prevents critical information loss at chunk boundaries.
Component Selection Comparison
1. Large Language Models (LLM)
Stages like intent classification, query expansion, and reranking generate massive requests, making response speed and token costs critical.
| Model | Pros | Cons | Recommendation |
|---|---|---|---|
| DeepSeek V4 Flash | Extremely cheap, lightning-fast, ultra-high cost-effectiveness | May require fallback for extremely heavy logic reasoning | Highly recommended for high-frequency stages (intent, rewrite, reranking) and general generation |
| GPT-4o-mini | Stable API, exceptionally rich toolchain ecosystem | Higher cost compared to DeepSeek | Backup option |
| Claude 3.5 Sonnet | Excellent at long-text logic assembly and complex instruction following | Expensive and slower generation | Only for complex final synthesis as a fallback |
2. Search Engine
Hybrid search is standard for knowledge bases; selection balances development speed and concurrency.
| Engine | Core Features | Pros | Cons |
|---|---|---|---|
| Qdrant | Native Dense + Sparse hybrid search | Designed for high-concurrency vector/hybrid search, ready out of the box | Requires managing an independent cluster |
| PGVector | Relational DB + Vector | Reuses PostgreSQL architecture, easily joins business data | Scaling ultra-large clusters is relatively difficult |
| Chroma | Lightweight vector search | Extremely lightweight, fast startup for local testing | Struggles with production-level high concurrency and hybrid architectures |
3. Embedding Models
Determines the quality of recall foundation and multilingual performance.
| Model | Deployment | Pros | Cons |
|---|---|---|---|
| BGE-m3 | Open-source local | Free, data remains on-premise; excellent native multilingual and long-text support | Requires local GPU resources and maintenance |
| text-embedding-3-small | Cloud API | Lowest integration cost, zero compute maintenance | Closed-source usage-based billing, potential data compliance concerns |
5. Open Access and MCP Integration (SSE MCP)
To allow other systems or agent frameworks to directly invoke this knowledge base node, the system adopts the Model Context Protocol (MCP) standard at the interface layer and exposes it via the Server-Sent Events (SSE) transport layer.
- SSE MCP Architecture Design:
The primary entry point of the knowledge base agent is encapsulated as a standard MCP Tool (e.g.,
search_knowledge_base). Compared to local execution via standard I/O (stdio), SSE allows external clients to remotely invoke this tool over an HTTP persistent connection, providing the capability to cross network boundaries and support independent distributed deployment. - Communication and Invocation Chain:
External applications establish a persistent listening stream via
GET /sse. When a query is needed, they send JSON-RPC standard invocation requests to thePOST /messagesendpoint. The server takes over and independently executes intent recognition, RAG, reranking, and thequeryLoopinternally, ultimately pushing the final answer back to the external application through the SSE event stream. - Integration Benefits: Any external client compatible with the MCP protocol (such as Claude Desktop or various open-source agent orchestration engines) can directly mount this knowledge base as an "external brain" just by providing the SSE endpoint URL. This requires writing zero custom integration code, bringing integration costs to near zero.