jianmu.rag¶
适用对象:知识库集成开发者 / 检索增强应用开发者
是否必读:按需
相关模块:jianmu.memory, jianmu.tool, jianmu.model
1. 模块职责¶
jianmu.rag 提供文档、嵌入、索引、检索、重排、知识管道和检索工具的公开 API。
这个模块覆盖的是“把外部知识变成可检索上下文”的完整链路。
2. 适合查什么¶
- 文档与检索结果:
Document、SearchResult - 搜索选项:
SearchOptions - 协议:
EmbedderProtocol、VectorStoreProtocol、RetrieverProtocol - 内置组件:
OpenAIEmbedder、DenseRetriever、HybridRetriever - store:
InMemoryStore、SQLiteStore、FaissStore、ChromaStore - ingest 组件:
DocumentLoaderProtocol、TextNormalizerProtocol、TextChunkerProtocol - 管道:
RAGPipeline、IngestPipeline、KnowledgeBase - 工具:
KnowledgeSearchTool
3. 使用建议¶
- 只想快速拼一个知识库时,先看
KnowledgeBase - 需要替换 embedding / store / retriever 时,再进入对应协议和实现层
- 想给 agent 暴露检索能力时,再看
KnowledgeSearchTool - 新的 ingest 设计更偏“组件策略拼装”,而不是把文档处理逻辑硬编码在单个函数里
4. 注意事项¶
- 不同 store / embedder 的依赖和运行要求不同
- 如果你只需要聊天历史或工作记忆,不必直接引入这个模块
Document、SearchResult、SearchOptions是 RAG 领域类型,应优先从jianmu.rag使用,而不是自己在业务层另造并行类型FaissStore与ChromaStore属于可选依赖路径;部署和测试环境需要提前安装对应依赖
5. 最小示例¶
from jianmu.rag import InMemoryStore, KnowledgeBase, SimpleEmbedder
store = InMemoryStore()
embedder = SimpleEmbedder()
kb = KnowledgeBase(store=store, embedder=embedder)
6. 常见入口¶
- 想定义文档与检索结果:看
Document、SearchResult、SearchOptions - 想选 embedder / store / retriever:看对应
*Protocol和内置实现 - 想做本地轻量存储:看
InMemoryStore/SQLiteStore - 想接外部向量后端:看
FaissStore/ChromaStore - 想直接组装知识库:看
KnowledgeBase - 想给 agent 暴露检索工具:看
KnowledgeSearchTool
7. API 参考¶
rag
¶
Knowledge (RAG) module: embeddings, ingestion, retrieval, stores, tools.
Document
dataclass
¶
Document(
id: str,
text: str,
metadata: Dict[str, Any] = dict(),
created_at: datetime = now_utc(),
embedding: Optional[List[float]] = None,
)
Represents a document with content, metadata, and optional embeddings.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
id |
str
|
Stable document identifier in the vector store. |
text |
str
|
Main textual content of the document. |
metadata |
Dict[str, Any]
|
Arbitrary metadata used for filtering and display. |
created_at |
datetime
|
Creation timestamp for the document record. |
embedding |
Optional[List[float]]
|
Optional cached dense vector embedding for the document. |
SearchResult
dataclass
¶
A single search result containing a document and its relevance score.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
document |
Document
|
Matched document returned by the retriever. |
score |
float
|
Relevance score assigned by the retriever or reranker. |
SearchOptions
dataclass
¶
SearchOptions(
k: int = 5,
mode: str = "hybrid",
alpha: float = _get_default_alpha(),
filter: Optional[Dict[str, Any]] = None,
)
Configuration options for conducting search queries.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
k |
int
|
Default number of results to retrieve. |
mode |
str
|
Retrieval mode such as |
alpha |
float
|
Hybrid-search weighting factor between sparse and dense scores. |
filter |
Optional[Dict[str, Any]]
|
Optional metadata filter applied during retrieval. |
EmbedderProtocol
¶
Bases: Protocol
Protocol for embedding text documents and queries into vectors.
embed_documents
¶
Embed multiple documents into dense vectors.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
texts
|
Sequence[str]
|
A sequence of document texts to embed. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[List[float]]
|
A list of dense vector embeddings for each text. |
embed_query
¶
Embed one query into a dense vector.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
text
|
str
|
The query string to embed. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[float]
|
A dense vector embedding for the query. |
VectorStoreProtocol
¶
Bases: Protocol
Protocol defining vector storage and search interfaces.
add
¶
Persist documents and return their ids.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
docs
|
List[Document]
|
A list of documents to add to the store. |
必需 |
embeddings
|
List[List[float]]
|
Parallel list of vector embeddings for the documents. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[str]
|
A list of document IDs that were persisted. |
源代码位于: jianmu/rag/base.py
search
¶
Search nearest documents for one query vector.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query_vec
|
List[float]
|
The query vector to search with. |
必需 |
k
|
int
|
The number of top results to return. |
5
|
filter
|
Optional[dict]
|
Metadata filter criteria. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[SearchResult]
|
A list of search results matching the criteria. |
源代码位于: jianmu/rag/base.py
delete
¶
Delete documents by id.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
ids
|
List[str]
|
A list of document IDs to delete. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
bool
|
True if deletion was successful, False otherwise. |
RetrieverProtocol
¶
Bases: Protocol
Protocol for retrieving documents based on queries.
search
¶
Retrieve relevant documents for one query.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query
|
str
|
The search query text. |
必需 |
k
|
int
|
The number of documents to retrieve. |
5
|
options
|
Optional[SearchOptions]
|
Optional search configuration overrides. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[SearchResult]
|
A list of retrieved search results. |
源代码位于: jianmu/rag/base.py
RerankerProtocol
¶
Bases: Protocol
Protocol for re-ranking search results to improve relevance.
rerank
¶
Reorder retrieved results for one query.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query
|
str
|
The search query text. |
必需 |
results
|
List[SearchResult]
|
A list of raw search results to re-rank. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[SearchResult]
|
A re-ranked list of search results. |
源代码位于: jianmu/rag/base.py
DocumentLoaderProtocol
¶
Bases: Protocol
Load raw text from one source path.
load
¶
Load text content from a source path.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
path
|
str
|
Source path to read from. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
str
|
Extracted plain text content. |
TextNormalizerProtocol
¶
Bases: Protocol
Normalize extracted raw text before chunking.
normalize
¶
Normalize raw text.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
text
|
str
|
Raw text to normalize. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
str
|
Normalized text content. |
TextChunkerProtocol
¶
Bases: Protocol
Split normalized text into ingestion chunks.
chunk
¶
Split text into chunks.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
text
|
str
|
Normalized text to split. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[str]
|
A list of text chunks. |
SimpleEmbedder
¶
Simple character-based embedder for testing.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
dim |
Dimensionality of generated embedding vectors. |
Configure embedding dimensionality for the fallback embedder.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
dim
|
int
|
Dimensionality of the generated vectors. |
64
|
源代码位于: jianmu/rag/embedder.py
embed_documents
¶
Embed multiple documents with the deterministic fallback embedder.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
texts
|
Sequence[str]
|
A sequence of document texts to embed. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[List[float]]
|
A list of dense vector embeddings for each document. |
源代码位于: jianmu/rag/embedder.py
embed_query
¶
Embed one query with the deterministic fallback embedder.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
text
|
str
|
Query text to embed. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[float]
|
A dense vector embedding for the query. |
OpenAIEmbedder
¶
OpenAIEmbedder(
api_key: Optional[str] = None,
base_url: Optional[str] = None,
model: Optional[str] = None,
)
OpenAI embeddings via official SDK.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
_client |
OpenAI SDK client or compatibility shim. |
|
_use_classic_client |
Whether the legacy SDK compatibility path is active. |
|
model |
Embedding model name used for requests. |
Initialize an OpenAI embedding client.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
api_key
|
Optional[str]
|
Optional API key. If omitted, fetched from environment. |
None
|
base_url
|
Optional[str]
|
Optional custom base URL for the API requests. |
None
|
model
|
Optional[str]
|
Optional embedding model name to use. |
None
|
引发:
| 类型 | 描述 |
|---|---|
RuntimeError
|
If no OpenAI API key is found. |
源代码位于: jianmu/rag/embedder.py
embed_documents
¶
Embed multiple documents with the OpenAI API.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
texts
|
Sequence[str]
|
A sequence of document texts to embed. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[List[float]]
|
A list of dense vector embeddings for each document. |
源代码位于: jianmu/rag/embedder.py
embed_query
¶
Embed one query with the OpenAI API.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
text
|
str
|
Query text to embed. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[float]
|
A dense vector embedding for the query. |
GeminiEmbedder
¶
GeminiEmbedder(
api_key: Optional[str] = None,
base_url: Optional[str] = None,
model: Optional[str] = None,
)
Google Gemini embeddings.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
_client |
Gemini SDK client used for embedding requests. |
|
model |
Embedding model name used for requests. |
Initialize a Gemini embedding client.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
api_key
|
Optional[str]
|
Optional API key. If omitted, fetched from environment. |
None
|
base_url
|
Optional[str]
|
Optional custom base URL for the API requests. |
None
|
model
|
Optional[str]
|
Optional embedding model name to use. |
None
|
引发:
| 类型 | 描述 |
|---|---|
RuntimeError
|
If genai package is missing or API key is not found. |
源代码位于: jianmu/rag/embedder.py
embed_documents
¶
Embed multiple documents with the Gemini API.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
texts
|
Sequence[str]
|
A sequence of document texts to embed. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[List[float]]
|
A list of dense vector embeddings for each document. |
源代码位于: jianmu/rag/embedder.py
embed_query
¶
Embed one query with the Gemini API.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
text
|
str
|
Query text to embed. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[float]
|
A dense vector embedding for the query. |
InMemoryStore
¶
Bases: VectorStoreProtocol
An in-memory implementation of VectorStoreProtocol with optional max size constraint.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
max_size |
Optional retention limit for stored documents. |
|
docs |
Dict[str, Document]
|
Document map keyed by document ID. |
order |
List[str]
|
Insertion-order list of document IDs. |
Initialize the in-memory vector store with optional retention.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
max_size
|
Optional[int]
|
Optional maximum number of documents to keep in memory. |
None
|
源代码位于: jianmu/rag/store.py
add
¶
Insert documents and embeddings into the in-memory store.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
docs
|
List[Document]
|
A list of Document objects to insert. |
必需 |
embeddings
|
List[List[float]]
|
Parallel list of vector embeddings for the documents. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[str]
|
A list of inserted document IDs. |
源代码位于: jianmu/rag/store.py
search
¶
Search the in-memory store with optional metadata filtering.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query_vec
|
List[float]
|
Dense query vector. |
必需 |
k
|
int
|
The number of nearest neighbors to retrieve. |
5
|
filter
|
Optional[dict]
|
Optional metadata equality filters. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[SearchResult]
|
A list of SearchResult objects sorted by descending similarity. |
源代码位于: jianmu/rag/store.py
list_all
¶
Return all documents in insertion order.
返回:
| 类型 | 描述 |
|---|---|
List[Document]
|
A list of all documents stored in memory. |
delete
¶
Delete documents from the in-memory store.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
ids
|
List[str]
|
List of document IDs to remove. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
bool
|
True if deletion succeeded. |
源代码位于: jianmu/rag/store.py
SQLiteStore
¶
Bases: VectorStoreProtocol
A SQLite-backed implementation of VectorStoreProtocol for persistent storage.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
path |
Filesystem path to the SQLite database file. |
Initialize the SQLite-backed vector store.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
path
|
Optional[str]
|
Optional filesystem path to the SQLite database file. |
None
|
源代码位于: jianmu/rag/store.py
add
¶
Insert or replace documents in the SQLite store.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
docs
|
List[Document]
|
A list of Document objects to store. |
必需 |
embeddings
|
List[List[float]]
|
Parallel list of vector embeddings for the documents. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[str]
|
A list of stored document IDs. |
源代码位于: jianmu/rag/store.py
search
¶
Search the SQLite store with optional metadata filtering.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query_vec
|
List[float]
|
Dense query vector. |
必需 |
k
|
int
|
The number of nearest neighbors to retrieve. |
5
|
filter
|
Optional[dict]
|
Optional metadata equality filters. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[SearchResult]
|
A list of SearchResult objects sorted by descending similarity. |
源代码位于: jianmu/rag/store.py
list_all
¶
Return all stored documents in insertion order.
返回:
| 类型 | 描述 |
|---|---|
List[Document]
|
A list of all documents stored in the database. |
源代码位于: jianmu/rag/store.py
delete
¶
Delete documents from the SQLite store.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
ids
|
List[str]
|
List of document IDs to remove. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
bool
|
True if deletion succeeded. |
源代码位于: jianmu/rag/store.py
FaissStore
¶
FaissStore(
dim: int,
index_factory: str = "Flat",
normalize: bool = True,
metric: Optional[str] = None,
)
Bases: VectorStoreProtocol
FAISS-backed vector store (optional dependency).
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
np |
Imported NumPy module used for vector preparation. |
|
faiss |
Imported FAISS module used for index operations. |
|
dim |
Embedding dimensionality expected by the index. |
|
normalize |
Whether vectors are unit-normalized before indexing/search. |
|
index |
Backing FAISS index instance. |
|
id_map |
Dict[int, Document]
|
Mapping from internal FAISS IDs to documents. |
next_id |
Next internal numeric ID assigned to inserted documents. |
Initialize a FAISS index and in-memory id map.
源代码位于: jianmu/rag/store.py
add
¶
Insert documents into the FAISS-backed store.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
docs
|
List[Document]
|
A list of Document objects to insert. |
必需 |
embeddings
|
List[List[float]]
|
Parallel list of vector embeddings for the documents. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[str]
|
A list of inserted document IDs. |
源代码位于: jianmu/rag/store.py
search
¶
Search the FAISS-backed store with optional metadata filtering.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query_vec
|
List[float]
|
Dense query vector. |
必需 |
k
|
int
|
The number of nearest neighbors to retrieve. |
5
|
filter
|
Optional[dict]
|
Optional metadata equality filters. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[SearchResult]
|
A list of SearchResult objects sorted by descending similarity. |
源代码位于: jianmu/rag/store.py
delete
¶
Delete documents by rebuilding the FAISS index without them.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
ids
|
List[str]
|
List of document IDs to remove. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
bool
|
True if deletion was successful, False otherwise. |
源代码位于: jianmu/rag/store.py
ChromaStore
¶
Bases: VectorStoreProtocol
Chroma vector store (optional dependency).
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
chromadb |
Imported ChromaDB module. |
|
client |
Persistent Chroma client instance. |
|
collection |
Backing Chroma collection handle. |
Initialize a persistent Chroma collection.
源代码位于: jianmu/rag/store.py
add
¶
Insert documents into the Chroma collection.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
docs
|
List[Document]
|
A list of Document objects to insert. |
必需 |
embeddings
|
List[List[float]]
|
Parallel list of vector embeddings for the documents. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[str]
|
A list of inserted document IDs. |
源代码位于: jianmu/rag/store.py
search
¶
Search the Chroma collection with optional metadata filtering.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query_vec
|
List[float]
|
Dense query vector. |
必需 |
k
|
int
|
The number of nearest neighbors to retrieve. |
5
|
filter
|
Optional[dict]
|
Optional metadata equality filters. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[SearchResult]
|
A list of SearchResult objects sorted by descending similarity. |
源代码位于: jianmu/rag/store.py
delete
¶
Delete documents from the Chroma collection.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
ids
|
List[str]
|
List of document IDs to remove. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
bool
|
True if deletion was successful, False otherwise. |
源代码位于: jianmu/rag/store.py
list_all
¶
Return all documents stored in the Chroma collection.
返回:
| 类型 | 描述 |
|---|---|
List[Document]
|
A list of all documents stored in the Chroma collection. |
源代码位于: jianmu/rag/store.py
BM25Retriever
¶
BM25 keyword-based retriever (lightweight implementation).
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
store |
Document store queried for keyword retrieval. |
|
k1 |
BM25 term-frequency saturation parameter. |
|
b |
BM25 document-length normalization parameter. |
Configure BM25 retrieval hyperparameters and backing store.
源代码位于: jianmu/rag/retriever.py
search
¶
Run BM25-style keyword retrieval over stored documents.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query
|
str
|
Input search query string. |
必需 |
k
|
int
|
The default number of results to retrieve. |
5
|
options
|
Optional[SearchOptions]
|
Optional SearchOptions overrides. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[SearchResult]
|
A list of SearchResult objects matching the keyword query. |
源代码位于: jianmu/rag/retriever.py
DenseRetriever
¶
Dense vector retriever using embeddings.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
store |
Vector store queried for candidate documents. |
|
embedder |
Embedding provider used to encode queries. |
Bind dense retrieval to a store and embedder.
源代码位于: jianmu/rag/retriever.py
search
¶
Run dense retrieval using the configured embedder and vector store.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query
|
str
|
Input search query string. |
必需 |
k
|
int
|
The default number of results to retrieve. |
5
|
options
|
Optional[SearchOptions]
|
Optional SearchOptions overrides. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[SearchResult]
|
A list of SearchResult objects matching the semantic query. |
源代码位于: jianmu/rag/retriever.py
HybridRetriever
¶
HybridRetriever(
store: VectorStoreProtocol,
embedder: Optional[EmbedderProtocol] = None,
alpha: Optional[float] = None,
)
Hybrid retriever combining BM25 and dense retrieval.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
store |
Shared document store used by both retrieval strategies. |
|
embedder |
Embedding provider used for dense query encoding. |
|
alpha |
Dense-score weight used for hybrid score blending. |
Configure hybrid sparse+dense retrieval.
源代码位于: jianmu/rag/retriever.py
search
¶
Blend dense and sparse retrieval scores into one ranking.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query
|
str
|
Input search query string. |
必需 |
k
|
int
|
The default number of results to retrieve. |
5
|
options
|
Optional[SearchOptions]
|
Optional SearchOptions overrides. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[SearchResult]
|
A list of SearchResult objects blending dense and sparse signals. |
源代码位于: jianmu/rag/retriever.py
PassthroughReranker
¶
Bases: RerankerProtocol
A dummy reranker that returns search results in their original order.
rerank
¶
Return retrieved results unchanged.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query
|
str
|
The search query text. |
必需 |
results
|
List[SearchResult]
|
A list of SearchResult objects to pass through. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[SearchResult]
|
The unmodified list of SearchResult objects. |
源代码位于: jianmu/rag/reranker.py
DefaultDocumentLoader
¶
Bases: DocumentLoaderProtocol
Load plain text from local text, PDF, and DOCX files.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
encoding |
Default text encoding for plain-text file reads. |
Configure the default text-file encoding.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
encoding
|
str
|
File character encoding for plain text files. |
'utf-8'
|
源代码位于: jianmu/rag/ingest.py
load
¶
Load plain text from a supported local file format.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
path
|
str
|
Path to the local file to load. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
str
|
The extracted plain text content. |
引发:
| 类型 | 描述 |
|---|---|
RuntimeError
|
If dependencies for PDF/DOCX are missing or for unsupported formats. |
源代码位于: jianmu/rag/ingest.py
WhitespaceNormalizer
¶
Bases: TextNormalizerProtocol
Normalize whitespace in raw extracted text.
normalize
¶
Normalize whitespace in raw extracted text.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
text
|
str
|
Raw text to normalize. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
str
|
The normalized text with collapsed whitespace. |
源代码位于: jianmu/rag/ingest.py
SimpleTextChunker
¶
Bases: TextChunkerProtocol
Split text into fixed-size overlapping character chunks.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
chunk_size |
Maximum character count emitted for each chunk. |
|
overlap |
Character overlap retained between adjacent chunks. |
Configure chunk size and overlap.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
chunk_size
|
int
|
Maximum character count for each chunk. |
500
|
overlap
|
int
|
Character count overlap between adjacent chunks. |
0
|
源代码位于: jianmu/rag/ingest.py
chunk
¶
Split text into overlapping chunks for ingestion.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
text
|
str
|
Raw or normalized text to split. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
List[str]
|
A list of split text chunks. |
源代码位于: jianmu/rag/ingest.py
RAGPipeline
¶
Pipeline for retrieval-augmented generation.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
retriever |
Retriever used to fetch candidate documents. |
|
reranker |
Optional reranker applied after retrieval. |
Bind retrieval and optional reranking into one pipeline.
源代码位于: jianmu/rag/pipeline.py
search
¶
Search and optionally rerank results.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query
|
str
|
Query search query text. |
必需 |
options
|
Optional[SearchOptions]
|
Optional SearchOptions overrides. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[SearchResult]
|
A list of retrieval search results. |
源代码位于: jianmu/rag/pipeline.py
IngestPipeline
¶
IngestPipeline(
embedder: Optional[EmbedderProtocol] = None,
store: Optional[VectorStoreProtocol] = None,
loader: DocumentLoaderProtocol | None = None,
normalizer: TextNormalizerProtocol | None = None,
chunker: TextChunkerProtocol | None = None,
)
Pipeline for ingesting documents into a vector store.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
embedder |
Embedding provider used to encode document chunks. |
|
store |
Vector store that persists documents and embeddings. |
|
loader |
Document loader used for file-based ingestion. |
|
normalizer |
Text normalizer applied before chunking. |
|
chunker |
Text chunker used to split normalized content. |
Bind ingestion helpers to an embedder, store, and ingestion strategies.
源代码位于: jianmu/rag/pipeline.py
ingest_text
¶
Chunk text and add to store.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
text
|
str
|
Input text content to ingest. |
必需 |
metadata
|
Optional[dict]
|
Optional dictionary of metadata to associate with the chunks. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[str]
|
A list of successfully ingested document IDs. |
源代码位于: jianmu/rag/pipeline.py
ingest_file
¶
Load file, chunk, and add to store.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
path
|
str
|
Target filesystem path to the file. |
必需 |
metadata
|
Optional[dict]
|
Optional dictionary of metadata to associate with the chunks. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[str]
|
A list of successfully ingested document IDs. |
源代码位于: jianmu/rag/pipeline.py
KnowledgeBase
¶
KnowledgeBase(
store: Optional[VectorStoreProtocol] = None,
embedder: Optional[EmbedderProtocol] = None,
retriever: Optional[RetrieverProtocol] = None,
reranker: Optional[RerankerProtocol] = None,
loader: DocumentLoaderProtocol | None = None,
normalizer: TextNormalizerProtocol | None = None,
chunker: TextChunkerProtocol | None = None,
)
Convenience facade combining store, embedder, retriever, and pipelines.
Example
kb = KnowledgeBase() kb.add("Hello world") kb.add("Python is great") results = kb.search("hello")
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
store |
Backing vector store for persisted documents. |
|
embedder |
Embedding provider used for add and ingest operations. |
|
retriever |
Retriever used to resolve search requests. |
|
reranker |
Optional reranker applied during search. |
|
_ingest |
Internal ingestion pipeline facade. |
|
_rag |
Internal retrieval pipeline facade. |
Assemble a convenience knowledge-base facade from core components.
源代码位于: jianmu/rag/pipeline.py
add
¶
Add a single document.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
text
|
str
|
Input document text. |
必需 |
metadata
|
Optional[dict]
|
Optional metadata dict. |
None
|
返回:
| 类型 | 描述 |
|---|---|
str
|
The generated ID of the added document. |
源代码位于: jianmu/rag/pipeline.py
ingest_text
¶
Chunk and ingest text.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
text
|
str
|
Input text content to ingest. |
必需 |
metadata
|
Optional[dict]
|
Optional dictionary of metadata to associate with the chunks. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[str]
|
A list of successfully ingested document IDs. |
源代码位于: jianmu/rag/pipeline.py
ingest_file
¶
Load and ingest a file.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
path
|
str
|
Target filesystem path to the file. |
必需 |
metadata
|
Optional[dict]
|
Optional dictionary of metadata to associate with the chunks. |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[str]
|
A list of successfully ingested document IDs. |
源代码位于: jianmu/rag/pipeline.py
search
¶
Search the knowledge base.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query
|
str
|
Query search query text. |
必需 |
k
|
int
|
Number of nearest neighbors to retrieve. |
5
|
mode
|
str
|
Search modality ("dense", "sparse", or "hybrid"). |
'hybrid'
|
返回:
| 类型 | 描述 |
|---|---|
List[SearchResult]
|
A list of SearchResult objects matching the query. |
源代码位于: jianmu/rag/pipeline.py
as_tool
¶
Return a KnowledgeSearchTool for this knowledge base.
返回:
| 类型 | 描述 |
|---|---|
'KnowledgeSearchTool'
|
The resulting |
源代码位于: jianmu/rag/pipeline.py
KnowledgeSearchTool
¶
Bases: Tool
Search external knowledge via a RAG pipeline.
属性:
| 名称 | 类型 | 描述 |
|---|---|---|
pipeline |
RAG pipeline used to execute search requests. |
Create a knowledge-search tool backed by one RAG pipeline.
源代码位于: jianmu/rag/tools.py
run
async
¶
Search the knowledge base and render results as plain text.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
query
|
str | None
|
The search query text. |
None
|
k
|
int
|
The number of search results to retrieve. |
3
|
mode
|
str | None
|
Search mode: semantic, keyword, or hybrid. |
None
|
**kwargs
|
Any
|
Additional keyword arguments. |
{}
|
返回:
| 类型 | 描述 |
|---|---|
str
|
A newline-separated string of the retrieved document texts. |
源代码位于: jianmu/rag/tools.py
resolve_embedder
¶
resolve_embedder(
preference: Optional[List[str]] = None,
model: Optional[str] = None,
base_url: Optional[str] = None,
) -> EmbedderProtocol
Auto-resolve embedder based on available API keys.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
preference
|
Optional[List[str]]
|
Optional ordered list of preferred providers. |
None
|
model
|
Optional[str]
|
Optional embedding model override. |
None
|
base_url
|
Optional[str]
|
Optional custom base URL override. |
None
|
返回:
| 类型 | 描述 |
|---|---|
EmbedderProtocol
|
An instantiated embedding provider conforming to EmbedderProtocol. |
源代码位于: jianmu/rag/embedder.py
simple_embedding
¶
Simple character-based embedding for testing/fallback.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
text
|
str
|
Input string to embed. |
必需 |
dim
|
int
|
Dimensionality of the output embedding vector. |
64
|
返回:
| 类型 | 描述 |
|---|---|
List[float]
|
A list of floats representing the normalized embedding vector. |
源代码位于: jianmu/rag/embedder.py
cosine_similarity
¶
Compute cosine similarity for two dense vectors.
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
a
|
List[float]
|
The first input vector. |
必需 |
b
|
List[float]
|
The second input vector. |
必需 |
返回:
| 类型 | 描述 |
|---|---|
float
|
The cosine similarity score between the two vectors. |