跳转至

jianmu.rag

适用对象:知识库集成开发者 / 检索增强应用开发者 是否必读:按需 相关模块:jianmu.memory, jianmu.tool, jianmu.model

1. 模块职责

jianmu.rag 提供文档、嵌入、索引、检索、重排、知识管道和检索工具的公开 API。

这个模块覆盖的是“把外部知识变成可检索上下文”的完整链路。

2. 适合查什么

  • 文档与检索结果:Document、SearchResult
  • 搜索选项:SearchOptions
  • 协议:EmbedderProtocol、VectorStoreProtocol、RetrieverProtocol
  • 内置组件:OpenAIEmbedder、DenseRetriever、HybridRetriever
  • store:InMemoryStore、SQLiteStore、FaissStore、ChromaStore
  • ingest 组件:DocumentLoaderProtocol、TextNormalizerProtocol、TextChunkerProtocol
  • 管道:RAGPipeline、IngestPipeline、KnowledgeBase
  • 工具:KnowledgeSearchTool

3. 使用建议

  • 只想快速拼一个知识库时,先看 KnowledgeBase
  • 需要替换 embedding / store / retriever 时,再进入对应协议和实现层
  • 想给 agent 暴露检索能力时,再看 KnowledgeSearchTool
  • 新的 ingest 设计更偏“组件策略拼装”,而不是把文档处理逻辑硬编码在单个函数里

4. 注意事项

  • 不同 store / embedder 的依赖和运行要求不同
  • 如果你只需要聊天历史或工作记忆,不必直接引入这个模块
  • Document、SearchResult、SearchOptions 是 RAG 领域类型,应优先从 jianmu.rag 使用,而不是自己在业务层另造并行类型
  • FaissStore 与 ChromaStore 属于可选依赖路径;部署和测试环境需要提前安装对应依赖

5. 最小示例

from jianmu.rag import InMemoryStore, KnowledgeBase, SimpleEmbedder

store = InMemoryStore()
embedder = SimpleEmbedder()
kb = KnowledgeBase(store=store, embedder=embedder)

6. 常见入口

  • 想定义文档与检索结果:看 Document、SearchResult、SearchOptions
  • 想选 embedder / store / retriever:看对应 *Protocol 和内置实现
  • 想做本地轻量存储:看 InMemoryStore / SQLiteStore
  • 想接外部向量后端:看 FaissStore / ChromaStore
  • 想直接组装知识库:看 KnowledgeBase
  • 想给 agent 暴露检索工具:看 KnowledgeSearchTool

7. API 参考

rag

Knowledge (RAG) module: embeddings, ingestion, retrieval, stores, tools.

Document dataclass

Document(
    id: str,
    text: str,
    metadata: Dict[str, Any] = dict(),
    created_at: datetime = now_utc(),
    embedding: Optional[List[float]] = None,
)

Represents a document with content, metadata, and optional embeddings.

属性:

名称 类型 描述
id str

Stable document identifier in the vector store.

text str

Main textual content of the document.

metadata Dict[str, Any]

Arbitrary metadata used for filtering and display.

created_at datetime

Creation timestamp for the document record.

embedding Optional[List[float]]

Optional cached dense vector embedding for the document.

SearchResult dataclass

SearchResult(document: Document, score: float)

A single search result containing a document and its relevance score.

属性:

名称 类型 描述
document Document

Matched document returned by the retriever.

score float

Relevance score assigned by the retriever or reranker.

SearchOptions dataclass

SearchOptions(
    k: int = 5,
    mode: str = "hybrid",
    alpha: float = _get_default_alpha(),
    filter: Optional[Dict[str, Any]] = None,
)

Configuration options for conducting search queries.

属性:

名称 类型 描述
k int

Default number of results to retrieve.

mode str

Retrieval mode such as vector, keyword, or hybrid.

alpha float

Hybrid-search weighting factor between sparse and dense scores.

filter Optional[Dict[str, Any]]

Optional metadata filter applied during retrieval.

EmbedderProtocol

Bases: Protocol

Protocol for embedding text documents and queries into vectors.

embed_documents

embed_documents(texts: Sequence[str]) -> List[List[float]]

Embed multiple documents into dense vectors.

参数:

名称 类型 描述 默认
texts Sequence[str]

A sequence of document texts to embed.

必需

返回:

类型 描述
List[List[float]]

A list of dense vector embeddings for each text.

源代码位于: jianmu/rag/base.py
def embed_documents(self, texts: Sequence[str]) -> List[List[float]]:
    """Embed multiple documents into dense vectors.

    Args:
        texts: A sequence of document texts to embed.

    Returns:
        A list of dense vector embeddings for each text.
    """
    ...

embed_query

embed_query(text: str) -> List[float]

Embed one query into a dense vector.

参数:

名称 类型 描述 默认
text str

The query string to embed.

必需

返回:

类型 描述
List[float]

A dense vector embedding for the query.

源代码位于: jianmu/rag/base.py
def embed_query(self, text: str) -> List[float]:
    """Embed one query into a dense vector.

    Args:
        text: The query string to embed.

    Returns:
        A dense vector embedding for the query.
    """
    ...

VectorStoreProtocol

Bases: Protocol

Protocol defining vector storage and search interfaces.

add

add(
    docs: List[Document], embeddings: List[List[float]]
) -> List[str]

Persist documents and return their ids.

参数:

名称 类型 描述 默认
docs List[Document]

A list of documents to add to the store.

必需
embeddings List[List[float]]

Parallel list of vector embeddings for the documents.

必需

返回:

类型 描述
List[str]

A list of document IDs that were persisted.

源代码位于: jianmu/rag/base.py
def add(self, docs: List[Document], embeddings: List[List[float]]) -> List[str]:
    """Persist documents and return their ids.

    Args:
        docs: A list of documents to add to the store.
        embeddings: Parallel list of vector embeddings for the documents.

    Returns:
        A list of document IDs that were persisted.
    """
    ...

search

search(
    query_vec: List[float],
    k: int = 5,
    filter: Optional[dict] = None,
) -> List[SearchResult]

Search nearest documents for one query vector.

参数:

名称 类型 描述 默认
query_vec List[float]

The query vector to search with.

必需
k int

The number of top results to return.

5
filter Optional[dict]

Metadata filter criteria.

None

返回:

类型 描述
List[SearchResult]

A list of search results matching the criteria.

源代码位于: jianmu/rag/base.py
def search(self, query_vec: List[float], k: int = 5, filter: Optional[dict] = None) -> List[SearchResult]:
    """Search nearest documents for one query vector.

    Args:
        query_vec: The query vector to search with.
        k: The number of top results to return.
        filter: Metadata filter criteria.

    Returns:
        A list of search results matching the criteria.
    """
    ...

delete

delete(ids: List[str]) -> bool

Delete documents by id.

参数:

名称 类型 描述 默认
ids List[str]

A list of document IDs to delete.

必需

返回:

类型 描述
bool

True if deletion was successful, False otherwise.

源代码位于: jianmu/rag/base.py
def delete(self, ids: List[str]) -> bool:
    """Delete documents by id.

    Args:
        ids: A list of document IDs to delete.

    Returns:
        True if deletion was successful, False otherwise.
    """
    ...

RetrieverProtocol

Bases: Protocol

Protocol for retrieving documents based on queries.

search

search(
    query: str,
    k: int = 5,
    options: Optional[SearchOptions] = None,
) -> List[SearchResult]

Retrieve relevant documents for one query.

参数:

名称 类型 描述 默认
query str

The search query text.

必需
k int

The number of documents to retrieve.

5
options Optional[SearchOptions]

Optional search configuration overrides.

None

返回:

类型 描述
List[SearchResult]

A list of retrieved search results.

源代码位于: jianmu/rag/base.py
def search(self, query: str, k: int = 5, options: Optional[SearchOptions] = None) -> List[SearchResult]:
    """Retrieve relevant documents for one query.

    Args:
        query: The search query text.
        k: The number of documents to retrieve.
        options: Optional search configuration overrides.

    Returns:
        A list of retrieved search results.
    """
    ...

RerankerProtocol

Bases: Protocol

Protocol for re-ranking search results to improve relevance.

rerank

rerank(
    query: str, results: List[SearchResult]
) -> List[SearchResult]

Reorder retrieved results for one query.

参数:

名称 类型 描述 默认
query str

The search query text.

必需
results List[SearchResult]

A list of raw search results to re-rank.

必需

返回:

类型 描述
List[SearchResult]

A re-ranked list of search results.

源代码位于: jianmu/rag/base.py
def rerank(self, query: str, results: List[SearchResult]) -> List[SearchResult]:
    """Reorder retrieved results for one query.

    Args:
        query: The search query text.
        results: A list of raw search results to re-rank.

    Returns:
        A re-ranked list of search results.
    """
    ...

DocumentLoaderProtocol

Bases: Protocol

Load raw text from one source path.

load

load(path: str) -> str

Load text content from a source path.

参数:

名称 类型 描述 默认
path str

Source path to read from.

必需

返回:

类型 描述
str

Extracted plain text content.

源代码位于: jianmu/rag/base.py
def load(self, path: str) -> str:
    """Load text content from a source path.

    Args:
        path: Source path to read from.

    Returns:
        Extracted plain text content.
    """
    ...

TextNormalizerProtocol

Bases: Protocol

Normalize extracted raw text before chunking.

normalize

normalize(text: str) -> str

Normalize raw text.

参数:

名称 类型 描述 默认
text str

Raw text to normalize.

必需

返回:

类型 描述
str

Normalized text content.

源代码位于: jianmu/rag/base.py
def normalize(self, text: str) -> str:
    """Normalize raw text.

    Args:
        text: Raw text to normalize.

    Returns:
        Normalized text content.
    """
    ...

TextChunkerProtocol

Bases: Protocol

Split normalized text into ingestion chunks.

chunk

chunk(text: str) -> List[str]

Split text into chunks.

参数:

名称 类型 描述 默认
text str

Normalized text to split.

必需

返回:

类型 描述
List[str]

A list of text chunks.

源代码位于: jianmu/rag/base.py
def chunk(self, text: str) -> List[str]:
    """Split text into chunks.

    Args:
        text: Normalized text to split.

    Returns:
        A list of text chunks.
    """
    ...

SimpleEmbedder

SimpleEmbedder(dim: int = 64)

Simple character-based embedder for testing.

属性:

名称 类型 描述
dim

Dimensionality of generated embedding vectors.

Configure embedding dimensionality for the fallback embedder.

参数:

名称 类型 描述 默认
dim int

Dimensionality of the generated vectors.

64
源代码位于: jianmu/rag/embedder.py
def __init__(self, dim: int = 64):
    """Configure embedding dimensionality for the fallback embedder.

    Args:
        dim: Dimensionality of the generated vectors.
    """
    self.dim = dim

embed_documents

embed_documents(texts: Sequence[str]) -> List[List[float]]

Embed multiple documents with the deterministic fallback embedder.

参数:

名称 类型 描述 默认
texts Sequence[str]

A sequence of document texts to embed.

必需

返回:

类型 描述
List[List[float]]

A list of dense vector embeddings for each document.

源代码位于: jianmu/rag/embedder.py
def embed_documents(self, texts: Sequence[str]) -> List[List[float]]:
    """Embed multiple documents with the deterministic fallback embedder.

    Args:
        texts: A sequence of document texts to embed.

    Returns:
        A list of dense vector embeddings for each document.
    """
    return [simple_embedding(t, dim=self.dim) for t in texts]

embed_query

embed_query(text: str) -> List[float]

Embed one query with the deterministic fallback embedder.

参数:

名称 类型 描述 默认
text str

Query text to embed.

必需

返回:

类型 描述
List[float]

A dense vector embedding for the query.

源代码位于: jianmu/rag/embedder.py
def embed_query(self, text: str) -> List[float]:
    """Embed one query with the deterministic fallback embedder.

    Args:
        text: Query text to embed.

    Returns:
        A dense vector embedding for the query.
    """
    return simple_embedding(text, dim=self.dim)

OpenAIEmbedder

OpenAIEmbedder(
    api_key: Optional[str] = None,
    base_url: Optional[str] = None,
    model: Optional[str] = None,
)

OpenAI embeddings via official SDK.

属性:

名称 类型 描述
_client

OpenAI SDK client or compatibility shim.

_use_classic_client

Whether the legacy SDK compatibility path is active.

model

Embedding model name used for requests.

Initialize an OpenAI embedding client.

参数:

名称 类型 描述 默认
api_key Optional[str]

Optional API key. If omitted, fetched from environment.

None
base_url Optional[str]

Optional custom base URL for the API requests.

None
model Optional[str]

Optional embedding model name to use.

None

引发:

类型 描述
RuntimeError

If no OpenAI API key is found.

源代码位于: jianmu/rag/embedder.py
def __init__(
    self,
    api_key: Optional[str] = None,
    base_url: Optional[str] = None,
    model: Optional[str] = None,
):
    """Initialize an OpenAI embedding client.

    Args:
        api_key: Optional API key. If omitted, fetched from environment.
        base_url: Optional custom base URL for the API requests.
        model: Optional embedding model name to use.

    Raises:
        RuntimeError: If no OpenAI API key is found.
    """
    cfg = get_config().rag.embedding
    key = api_key or os.getenv("OPENAI_API_KEY") or os.getenv("API_KEY")
    if not key:
        raise RuntimeError("OpenAI API key not found (OPENAI_API_KEY/API_KEY)")

    resolved_base_url = (
        base_url
        or cfg.base_url
        or os.getenv("OPENAI_BASE_URL")
        or os.getenv("OPENAI_API_BASE")
        or os.getenv("BASE_URL")
    )

    try:
        from openai import OpenAI
        self._client = OpenAI(api_key=key, base_url=resolved_base_url)
        self._use_classic_client = False
    except Exception:
        import openai
        openai.api_key = key
        if resolved_base_url:
            openai.base_url = resolved_base_url
        self._client = openai
        self._use_classic_client = True

    self.model = model or cfg.model or os.getenv("OPENAI_EMBED_MODEL") or os.getenv("EMBEDDING_MODEL") or "text-embedding-3-small"

embed_documents

embed_documents(texts: Sequence[str]) -> List[List[float]]

Embed multiple documents with the OpenAI API.

参数:

名称 类型 描述 默认
texts Sequence[str]

A sequence of document texts to embed.

必需

返回:

类型 描述
List[List[float]]

A list of dense vector embeddings for each document.

源代码位于: jianmu/rag/embedder.py
def embed_documents(self, texts: Sequence[str]) -> List[List[float]]:
    """Embed multiple documents with the OpenAI API.

    Args:
        texts: A sequence of document texts to embed.

    Returns:
        A list of dense vector embeddings for each document.
    """
    return [self._embed_single(t) for t in texts]

embed_query

embed_query(text: str) -> List[float]

Embed one query with the OpenAI API.

参数:

名称 类型 描述 默认
text str

Query text to embed.

必需

返回:

类型 描述
List[float]

A dense vector embedding for the query.

源代码位于: jianmu/rag/embedder.py
def embed_query(self, text: str) -> List[float]:
    """Embed one query with the OpenAI API.

    Args:
        text: Query text to embed.

    Returns:
        A dense vector embedding for the query.
    """
    return self._embed_single(text)

GeminiEmbedder

GeminiEmbedder(
    api_key: Optional[str] = None,
    base_url: Optional[str] = None,
    model: Optional[str] = None,
)

Google Gemini embeddings.

属性:

名称 类型 描述
_client

Gemini SDK client used for embedding requests.

model

Embedding model name used for requests.

Initialize a Gemini embedding client.

参数:

名称 类型 描述 默认
api_key Optional[str]

Optional API key. If omitted, fetched from environment.

None
base_url Optional[str]

Optional custom base URL for the API requests.

None
model Optional[str]

Optional embedding model name to use.

None

引发:

类型 描述
RuntimeError

If genai package is missing or API key is not found.

源代码位于: jianmu/rag/embedder.py
def __init__(
    self,
    api_key: Optional[str] = None,
    base_url: Optional[str] = None,
    model: Optional[str] = None,
):
    """Initialize a Gemini embedding client.

    Args:
        api_key: Optional API key. If omitted, fetched from environment.
        base_url: Optional custom base URL for the API requests.
        model: Optional embedding model name to use.

    Raises:
        RuntimeError: If genai package is missing or API key is not found.
    """
    try:
        from google import genai
    except ImportError as e:
        raise RuntimeError(
            "google-genai package not installed. Run: pip install google-genai"
        ) from e

    cfg = get_config().rag.embedding
    key = api_key or os.getenv("GOOGLE_API_KEY") or os.getenv("GEMINI_API_KEY")
    if not key:
        raise RuntimeError("Gemini API key not found (GOOGLE_API_KEY/GEMINI_API_KEY)")

    resolved_base_url = base_url or cfg.base_url or os.getenv("EMBEDDING_BASE_URL") or os.getenv("BASE_URL")
    http_options = {"base_url": resolved_base_url} if resolved_base_url else None

    self._client = genai.Client(api_key=key, http_options=http_options)
    self.model = model or cfg.model or os.getenv("GEMINI_EMBED_MODEL") or os.getenv("EMBEDDING_MODEL") or "text-embedding-004"

embed_documents

embed_documents(texts: Sequence[str]) -> List[List[float]]

Embed multiple documents with the Gemini API.

参数:

名称 类型 描述 默认
texts Sequence[str]

A sequence of document texts to embed.

必需

返回:

类型 描述
List[List[float]]

A list of dense vector embeddings for each document.

源代码位于: jianmu/rag/embedder.py
def embed_documents(self, texts: Sequence[str]) -> List[List[float]]:
    """Embed multiple documents with the Gemini API.

    Args:
        texts: A sequence of document texts to embed.

    Returns:
        A list of dense vector embeddings for each document.
    """
    return [self._embed_single(t) for t in texts]

embed_query

embed_query(text: str) -> List[float]

Embed one query with the Gemini API.

参数:

名称 类型 描述 默认
text str

Query text to embed.

必需

返回:

类型 描述
List[float]

A dense vector embedding for the query.

源代码位于: jianmu/rag/embedder.py
def embed_query(self, text: str) -> List[float]:
    """Embed one query with the Gemini API.

    Args:
        text: Query text to embed.

    Returns:
        A dense vector embedding for the query.
    """
    return self._embed_single(text)

InMemoryStore

InMemoryStore(max_size: Optional[int] = None)

Bases: VectorStoreProtocol

An in-memory implementation of VectorStoreProtocol with optional max size constraint.

属性:

名称 类型 描述
max_size

Optional retention limit for stored documents.

docs Dict[str, Document]

Document map keyed by document ID.

order List[str]

Insertion-order list of document IDs.

Initialize the in-memory vector store with optional retention.

参数:

名称 类型 描述 默认
max_size Optional[int]

Optional maximum number of documents to keep in memory.

None
源代码位于: jianmu/rag/store.py
def __init__(self, max_size: Optional[int] = None):
    """Initialize the in-memory vector store with optional retention.

    Args:
        max_size: Optional maximum number of documents to keep in memory.
    """
    self.max_size = max_size
    self.docs: Dict[str, Document] = {}
    self.order: List[str] = []

add

add(
    docs: List[Document], embeddings: List[List[float]]
) -> List[str]

Insert documents and embeddings into the in-memory store.

参数:

名称 类型 描述 默认
docs List[Document]

A list of Document objects to insert.

必需
embeddings List[List[float]]

Parallel list of vector embeddings for the documents.

必需

返回:

类型 描述
List[str]

A list of inserted document IDs.

源代码位于: jianmu/rag/store.py
def add(self, docs: List[Document], embeddings: List[List[float]]) -> List[str]:
    """Insert documents and embeddings into the in-memory store.

    Args:
        docs: A list of Document objects to insert.
        embeddings: Parallel list of vector embeddings for the documents.

    Returns:
        A list of inserted document IDs.
    """
    ids: List[str] = []
    for doc, emb in zip(docs, embeddings):
        doc.embedding = emb
        self.docs[doc.id] = doc
        self.order.append(doc.id)
        ids.append(doc.id)
    if self.max_size is not None and len(self.order) > self.max_size:
        overflow = len(self.order) - self.max_size
        for _ in range(overflow):
            old_id = self.order.pop(0)
            self.docs.pop(old_id, None)
    return ids

search

search(
    query_vec: List[float],
    k: int = 5,
    filter: Optional[dict] = None,
) -> List[SearchResult]

Search the in-memory store with optional metadata filtering.

参数:

名称 类型 描述 默认
query_vec List[float]

Dense query vector.

必需
k int

The number of nearest neighbors to retrieve.

5
filter Optional[dict]

Optional metadata equality filters.

None

返回:

类型 描述
List[SearchResult]

A list of SearchResult objects sorted by descending similarity.

源代码位于: jianmu/rag/store.py
def search(self, query_vec: List[float], k: int = 5, filter: Optional[dict] = None) -> List[SearchResult]:
    """Search the in-memory store with optional metadata filtering.

    Args:
        query_vec: Dense query vector.
        k: The number of nearest neighbors to retrieve.
        filter: Optional metadata equality filters.

    Returns:
        A list of SearchResult objects sorted by descending similarity.
    """
    ordered = [self.docs[i] for i in self.order if i in self.docs and _metadata_match(self.docs[i].metadata or {}, filter)]
    if not ordered:
        return []
    if not query_vec:
        sliced = ordered[-k:] if k else ordered
        return [SearchResult(document=d, score=1.0) for d in sliced]

    scored = []
    for doc in ordered:
        score = cosine_similarity(query_vec, doc.embedding or [])
        scored.append((score, doc))
    scored.sort(key=lambda x: x[0], reverse=True)
    top = scored[:k] if k else scored
    return [SearchResult(document=d, score=s) for s, d in top]

list_all

list_all() -> List[Document]

Return all documents in insertion order.

返回:

类型 描述
List[Document]

A list of all documents stored in memory.

源代码位于: jianmu/rag/store.py
def list_all(self) -> List[Document]:
    """Return all documents in insertion order.

    Returns:
        A list of all documents stored in memory.
    """
    return [self.docs[i] for i in self.order if i in self.docs]

delete

delete(ids: List[str]) -> bool

Delete documents from the in-memory store.

参数:

名称 类型 描述 默认
ids List[str]

List of document IDs to remove.

必需

返回:

类型 描述
bool

True if deletion succeeded.

源代码位于: jianmu/rag/store.py
def delete(self, ids: List[str]) -> bool:
    """Delete documents from the in-memory store.

    Args:
        ids: List of document IDs to remove.

    Returns:
        True if deletion succeeded.
    """
    for i in ids:
        self.docs.pop(i, None)
        self.order = [x for x in self.order if x != i]
    return True

SQLiteStore

SQLiteStore(path: Optional[str] = None)

Bases: VectorStoreProtocol

A SQLite-backed implementation of VectorStoreProtocol for persistent storage.

属性:

名称 类型 描述
path

Filesystem path to the SQLite database file.

Initialize the SQLite-backed vector store.

参数:

名称 类型 描述 默认
path Optional[str]

Optional filesystem path to the SQLite database file.

None
源代码位于: jianmu/rag/store.py
def __init__(self, path: Optional[str] = None):
    """Initialize the SQLite-backed vector store.

    Args:
        path: Optional filesystem path to the SQLite database file.
    """
    if path is None:
        from jianmu.config.loader import get_config
        path = get_config().paths.knowledge_db
    self.path = Path(path)
    self.path.parent.mkdir(parents=True, exist_ok=True)
    self._init_db()

add

add(
    docs: List[Document], embeddings: List[List[float]]
) -> List[str]

Insert or replace documents in the SQLite store.

参数:

名称 类型 描述 默认
docs List[Document]

A list of Document objects to store.

必需
embeddings List[List[float]]

Parallel list of vector embeddings for the documents.

必需

返回:

类型 描述
List[str]

A list of stored document IDs.

源代码位于: jianmu/rag/store.py
def add(self, docs: List[Document], embeddings: List[List[float]]) -> List[str]:
    """Insert or replace documents in the SQLite store.

    Args:
        docs: A list of Document objects to store.
        embeddings: Parallel list of vector embeddings for the documents.

    Returns:
        A list of stored document IDs.
    """
    with sqlite3.connect(self.path) as conn:
        for doc, emb in zip(docs, embeddings):
            conn.execute(
                "INSERT OR REPLACE INTO documents (id, text, metadata, embedding) VALUES (?, ?, ?, ?)",
                (doc.id, doc.text, json.dumps(doc.metadata), json.dumps(emb)),
            )
    return [d.id for d in docs]

search

search(
    query_vec: List[float],
    k: int = 5,
    filter: Optional[dict] = None,
) -> List[SearchResult]

Search the SQLite store with optional metadata filtering.

参数:

名称 类型 描述 默认
query_vec List[float]

Dense query vector.

必需
k int

The number of nearest neighbors to retrieve.

5
filter Optional[dict]

Optional metadata equality filters.

None

返回:

类型 描述
List[SearchResult]

A list of SearchResult objects sorted by descending similarity.

源代码位于: jianmu/rag/store.py
def search(self, query_vec: List[float], k: int = 5, filter: Optional[dict] = None) -> List[SearchResult]:
    """Search the SQLite store with optional metadata filtering.

    Args:
        query_vec: Dense query vector.
        k: The number of nearest neighbors to retrieve.
        filter: Optional metadata equality filters.

    Returns:
        A list of SearchResult objects sorted by descending similarity.
    """
    with sqlite3.connect(self.path) as conn:
        rows = conn.execute("SELECT id, text, metadata, embedding FROM documents").fetchall()

    docs: List[Document] = []
    for row in rows:
        meta = json.loads(row[2] or "{}")
        if not _metadata_match(meta, filter):
            continue
        emb = json.loads(row[3]) if row[3] else []
        docs.append(Document(id=row[0], text=row[1], metadata=meta, embedding=emb))

    if not docs:
        return []
    if not query_vec:
        sliced = docs[-k:] if k else docs
        return [SearchResult(document=d, score=1.0) for d in sliced]

    scored = []
    for doc in docs:
        score = cosine_similarity(query_vec, doc.embedding or [])
        scored.append((score, doc))
    scored.sort(key=lambda x: x[0], reverse=True)
    top = scored[:k] if k else scored
    return [SearchResult(document=d, score=s) for s, d in top]

list_all

list_all() -> List[Document]

Return all stored documents in insertion order.

返回:

类型 描述
List[Document]

A list of all documents stored in the database.

源代码位于: jianmu/rag/store.py
def list_all(self) -> List[Document]:
    """Return all stored documents in insertion order.

    Returns:
        A list of all documents stored in the database.
    """
    with sqlite3.connect(self.path) as conn:
        rows = conn.execute("SELECT id, text, metadata, embedding FROM documents ORDER BY rowid").fetchall()
    docs: List[Document] = []
    for row in rows:
        meta = json.loads(row[2] or "{}")
        emb = json.loads(row[3]) if row[3] else []
        docs.append(Document(id=row[0], text=row[1], metadata=meta, embedding=emb))
    return docs

delete

delete(ids: List[str]) -> bool

Delete documents from the SQLite store.

参数:

名称 类型 描述 默认
ids List[str]

List of document IDs to remove.

必需

返回:

类型 描述
bool

True if deletion succeeded.

源代码位于: jianmu/rag/store.py
def delete(self, ids: List[str]) -> bool:
    """Delete documents from the SQLite store.

    Args:
        ids: List of document IDs to remove.

    Returns:
        True if deletion succeeded.
    """
    with sqlite3.connect(self.path) as conn:
        conn.executemany("DELETE FROM documents WHERE id = ?", [(i,) for i in ids])
    return True

FaissStore

FaissStore(
    dim: int,
    index_factory: str = "Flat",
    normalize: bool = True,
    metric: Optional[str] = None,
)

Bases: VectorStoreProtocol

FAISS-backed vector store (optional dependency).

属性:

名称 类型 描述
np

Imported NumPy module used for vector preparation.

faiss

Imported FAISS module used for index operations.

dim

Embedding dimensionality expected by the index.

normalize

Whether vectors are unit-normalized before indexing/search.

index

Backing FAISS index instance.

id_map Dict[int, Document]

Mapping from internal FAISS IDs to documents.

next_id

Next internal numeric ID assigned to inserted documents.

Initialize a FAISS index and in-memory id map.

源代码位于: jianmu/rag/store.py
def __init__(self, dim: int, index_factory: str = "Flat", normalize: bool = True, metric: Optional[str] = None):
    """Initialize a FAISS index and in-memory id map."""
    try:
        import faiss  # type: ignore
    except Exception as e:
        raise RuntimeError("faiss is required for FaissStore. pip install faiss-cpu") from e

    import numpy as np  # type: ignore

    self.np = np
    self.faiss = faiss
    self.dim = dim
    self.normalize = normalize
    metric_type = faiss.METRIC_INNER_PRODUCT if (metric or "ip").lower().startswith("ip") else faiss.METRIC_L2
    base_index = faiss.index_factory(dim, index_factory, metric_type)
    self.index = faiss.IndexIDMap2(base_index)
    self.id_map: Dict[int, Document] = {}
    self.next_id = 0

add

add(
    docs: List[Document], embeddings: List[List[float]]
) -> List[str]

Insert documents into the FAISS-backed store.

参数:

名称 类型 描述 默认
docs List[Document]

A list of Document objects to insert.

必需
embeddings List[List[float]]

Parallel list of vector embeddings for the documents.

必需

返回:

类型 描述
List[str]

A list of inserted document IDs.

源代码位于: jianmu/rag/store.py
def add(self, docs: List[Document], embeddings: List[List[float]]) -> List[str]:
    """Insert documents into the FAISS-backed store.

    Args:
        docs: A list of Document objects to insert.
        embeddings: Parallel list of vector embeddings for the documents.

    Returns:
        A list of inserted document IDs.
    """
    if not embeddings:
        return []
    vecs = self._prep_vecs(embeddings)
    ids = []
    idx_ids = []
    for doc, vec in zip(docs, vecs):
        doc.embedding = vec.tolist()
        ids.append(doc.id)
        idx_ids.append(self.next_id)
        self.id_map[self.next_id] = doc
        self.next_id += 1
    self.index.add_with_ids(vecs, self.np.array(idx_ids))
    return ids

search

search(
    query_vec: List[float],
    k: int = 5,
    filter: Optional[dict] = None,
) -> List[SearchResult]

Search the FAISS-backed store with optional metadata filtering.

参数:

名称 类型 描述 默认
query_vec List[float]

Dense query vector.

必需
k int

The number of nearest neighbors to retrieve.

5
filter Optional[dict]

Optional metadata equality filters.

None

返回:

类型 描述
List[SearchResult]

A list of SearchResult objects sorted by descending similarity.

源代码位于: jianmu/rag/store.py
def search(self, query_vec: List[float], k: int = 5, filter: Optional[dict] = None) -> List[SearchResult]:
    """Search the FAISS-backed store with optional metadata filtering.

    Args:
        query_vec: Dense query vector.
        k: The number of nearest neighbors to retrieve.
        filter: Optional metadata equality filters.

    Returns:
        A list of SearchResult objects sorted by descending similarity.
    """
    if not query_vec:
        return []
    q = self._prep_vecs([query_vec])
    scores, idxs = self.index.search(q, k)
    results: List[SearchResult] = []
    for score, idx in zip(scores[0], idxs[0]):
        if idx == -1:
            continue
        doc = self.id_map.get(int(idx))
        if doc is None:
            continue
        if filter:
            meta = doc.metadata or {}
            match = all(meta.get(kf) == vf for kf, vf in filter.items())
            if not match:
                continue
        results.append(SearchResult(document=doc, score=float(score)))
    return results

delete

delete(ids: List[str]) -> bool

Delete documents by rebuilding the FAISS index without them.

参数:

名称 类型 描述 默认
ids List[str]

List of document IDs to remove.

必需

返回:

类型 描述
bool

True if deletion was successful, False otherwise.

源代码位于: jianmu/rag/store.py
def delete(self, ids: List[str]) -> bool:
    """Delete documents by rebuilding the FAISS index without them.

    Args:
        ids: List of document IDs to remove.

    Returns:
        True if deletion was successful, False otherwise.
    """
    # Rebuild index without deleted docs
    keep_docs = [doc for doc in self.id_map.values() if doc.id not in ids]
    if len(keep_docs) == len(self.id_map):
        return False
    self.id_map = {}
    self.index.reset()
    self.next_id = 0
    embs = [d.embedding or [] for d in keep_docs]
    if any(len(e) == 0 for e in embs):
        return True  # cannot rebuild without embeddings; keep index empty but state consistent
    self.add(keep_docs, embs)
    return True

list_all

list_all() -> List[Document]

Return all documents currently indexed by FAISS.

返回:

类型 描述
List[Document]

A list of all documents currently indexed in FAISS.

源代码位于: jianmu/rag/store.py
def list_all(self) -> List[Document]:
    """Return all documents currently indexed by FAISS.

    Returns:
        A list of all documents currently indexed in FAISS.
    """
    return list(self.id_map.values())

ChromaStore

ChromaStore(
    path: str = ".chroma", collection: str = "jianmu"
)

Bases: VectorStoreProtocol

Chroma vector store (optional dependency).

属性:

名称 类型 描述
chromadb

Imported ChromaDB module.

client

Persistent Chroma client instance.

collection

Backing Chroma collection handle.

Initialize a persistent Chroma collection.

源代码位于: jianmu/rag/store.py
def __init__(self, path: str = ".chroma", collection: str = "jianmu"):
    """Initialize a persistent Chroma collection."""
    try:
        import chromadb  # type: ignore
    except Exception as e:
        raise RuntimeError("chromadb is required for ChromaStore. pip install chromadb") from e
    self.chromadb = chromadb
    self.client = chromadb.PersistentClient(path=path)
    self.collection = self.client.get_or_create_collection(collection)

add

add(
    docs: List[Document], embeddings: List[List[float]]
) -> List[str]

Insert documents into the Chroma collection.

参数:

名称 类型 描述 默认
docs List[Document]

A list of Document objects to insert.

必需
embeddings List[List[float]]

Parallel list of vector embeddings for the documents.

必需

返回:

类型 描述
List[str]

A list of inserted document IDs.

源代码位于: jianmu/rag/store.py
def add(self, docs: List[Document], embeddings: List[List[float]]) -> List[str]:
    """Insert documents into the Chroma collection.

    Args:
        docs: A list of Document objects to insert.
        embeddings: Parallel list of vector embeddings for the documents.

    Returns:
        A list of inserted document IDs.
    """
    ids = [d.id for d in docs]
    metadatas = [d.metadata for d in docs]
    texts = [d.text for d in docs]
    self.collection.add(ids=ids, embeddings=embeddings, documents=texts, metadatas=metadatas)
    for d, emb in zip(docs, embeddings):
        d.embedding = emb
    return ids

search

search(
    query_vec: List[float],
    k: int = 5,
    filter: Optional[dict] = None,
) -> List[SearchResult]

Search the Chroma collection with optional metadata filtering.

参数:

名称 类型 描述 默认
query_vec List[float]

Dense query vector.

必需
k int

The number of nearest neighbors to retrieve.

5
filter Optional[dict]

Optional metadata equality filters.

None

返回:

类型 描述
List[SearchResult]

A list of SearchResult objects sorted by descending similarity.

源代码位于: jianmu/rag/store.py
def search(self, query_vec: List[float], k: int = 5, filter: Optional[dict] = None) -> List[SearchResult]:
    """Search the Chroma collection with optional metadata filtering.

    Args:
        query_vec: Dense query vector.
        k: The number of nearest neighbors to retrieve.
        filter: Optional metadata equality filters.

    Returns:
        A list of SearchResult objects sorted by descending similarity.
    """
    if not query_vec:
        return []
    res = self.collection.query(
        query_embeddings=[query_vec],
        n_results=k,
        where=filter if filter else None,
        include=["documents", "metadatas", "distances"],
    )
    results: List[SearchResult] = []
    ids = res.get("ids", [[]])[0]
    docs = res.get("documents", [[]])[0]
    metas = res.get("metadatas", [[]])[0]
    dists = res.get("distances", [[]])[0]
    for i, text in enumerate(docs):
        doc = Document(id=str(ids[i]), text=text or "", metadata=metas[i] or {})
        results.append(SearchResult(document=doc, score=float(1.0 - (dists[i] or 0.0))))
    return results

delete

delete(ids: List[str]) -> bool

Delete documents from the Chroma collection.

参数:

名称 类型 描述 默认
ids List[str]

List of document IDs to remove.

必需

返回:

类型 描述
bool

True if deletion was successful, False otherwise.

源代码位于: jianmu/rag/store.py
def delete(self, ids: List[str]) -> bool:
    """Delete documents from the Chroma collection.

    Args:
        ids: List of document IDs to remove.

    Returns:
        True if deletion was successful, False otherwise.
    """
    self.collection.delete(ids=ids)
    return True

list_all

list_all() -> List[Document]

Return all documents stored in the Chroma collection.

返回:

类型 描述
List[Document]

A list of all documents stored in the Chroma collection.

源代码位于: jianmu/rag/store.py
def list_all(self) -> List[Document]:
    """Return all documents stored in the Chroma collection.

    Returns:
        A list of all documents stored in the Chroma collection.
    """
    res = self.collection.get(include=["documents", "metadatas", "embeddings"])
    docs: List[Document] = []
    metadatas = res.get("metadatas")
    embeddings = res.get("embeddings")
    ids = res.get("ids") or []
    for i, text in enumerate(res.get("documents", []) or []):
        meta = metadatas[i] if metadatas is not None else {}
        emb = embeddings[i] if embeddings is not None else []
        if hasattr(emb, "tolist"):
            emb = emb.tolist()
        docs.append(Document(id=str(ids[i]), text=text or "", metadata=meta or {}, embedding=emb))
    return docs

BM25Retriever

BM25Retriever(
    store: VectorStoreProtocol,
    k1: float = 1.5,
    b: float = 0.75,
)

BM25 keyword-based retriever (lightweight implementation).

属性:

名称 类型 描述
store

Document store queried for keyword retrieval.

k1

BM25 term-frequency saturation parameter.

b

BM25 document-length normalization parameter.

Configure BM25 retrieval hyperparameters and backing store.

源代码位于: jianmu/rag/retriever.py
def __init__(self, store: VectorStoreProtocol, k1: float = 1.5, b: float = 0.75):
    """Configure BM25 retrieval hyperparameters and backing store."""
    self.store = store
    self.k1 = k1
    self.b = b

search

search(
    query: str,
    k: int = 5,
    options: Optional[SearchOptions] = None,
) -> List[SearchResult]

Run BM25-style keyword retrieval over stored documents.

参数:

名称 类型 描述 默认
query str

Input search query string.

必需
k int

The default number of results to retrieve.

5
options Optional[SearchOptions]

Optional SearchOptions overrides.

None

返回:

类型 描述
List[SearchResult]

A list of SearchResult objects matching the keyword query.

源代码位于: jianmu/rag/retriever.py
def search(self, query: str, k: int = 5, options: Optional[SearchOptions] = None) -> List[SearchResult]:
    """Run BM25-style keyword retrieval over stored documents.

    Args:
        query: Input search query string.
        k: The default number of results to retrieve.
        options: Optional SearchOptions overrides.

    Returns:
        A list of SearchResult objects matching the keyword query.
    """
    options = options or SearchOptions(k=k, mode="keyword")
    records = self.store.list_all() if hasattr(self.store, "list_all") else []
    if not records:
        return []

    query_tokens = _tokenize(query)
    if not query_tokens:
        return []

    # Build corpus stats
    doc_tokens_list = [_tokenize(doc.text) for doc in records]
    doc_lengths = [len(toks) or 1 for toks in doc_tokens_list]
    avg_dl = sum(doc_lengths) / len(doc_lengths)

    # document frequencies
    df = Counter()
    for toks in doc_tokens_list:
        df.update(set(toks))

    scores: List[float] = []
    N = len(records)
    for toks, dl in zip(doc_tokens_list, doc_lengths):
        tf = Counter(toks)
        score = 0.0
        for term in query_tokens:
            freq = tf.get(term, 0)
            if freq == 0:
                continue
            idf = math.log((N - df[term] + 0.5) / (df[term] + 0.5) + 1)
            denom = freq + self.k1 * (1 - self.b + self.b * dl / avg_dl)
            score += idf * (freq * (self.k1 + 1)) / denom
        scores.append(score)

    paired = sorted(zip(records, scores), key=lambda x: x[1], reverse=True)
    return [SearchResult(document=doc, score=s) for doc, s in paired[: options.k]]

DenseRetriever

DenseRetriever(
    store: VectorStoreProtocol, embedder: EmbedderProtocol
)

Dense vector retriever using embeddings.

属性:

名称 类型 描述
store

Vector store queried for candidate documents.

embedder

Embedding provider used to encode queries.

Bind dense retrieval to a store and embedder.

源代码位于: jianmu/rag/retriever.py
def __init__(self, store: VectorStoreProtocol, embedder: EmbedderProtocol):
    """Bind dense retrieval to a store and embedder."""
    self.store = store
    self.embedder = embedder

search

search(
    query: str,
    k: int = 5,
    options: Optional[SearchOptions] = None,
) -> List[SearchResult]

Run dense retrieval using the configured embedder and vector store.

参数:

名称 类型 描述 默认
query str

Input search query string.

必需
k int

The default number of results to retrieve.

5
options Optional[SearchOptions]

Optional SearchOptions overrides.

None

返回:

类型 描述
List[SearchResult]

A list of SearchResult objects matching the semantic query.

源代码位于: jianmu/rag/retriever.py
def search(self, query: str, k: int = 5, options: Optional[SearchOptions] = None) -> List[SearchResult]:
    """Run dense retrieval using the configured embedder and vector store.

    Args:
        query: Input search query string.
        k: The default number of results to retrieve.
        options: Optional SearchOptions overrides.

    Returns:
        A list of SearchResult objects matching the semantic query.
    """
    options = options or SearchOptions(k=k, mode="semantic")
    q_vec = _coerce_embedding(self.embedder.embed_query(query)) or []
    q_vec = _normalize_vector(q_vec) if q_vec else []

    # Prefer store-side vector search if provided
    try:
        results = self.store.search(q_vec, k=options.k, filter=options.filter)
        if results is not None:
            return results
    except Exception:
        pass

    records = self.store.list_all() if hasattr(self.store, "list_all") else []
    if not records:
        return []

    scores = []
    for doc in records:
        emb = doc.embedding or []
        scores.append(cosine_similarity(q_vec, emb) if emb else 0.0)

    paired = sorted(zip(records, scores), key=lambda x: x[1], reverse=True)
    return [SearchResult(document=doc, score=s) for doc, s in paired[: options.k]]

HybridRetriever

HybridRetriever(
    store: VectorStoreProtocol,
    embedder: Optional[EmbedderProtocol] = None,
    alpha: Optional[float] = None,
)

Hybrid retriever combining BM25 and dense retrieval.

属性:

名称 类型 描述
store

Shared document store used by both retrieval strategies.

embedder

Embedding provider used for dense query encoding.

alpha

Dense-score weight used for hybrid score blending.

Configure hybrid sparse+dense retrieval.

源代码位于: jianmu/rag/retriever.py
def __init__(self, store: VectorStoreProtocol, embedder: Optional[EmbedderProtocol] = None, alpha: Optional[float] = None):
    """Configure hybrid sparse+dense retrieval."""
    from jianmu.config.loader import get_config
    self.store = store
    self.embedder = embedder or SimpleEmbedder()
    self.alpha = alpha if alpha is not None else get_config().limits.rerank_alpha

search

search(
    query: str,
    k: int = 5,
    options: Optional[SearchOptions] = None,
) -> List[SearchResult]

Blend dense and sparse retrieval scores into one ranking.

参数:

名称 类型 描述 默认
query str

Input search query string.

必需
k int

The default number of results to retrieve.

5
options Optional[SearchOptions]

Optional SearchOptions overrides.

None

返回:

类型 描述
List[SearchResult]

A list of SearchResult objects blending dense and sparse signals.

源代码位于: jianmu/rag/retriever.py
def search(self, query: str, k: int = 5, options: Optional[SearchOptions] = None) -> List[SearchResult]:
    """Blend dense and sparse retrieval scores into one ranking.

    Args:
        query: Input search query string.
        k: The default number of results to retrieve.
        options: Optional SearchOptions overrides.

    Returns:
        A list of SearchResult objects blending dense and sparse signals.
    """
    opts = options or SearchOptions(k=k)
    mode = (opts.mode or "hybrid").lower()
    alpha = opts.alpha if opts.alpha is not None else self.alpha

    # Run both retrievers
    bm25_results = []
    dense_results = []
    if mode in ("keyword", "hybrid"):
        bm25_results = BM25Retriever(self.store).search(query, k=opts.k * 2, options=SearchOptions(k=opts.k))
    if mode in ("semantic", "hybrid"):
        dense_results = DenseRetriever(self.store, self.embedder).search(query, k=opts.k * 2, options=SearchOptions(k=opts.k))

    # Normalize scores
    def _normalize(res: List[SearchResult]) -> dict:
        """Scale result scores into the [0, 1] range."""
        if not res:
            return {}
        max_s = max(r.score for r in res) or 1.0
        return {r.document.id: r.score / max_s for r in res}

    bm25_map = _normalize(bm25_results)
    dense_map = _normalize(dense_results)

    # Merge scores
    ids = set(bm25_map.keys()) | set(dense_map.keys())
    merged: List[SearchResult] = []
    records_map = {}
    for r in bm25_results + dense_results:
        records_map[r.document.id] = r.document
    for _id in ids:
        b = bm25_map.get(_id, 0.0)
        d = dense_map.get(_id, 0.0)
        if mode == "keyword":
            score = b
        elif mode == "semantic":
            score = d
        else:
            score = alpha * d + (1 - alpha) * b
        merged.append(SearchResult(document=records_map[_id], score=score))

    merged.sort(key=lambda x: x.score, reverse=True)
    return merged[: opts.k]

PassthroughReranker

Bases: RerankerProtocol

A dummy reranker that returns search results in their original order.

rerank

rerank(
    query: str, results: List[SearchResult]
) -> List[SearchResult]

Return retrieved results unchanged.

参数:

名称 类型 描述 默认
query str

The search query text.

必需
results List[SearchResult]

A list of SearchResult objects to pass through.

必需

返回:

类型 描述
List[SearchResult]

The unmodified list of SearchResult objects.

源代码位于: jianmu/rag/reranker.py
def rerank(self, query: str, results: List[SearchResult]) -> List[SearchResult]:
    """Return retrieved results unchanged.

    Args:
        query: The search query text.
        results: A list of SearchResult objects to pass through.

    Returns:
        The unmodified list of SearchResult objects.
    """
    return results

DefaultDocumentLoader

DefaultDocumentLoader(encoding: str = 'utf-8')

Bases: DocumentLoaderProtocol

Load plain text from local text, PDF, and DOCX files.

属性:

名称 类型 描述
encoding

Default text encoding for plain-text file reads.

Configure the default text-file encoding.

参数:

名称 类型 描述 默认
encoding str

File character encoding for plain text files.

'utf-8'
源代码位于: jianmu/rag/ingest.py
def __init__(self, encoding: str = "utf-8"):
    """Configure the default text-file encoding.

    Args:
        encoding: File character encoding for plain text files.
    """
    self.encoding = encoding

load

load(path: str) -> str

Load plain text from a supported local file format.

参数:

名称 类型 描述 默认
path str

Path to the local file to load.

必需

返回:

类型 描述
str

The extracted plain text content.

引发:

类型 描述
RuntimeError

If dependencies for PDF/DOCX are missing or for unsupported formats.

源代码位于: jianmu/rag/ingest.py
def load(self, path: str) -> str:
    """Load plain text from a supported local file format.

    Args:
        path: Path to the local file to load.

    Returns:
        The extracted plain text content.

    Raises:
        RuntimeError: If dependencies for PDF/DOCX are missing or for unsupported formats.
    """
    file_path = Path(path)
    suffix = file_path.suffix.lower()

    if suffix == ".pdf":
        if PdfReader is None:
            raise RuntimeError("pypdf not installed. Please run: pip install pypdf")
        reader = PdfReader(str(file_path))
        pages = []
        for page in reader.pages:
            text = page.extract_text() or ""
            if text:
                pages.append(text)
        return "\n".join(pages)

    if suffix == ".docx":
        if docx is None:
            raise RuntimeError("python-docx not installed. Please run: pip install python-docx")
        doc = docx.Document(str(file_path))
        return "\n".join(p.text for p in doc.paragraphs if p.text)

    if suffix == ".doc":
        raise RuntimeError(
            f"Binary .doc format is not directly supported: {file_path}. "
            "Please convert it to .docx or use a library like 'textract'."
        )

    return file_path.read_text(encoding=self.encoding)

WhitespaceNormalizer

Bases: TextNormalizerProtocol

Normalize whitespace in raw extracted text.

normalize

normalize(text: str) -> str

Normalize whitespace in raw extracted text.

参数:

名称 类型 描述 默认
text str

Raw text to normalize.

必需

返回:

类型 描述
str

The normalized text with collapsed whitespace.

源代码位于: jianmu/rag/ingest.py
def normalize(self, text: str) -> str:
    """Normalize whitespace in raw extracted text.

    Args:
        text: Raw text to normalize.

    Returns:
        The normalized text with collapsed whitespace.
    """
    import re

    normalized = text.strip()
    normalized = re.sub(r"\s+", " ", normalized)
    return normalized

SimpleTextChunker

SimpleTextChunker(chunk_size: int = 500, overlap: int = 0)

Bases: TextChunkerProtocol

Split text into fixed-size overlapping character chunks.

属性:

名称 类型 描述
chunk_size

Maximum character count emitted for each chunk.

overlap

Character overlap retained between adjacent chunks.

Configure chunk size and overlap.

参数:

名称 类型 描述 默认
chunk_size int

Maximum character count for each chunk.

500
overlap int

Character count overlap between adjacent chunks.

0
源代码位于: jianmu/rag/ingest.py
def __init__(self, chunk_size: int = 500, overlap: int = 0):
    """Configure chunk size and overlap.

    Args:
        chunk_size: Maximum character count for each chunk.
        overlap: Character count overlap between adjacent chunks.
    """
    self.chunk_size = chunk_size
    self.overlap = overlap

chunk

chunk(text: str) -> List[str]

Split text into overlapping chunks for ingestion.

参数:

名称 类型 描述 默认
text str

Raw or normalized text to split.

必需

返回:

类型 描述
List[str]

A list of split text chunks.

源代码位于: jianmu/rag/ingest.py
def chunk(self, text: str) -> List[str]:
    """Split text into overlapping chunks for ingestion.

    Args:
        text: Raw or normalized text to split.

    Returns:
        A list of split text chunks.
    """
    if self.chunk_size <= 0:
        return [text]
    overlap = max(0, self.overlap)
    chunks: List[str] = []
    start = 0
    while start < len(text):
        end = start + self.chunk_size
        chunks.append(text[start:end])
        start = end - overlap
    return chunks

RAGPipeline

RAGPipeline(
    retriever: RetrieverProtocol,
    reranker: Optional[RerankerProtocol] = None,
)

Pipeline for retrieval-augmented generation.

属性:

名称 类型 描述
retriever

Retriever used to fetch candidate documents.

reranker

Optional reranker applied after retrieval.

Bind retrieval and optional reranking into one pipeline.

源代码位于: jianmu/rag/pipeline.py
def __init__(
    self,
    retriever: RetrieverProtocol,
    reranker: Optional[RerankerProtocol] = None,
):
    """Bind retrieval and optional reranking into one pipeline."""
    self.retriever = retriever
    self.reranker = reranker

search

search(
    query: str, options: Optional[SearchOptions] = None
) -> List[SearchResult]

Search and optionally rerank results.

参数:

名称 类型 描述 默认
query str

Query search query text.

必需
options Optional[SearchOptions]

Optional SearchOptions overrides.

None

返回:

类型 描述
List[SearchResult]

A list of retrieval search results.

源代码位于: jianmu/rag/pipeline.py
def search(self, query: str, options: Optional[SearchOptions] = None) -> List[SearchResult]:
    """Search and optionally rerank results.

    Args:
        query: Query search query text.
        options: Optional SearchOptions overrides.

    Returns:
        A list of retrieval search results.
    """
    opts = options or SearchOptions()
    fetch_k = opts.k * 2 if self.reranker else opts.k
    fetch_opts = SearchOptions(k=fetch_k, mode=opts.mode, alpha=opts.alpha, filter=opts.filter)
    results = self.retriever.search(query=query, k=fetch_k, options=fetch_opts)
    if self.reranker:
        results = self.reranker.rerank(query, results)
        results = results[: opts.k]
    return results

IngestPipeline

IngestPipeline(
    embedder: Optional[EmbedderProtocol] = None,
    store: Optional[VectorStoreProtocol] = None,
    loader: DocumentLoaderProtocol | None = None,
    normalizer: TextNormalizerProtocol | None = None,
    chunker: TextChunkerProtocol | None = None,
)

Pipeline for ingesting documents into a vector store.

属性:

名称 类型 描述
embedder

Embedding provider used to encode document chunks.

store

Vector store that persists documents and embeddings.

loader

Document loader used for file-based ingestion.

normalizer

Text normalizer applied before chunking.

chunker

Text chunker used to split normalized content.

Bind ingestion helpers to an embedder, store, and ingestion strategies.

源代码位于: jianmu/rag/pipeline.py
def __init__(
    self,
    embedder: Optional[EmbedderProtocol] = None,
    store: Optional[VectorStoreProtocol] = None,
    loader: DocumentLoaderProtocol | None = None,
    normalizer: TextNormalizerProtocol | None = None,
    chunker: TextChunkerProtocol | None = None,
):
    """Bind ingestion helpers to an embedder, store, and ingestion strategies."""
    self.embedder = embedder or SimpleEmbedder()
    self.store = store or InMemoryStore()
    self.loader = loader or DefaultDocumentLoader()
    self.normalizer = normalizer or WhitespaceNormalizer()
    self.chunker = chunker or _default_chunker()

ingest_text

ingest_text(
    text: str, metadata: Optional[dict] = None
) -> List[str]

Chunk text and add to store.

参数:

名称 类型 描述 默认
text str

Input text content to ingest.

必需
metadata Optional[dict]

Optional dictionary of metadata to associate with the chunks.

None

返回:

类型 描述
List[str]

A list of successfully ingested document IDs.

源代码位于: jianmu/rag/pipeline.py
def ingest_text(
    self,
    text: str,
    metadata: Optional[dict] = None,
) -> List[str]:
    """Chunk text and add to store.

    Args:
        text: Input text content to ingest.
        metadata: Optional dictionary of metadata to associate with the chunks.

    Returns:
        A list of successfully ingested document IDs.
    """
    normalized = self.normalizer.normalize(text)
    chunks = self.chunker.chunk(normalized)
    docs = [
        Document(
            id=str(uuid.uuid4()),
            text=c,
            metadata=dict(metadata or {}, chunk_index=i),
        )
        for i, c in enumerate(chunks)
    ]
    embs = self.embedder.embed_documents([d.text for d in docs])
    for doc, emb in zip(docs, embs):
        doc.embedding = emb
    return self.store.add(docs, embs)

ingest_file

ingest_file(
    path: str, metadata: Optional[dict] = None
) -> List[str]

Load file, chunk, and add to store.

参数:

名称 类型 描述 默认
path str

Target filesystem path to the file.

必需
metadata Optional[dict]

Optional dictionary of metadata to associate with the chunks.

None

返回:

类型 描述
List[str]

A list of successfully ingested document IDs.

源代码位于: jianmu/rag/pipeline.py
def ingest_file(
    self,
    path: str,
    metadata: Optional[dict] = None,
) -> List[str]:
    """Load file, chunk, and add to store.

    Args:
        path: Target filesystem path to the file.
        metadata: Optional dictionary of metadata to associate with the chunks.

    Returns:
        A list of successfully ingested document IDs.
    """
    text = self.loader.load(path)
    meta = dict(metadata or {}, source=path)
    return self.ingest_text(text, metadata=meta)

KnowledgeBase

KnowledgeBase(
    store: Optional[VectorStoreProtocol] = None,
    embedder: Optional[EmbedderProtocol] = None,
    retriever: Optional[RetrieverProtocol] = None,
    reranker: Optional[RerankerProtocol] = None,
    loader: DocumentLoaderProtocol | None = None,
    normalizer: TextNormalizerProtocol | None = None,
    chunker: TextChunkerProtocol | None = None,
)

Convenience facade combining store, embedder, retriever, and pipelines.

Example

kb = KnowledgeBase() kb.add("Hello world") kb.add("Python is great") results = kb.search("hello")

属性:

名称 类型 描述
store

Backing vector store for persisted documents.

embedder

Embedding provider used for add and ingest operations.

retriever

Retriever used to resolve search requests.

reranker

Optional reranker applied during search.

_ingest

Internal ingestion pipeline facade.

_rag

Internal retrieval pipeline facade.

Assemble a convenience knowledge-base facade from core components.

源代码位于: jianmu/rag/pipeline.py
def __init__(
    self,
    store: Optional[VectorStoreProtocol] = None,
    embedder: Optional[EmbedderProtocol] = None,
    retriever: Optional[RetrieverProtocol] = None,
    reranker: Optional[RerankerProtocol] = None,
    loader: DocumentLoaderProtocol | None = None,
    normalizer: TextNormalizerProtocol | None = None,
    chunker: TextChunkerProtocol | None = None,
):
    """Assemble a convenience knowledge-base facade from core components."""
    self.store = store or InMemoryStore()
    self.embedder = embedder or SimpleEmbedder()
    # Retriever needs store reference to fetch documents
    self.retriever = retriever or HybridRetriever(store=self.store, embedder=self.embedder)
    self.reranker = reranker

    self._ingest = IngestPipeline(
        embedder=self.embedder,
        store=self.store,
        loader=loader,
        normalizer=normalizer,
        chunker=chunker,
    )
    self._rag = RAGPipeline(retriever=self.retriever, reranker=self.reranker)

add

add(text: str, metadata: Optional[dict] = None) -> str

Add a single document.

参数:

名称 类型 描述 默认
text str

Input document text.

必需
metadata Optional[dict]

Optional metadata dict.

None

返回:

类型 描述
str

The generated ID of the added document.

源代码位于: jianmu/rag/pipeline.py
def add(self, text: str, metadata: Optional[dict] = None) -> str:
    """Add a single document.

    Args:
        text: Input document text.
        metadata: Optional metadata dict.

    Returns:
        The generated ID of the added document.
    """
    doc = Document(id=str(uuid.uuid4()), text=text, metadata=metadata or {})
    emb = self.embedder.embed_query(text)
    doc.embedding = emb
    ids = self.store.add([doc], [emb])
    return ids[0] if ids else doc.id

ingest_text

ingest_text(
    text: str, metadata: Optional[dict] = None
) -> List[str]

Chunk and ingest text.

参数:

名称 类型 描述 默认
text str

Input text content to ingest.

必需
metadata Optional[dict]

Optional dictionary of metadata to associate with the chunks.

None

返回:

类型 描述
List[str]

A list of successfully ingested document IDs.

源代码位于: jianmu/rag/pipeline.py
def ingest_text(self, text: str, metadata: Optional[dict] = None) -> List[str]:
    """Chunk and ingest text.

    Args:
        text: Input text content to ingest.
        metadata: Optional dictionary of metadata to associate with the chunks.

    Returns:
        A list of successfully ingested document IDs.
    """
    return self._ingest.ingest_text(text, metadata=metadata)

ingest_file

ingest_file(
    path: str, metadata: Optional[dict] = None
) -> List[str]

Load and ingest a file.

参数:

名称 类型 描述 默认
path str

Target filesystem path to the file.

必需
metadata Optional[dict]

Optional dictionary of metadata to associate with the chunks.

None

返回:

类型 描述
List[str]

A list of successfully ingested document IDs.

源代码位于: jianmu/rag/pipeline.py
def ingest_file(
    self,
    path: str,
    metadata: Optional[dict] = None,
) -> List[str]:
    """Load and ingest a file.

    Args:
        path: Target filesystem path to the file.
        metadata: Optional dictionary of metadata to associate with the chunks.

    Returns:
        A list of successfully ingested document IDs.
    """
    return self._ingest.ingest_file(path, metadata=metadata)

search

search(
    query: str, k: int = 5, mode: str = "hybrid"
) -> List[SearchResult]

Search the knowledge base.

参数:

名称 类型 描述 默认
query str

Query search query text.

必需
k int

Number of nearest neighbors to retrieve.

5
mode str

Search modality ("dense", "sparse", or "hybrid").

'hybrid'

返回:

类型 描述
List[SearchResult]

A list of SearchResult objects matching the query.

源代码位于: jianmu/rag/pipeline.py
def search(self, query: str, k: int = 5, mode: str = "hybrid") -> List[SearchResult]:
    """Search the knowledge base.

    Args:
        query: Query search query text.
        k: Number of nearest neighbors to retrieve.
        mode: Search modality ("dense", "sparse", or "hybrid").

    Returns:
        A list of SearchResult objects matching the query.
    """
    options = SearchOptions(k=k, mode=mode)
    return self._rag.search(query, options=options)

as_tool

as_tool() -> 'KnowledgeSearchTool'

Return a KnowledgeSearchTool for this knowledge base.

返回:

类型 描述
'KnowledgeSearchTool'

The resulting 'KnowledgeSearchTool' value.

源代码位于: jianmu/rag/pipeline.py
def as_tool(self) -> "KnowledgeSearchTool":
    """Return a KnowledgeSearchTool for this knowledge base.

    Returns:
        The resulting `'KnowledgeSearchTool'` value.
    """
    from jianmu.rag.tools import KnowledgeSearchTool
    return KnowledgeSearchTool(pipeline=self._rag)

KnowledgeSearchTool

KnowledgeSearchTool(pipeline: RAGPipeline)

Bases: Tool

Search external knowledge via a RAG pipeline.

属性:

名称 类型 描述
pipeline

RAG pipeline used to execute search requests.

Create a knowledge-search tool backed by one RAG pipeline.

源代码位于: jianmu/rag/tools.py
def __init__(self, pipeline: RAGPipeline):
    """Create a knowledge-search tool backed by one RAG pipeline."""
    self.pipeline = pipeline

run async

run(
    query: str | None = None,
    k: int = 3,
    mode: str | None = None,
    **kwargs: Any,
) -> str

Search the knowledge base and render results as plain text.

参数:

名称 类型 描述 默认
query str | None

The search query text.

None
k int

The number of search results to retrieve.

3
mode str | None

Search mode: semantic, keyword, or hybrid.

None
**kwargs Any

Additional keyword arguments.

{}

返回:

类型 描述
str

A newline-separated string of the retrieved document texts.

源代码位于: jianmu/rag/tools.py
async def run(
    self,
    query: str | None = None,
    k: int = 3,
    mode: str | None = None,
    **kwargs: Any,
) -> str:
    """Search the knowledge base and render results as plain text.

    Args:
        query: The search query text.
        k: The number of search results to retrieve.
        mode: Search mode: semantic, keyword, or hybrid.
        **kwargs: Additional keyword arguments.

    Returns:
        A newline-separated string of the retrieved document texts.
    """
    if query is None:
        query = kwargs.get("input", "")
    if not query:
        return "Error: No query provided"
    options = SearchOptions(k=k, mode=mode or "hybrid")
    items = self.pipeline.search(query, options=options)
    if not items:
        return "No relevant knowledge found."
    results: List[str] = []
    for i, item in enumerate(items, 1):
        doc = getattr(item, "document", None) or item
        text = getattr(doc, "text", "")
        results.append(f"{i}. {text}")
    return "\n".join(results)

resolve_embedder

resolve_embedder(
    preference: Optional[List[str]] = None,
    model: Optional[str] = None,
    base_url: Optional[str] = None,
) -> EmbedderProtocol

Auto-resolve embedder based on available API keys.

参数:

名称 类型 描述 默认
preference Optional[List[str]]

Optional ordered list of preferred providers.

None
model Optional[str]

Optional embedding model override.

None
base_url Optional[str]

Optional custom base URL override.

None

返回:

类型 描述
EmbedderProtocol

An instantiated embedding provider conforming to EmbedderProtocol.

源代码位于: jianmu/rag/embedder.py
def resolve_embedder(
    preference: Optional[List[str]] = None,
    model: Optional[str] = None,
    base_url: Optional[str] = None,
) -> EmbedderProtocol:
    """Auto-resolve embedder based on available API keys.

    Args:
        preference: Optional ordered list of preferred providers.
        model: Optional embedding model override.
        base_url: Optional custom base URL override.

    Returns:
        An instantiated embedding provider conforming to EmbedderProtocol.
    """
    cfg = get_config().rag.embedding
    order = preference or ([cfg.provider] if cfg.provider else ["gemini", "openai", "simple"])

    for name in order:
        name = name.lower()
        if name == "gemini":
            if os.getenv("GOOGLE_API_KEY") or os.getenv("GEMINI_API_KEY"):
                try:
                    return GeminiEmbedder(api_key=None, base_url=base_url, model=model)
                except Exception:
                    pass
        elif name == "openai":
            if os.getenv("OPENAI_API_KEY") or os.getenv("API_KEY"):
                try:
                    return OpenAIEmbedder(api_key=None, base_url=base_url, model=model)
                except Exception:
                    pass
        elif name == "simple":
            return SimpleEmbedder()

    return SimpleEmbedder()

simple_embedding

simple_embedding(text: str, dim: int = 64) -> List[float]

Simple character-based embedding for testing/fallback.

参数:

名称 类型 描述 默认
text str

Input string to embed.

必需
dim int

Dimensionality of the output embedding vector.

64

返回:

类型 描述
List[float]

A list of floats representing the normalized embedding vector.

源代码位于: jianmu/rag/embedder.py
def simple_embedding(text: str, dim: int = 64) -> List[float]:
    """Simple character-based embedding for testing/fallback.

    Args:
        text: Input string to embed.
        dim: Dimensionality of the output embedding vector.

    Returns:
        A list of floats representing the normalized embedding vector.
    """
    import math
    vec = [0.0] * dim
    for i, ch in enumerate(text.lower()):
        vec[ord(ch) % dim] += 1.0 / (1 + i * 0.01)
    norm = math.sqrt(sum(x * x for x in vec)) or 1.0
    return [x / norm for x in vec]

cosine_similarity

cosine_similarity(a: List[float], b: List[float]) -> float

Compute cosine similarity for two dense vectors.

参数:

名称 类型 描述 默认
a List[float]

The first input vector.

必需
b List[float]

The second input vector.

必需

返回:

类型 描述
float

The cosine similarity score between the two vectors.

源代码位于: jianmu/rag/utils.py
def cosine_similarity(a: List[float], b: List[float]) -> float:
    """Compute cosine similarity for two dense vectors.

    Args:
        a: The first input vector.
        b: The second input vector.

    Returns:
        The cosine similarity score between the two vectors.
    """
    if not a or not b:
        return 0.0
    dot = sum(x * y for x, y in zip(a, b))
    denom = math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b))
    if denom == 0.0:
        return 0.0
    return dot / denom