Docs / Python SDK deep dive
Python SDK — full method reference
Single scrollable page with every method card, pipeline options, and error tables. Prefer the split SDK pages in the sidebar for quicker navigation.
LLM provider keys
Ephemeral vs stored keys are documented on Native → LLM provider keys. Links below that point here scroll to this note.
Python SDK reference
The tokensaver-sdk package is a thin client over the TokenSaver HTTP API. For ephemeral provider_api_key vs keys stored in Settings, read LLM provider keys at the top of this page. Then use the snippet below and jump to any method for its purpose, return type, and signature.
Test a basic request
Create a client with your API key and call ask — the smallest path from install to a model response. You do not need base_url unless you target a non-default host (see Client).
from tokensaver_sdk import TokenSaver
ts = TokenSaver(api_key="ts_...")
out = ts.ask(
"Say hello in one sentence.",
provider="openai",
model="gpt-4o",
)
print(out.text)
Run with: python main.py (after pip install tokensaver-sdk).
Methods overview
Jump to the group you need. Each card below explains one callable with a short description and a copy-pasteable signature.
Client
TokenSaver (alias TokenSaverClient) holds your API key, base URL, and chat for session helpers.
TokenSaver.__init__(base_url=None, api_key=None, *, provider_api_key=None, timeout_total=30, connect_timeout=5, read_timeout=25, max_retries=2, headers=None)
Builds the client. Omit base_url to use the built-in production API root (https://api.tokensaver.fr/api/v1). Optional provider_api_key: default LLM provider secret sent on every pipeline run (overrides org keys for that run only; never stored). Raises ValueError if api_key is missing.
Returns — TokenSaver
ts = TokenSaver(api_key="ts_...")
ts = TokenSaver(api_key="ts_...", provider_api_key="sk-...")
ts = TokenSaver("https://api.example.com/api/v1", "ts_...")
ts = TokenSaver(base_url="http://localhost:8000/api/v1", api_key="ts_...")
Attributes
base_url— Normalized API root (no trailing slash).api_key— Same string you passed in (sent as Bearer).chat— Usets.chat.session(...)for chat flows (see Chat sessions).
Governance policies
Module on/off (cache, RAG, compression, PII) is controlled by governance policies on the API key — not use_cache / use_rag on ask(). Read runtime state via get_pipeline_settings() or effective_modules(). Create or toggle policies with the methods below (scope api_key only via SDK; org/workspace policies are inherited and managed in the console).
Plan entitlements
Before creating an enabled module policy, the SDK checks plan_features (has_cache, has_rag, …). Mismatch raises ValidationError (PLAN_CACHE_DISABLED, …). Pass validate_plan=False to skip the client check.
from tokensaver_sdk import TokenSaver
ts = TokenSaver(api_key="ts_...")
# Module on/off = governance policies (not use_cache on ask())
policy = ts.create_governance_policy(
"Production cache",
kind="cache",
config={
"exact_cache": True,
"semantic_cache": True,
"similarity_threshold": 0.85,
},
)
assert ts.effective_modules()["use_cache"] is True
# Per-run threshold only (does not enable the module)
out = ts.ask(
"Hello again",
provider="openai",
model="gpt-4o",
cache_similarity_threshold=0.85,
)
get_pipeline_settings() → dict
Thresholds on the key, effective_modules (policy-driven gates), plan_features, server_defaults.
HTTP — GET /sdk/pipeline-settings
Returns — dict
effective_modules() → dict[str, bool]
Policy-driven module gates: use_cache, use_rag, use_compression, use_pii_filter (read-only). Uses cached settings when available.
Returns — dict
patch_pipeline_settings(*, pipeline_settings=None, clear_fields=None, reset=False) → dict
Merge thresholds, pii_options, default_model, etc. Does not toggle modules.
HTTP — PATCH /sdk/pipeline-settings
Returns — dict
list_governance_policies(*, kind=None) → dict
Own api_key policies plus inherited org/workspace rows (read-only). Includes effective_gates and pii_disable_hint when relevant.
HTTP — GET /sdk/governance/policies
Returns — dict
get_governance_policy(policy_id) → dict
Fetch one policy owned by this API key.
HTTP — GET /sdk/governance/policies/{id}
Returns — dict
create_governance_policy(name, *, kind=‘pii’, config=None, enabled=True, priority=100, validate_plan=True) → dict
Create an api_key-scoped policy. Kinds: pii, cache, rag, compression, prompt_injection, decorator. Typed configs: CachePolicyConfig, RagPolicyConfig, CompressionPolicyConfig, PiiPolicyConfig.
HTTP — POST /sdk/governance/policies
Returns — dict
update_governance_policy(policy_id, *, name=None, config=None, enabled=None, priority=None, validate_plan=True) → dict
Patch a key-owned policy. Disabling PII may still leave PII active if inherited org/workspace policies apply (see warning in response).
HTTP — PATCH /sdk/governance/policies/{id}
Returns — dict
delete_governance_policy(policy_id) → None
Delete a key-owned policy.
HTTP — DELETE /sdk/governance/policies/{id}
compress_context_preview(blocks, *, provider=‘openai’, model=‘gpt-4o-mini’, policy=None, compare_headroom=False) → dict
Dry-run content-aware compression without an LLM call. Returns per-block strategy, token counts, optional policy_applied snapshot, and best-of-both metadata (alternate_strategy, headroom_rejected, tie_broken_prefer_native). Set compare_headroom=True with headroom in enabled_strategies for native vs TokenSaver+Headroom lanes.
HTTP — POST /compression/preview
retrieve_ccr(content_id, *, filter_hint=None, max_chars=None) → dict
Fetch the original payload stored for a TokenSaver CCR content_id (after compression). Optional filter_hint (errors, row:N, path) and max_chars cap match the MCP retrieve tool.
HTTP — GET /compression/ccr/{content_id}
Returns — dict with content_id, text, chars
Pipeline & responses
These methods call the governed pipeline. Which modules run (cache, RAG, compression, PII) is resolved from governance policies and your organisation plan — not boolean flags on the request body.
Hosted API (default base URL)
On https://api.tokensaver.fr/api/v1, use provider in HOSTED_SAAS_LLM_PROVIDERS (OpenAI, Anthropic, Google, Mistral, Grok, DeepSeek) — same as the hosted console. Other codes raise ValidationError (HOSTED_LLM_PROVIDER). Point base_url at a self-hosted stack for additional providers where the backend allows them.
ask(…)
Recommended entry point: same JSON body as run_pipeline. Returns RunResult (.text, metrics, trace). Optional provider_api_key per call. History: chat_id with HISTORY_SERVER or chat_history with HISTORY_LOCAL. Per-run tuning: temperature; rag_similarity_threshold; cache_similarity_threshold; compression_level; rag_options; pii_options; context_layers or legacy instruction fields. Module on/off: governance policies only.
HTTP — POST /pipelines/run
Returns — RunResult
result = ts.ask(
"Summarize this in 3 bullets.",
provider="openai",
model="gpt-4o",
rag_similarity_threshold=0.55,
rag_options={"document_ids": ["<document_id>"], "top_k": 8},
)
print(result.text, result.metrics.cost_usd)
run_pipeline(…)
Lower-level: raw API JSON. Same optional threshold/option fields as ask (not use_* module flags).
HTTP — POST /pipelines/run
Returns — dict
estimate_cost(prompt_tokens, completion_tokens, provider, model)
Ask TokenSaver for an estimated cost from token counts — no LLM call.
HTTP — POST /pricing/estimate
Returns — dict
Pipeline request options (console parity)
ask and run_pipeline send the same optional JSON fields as the TokenSaver console for POST /pipelines/run (thresholds and options — not module on/off). Omit a field to use server defaults.
| Parameter | Purpose |
|---|---|
| temperature | LLM temperature (0–2). |
| use_cache, use_rag, use_compression, use_pii_filter | Removed from the client schema — use governance policies. Sending them returns 422 MODULE_GATE_POLICY_ONLY. |
| stream | Reserved on native run (sync today). |
| rag_similarity_threshold | RAG retrieval similarity floor (0–1). |
| cache_similarity_threshold | Semantic cache similarity (0–1). |
| compression_level | Compression strength 1–5 when compression is on. |
| provider_api_key | SDK only. Per-request LLM provider secret; takes precedence over organisation keys in the database for that run; never stored. Set on TokenSaver(..., provider_api_key=...) or pass to ask / run_pipeline. |
| rag_options | Dict: document_ids, top_k, query_image_url. |
| pii_options | Dict: engine, strategy, confidence_threshold, entity_types, language, regex_fallback. |
| context_layers | Structured instruction / knowledge / interaction (canonical API). |
| system_prompt, profile_context, workspace_instructions | Legacy flat instruction fields (if not using context_layers). |
| chat_id, chat_history | Session routing (see history modes). |
Types RagOptions, PiiOptions, CachePolicyConfig, … are exported from tokensaver_sdk for editor hints.
HTTP helpers
Authenticated httpx calls with retries on transient errors. Paths are relative to base_url (e.g. "rag/documents").
get(path, **kwargs) → Response
GET with Authorization header and JSON Accept.
post(path, **kwargs) → Response
POST JSON by default; use httpx kwargs for custom bodies.
patch(path, **kwargs) → Response
PATCH JSON (governance policies, pipeline-settings).
delete(path, **kwargs) → Response
Used by ChatSession.close() and delete_governance_policy().
RAG documents
Upload supported files to your workspace, wait for ingestion, then pass document_ids in rag_options on ask(...) when a RAG governance policy is active — or use ChatSession.attach_knowledge on server chats. Allowed extensions match the console and RAG_UPLOAD_EXTENSIONS (pdf, txt, md, csv, json, docx). Wrong extension → ValidationError (RAG_UNSUPPORTED_FILE_TYPE) before any HTTP call.
rag_list_documents()
Lists ingested documents for the current API key / workspace (newest first). Maps to GET /rag/documents — each row may include embedding_key, embedding_dim, embedding_compute, chunk_config, metadata.
HTTP — GET /rag/documents
Returns — dict with keys documents, chunk_config
rag_upload_document(file_path, *, name=None, description=None)
Multipart upload only; does not wait for chunking. Raises ValidationError (RAG_FILE_NOT_FOUND) if the path is missing or not a file; RAG_UNSUPPORTED_FILE_TYPE if the extension is not allowed.
HTTP — POST /rag/documents
Returns — dict (includes document_id)
rag_get_document(document_id)
Fetch status, chunk counts, metadata, chunk_config, and chunks_by_type (text/image/table) for one document.
HTTP — GET /rag/documents/{id}
Returns — dict
rag_wait_document_ready(document_id, *, timeout_seconds=90, poll_interval_seconds=2)
Polls until status is done or ingested, or raises on error / timeout.
Returns — dict
rag_upload_and_wait(file_path, *, name=None, description=None, timeout_seconds=90, poll_interval_seconds=2)
Upload then block until ingestion completes. Same ValidationError (RAG_FILE_NOT_FOUND) as rag_upload_document if the local file is missing.
Returns — dict
rag_ensure_document(file_path, *, reuse_existing=True, name=None, description=None, timeout_seconds=90, poll_interval_seconds=2)
If a document with the same filename already exists and is ready, returns it without re-uploading. If pending, waits. Otherwise uploads and waits (missing local file → RAG_FILE_NOT_FOUND; bad extension → RAG_UNSUPPORTED_FILE_TYPE). Set reuse_existing=False to always send a new file.
Returns — dict
Additional REST endpoints (console / integrations, no dedicated SDK helper yet): GET /rag/embedding-spaces, GET /rag/documents/{id}/chunks?full=true, GET /rag/documents/{id}/preview, GET /rag/queue-stats. Full path and response shapes: Native API → RAG documents in the API reference navigation.
ChatSession.attach_knowledge
On a HISTORY_SERVER session, remember one or more RAG document_id strings. Each ask merges them into rag_options["document_ids"] (deduped with any IDs you pass explicitly). Same workflow as attaching knowledge from the console chat “+” menu. Available in SDK 0.1.9+.
from tokensaver_sdk import HISTORY_SERVER, TokenSaver
ts = TokenSaver(api_key="ts_...")
doc_id = str(ts.rag_ensure_document("./handbook.pdf")["document_id"])
session = ts.chat.session(history=HISTORY_SERVER, name="Docs Q&A")
session.attach_knowledge(doc_id)
out = session.ask(
"What is the refund policy?",
provider="openai",
model="gpt-4o",
rag_options={"top_k": 6},
)
print(out.text)
ChatSession.attach_knowledge(*document_ids: str) → None
Append non-empty document IDs to the session list (no HTTP call). IDs are sent on subsequent ask() calls via merged rag_options.
ChatSession.clear_knowledge() → None
Remove all IDs added with attach_knowledge (in-memory only).
Chat sessions
Use HISTORY_SERVER for chats stored in TokenSaver, or HISTORY_LOCAL for in-process memory.
ts.chat.session(*, history=HISTORY_NONE, name=‘New Chat’, provider=None, model=None) → ChatSession
When history is HISTORY_SERVER, creates a chat via POST /sdk/chats (Idempotency-Key set automatically) and returns a ChatSession with chat_id.
HTTP — POST /sdk/chats (server mode)
Returns — ChatSession
ChatSession
ask(prompt, **kwargs) → RunResult— forwards to TokenSaver.ask; in HISTORY_LOCAL, appends turns to memory. Merges attach_knowledge IDs intorag_optionsbefore each call.attach_knowledge(*document_ids),clear_knowledge()— RAG document IDs for server/local/none sessions (merge behaviour applies whenever ask runs).messages(limit=50, cursor=None, order="asc") → dict— server transcript when HISTORY_SERVER.close()— deletes the server chat when applicable.
Return types
Dataclasses returned by ask and attached to session calls.
RunResult
text, raw, metrics (MetricsView), trace (TraceView), context (ContextView). Method to_dict().
MetricsView
cost_usd, latency_ms, cache_hit, tokens_*, savings_ratio (all optional).
TraceView
request_id, provider, model.
ContextView
history_mode, chat_id, layers_used.
Errors
Import from tokensaver_sdk.errors. All subclasses expose code, message, status_code, request_id, raw.
TokenSaverError— base class.AuthenticationError,ValidationError,ServerError,TimeoutError- Transport / DNS —
ServerErrorwithcode=NETWORK_ERRORwhen the host cannot be reached (wrongbase_url, DNS failure, TLS, connection refused). Not an HTTP response from TokenSaver. ProviderKeyMissingError— extra fieldprovider- Hosted default URL —
ValidationErrorwithcode=HOSTED_LLM_PROVIDER(ERROR_HOSTED_LLM_PROVIDER) ifprovideris not OpenAI on the public API root (see Pipeline above). - Model catalogue —
ValidationErrorwithcode=LLM_MODEL_NOT_SUPPORTED(ERROR_LLM_MODEL_NOT_SUPPORTED) ifprovider/modelare not an active row in the platform LLM reference (same rule asGET /api/v1/llm-reference/models). QuotaExceededError— quota_dimension, limit, current_usage, retry_after_secondsRateLimitError— retry_after_secondsValidationErrorwithcode=PLAN_CACHE_DISABLED(andPLAN_RAG_DISABLED, …) when creating enabled module policies on a plan that excludes the feature.- Client-side RAG path —
rag_upload_document/rag_upload_and_wait/rag_ensure_documentraiseValidationErrorwithcode=RAG_FILE_NOT_FOUNDif the path is missing, orcode=RAG_UNSUPPORTED_FILE_TYPE(ERROR_RAG_UNSUPPORTED_FILE_TYPE) if the extension is not inRAG_UPLOAD_EXTENSIONS. Compare withERROR_RAG_FILE_NOT_FOUND;raw["path"]may hold the resolved path.
Package imports
Public surface matches tokensaver_sdk.__all__ (stable imports for docs and IDEs).
from tokensaver_sdk import (
TokenSaver,
TokenSaverClient,
API_PIPELINE_LLM_PROVIDERS,
DEFAULT_PUBLIC_API_BASE_URL,
HOSTED_SAAS_LLM_PROVIDERS,
RAG_UPLOAD_EXTENSIONS,
mime_type_for_rag_filename,
ERROR_HOSTED_LLM_PROVIDER,
ERROR_LLM_MODEL_NOT_SUPPORTED,
ERROR_RAG_FILE_NOT_FOUND,
ERROR_RAG_UNSUPPORTED_FILE_TYPE,
HISTORY_NONE,
HISTORY_LOCAL,
HISTORY_SERVER,
HistoryMode,
RunResult,
MetricsView,
TraceView,
ContextView,
GOVERNANCE_POLICY_KINDS,
MODULE_POLICY_KINDS,
POLICY_KIND_TO_PLAN_FEATURE,
RagOptions,
PiiOptions,
CachePolicyConfig,
RagPolicyConfig,
CompressionPolicyConfig,
PiiPolicyConfig,
PolicyKind,
ModulePolicyKind,
TokenSaverError,
AuthenticationError,
ProviderKeyMissingError,
QuotaExceededError,
RateLimitError,
ValidationError,
ServerError,
TimeoutError,
)
import tokensaver_sdk
print(tokensaver_sdk.__version__)
History constants
HISTORY_NONE # "none" — no memory
HISTORY_LOCAL # "local" — SDK process memory
HISTORY_SERVER # "server" — persisted on TokenSaver