Docs / Python SDK deep dive

Python SDK — full method reference

Single scrollable page with every method card, pipeline options, and error tables. Prefer the split SDK pages in the sidebar for quicker navigation.

LLM provider keys

Ephemeral vs stored keys are documented on Native → LLM provider keys. Links below that point here scroll to this note.

Python SDK reference

The tokensaver-sdk package is a thin client over the TokenSaver HTTP API. For ephemeral provider_api_key vs keys stored in Settings, read LLM provider keys at the top of this page. Then use the snippet below and jump to any method for its purpose, return type, and signature.

Test a basic request

Create a client with your API key and call ask — the smallest path from install to a model response. You do not need base_url unless you target a non-default host (see Client).

from tokensaver_sdk import TokenSaver

ts = TokenSaver(api_key="ts_...")
out = ts.ask(
    "Say hello in one sentence.",
    provider="openai",
    model="gpt-4o",
)
print(out.text)

Run with: python main.py (after pip install tokensaver-sdk).

Methods overview

Jump to the group you need. Each card below explains one callable with a short description and a copy-pasteable signature.

TokenSaverConstructor, timeouts, chat facade
GovernancePolicies, effective_modules, plan gates
askTyped prompt → RunResult
run_pipelineRaw POST /pipelines/run JSON
estimate_costPrice estimate without LLM call
Pipeline optionsThresholds & module params (console parity)
get / post / deleteAuthenticated httpx wrappers
rag_ensure_documentUpload or reuse by filename
attach_knowledgeServer chat + RAG document IDs (console “+” parity)
chat.sessionServer or local chat sessions
RunResult & viewsMetrics, trace, context
ExceptionsTyped errors from the API

Client

TokenSaver (alias TokenSaverClient) holds your API key, base URL, and chat for session helpers.

TokenSaver.__init__(base_url=None, api_key=None, *, provider_api_key=None, timeout_total=30, connect_timeout=5, read_timeout=25, max_retries=2, headers=None)

Builds the client. Omit base_url to use the built-in production API root (https://api.tokensaver.fr/api/v1). Optional provider_api_key: default LLM provider secret sent on every pipeline run (overrides org keys for that run only; never stored). Raises ValueError if api_key is missing.

Returns — TokenSaver

ts = TokenSaver(api_key="ts_...")
ts = TokenSaver(api_key="ts_...", provider_api_key="sk-...")
ts = TokenSaver("https://api.example.com/api/v1", "ts_...")
ts = TokenSaver(base_url="http://localhost:8000/api/v1", api_key="ts_...")

Attributes

  • base_url — Normalized API root (no trailing slash).
  • api_key — Same string you passed in (sent as Bearer).
  • chat — Use ts.chat.session(...) for chat flows (see Chat sessions).

Governance policies

Module on/off (cache, RAG, compression, PII) is controlled by governance policies on the API key — not use_cache / use_rag on ask(). Read runtime state via get_pipeline_settings() or effective_modules(). Create or toggle policies with the methods below (scope api_key only via SDK; org/workspace policies are inherited and managed in the console).

Plan entitlements

Before creating an enabled module policy, the SDK checks plan_features (has_cache, has_rag, …). Mismatch raises ValidationError (PLAN_CACHE_DISABLED, …). Pass validate_plan=False to skip the client check.

from tokensaver_sdk import TokenSaver

ts = TokenSaver(api_key="ts_...")

# Module on/off = governance policies (not use_cache on ask())
policy = ts.create_governance_policy(
    "Production cache",
    kind="cache",
    config={
        "exact_cache": True,
        "semantic_cache": True,
        "similarity_threshold": 0.85,
    },
)
assert ts.effective_modules()["use_cache"] is True

# Per-run threshold only (does not enable the module)
out = ts.ask(
    "Hello again",
    provider="openai",
    model="gpt-4o",
    cache_similarity_threshold=0.85,
)

get_pipeline_settings() → dict

Thresholds on the key, effective_modules (policy-driven gates), plan_features, server_defaults.

HTTP — GET /sdk/pipeline-settings

Returns — dict

effective_modules() → dict[str, bool]

Policy-driven module gates: use_cache, use_rag, use_compression, use_pii_filter (read-only). Uses cached settings when available.

Returns — dict

patch_pipeline_settings(*, pipeline_settings=None, clear_fields=None, reset=False) → dict

Merge thresholds, pii_options, default_model, etc. Does not toggle modules.

HTTP — PATCH /sdk/pipeline-settings

Returns — dict

list_governance_policies(*, kind=None) → dict

Own api_key policies plus inherited org/workspace rows (read-only). Includes effective_gates and pii_disable_hint when relevant.

HTTP — GET /sdk/governance/policies

Returns — dict

get_governance_policy(policy_id) → dict

Fetch one policy owned by this API key.

HTTP — GET /sdk/governance/policies/{id}

Returns — dict

create_governance_policy(name, *, kind=‘pii’, config=None, enabled=True, priority=100, validate_plan=True) → dict

Create an api_key-scoped policy. Kinds: pii, cache, rag, compression, prompt_injection, decorator. Typed configs: CachePolicyConfig, RagPolicyConfig, CompressionPolicyConfig, PiiPolicyConfig.

HTTP — POST /sdk/governance/policies

Returns — dict

update_governance_policy(policy_id, *, name=None, config=None, enabled=None, priority=None, validate_plan=True) → dict

Patch a key-owned policy. Disabling PII may still leave PII active if inherited org/workspace policies apply (see warning in response).

HTTP — PATCH /sdk/governance/policies/{id}

Returns — dict

delete_governance_policy(policy_id) → None

Delete a key-owned policy.

HTTP — DELETE /sdk/governance/policies/{id}

compress_context_preview(blocks, *, provider=‘openai’, model=‘gpt-4o-mini’, policy=None, compare_headroom=False) → dict

Dry-run content-aware compression without an LLM call. Returns per-block strategy, token counts, optional policy_applied snapshot, and best-of-both metadata (alternate_strategy, headroom_rejected, tie_broken_prefer_native). Set compare_headroom=True with headroom in enabled_strategies for native vs TokenSaver+Headroom lanes.

HTTP — POST /compression/preview

retrieve_ccr(content_id, *, filter_hint=None, max_chars=None) → dict

Fetch the original payload stored for a TokenSaver CCR content_id (after compression). Optional filter_hint (errors, row:N, path) and max_chars cap match the MCP retrieve tool.

HTTP — GET /compression/ccr/{content_id}

Returns — dict with content_id, text, chars

Pipeline & responses

These methods call the governed pipeline. Which modules run (cache, RAG, compression, PII) is resolved from governance policies and your organisation plan — not boolean flags on the request body.

Hosted API (default base URL)

On https://api.tokensaver.fr/api/v1, use provider in HOSTED_SAAS_LLM_PROVIDERS (OpenAI, Anthropic, Google, Mistral, Grok, DeepSeek) — same as the hosted console. Other codes raise ValidationError (HOSTED_LLM_PROVIDER). Point base_url at a self-hosted stack for additional providers where the backend allows them.

ask(…)

Recommended entry point: same JSON body as run_pipeline. Returns RunResult (.text, metrics, trace). Optional provider_api_key per call. History: chat_id with HISTORY_SERVER or chat_history with HISTORY_LOCAL. Per-run tuning: temperature; rag_similarity_threshold; cache_similarity_threshold; compression_level; rag_options; pii_options; context_layers or legacy instruction fields. Module on/off: governance policies only.

HTTP — POST /pipelines/run

Returns — RunResult

result = ts.ask(
    "Summarize this in 3 bullets.",
    provider="openai",
    model="gpt-4o",
    rag_similarity_threshold=0.55,
    rag_options={"document_ids": ["<document_id>"], "top_k": 8},
)
print(result.text, result.metrics.cost_usd)

run_pipeline(…)

Lower-level: raw API JSON. Same optional threshold/option fields as ask (not use_* module flags).

HTTP — POST /pipelines/run

Returns — dict

estimate_cost(prompt_tokens, completion_tokens, provider, model)

Ask TokenSaver for an estimated cost from token counts — no LLM call.

HTTP — POST /pricing/estimate

Returns — dict

Pipeline request options (console parity)

ask and run_pipeline send the same optional JSON fields as the TokenSaver console for POST /pipelines/run (thresholds and options — not module on/off). Omit a field to use server defaults.

Parameter Purpose
temperature LLM temperature (0–2).
use_cache, use_rag, use_compression, use_pii_filter Removed from the client schema — use governance policies. Sending them returns 422 MODULE_GATE_POLICY_ONLY.
stream Reserved on native run (sync today).
rag_similarity_threshold RAG retrieval similarity floor (0–1).
cache_similarity_threshold Semantic cache similarity (0–1).
compression_level Compression strength 1–5 when compression is on.
provider_api_key SDK only. Per-request LLM provider secret; takes precedence over organisation keys in the database for that run; never stored. Set on TokenSaver(..., provider_api_key=...) or pass to ask / run_pipeline.
rag_options Dict: document_ids, top_k, query_image_url.
pii_options Dict: engine, strategy, confidence_threshold, entity_types, language, regex_fallback.
context_layers Structured instruction / knowledge / interaction (canonical API).
system_prompt, profile_context, workspace_instructions Legacy flat instruction fields (if not using context_layers).
chat_id, chat_history Session routing (see history modes).

Types RagOptions, PiiOptions, CachePolicyConfig, … are exported from tokensaver_sdk for editor hints.

HTTP helpers

Authenticated httpx calls with retries on transient errors. Paths are relative to base_url (e.g. "rag/documents").

get(path, **kwargs) → Response

GET with Authorization header and JSON Accept.

post(path, **kwargs) → Response

POST JSON by default; use httpx kwargs for custom bodies.

patch(path, **kwargs) → Response

PATCH JSON (governance policies, pipeline-settings).

delete(path, **kwargs) → Response

Used by ChatSession.close() and delete_governance_policy().

RAG documents

Upload supported files to your workspace, wait for ingestion, then pass document_ids in rag_options on ask(...) when a RAG governance policy is active — or use ChatSession.attach_knowledge on server chats. Allowed extensions match the console and RAG_UPLOAD_EXTENSIONS (pdf, txt, md, csv, json, docx). Wrong extension → ValidationError (RAG_UNSUPPORTED_FILE_TYPE) before any HTTP call.

rag_list_documents()

Lists ingested documents for the current API key / workspace (newest first). Maps to GET /rag/documents — each row may include embedding_key, embedding_dim, embedding_compute, chunk_config, metadata.

HTTP — GET /rag/documents

Returns — dict with keys documents, chunk_config

rag_upload_document(file_path, *, name=None, description=None)

Multipart upload only; does not wait for chunking. Raises ValidationError (RAG_FILE_NOT_FOUND) if the path is missing or not a file; RAG_UNSUPPORTED_FILE_TYPE if the extension is not allowed.

HTTP — POST /rag/documents

Returns — dict (includes document_id)

rag_get_document(document_id)

Fetch status, chunk counts, metadata, chunk_config, and chunks_by_type (text/image/table) for one document.

HTTP — GET /rag/documents/{id}

Returns — dict

rag_wait_document_ready(document_id, *, timeout_seconds=90, poll_interval_seconds=2)

Polls until status is done or ingested, or raises on error / timeout.

Returns — dict

rag_upload_and_wait(file_path, *, name=None, description=None, timeout_seconds=90, poll_interval_seconds=2)

Upload then block until ingestion completes. Same ValidationError (RAG_FILE_NOT_FOUND) as rag_upload_document if the local file is missing.

Returns — dict

rag_ensure_document(file_path, *, reuse_existing=True, name=None, description=None, timeout_seconds=90, poll_interval_seconds=2)

If a document with the same filename already exists and is ready, returns it without re-uploading. If pending, waits. Otherwise uploads and waits (missing local file → RAG_FILE_NOT_FOUND; bad extension → RAG_UNSUPPORTED_FILE_TYPE). Set reuse_existing=False to always send a new file.

Returns — dict

Additional REST endpoints (console / integrations, no dedicated SDK helper yet): GET /rag/embedding-spaces, GET /rag/documents/{id}/chunks?full=true, GET /rag/documents/{id}/preview, GET /rag/queue-stats. Full path and response shapes: Native API → RAG documents in the API reference navigation.

ChatSession.attach_knowledge

On a HISTORY_SERVER session, remember one or more RAG document_id strings. Each ask merges them into rag_options["document_ids"] (deduped with any IDs you pass explicitly). Same workflow as attaching knowledge from the console chat “+” menu. Available in SDK 0.1.9+.

from tokensaver_sdk import HISTORY_SERVER, TokenSaver

ts = TokenSaver(api_key="ts_...")
doc_id = str(ts.rag_ensure_document("./handbook.pdf")["document_id"])

session = ts.chat.session(history=HISTORY_SERVER, name="Docs Q&A")
session.attach_knowledge(doc_id)
out = session.ask(
    "What is the refund policy?",
    provider="openai",
    model="gpt-4o",
    rag_options={"top_k": 6},
)
print(out.text)

ChatSession.attach_knowledge(*document_ids: str) → None

Append non-empty document IDs to the session list (no HTTP call). IDs are sent on subsequent ask() calls via merged rag_options.

ChatSession.clear_knowledge() → None

Remove all IDs added with attach_knowledge (in-memory only).

Chat sessions

Use HISTORY_SERVER for chats stored in TokenSaver, or HISTORY_LOCAL for in-process memory.

ts.chat.session(*, history=HISTORY_NONE, name=‘New Chat’, provider=None, model=None) → ChatSession

When history is HISTORY_SERVER, creates a chat via POST /sdk/chats (Idempotency-Key set automatically) and returns a ChatSession with chat_id.

HTTP — POST /sdk/chats (server mode)

Returns — ChatSession

ChatSession

  • ask(prompt, **kwargs) → RunResult — forwards to TokenSaver.ask; in HISTORY_LOCAL, appends turns to memory. Merges attach_knowledge IDs into rag_options before each call.
  • attach_knowledge(*document_ids), clear_knowledge() — RAG document IDs for server/local/none sessions (merge behaviour applies whenever ask runs).
  • messages(limit=50, cursor=None, order="asc") → dict — server transcript when HISTORY_SERVER.
  • close() — deletes the server chat when applicable.

Return types

Dataclasses returned by ask and attached to session calls.

RunResult

text, raw, metrics (MetricsView), trace (TraceView), context (ContextView). Method to_dict().

MetricsView

cost_usd, latency_ms, cache_hit, tokens_*, savings_ratio (all optional).

TraceView

request_id, provider, model.

ContextView

history_mode, chat_id, layers_used.

Errors

Import from tokensaver_sdk.errors. All subclasses expose code, message, status_code, request_id, raw.

  • TokenSaverError — base class.
  • AuthenticationError, ValidationError, ServerError, TimeoutError
  • Transport / DNS — ServerError with code=NETWORK_ERROR when the host cannot be reached (wrong base_url, DNS failure, TLS, connection refused). Not an HTTP response from TokenSaver.
  • ProviderKeyMissingError — extra field provider
  • Hosted default URL — ValidationError with code=HOSTED_LLM_PROVIDER (ERROR_HOSTED_LLM_PROVIDER) if provider is not OpenAI on the public API root (see Pipeline above).
  • Model catalogue — ValidationError with code=LLM_MODEL_NOT_SUPPORTED (ERROR_LLM_MODEL_NOT_SUPPORTED) if provider / model are not an active row in the platform LLM reference (same rule as GET /api/v1/llm-reference/models).
  • QuotaExceededError — quota_dimension, limit, current_usage, retry_after_seconds
  • RateLimitError — retry_after_seconds
  • ValidationError with code=PLAN_CACHE_DISABLED (and PLAN_RAG_DISABLED, …) when creating enabled module policies on a plan that excludes the feature.
  • Client-side RAG path — rag_upload_document / rag_upload_and_wait / rag_ensure_document raise ValidationError with code=RAG_FILE_NOT_FOUND if the path is missing, or code=RAG_UNSUPPORTED_FILE_TYPE (ERROR_RAG_UNSUPPORTED_FILE_TYPE) if the extension is not in RAG_UPLOAD_EXTENSIONS. Compare with ERROR_RAG_FILE_NOT_FOUND; raw["path"] may hold the resolved path.

Package imports

Public surface matches tokensaver_sdk.__all__ (stable imports for docs and IDEs).

from tokensaver_sdk import (
    TokenSaver,
    TokenSaverClient,
    API_PIPELINE_LLM_PROVIDERS,
    DEFAULT_PUBLIC_API_BASE_URL,
    HOSTED_SAAS_LLM_PROVIDERS,
    RAG_UPLOAD_EXTENSIONS,
    mime_type_for_rag_filename,
    ERROR_HOSTED_LLM_PROVIDER,
    ERROR_LLM_MODEL_NOT_SUPPORTED,
    ERROR_RAG_FILE_NOT_FOUND,
    ERROR_RAG_UNSUPPORTED_FILE_TYPE,
    HISTORY_NONE,
    HISTORY_LOCAL,
    HISTORY_SERVER,
    HistoryMode,
    RunResult,
    MetricsView,
    TraceView,
    ContextView,
    GOVERNANCE_POLICY_KINDS,
    MODULE_POLICY_KINDS,
    POLICY_KIND_TO_PLAN_FEATURE,
    RagOptions,
    PiiOptions,
    CachePolicyConfig,
    RagPolicyConfig,
    CompressionPolicyConfig,
    PiiPolicyConfig,
    PolicyKind,
    ModulePolicyKind,
    TokenSaverError,
    AuthenticationError,
    ProviderKeyMissingError,
    QuotaExceededError,
    RateLimitError,
    ValidationError,
    ServerError,
    TimeoutError,
)
import tokensaver_sdk

print(tokensaver_sdk.__version__)

History constants

HISTORY_NONE    # "none"   — no memory
HISTORY_LOCAL   # "local"  — SDK process memory
HISTORY_SERVER  # "server" — persisted on TokenSaver