La documentation est disponible en anglais.
Documentation / OpenAI-compatible
Headers & body extensions
Authentication, TokenSaver pipeline extensions, and JSON merges. Chat and responses: module on/off from governance policies; thresholds via tokensaver / X-Tokensaver-Options and optional X-Tokensaver-Apply-Key-Pipeline-Defaults. Embeddings: POST /openai/v1/embeddings uses X-Tokensaver-Provider, optional tokensaver / X-Tokensaver-Options, and X-Tokensaver-Use-Embedding-Cache for vector cache. Per-request headers and body win over merged defaults.
POST/openai/v1/chat/completions
POST/openai/v1/responses
POST/openai/v1/embeddings
Authentication
| Header | Role |
|---|---|
| Authorization: Bearer <ts_…> | Primary: TokenSaver API key (not the LLM vendor secret alone). |
| X-TokenSaver-Key: <ts_…> | Alternative if you cannot use Bearer (same key). |
curl -H "Authorization: Bearer ts_..." "https://api.tokensaver.fr/openai/v1/models"
# or
curl -H "X-TokenSaver-Key: ts_..." "https://api.tokensaver.fr/openai/v1/models"
Module on/off — governance policies only
You can still pass thresholds and options in JSON headers: cache_similarity_threshold, rag_options, pii_options, compression_level, etc.
Request headers — apply key pipeline defaults (chat / responses)
For minimal OpenAI clients: when the client sends no explicit pipeline control (thresholds/options in JSON — not use_*), you may set X-Tokensaver-Apply-Key-Pipeline-Defaults: true so the server merges api_keys.pipeline_settings (thresholds, default_model, … — not module on/off). Module gates always come from governance policies.
| HTTP header | Routes | Role |
|---|---|---|
| X-Tokensaver-Apply-Key-Pipeline-Defaults | /chat/completions, /responses |
When true and no explicit module control: merge key pipeline_settings. Ignored if the client already fixed module config via body or headers. |
tokensaver & JSON headers — module options (chat / responses)
Put these keys in the JSON body tokensaver object, or in X-Tokensaver-Extensions then X-Tokensaver-Options (same keys; Options overwrites Extensions on duplicates). Only keys that exist on PipelineRunRequest (other than prompt, provider, model) are merged. Invalid JSON in the headers → 400.
| Key | Type | Module | Description |
|---|---|---|---|
| use_cache, use_rag, use_compression, use_pii_filter | boolean | — | Rejected — returns 422 MODULE_GATE_POLICY_ONLY. Use governance policies. |
| cache_similarity_threshold | float 0–1 | LLM response cache | Minimum similarity for “near” questions. Identical prompts hit without this threshold. |
| cache_embedding_compute | provider | local |
Embeddings for cache | Where to compute question embeddings for semantic response cache: OpenAI catalogue vs internal EMBEDDING_SERVICE_URL. Pair with cache_embedding_model. |
| cache_embedding_model | string | Embeddings for cache | Catalogue embedding id (e.g. openai/text-embedding-3-small) when compute is provider; fixed/local model id when local. |
| rag_similarity_threshold | float 0–1 | RAG | Minimum chunk similarity to the query. |
| rag_options | object | RAG | document_ids, top_k (1–100), optional query_image_url. See the dedicated table below. |
| compression_level | int 1–5 | Compression | Higher = stronger compression (and optional model phase when budget is configured server-side). |
| pii_options | object | PII | engine, strategy, confidence_threshold, entity_types, language, regex_fallback — see table below. |
| chat_id | string | Session | Bind the run to a server-side chat (history + instructions loaded by the backend). |
| context_layers | object | Context | Structured instruction / knowledge / interaction layers (overrides ad-hoc when provided). |
| temperature | float 0–2 | LLM | Also sent as top-level OpenAI temperature; JSON merge can override. |
| provider_api_key | string | Vendor | Per-request LLM provider secret (not persisted); overrides org keys for that call only. |
| pipeline_id | string | Pipeline | Select a named pipeline when your workspace defines several. |
Request headers — JSON & provider
X-Tokensaver-Extensions and X-Tokensaver-Options must be a single-line JSON object. Invalid JSON → 400.
| Header | Format | Role |
|---|---|---|
| X-Tokensaver-Provider | Plain string | Provider code (openai, anthropic, …) when model is ambiguous. Chat, responses, embeddings. |
| X-Tokensaver-Extensions | JSON object | Merges into the pipeline request: any PipelineRunRequest key except prompt, provider, model. |
| X-Tokensaver-Options | JSON object | Same allowed keys as Extensions. Merged after Extensions; duplicate keys are overwritten by Options. |
| X-Tokensaver-Rag-Options | JSON (CORS) | Listed for browser preflight. On /openai/v1/*, put RAG parameters in rag_options inside Options or tokensaver so they merge automatically. |
| X-Tokensaver-Context-Layers | JSON (CORS) | Listed for preflight. Prefer context_layers inside Options or tokensaver. |
OpenAI-compat: keys that count as “pipeline control”
Thresholds and options below merge into the pipeline request. Legacy use_* keys and X-Tokensaver-Use-* headers are rejected. Module on/off is always from governance policies (effective_modules).
| Key | Meaning |
|---|---|
| use_* | Rejected — use governance policies. |
| pii_options | Counts as explicit PII configuration. |
| rag_options | Counts as explicit RAG configuration. |
| cache_similarity_threshold | Counts as explicit cache tuning. |
| rag_similarity_threshold | Counts as explicit RAG tuning. |
| compression_level | Counts as explicit compression tuning. |
| context_layers | Counts as explicit pipeline context control. |
| pipeline_id | Counts as explicit pipeline selection. |
| apply_key_pipeline_defaults | Reserved control flag only — does not count as explicit module control; when true (and no explicit control), triggers merge of api_keys.pipeline_settings. |
Merge order (end state on PipelineRunRequest)
- Body built from OpenAI fields (messages → prompt, temperature, max tokens, tools, …).
tokensaverobject in the JSON body (if present) merged into the pipeline request.X-Tokensaver-ExtensionsJSON, thenX-Tokensaver-OptionsJSON (Options wins on duplicate keys). Invalid JSON →400. Only keys present onPipelineRunRequestare applied from these payloads;apply_key_pipeline_defaultsis read separately for the step below.- Legacy
X-Tokensaver-Use-*headers anduse_*keys in JSON →422 MODULE_GATE_POLICY_ONLY(not merged). - Module gates: resolved from governance policies on the API key (and plan entitlements). Optional merge of
api_keys.pipeline_settingswhenX-Tokensaver-Apply-Key-Pipeline-Defaultsis set — thresholds only, not module on/off.
rag_options (object)
| Key | Type | Description |
|---|---|---|
| document_ids | string[] | null | Restrict retrieval to these workspace document UUIDs. |
| top_k | int | null | Max chunks (1–100); overrides server default when set. |
| query_image_url | string | null | Optional image URL for multimodal RAG queries when supported. |
pii_options (object)
| Key | Type / values | Description |
|---|---|---|
| engine | gliner | spacy |
Default gliner. |
| strategy | mask | replace | remove |
How to apply detections to the text. |
| confidence_threshold | float 0–1 | Default 0.5. |
| entity_types | string[] | Presidio entity types to keep; empty = all supported. |
| language | fr | en |
For spaCy engine; default fr. |
| regex_fallback | boolean | Default true; extra regex for emails/phones, etc. |
context_layers (object)
Prefer OpenAI system + messages for the common case; use this when you need explicit layer control from integrations.
| Key | Description |
|---|---|
| instruction_context | Object with workspace_instruction, user_profile_instruction, chat_instruction (strings). |
| knowledge_context | rag_documents (string[]), tool_outputs (string[]) — usually filled by the pipeline; optional input for advanced flows. |
| interaction_context | chat_history: array of messages with role user | assistant | tool, content, optional tool_calls / tool_call_id / name. |
| token_budget | Optional caps: instructions, rag, history (ints ≥ 0). |
OpenAI JSON body (outside tokensaver)
Standard fields on POST /openai/v1/chat/completions that the adapter maps before merges:
| Field | Role |
|---|---|
| model | Resolved to provider + catalogue model (prefer provider/model_id). |
| messages | Last user text → prompt; system → instructions; prior turns → history / context layers. |
| temperature | Mapped to pipeline temperature (0–2). |
| stream | If true → SSE chat.completion.chunk stream. |
| tools | OpenAI function definitions → openai_tools (OpenAI provider only). |
| tool_choice | Mapped to openai_tool_choice. |
| parallel_tool_calls | Mapped to openai_parallel_tool_calls. |
| user | Optional end-user id for logging (OpenAI field). |
Snippets: headers vs body
httpx (Python)
import httpx, json
url = "https://api.tokensaver.fr/openai/v1/chat/completions"
headers = {
"Authorization": "Bearer ts_...",
"Content-Type": "application/json",
"X-Tokensaver-Options": json.dumps({"cache_similarity_threshold": 0.9}),
}
body = {"model": "openai/gpt-4o", "messages": [{"role": "user", "content": "Hi"}]}
r = httpx.post(url, headers=headers, json=body, timeout=120.0)
r.raise_for_status()
print(r.json()["choices"][0]["message"]["content"])
OpenAI SDK + default_headers
import json
from openai import OpenAI
client = OpenAI(
api_key="ts_...",
base_url="https://api.tokensaver.fr/openai/v1",
default_headers={
"X-Tokensaver-Options": json.dumps({
"rag_similarity_threshold": 0.55,
"rag_options": {"document_ids": ["<uuid>"], "top_k": 8},
}),
},
)
print(client.chat.completions.create(
model="openai/gpt-4o",
messages=[{"role": "user", "content": "What does the doc say?"}],
).choices[0].message.content)
Node (fetch)
const opts = JSON.stringify({
compression_level: 4,
});
const r = await fetch("https://api.tokensaver.fr/openai/v1/chat/completions", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.TS_KEY,
"Content-Type": "application/json",
"X-Tokensaver-Options": opts,
},
body: JSON.stringify({
model: "openai/gpt-4o",
messages: [{ role: "user", content: "Short summary." }],
}),
});
console.log(await r.json());
LangChain (default_headers)
import json
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="openai/gpt-4o",
api_key="ts_...",
base_url="https://api.tokensaver.fr/openai/v1",
default_headers={
"X-Tokensaver-Options": json.dumps({
"cache_similarity_threshold": 0.88,
"rag_similarity_threshold": 0.55,
"rag_options": {"document_ids": ["<uuid>"], "top_k": 6},
}),
},
)
Response headers
Typical HTTP API responses (including /openai/v1/* JSON) include correlation and version headers. Some proxy or chat paths also forward pipeline diagnostics for UIs.
| Header | Typical use |
|---|---|
| X-Tokensaver-Api-Version | Backend API version string. |
| X-Request-Id | Correlation id for support and logs (also on errors). |
| X-Tokensaver-Cache-Hit | true / false when the response path exposes cache outcome (e.g. some console proxies). |
| X-Tokensaver-Token-Metrics | Structured token / cost metrics (optional encoding via X-Tokensaver-Token-Metrics-Encoding). |
| X-Tokensaver-RAG-Sources | JSON list of RAG source snippets when exposed by the integration path. |
JSON chat.completion may include a tokensaver object (e.g. model_resolved, metadata) when the pipeline returns metadata.
POST /openai/v1/embeddings (body + TokenSaver options)
No chat pipeline. Merge order for TokenSaver-specific options: defaults from the API key’s pipeline_settings (console), then X-Tokensaver-Extensions, then X-Tokensaver-Options, then body tokensaver (later wins). Header X-Tokensaver-Use-Embedding-Cache, when present, overrides the resolved use_embedding_cache boolean for this request.
| tokensaver / JSON key | Type | Description |
|---|---|---|
| use_embedding_cache | boolean | Enable Redis exact cache of embedding vectors. Same effect as X-Tokensaver-Use-Embedding-Cache when the header is set (header wins if both are sent). |
| embedding_compute | local | other |
local → internal embedding service (EMBEDDING_SERVICE_URL). Otherwise default OpenAI (or org) path. Aliases such as embedding_compute_backend / use_internal_embedding_service are also recognised by the server. |
| (from key defaults) | — | Per-key cache_embedding_compute and use_embedding_cache from the console apply when the request does not override them. |
OpenAI-shaped body (required fields):
| Field | Description |
|---|---|
| model | Embedding catalogue id (provider/model_id). |
| input | String or array of strings to embed. |
| tokensaver | Optional object; merged last with the rules above. |
| encoding_format | Optional; default float (extra fields ignored by schema). |
| dimensions | Optional; ignored if not applicable to the internal embedding service. |
CORS
Preflight allows the X-Tokensaver-* headers listed above, including X-Tokensaver-Apply-Key-Pipeline-Defaults (plus X-TokenSaver-Key) so browser apps can send extensions from another origin when the API CORS policy permits.