Docs / OpenAI-compatible

Headers & body extensions

Authentication, TokenSaver pipeline extensions, and JSON merges. Chat and responses: module on/off from governance policies; thresholds via tokensaver / X-Tokensaver-Options and optional X-Tokensaver-Apply-Key-Pipeline-Defaults. Embeddings: POST /openai/v1/embeddings uses X-Tokensaver-Provider, optional tokensaver / X-Tokensaver-Options, and X-Tokensaver-Use-Embedding-Cache for vector cache. Per-request headers and body win over merged defaults.

POST/openai/v1/chat/completions

POST/openai/v1/responses

POST/openai/v1/embeddings

Authentication

Header Role
Authorization: Bearer <ts_…> Primary: TokenSaver API key (not the LLM vendor secret alone).
X-TokenSaver-Key: <ts_…> Alternative if you cannot use Bearer (same key).
curl -H "Authorization: Bearer ts_..." "https://api.tokensaver.fr/openai/v1/models"
# or
curl -H "X-TokenSaver-Key: ts_..." "https://api.tokensaver.fr/openai/v1/models"

Module on/off — governance policies only

You can still pass thresholds and options in JSON headers: cache_similarity_threshold, rag_options, pii_options, compression_level, etc.

Request headers — apply key pipeline defaults (chat / responses)

For minimal OpenAI clients: when the client sends no explicit pipeline control (thresholds/options in JSON — not use_*), you may set X-Tokensaver-Apply-Key-Pipeline-Defaults: true so the server merges api_keys.pipeline_settings (thresholds, default_model, … — not module on/off). Module gates always come from governance policies.

HTTP header Routes Role
X-Tokensaver-Apply-Key-Pipeline-Defaults /chat/completions, /responses When true and no explicit module control: merge key pipeline_settings. Ignored if the client already fixed module config via body or headers.

tokensaver & JSON headers — module options (chat / responses)

Put these keys in the JSON body tokensaver object, or in X-Tokensaver-Extensions then X-Tokensaver-Options (same keys; Options overwrites Extensions on duplicates). Only keys that exist on PipelineRunRequest (other than prompt, provider, model) are merged. Invalid JSON in the headers → 400.

Key Type Module Description
use_cache, use_rag, use_compression, use_pii_filter boolean — Rejected — returns 422 MODULE_GATE_POLICY_ONLY. Use governance policies.
cache_similarity_threshold float 0–1 LLM response cache Minimum similarity for “near” questions. Identical prompts hit without this threshold.
cache_embedding_compute provider | local Embeddings for cache Where to compute question embeddings for semantic response cache: OpenAI catalogue vs internal EMBEDDING_SERVICE_URL. Pair with cache_embedding_model.
cache_embedding_model string Embeddings for cache Catalogue embedding id (e.g. openai/text-embedding-3-small) when compute is provider; fixed/local model id when local.
rag_similarity_threshold float 0–1 RAG Minimum chunk similarity to the query.
rag_options object RAG document_ids, top_k (1–100), optional query_image_url. See the dedicated table below.
compression_level int 1–5 Compression Higher = stronger compression (and optional model phase when budget is configured server-side).
pii_options object PII engine, strategy, confidence_threshold, entity_types, language, regex_fallback — see table below.
chat_id string Session Bind the run to a server-side chat (history + instructions loaded by the backend).
context_layers object Context Structured instruction / knowledge / interaction layers (overrides ad-hoc when provided).
temperature float 0–2 LLM Also sent as top-level OpenAI temperature; JSON merge can override.
provider_api_key string Vendor Per-request LLM provider secret (not persisted); overrides org keys for that call only.
pipeline_id string Pipeline Select a named pipeline when your workspace defines several.

Request headers — JSON & provider

X-Tokensaver-Extensions and X-Tokensaver-Options must be a single-line JSON object. Invalid JSON → 400.

Header Format Role
X-Tokensaver-Provider Plain string Provider code (openai, anthropic, …) when model is ambiguous. Chat, responses, embeddings.
X-Tokensaver-Extensions JSON object Merges into the pipeline request: any PipelineRunRequest key except prompt, provider, model.
X-Tokensaver-Options JSON object Same allowed keys as Extensions. Merged after Extensions; duplicate keys are overwritten by Options.
X-Tokensaver-Rag-Options JSON (CORS) Listed for browser preflight. On /openai/v1/*, put RAG parameters in rag_options inside Options or tokensaver so they merge automatically.
X-Tokensaver-Context-Layers JSON (CORS) Listed for preflight. Prefer context_layers inside Options or tokensaver.

OpenAI-compat: keys that count as “pipeline control”

Thresholds and options below merge into the pipeline request. Legacy use_* keys and X-Tokensaver-Use-* headers are rejected. Module on/off is always from governance policies (effective_modules).

Key Meaning
use_* Rejected — use governance policies.
pii_options Counts as explicit PII configuration.
rag_options Counts as explicit RAG configuration.
cache_similarity_threshold Counts as explicit cache tuning.
rag_similarity_threshold Counts as explicit RAG tuning.
compression_level Counts as explicit compression tuning.
context_layers Counts as explicit pipeline context control.
pipeline_id Counts as explicit pipeline selection.
apply_key_pipeline_defaults Reserved control flag only — does not count as explicit module control; when true (and no explicit control), triggers merge of api_keys.pipeline_settings.

Merge order (end state on PipelineRunRequest)

  1. Body built from OpenAI fields (messages → prompt, temperature, max tokens, tools, …).
  2. tokensaver object in the JSON body (if present) merged into the pipeline request.
  3. X-Tokensaver-Extensions JSON, then X-Tokensaver-Options JSON (Options wins on duplicate keys). Invalid JSON → 400. Only keys present on PipelineRunRequest are applied from these payloads; apply_key_pipeline_defaults is read separately for the step below.
  4. Legacy X-Tokensaver-Use-* headers and use_* keys in JSON → 422 MODULE_GATE_POLICY_ONLY (not merged).
  5. Module gates: resolved from governance policies on the API key (and plan entitlements). Optional merge of api_keys.pipeline_settings when X-Tokensaver-Apply-Key-Pipeline-Defaults is set — thresholds only, not module on/off.

rag_options (object)

Key Type Description
document_ids string[] | null Restrict retrieval to these workspace document UUIDs.
top_k int | null Max chunks (1–100); overrides server default when set.
query_image_url string | null Optional image URL for multimodal RAG queries when supported.

pii_options (object)

Key Type / values Description
engine gliner | spacy Default gliner.
strategy mask | replace | remove How to apply detections to the text.
confidence_threshold float 0–1 Default 0.5.
entity_types string[] Presidio entity types to keep; empty = all supported.
language fr | en For spaCy engine; default fr.
regex_fallback boolean Default true; extra regex for emails/phones, etc.

context_layers (object)

Prefer OpenAI system + messages for the common case; use this when you need explicit layer control from integrations.

Key Description
instruction_context Object with workspace_instruction, user_profile_instruction, chat_instruction (strings).
knowledge_context rag_documents (string[]), tool_outputs (string[]) — usually filled by the pipeline; optional input for advanced flows.
interaction_context chat_history: array of messages with role user | assistant | tool, content, optional tool_calls / tool_call_id / name.
token_budget Optional caps: instructions, rag, history (ints ≥ 0).

OpenAI JSON body (outside tokensaver)

Standard fields on POST /openai/v1/chat/completions that the adapter maps before merges:

Field Role
model Resolved to provider + catalogue model (prefer provider/model_id).
messages Last user text → prompt; system → instructions; prior turns → history / context layers.
temperature Mapped to pipeline temperature (0–2).
stream If true → SSE chat.completion.chunk stream.
tools OpenAI function definitions → openai_tools (OpenAI provider only).
tool_choice Mapped to openai_tool_choice.
parallel_tool_calls Mapped to openai_parallel_tool_calls.
user Optional end-user id for logging (OpenAI field).

Snippets: headers vs body

httpx (Python)

import httpx, json

url = "https://api.tokensaver.fr/openai/v1/chat/completions"
headers = {
    "Authorization": "Bearer ts_...",
    "Content-Type": "application/json",
    "X-Tokensaver-Options": json.dumps({"cache_similarity_threshold": 0.9}),
}
body = {"model": "openai/gpt-4o", "messages": [{"role": "user", "content": "Hi"}]}
r = httpx.post(url, headers=headers, json=body, timeout=120.0)
r.raise_for_status()
print(r.json()["choices"][0]["message"]["content"])

OpenAI SDK + default_headers

import json
from openai import OpenAI

client = OpenAI(
    api_key="ts_...",
    base_url="https://api.tokensaver.fr/openai/v1",
    default_headers={
        "X-Tokensaver-Options": json.dumps({
            "rag_similarity_threshold": 0.55,
            "rag_options": {"document_ids": ["<uuid>"], "top_k": 8},
        }),
    },
)
print(client.chat.completions.create(
    model="openai/gpt-4o",
    messages=[{"role": "user", "content": "What does the doc say?"}],
).choices[0].message.content)

Node (fetch)

const opts = JSON.stringify({
  compression_level: 4,
});
const r = await fetch("https://api.tokensaver.fr/openai/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: "Bearer " + process.env.TS_KEY,
    "Content-Type": "application/json",
    "X-Tokensaver-Options": opts,
  },
  body: JSON.stringify({
    model: "openai/gpt-4o",
    messages: [{ role: "user", content: "Short summary." }],
  }),
});
console.log(await r.json());

LangChain (default_headers)

import json
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="openai/gpt-4o",
    api_key="ts_...",
    base_url="https://api.tokensaver.fr/openai/v1",
    default_headers={
        "X-Tokensaver-Options": json.dumps({
            "cache_similarity_threshold": 0.88,
            "rag_similarity_threshold": 0.55,
            "rag_options": {"document_ids": ["<uuid>"], "top_k": 6},
        }),
    },
)

Response headers

Typical HTTP API responses (including /openai/v1/* JSON) include correlation and version headers. Some proxy or chat paths also forward pipeline diagnostics for UIs.

Header Typical use
X-Tokensaver-Api-Version Backend API version string.
X-Request-Id Correlation id for support and logs (also on errors).
X-Tokensaver-Cache-Hit true / false when the response path exposes cache outcome (e.g. some console proxies).
X-Tokensaver-Token-Metrics Structured token / cost metrics (optional encoding via X-Tokensaver-Token-Metrics-Encoding).
X-Tokensaver-RAG-Sources JSON list of RAG source snippets when exposed by the integration path.

JSON chat.completion may include a tokensaver object (e.g. model_resolved, metadata) when the pipeline returns metadata.

POST /openai/v1/embeddings (body + TokenSaver options)

No chat pipeline. Merge order for TokenSaver-specific options: defaults from the API key’s pipeline_settings (console), then X-Tokensaver-Extensions, then X-Tokensaver-Options, then body tokensaver (later wins). Header X-Tokensaver-Use-Embedding-Cache, when present, overrides the resolved use_embedding_cache boolean for this request.

tokensaver / JSON key Type Description
use_embedding_cache boolean Enable Redis exact cache of embedding vectors. Same effect as X-Tokensaver-Use-Embedding-Cache when the header is set (header wins if both are sent).
embedding_compute local | other local → internal embedding service (EMBEDDING_SERVICE_URL). Otherwise default OpenAI (or org) path. Aliases such as embedding_compute_backend / use_internal_embedding_service are also recognised by the server.
(from key defaults) — Per-key cache_embedding_compute and use_embedding_cache from the console apply when the request does not override them.

OpenAI-shaped body (required fields):

Field Description
model Embedding catalogue id (provider/model_id).
input String or array of strings to embed.
tokensaver Optional object; merged last with the rules above.
encoding_format Optional; default float (extra fields ignored by schema).
dimensions Optional; ignored if not applicable to the internal embedding service.

CORS

Preflight allows the X-Tokensaver-* headers listed above, including X-Tokensaver-Apply-Key-Pipeline-Defaults (plus X-TokenSaver-Key) so browser apps can send extensions from another origin when the API CORS policy permits.