La documentation est disponible en anglais.

Documentation / Native API

RAG documents

Upload and manage workspace documents (Connaissance), then pass document_ids in rag_options when a RAG governance policy is active. Auth: Bearer API key or JWT (console).

Base path: /api/v1/rag. Supported upload extensions: pdf, txt, md, csv, json, docx. user_id and workspace_id are derived from auth — never sent in the body.

Endpoints

Method Path Purpose
POST /rag/documents Multipart upload (file, optional name, description). 201 sync or 202 async.
GET /rag/documents List documents for the authenticated user / workspace.
GET /rag/documents/{id} Document status, metadata, chunk_config, chunks_by_type.
DELETE /rag/documents/{id} Remove document and indexed chunks.
GET /rag/documents/{id}/chunks Chunk list; ?full=true returns full text (default: 200-char previews).
GET /rag/documents/{id}/preview Stream file inline (iframe-safe); ?download=true forces attachment.
GET /rag/documents/{id}/preview-url Presigned MinIO URL (legacy / direct download).
GET /rag/embedding-spaces Chunk count per embedding space (multi-model RAG index).
GET /rag/queue-stats ARQ ingestion queue length ({ queued_jobs }).
POST /rag/queue/clear Clear pending jobs (?scope=mine|all; all = workspace admin).
GET /rag/org-storage-summary Org storage breakdown (JWT, workspace/org admin).

List response (GET /rag/documents)

Each item in documents[] includes: document_id, filename, mime_type, status (pending | processing | done | ingested | failed), chunks_count, size_bytes, metadata ( display_name, description), plus per-document embedding summary:

  • embedding_key — e.g. local:… or provider:openai/text-embedding-3-large (dominant space for the document’s chunks; may be null for legacy rows).
  • embedding_dim, embedding_compute
  • chunk_config — { size_chars, overlap_chars } (ingestion defaults: 1200 / 150).

Top-level chunk_config mirrors the same global chunking settings.

Embedding spaces (GET /rag/embedding-spaces)

Returns workspace-level RAG index occupancy (isolated vector spaces per model/dimension):

{
  "workspace_id": "uuid",
  "total_chunks": 27,
  "chunk_config": { "size_chars": 1200, "overlap_chars": 150 },
  "spaces": [
    {
      "embedding_key": "local:minilm",
      "embedding_dim": 384,
      "embedding_compute": "local",
      "entries": 18
    }
  ]
}

Document detail (GET /rag/documents/{id})

Adds chunks_by_type (text, image, table) — useful when PDF image extraction differs between environments.

Preview vs preview-url

Prefer GET …/preview for browser embedding: streams bytes with Content-Disposition: inline (same-origin console proxy). Use preview-url when you need a time-limited presigned MinIO URL.

Python SDK wrappers: rag_list_documents(), rag_upload_document(), rag_get_document(), rag_ensure_document(). Console UI: page Connaissance (/files).