# RAG (/docs/ai/rag)



`Files & RAG` in AI is the capability chain that turns uploaded documents into grounded context for chat, assistants, and other retrieval-enabled workflows.

At a high level, the flow is:

1. a user uploads a file through the platform
2. the backend stores the source file and creates a library record
3. a durable background workflow parses and chunks the document
4. embeddings are generated and written to `Milvus` in bounded batches
5. chat and assistant flows retrieve matching chunks later through RAG

## What This Capability Covers [#what-this-capability-covers]

This capability is broader than file upload alone.

It includes:

* the user-facing file library in the `Portal`
* file metadata and status tracking in the backend
* source-file storage in the configured `S3`-compatible object store
* background indexing through `Temporal`
* document chunking and normalization through `Docling`
* vector storage and similarity search in `Milvus`
* retrieval use inside chat and assistant flows

## User-Facing Entry Points [#user-facing-entry-points]

The main user-facing entry points are:

* the `Files` section in the `Portal`
* the chat `RAG` picker in the portal composer

The Files area is where users upload, organize, rename, delete, reprocess, and share content.

The chat `RAG` control is where users decide which files or folders should be used as retrieval context for a conversation.

## Core Processing Flow [#core-processing-flow]

The backend implementation currently spans multiple runtime layers.

<Mermaid
  chart="flowchart TD
    A[Portal Files UI] --> B[`/v1/files/upload` or `/v1/files/upload-stream`]
    B --> C[Store source file in S3-compatible object storage]
    C --> D[Create `library_files` row in Postgres with `rag_pending`]
    D --> E[Start Temporal workflow on `file-events-queue`]

    E --> F[file-event worker]
    F --> G[Set status to `rag_processing`]
    G --> H[Remove older vectors for the same source]
    H --> I[Download source file]
    I --> J[Parse and chunk the document]
    J --> K[Docling API and Redis-backed RQ job queue for supported formats]
    K --> L[Write chunk and artifact files to the artifacts bucket]
    L --> M[Generate embeddings and insert them into the tenant collection]
    M --> N[Set status to `rag_available`]

    N --> O[Later chat retrieval]
    O --> P[Use selected file or folder scope]
    P --> Q[Resolve target document IDs]
    Q --> R[Batch authorization checks in OpenFGA]
    R --> S[Search the tenant's active Milvus collection]
    S --> T[Pass matching chunks to the model as grounded context]"
/>

## Where The Data Lives [#where-the-data-lives]

Different parts of the capability live in different systems.

| Data type                                 | Main system                                                       |
| ----------------------------------------- | ----------------------------------------------------------------- |
| Original uploaded file                    | Configured `S3`-compatible storage, either in-cluster or external |
| File library metadata and status          | `PostgreSQL`                                                      |
| Chunk JSON and derived artifacts          | `{bucket}-artifacts` object-storage bucket                        |
| Searchable vectors and retrieval payloads | `Milvus`                                                          |

This boundary matters.

The file row shown in the portal is not the retrieval index itself. The portal library is driven by database metadata, while actual retrieval depends on vectorized chunks in `Milvus`.

## Tenant-scoped vector collections [#tenant-scoped-vector-collections]

BullSequana AI uses one shared Milvus service and separate platform-managed collections for each tenant. It does not deploy a separate Milvus instance for every tenant.

The backend derives a collection name from the canonical tenant UUID, embedding model, and embedding dimension. The first ingestion that needs the collection ensures that it exists. Changing the embedding model or dimension creates a new tenant-scoped target collection through the reindex workflow.

Collection listing, statistics, document inspection, and deletion operations expose only names in the active tenant's collection namespace. Collection search and Responses API file search accept only the tenant's active collection. A foreign or inactive collection identifier is returned as not found, and Responses API file search requires exactly one vector store identifier.

An upgraded tenant can continue using its previously configured legacy collection until a reindex completes. The reindex switches that tenant to the derived collection and preserves the legacy collection instead of deleting a resource that another tenant might still use.

## Status Lifecycle [#status-lifecycle]

Files move through a visible RAG lifecycle.

* `rag_pending`: upload succeeded and background indexing is queued
* `rag_processing`: the worker is actively chunking, embedding, and indexing the file
* `rag_available`: the file is indexed and can be used for retrieval
* `rag_failed`: indexing failed
* `rag_removed`: the file was removed from the retrieval lifecycle

This is why upload success is not the same as retrieval readiness.

A file can appear in the portal library before it is usable in grounded chat.

## How Uploads Become Retrieval-Ready [#how-uploads-become-retrieval-ready]

The backend upload path currently uses:

* `POST /v1/files/upload`
* `POST /v1/files/upload-stream`

These endpoints do more than store bytes.

They also:

* compute a deterministic object key
* create or update the corresponding library record
* set the initial status to `rag_pending`
* start a `Temporal` workflow on `file-events-queue`

The actual indexing work does not happen inside the upload request itself.

Instead, a separate file-event worker performs the expensive background steps later. If that worker is not running, files can remain stuck in `rag_pending` even though the upload returned success.

## Docling, Chunking, And Embeddings [#docling-chunking-and-embeddings]

Supported file types include documents (`pdf`, `docx`, `md`, `pptx`, `txt`, `vtt`), data files (`json`, `csv`, `xlsx`), code files (`py`, `ts`, `js`, `java`, `go`, `rs`, and others), and config files (`yaml`, `yml`, `toml`, `ini`). OCR is enabled by default with automatic engine selection.

For standard document formats, the worker uses `Docling`-based processing to chunk and normalize the source material.

That processing can also produce additional artifacts such as chunk JSON and, where applicable, metadata, images, or tables.

The 1.3 processing path includes several safeguards for large documents, busy queues, and transient service failures:

* Docling conversion runs as a queued job, so processing is not tied to one API request or backend pod
* the worker sends liveness heartbeats while it polls Docling and while it embeds batches
* queue wait is treated as normal load, with separate execution and end-to-end workflow budgets
* missing Docling task IDs fail immediately, while a small number of transient polling failures are tolerated before the job fails
* OCR selection is resolved once per document and reused for parsing and artifact extraction
* the chunk tokenizer follows the active embedding model, avoiding tokenizer drift after an administrator changes models
* text beyond the embedding model's input limit is split safely and combined into one vector for the source chunk
* embeddings and vector inserts run in bounded batches to reduce peak memory use on large or scanned files

Artifact extraction is best-effort: a failed image or table extraction does not discard otherwise valid text chunks.

After chunking, embeddings are generated for the chunks and inserted into `Milvus`. If ingestion ultimately fails, partial vectors from that attempt are removed and the file moves to `rag_failed`. A document that produces no chunks also fails without repeatedly rerunning the same conversion.

Those vectors are what later power retrieval.

## How Chat Uses Files Later [#how-chat-uses-files-later]

The portal chat composer has a `RAG` control that searches the file library and lets users pick:

* individual files
* folders
* all available retrieval-ready content

Only files with `rag_available` status are surfaced as directly selectable retrieval files in that picker.

When a user submits a chat request with retrieval enabled, the UI sends:

* `fileSearch`
* optional `documentIds`
* optional `folderIds`

The backend then resolves the effective retrieval scope.

Folder selections are expanded on the backend into the underlying file IDs for that user, and the final retrieval step uses those document IDs when running similarity search.

This makes the flow user-friendly in the portal while still keeping retrieval scoped to concrete indexed documents under the hood.

The selected files, folders, and file-search state are stored with the conversation. Reopening a chat restores the last RAG selection instead of silently widening or losing its retrieval scope.

## Retrieval Boundary [#retrieval-boundary]

The retrieval path is currently centered around the backend `RetrievalService`.

In practice, it:

* resolves and searches the active tenant's `Milvus` collection
* optionally filters by document IDs
* can apply a user-specific filter
* returns chunk text that is assembled into model context

The chat-oriented `file_search` tool uses this retrieval layer so that the assistant can ground responses in the selected content instead of relying only on base-model knowledge.

The same file-and-folder scoping is available to portal chat, the Responses API `file_search` tool, and collection search. Explicit file and folder selections are checked through batched `OpenFGA` authorization calls, then intersected with records in the active tenant before retrieval. If no readable documents remain, the result is a deny-all scope rather than an unfiltered search.

## Files, Folders, And Scope [#files-folders-and-scope]

The platform treats files and folders differently:

* files are the actual indexed retrieval units
* folders are a user-facing grouping and selection mechanism

That means selecting a folder in chat is really a convenient way to include the files contained in that folder hierarchy.

This also explains why the Files area and the chat `RAG` picker are so closely connected.

## Sharing And Access Scopes [#sharing-and-access-scopes]

Files and folders can be shared with other users or Keycloak groups through the portal Files area.

The backend supports three listing scopes:

| Scope            | What it returns                                                            |
| ---------------- | -------------------------------------------------------------------------- |
| `mine`           | Files and folders owned by the current user                                |
| `shared_with_me` | Files and folders accessible to the current user but owned by someone else |
| `all`            | Combined scope used internally by the chat RAG search                      |

The `all` scope ensures that the chat `RAG` picker surfaces both the user's own files and all files shared with them.

Sharing is enforced through `OpenFGA`. Org admins receive implicit access to all files in the system through an organization-level tuple written at upload time. Regular users only see files explicitly shared with them by user or group assignment.

## Reprocess And Delete [#reprocess-and-delete]

The Files area is not only for initial upload.

It also supports later lifecycle actions.

### Reprocess [#reprocess]

Reprocess reruns indexing for an already stored source file.

This is useful when:

* indexing previously failed
* processing settings changed
* the vector representation needs to be rebuilt

Reprocess does not require the user to upload the file again.

### Delete [#delete]

Delete removes the file from the library lifecycle and triggers cleanup behavior for the indexed content.

From a user perspective, the important point is that delete is not only a UI metadata operation. It also participates in the broader storage and retrieval cleanup path.

## Operational Realities [#operational-realities]

The current implementation has a few practical behaviors that matter to operators and advanced users:

* indexing is asynchronous
* upload success does not guarantee retrieval success
* `Temporal`, `Docling`, object storage, and `Milvus` all need to be available for the full pipeline to complete
* only `/v1/files/upload` and `/v1/files/upload-stream` enter the main RAG indexing path
* long conversions and temporary queue backlogs are tolerated, but bounded timeouts still turn stalled work into a visible `rag_failed` state
* ingestion, file metadata, folder expansion, and vector operations carry the active tenant context, and new collections include the tenant identity in their derived name
* a failed reprocess can temporarily leave a file without active vectors until indexing succeeds again

For day-to-day portal use, the most important takeaway is simple: wait for `RAG Available` before expecting grounded chat answers from a newly uploaded document.

## Why This Matters In AI [#why-this-matters-in-ai]

This capability is one of the main reasons AI can support grounded assistants and enterprise document workflows.

Without it, chat would only have direct model generation.

With it, the platform can:

* ground answers in uploaded content
* scope retrieval to specific files or folders
* connect document management with conversational AI
* support more governed and inspectable knowledge workflows

## Related Pages [#related-pages]

* [Portal Guide](/docs/use/portal)
* [Milvus](/docs/ai/components/milvus)
* [Docling](/docs/ai/components/docling)
* [BSQAI API](/docs/ai/components/bsqai-api)
