From 9d7ae179b3ff11fb837c43c8d0a89639375b093f Mon Sep 17 00:00:00 2001 From: "claude[bot]" <41898282+claude[bot]@users.noreply.github.com> Date: Fri, 2 Oct 2026 13:49:54 +0000 Subject: [PATCH 1/2] update changelog and docs from OCS PR #4601 Co-Authored-By: Claude Opus 5 (1M context) --- docs/changelog.md | 1 + docs/concepts/collections/indexed.md | 20 ++++++ docs/how-to/import_csv_tsv_rows.md | 76 +++++++++++++++++++++++ docs/tech-hub/local-index-optimization.md | 16 +++++ mkdocs.yml | 1 + 5 files changed, 114 insertions(+) create mode 100644 docs/how-to/import_csv_tsv_rows.md diff --git a/docs/changelog.md b/docs/changelog.md index 039606e9..eab1c0fd 100644 --- a/docs/changelog.md +++ b/docs/changelog.md @@ -12,6 +12,7 @@ hide: Looking for older entries? See the [GitHub release notes](https://github.com/dimagi/open-chat-studio-docs/releases). ## Oct 2, 2026 +* **NEW** An [indexed collection](concepts/collections/indexed.md#importing-csvtsv-rows) on a local index can take a `.csv` or `.tsv` through the new **Import CSV/TSV rows** entry under **Add Files**, indexing each row as its own searchable record rather than chunking the whole sheet as text. You tick the columns to keep as metadata on each row — up to 16 — and the dialog shows the detected headers, the row count and the first five rows as they will be indexed before anything is uploaded. The chunks page and search results both show a row's number and its metadata, so a chatbot can look up individual records in the sheet. An import can hold up to 10,000 rows, each under 2,000 tokens once rendered, and both limits are checked at preview. Remote (OpenAI) indexes do their own chunking and don't accept row import. See [Import CSV/TSV Rows into a Collection](how-to/import_csv_tsv_rows.md). * **NEW** **Generate Chat Export** on a chatbot's **Sessions** tab now opens a dialog where you choose which columns go into the CSV. Every column is selected to begin with, and **Select all** and **Clear all** change them all at once. Message ID and Message Type are always included. The file keeps the export's usual column order whatever order you tick the columns in. * **NEW** The **Migration** card in the **Data** section of [Team Settings](concepts/team/index.md) now lets Team Admins choose what a [team migration](tech-hub/migrate_team.md) exports: the whole team, or selected chatbots. Selecting a chatbot includes all of its versions, and team members, tags, pricing rules and notifications are always exported for the whole team. When a selection is active, the migration banner says how many chatbots are being migrated, and migration mode stops triggers and scheduled messages only for those chatbots — the rest of the team keeps running. The card warns you when a change would leave a chatbot firing on both servers. * **CHANGE** `sync_team` now names the chatbots it migrated and states what it did not migrate, and a partial sync no longer requires the team's files bundle — it fetches any missing files from the source server as it goes. `sync_team --force-delete` still deletes the whole local team before re-importing, whether the sync is partial or not. diff --git a/docs/concepts/collections/indexed.md b/docs/concepts/collections/indexed.md index 9c060e69..01e7e767 100644 --- a/docs/concepts/collections/indexed.md +++ b/docs/concepts/collections/indexed.md @@ -79,12 +79,32 @@ Local indexes are hosted and managed by OCS. When you create a local index, you - **Supported file types**: pdf, txt, csv, docx - **Supported embedding models**: You can see the list of embedding models for the LLM provider you have selected. +`.tsv` files are not accepted through regular file upload — only through [Importing CSV/TSV rows](#importing-csvtsv-rows). + ### Chunking and Optimization When you upload a document to a local index, OCS breaks it into smaller parts called **chunks** and stores them in the index. The default chunking settings work well for most use cases. For advanced configuration — including chunk size, chunk overlap, and embedding model selection — see [Local Index Optimization](../../tech-hub/local-index-optimization.md). +### Importing CSV/TSV rows + +You can import a `.csv` or `.tsv` file so each row becomes its own searchable record, instead of the whole file being chunked as plain text. +Use **Add Files → Import CSV/TSV rows** on a local-index collection. +Uploading a CSV through regular file upload still indexes it as text chunks, so row import is an alternative rather than a change to how uploads behave. + +Each row is indexed as one chunk. +It renders as the sheet name, followed by `column: value` lines for the metadata columns you chose. +Search results include the row number and chosen metadata next to the row's text. +This lets a chatbot answer questions like "find the record like this one" against the sheet. + +!!! note "Local indexes only" + Row import works for local indexes only. + Remote (OpenAI) indexes do their own chunking and don't accept CSV/TSV row import. + Uploading a CSV to a remote index still indexes it as regular text chunks. + +See [Import CSV/TSV Rows into a Collection](../../how-to/import_csv_tsv_rows.md) for step-by-step instructions. + ## Document Sources for Indexed Collections Instead of uploading files manually, you can connect OCS to an external document source — such as a Confluence space or GitHub repository — and have it fetch and index content automatically on a schedule. This keeps your OCS indexed collection (for both remote and local indexes) current without manual uploads. diff --git a/docs/how-to/import_csv_tsv_rows.md b/docs/how-to/import_csv_tsv_rows.md new file mode 100644 index 00000000..ed6389f9 --- /dev/null +++ b/docs/how-to/import_csv_tsv_rows.md @@ -0,0 +1,76 @@ +--- +title: Import CSV/TSV Rows into a Collection +--- + +# Import CSV/TSV Rows into a Collection + +This guide shows you how to import a `.csv` or `.tsv` file into an indexed collection. +Each row becomes its own searchable record, instead of the whole file being chunked as plain text. +Use this when your chatbot needs to look up individual rows, such as matching a participant's input against a product list or reference table. + +For a conceptual overview, see [Importing CSV/TSV rows](../concepts/collections/indexed.md#importing-csvtsv-rows). + +## Prerequisites + +- An [indexed collection](../concepts/collections/indexed.md) using a **[Local Index](../concepts/collections/indexed.md#local-index)**. Remote (OpenAI) indexes do their own chunking and do not accept CSV/TSV row import. +- A `.csv` or `.tsv` file, within the standard file upload size limit. + +## Import a file + +1. Open your indexed collection and click **Add Files**. +2. Select **Import CSV/TSV rows**. +3. Choose your `.csv` or `.tsv` file. +4. In the dialog, tick the columns you want to keep as metadata on each row's record. You can keep up to 16 metadata columns. +5. Review the preview — it shows the detected headers, the number of rows found, and the first five rows rendered as they will be indexed. +6. Click **Import** to upload the file and start indexing. + +Each row is indexed as one chunk, rendered as the sheet name followed by `column: value` lines for the metadata columns you chose. + +!!! note "Row numbers count from the first data row" + Row numbers in the preview, the chunks page, and search results count data rows starting at 1. + "Row 4" is the fifth line of the file once the header row is counted. + +## Viewing imported rows + +On the file's chunks page, each imported row appears as its own chunk, showing its row number and the metadata you chose to keep. + +When your chatbot searches the collection, search results include the row number and the chosen metadata alongside the row's text. +This lets a chatbot answer questions like "find the record like this one" against the sheet. + +## File requirements + +- Headers must be unique and non-blank. +- Every row must have the same number of cells as the header row. A row with a different number of cells (a ragged row) is rejected, and its row number is reported. +- Files must be UTF-8 encoded. A byte-order mark at the start of the file is handled automatically. +- The delimiter for `.csv` files is detected automatically. `.tsv` files must be tab-separated. + +## Limits + +- A single import can hold up to 10,000 rows. +- Each row must be under 2,000 tokens once rendered as `column: value` lines. + +Both limits are checked when you preview the file, so an oversized row is reported before indexing starts rather than partway through. +Self-hosted operators who need different limits can override the application settings — see [Local Index Optimization](../tech-hub/local-index-optimization.md#csvtsv-row-import-limits). + +## Common issues + +### Some rows did not import + +If some rows in the file fail to embed, the rest of the sheet still indexes. +The failed rows are listed on the file's status tooltip in the collection's file list. + +The file is still marked **completed**, since most of its rows indexed successfully, so **Retry Failed Uploads** does not pick it up. +To recover the missing rows, delete the file and import it again. + +If every row in a batch fails, the whole file fails and stays retryable, so **Retry Failed Uploads** picks it up as usual. + +### The file was rejected before importing + +Check the [file requirements](#file-requirements) above. +A common cause is a ragged row, a non-unique or blank header, or a file that isn't UTF-8 encoded. +The error message reports the row number where the problem was found. + +## See also + +- [Indexed Collection for RAG](../concepts/collections/indexed.md) — how indexed collections and local indexes work +- [Local Index Optimization](../tech-hub/local-index-optimization.md) — chunking configuration and row import limits diff --git a/docs/tech-hub/local-index-optimization.md b/docs/tech-hub/local-index-optimization.md index e70d5636..6692db20 100644 --- a/docs/tech-hub/local-index-optimization.md +++ b/docs/tech-hub/local-index-optimization.md @@ -40,3 +40,19 @@ In most cases the default chunking strategy works well. You can customise it per - **Structured data (tables, forms)**: Experiment with overlap settings — tables often lose meaning when split mid-row. !!! warning "Changing the chunking strategy after upload requires re-indexing your files." + +## CSV/TSV row import limits + +Importing a `.csv` or `.tsv` file as individual rows is an alternative to chunking. +It's available for local indexes only. +See [Importing CSV/TSV rows](../concepts/collections/indexed.md#importing-csvtsv-rows) for the concept, and [Import CSV/TSV Rows into a Collection](../how-to/import_csv_tsv_rows.md) for the steps. + +Three limits apply to a row import: + +| Limit | Default | What it controls | +|---|---|---| +| `COLLECTION_ROW_IMPORT_MAX_ROWS` | 10,000 | Maximum number of rows a single import can hold | +| `COLLECTION_ROW_IMPORT_MAX_ROW_TOKENS` | 2,000 | Maximum tokens allowed per rendered row before it's rejected | +| `COLLECTION_FILE_MAX_METADATA_COLUMNS` | 16 | Maximum number of columns that can be kept as metadata per row | + +These are constants in the application settings, not environment variables, so self-hosted operators who need different limits must override the settings values in their deployment. diff --git a/mkdocs.yml b/mkdocs.yml index 6aae027a..8b177ea5 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -105,6 +105,7 @@ nav: - Add a Custom LLM Model: how-to/add_custom_llm_model.md - how-to/workflow_cookbook.md - how-to/add_a_knowledge_base.md + - Import CSV/TSV Rows into a Collection: how-to/import_csv_tsv_rows.md - Router Nodes: - how-to/routers/index.md - Configure LLM Router: how-to/routers/llm_router.md From fd5b17766934fccfe9c3bbc4d36e9358856fadf7 Mon Sep 17 00:00:00 2001 From: barry47products Date: Wed, 7 Oct 2026 09:37:45 +0200 Subject: [PATCH 2/2] Correct CSV row import details against the shipped behaviour --- docs/changelog.md | 2 +- docs/concepts/collections/indexed.md | 15 ++++++------ docs/how-to/import_csv_tsv_rows.md | 36 +++++++++++++++------------- 3 files changed, 28 insertions(+), 25 deletions(-) diff --git a/docs/changelog.md b/docs/changelog.md index eab1c0fd..a9e94905 100644 --- a/docs/changelog.md +++ b/docs/changelog.md @@ -12,7 +12,7 @@ hide: Looking for older entries? See the [GitHub release notes](https://github.com/dimagi/open-chat-studio-docs/releases). ## Oct 2, 2026 -* **NEW** An [indexed collection](concepts/collections/indexed.md#importing-csvtsv-rows) on a local index can take a `.csv` or `.tsv` through the new **Import CSV/TSV rows** entry under **Add Files**, indexing each row as its own searchable record rather than chunking the whole sheet as text. You tick the columns to keep as metadata on each row — up to 16 — and the dialog shows the detected headers, the row count and the first five rows as they will be indexed before anything is uploaded. The chunks page and search results both show a row's number and its metadata, so a chatbot can look up individual records in the sheet. An import can hold up to 10,000 rows, each under 2,000 tokens once rendered, and both limits are checked at preview. Remote (OpenAI) indexes do their own chunking and don't accept row import. See [Import CSV/TSV Rows into a Collection](how-to/import_csv_tsv_rows.md). +* **NEW** An [indexed collection](concepts/collections/indexed.md#importing-csvtsv-rows) on a local index can take a `.csv` or `.tsv` through the new **Import CSV/TSV rows** entry under **Add Files**, indexing each row as its own searchable record. Regular file upload doesn't accept CSV or TSV files, so this is how to add them. Choosing a file shows a preview with the row count and the first five rows as they will be indexed, and you tick up to 16 columns to store as metadata on each row. The chunks page shows each row's number and metadata, and a chatbot's collection search returns them with each matching row. An import can hold up to 10,000 rows, each at most 2,000 tokens once rendered, and both limits are checked at preview. Remote (OpenAI) indexes don't offer row import. See [Import CSV/TSV Rows into a Collection](how-to/import_csv_tsv_rows.md). * **NEW** **Generate Chat Export** on a chatbot's **Sessions** tab now opens a dialog where you choose which columns go into the CSV. Every column is selected to begin with, and **Select all** and **Clear all** change them all at once. Message ID and Message Type are always included. The file keeps the export's usual column order whatever order you tick the columns in. * **NEW** The **Migration** card in the **Data** section of [Team Settings](concepts/team/index.md) now lets Team Admins choose what a [team migration](tech-hub/migrate_team.md) exports: the whole team, or selected chatbots. Selecting a chatbot includes all of its versions, and team members, tags, pricing rules and notifications are always exported for the whole team. When a selection is active, the migration banner says how many chatbots are being migrated, and migration mode stops triggers and scheduled messages only for those chatbots — the rest of the team keeps running. The card warns you when a change would leave a chatbot firing on both servers. * **CHANGE** `sync_team` now names the chatbots it migrated and states what it did not migrate, and a partial sync no longer requires the team's files bundle — it fetches any missing files from the source server as it goes. `sync_team --force-delete` still deletes the whole local team before re-importing, whether the sync is partial or not. diff --git a/docs/concepts/collections/indexed.md b/docs/concepts/collections/indexed.md index 01e7e767..bf68ecf5 100644 --- a/docs/concepts/collections/indexed.md +++ b/docs/concepts/collections/indexed.md @@ -79,7 +79,7 @@ Local indexes are hosted and managed by OCS. When you create a local index, you - **Supported file types**: pdf, txt, csv, docx - **Supported embedding models**: You can see the list of embedding models for the LLM provider you have selected. -`.tsv` files are not accepted through regular file upload — only through [Importing CSV/TSV rows](#importing-csvtsv-rows). +`.csv` and `.tsv` files are not accepted through regular file upload. Add them through [Importing CSV/TSV rows](#importing-csvtsv-rows) instead. ### Chunking and Optimization @@ -89,19 +89,18 @@ For advanced configuration — including chunk size, chunk overlap, and embeddin ### Importing CSV/TSV rows -You can import a `.csv` or `.tsv` file so each row becomes its own searchable record, instead of the whole file being chunked as plain text. +You can import a `.csv` or `.tsv` file so each row becomes its own searchable record. Use **Add Files → Import CSV/TSV rows** on a local-index collection. -Uploading a CSV through regular file upload still indexes it as text chunks, so row import is an alternative rather than a change to how uploads behave. +This is the only way to add a CSV or TSV file to an indexed collection, because regular file upload doesn't accept them. Each row is indexed as one chunk. -It renders as the sheet name, followed by `column: value` lines for the metadata columns you chose. -Search results include the row number and chosen metadata next to the row's text. -This lets a chatbot answer questions like "find the record like this one" against the sheet. +The chunk text is the file name followed by one `column: value` line for every column in the row. +The columns you tick are also stored as metadata on the row's chunk. +When a chatbot searches the collection, each matching row comes back with its row number and its metadata. !!! note "Local indexes only" Row import works for local indexes only. - Remote (OpenAI) indexes do their own chunking and don't accept CSV/TSV row import. - Uploading a CSV to a remote index still indexes it as regular text chunks. + Remote (OpenAI) indexes don't offer the **Import CSV/TSV rows** option. See [Import CSV/TSV Rows into a Collection](../../how-to/import_csv_tsv_rows.md) for step-by-step instructions. diff --git a/docs/how-to/import_csv_tsv_rows.md b/docs/how-to/import_csv_tsv_rows.md index ed6389f9..ff122e50 100644 --- a/docs/how-to/import_csv_tsv_rows.md +++ b/docs/how-to/import_csv_tsv_rows.md @@ -5,51 +5,55 @@ title: Import CSV/TSV Rows into a Collection # Import CSV/TSV Rows into a Collection This guide shows you how to import a `.csv` or `.tsv` file into an indexed collection. -Each row becomes its own searchable record, instead of the whole file being chunked as plain text. +Each row becomes its own searchable record. +Regular file upload doesn't accept CSV or TSV files, so row import is how you add them. Use this when your chatbot needs to look up individual rows, such as matching a participant's input against a product list or reference table. For a conceptual overview, see [Importing CSV/TSV rows](../concepts/collections/indexed.md#importing-csvtsv-rows). ## Prerequisites -- An [indexed collection](../concepts/collections/indexed.md) using a **[Local Index](../concepts/collections/indexed.md#local-index)**. Remote (OpenAI) indexes do their own chunking and do not accept CSV/TSV row import. +- An [indexed collection](../concepts/collections/indexed.md) using a **[Local Index](../concepts/collections/indexed.md#local-index)**. Remote (OpenAI) indexes don't offer row import. - A `.csv` or `.tsv` file, within the standard file upload size limit. ## Import a file 1. Open your indexed collection and click **Add Files**. 2. Select **Import CSV/TSV rows**. -3. Choose your `.csv` or `.tsv` file. -4. In the dialog, tick the columns you want to keep as metadata on each row's record. You can keep up to 16 metadata columns. -5. Review the preview — it shows the detected headers, the number of rows found, and the first five rows rendered as they will be indexed. +3. Choose your `.csv` or `.tsv` file. A preview loads as soon as you choose it. +4. Check the preview. It shows the number of rows found and the first five rows as they will be indexed. If a row is too long, the preview names it and **Import** is disabled. +5. Under **Store as metadata**, tick the columns you want stored with each row. You can tick up to 16. 6. Click **Import** to upload the file and start indexing. -Each row is indexed as one chunk, rendered as the sheet name followed by `column: value` lines for the metadata columns you chose. +Each row is indexed as one chunk. +The chunk text is the file name followed by one `column: value` line for every column, whether or not you ticked it. !!! note "Row numbers count from the first data row" Row numbers in the preview, the chunks page, and search results count data rows starting at 1. "Row 4" is the fifth line of the file once the header row is counted. + Blank rows are skipped but keep their place in the count, so the numbers can have gaps. ## Viewing imported rows On the file's chunks page, each imported row appears as its own chunk, showing its row number and the metadata you chose to keep. -When your chatbot searches the collection, search results include the row number and the chosen metadata alongside the row's text. -This lets a chatbot answer questions like "find the record like this one" against the sheet. +When your chatbot searches the collection, each matching row comes back with its row number and the metadata you chose. +The collection's **Index Inspector** page (the search icon) doesn't show the row number or metadata. ## File requirements - Headers must be unique and non-blank. - Every row must have the same number of cells as the header row. A row with a different number of cells (a ragged row) is rejected, and its row number is reported. -- Files must be UTF-8 encoded. A byte-order mark at the start of the file is handled automatically. -- The delimiter for `.csv` files is detected automatically. `.tsv` files must be tab-separated. +- The file must have at least one data row. +- UTF-8 is tried first, with or without a byte-order mark. If the file isn't UTF-8, OCS tries to detect its encoding and rejects the file only if detection fails. +- The delimiter for `.csv` files is detected automatically from comma, semicolon, tab and pipe. `.tsv` files must be tab-separated. ## Limits - A single import can hold up to 10,000 rows. -- Each row must be under 2,000 tokens once rendered as `column: value` lines. +- Each row can be at most 2,000 tokens once rendered as `column: value` lines. -Both limits are checked when you preview the file, so an oversized row is reported before indexing starts rather than partway through. +Both limits are checked when you preview the file, so an oversized row is reported before indexing starts. Self-hosted operators who need different limits can override the application settings — see [Local Index Optimization](../tech-hub/local-index-optimization.md#csvtsv-row-import-limits). ## Common issues @@ -57,9 +61,9 @@ Self-hosted operators who need different limits can override the application set ### Some rows did not import If some rows in the file fail to embed, the rest of the sheet still indexes. -The failed rows are listed on the file's status tooltip in the collection's file list. +The failed rows are listed on the file's status tooltip in the collection's file list. A long list is cut short. -The file is still marked **completed**, since most of its rows indexed successfully, so **Retry Failed Uploads** does not pick it up. +The file is still marked **completed** if at least one row indexed, so **Retry Failed Uploads** does not pick it up. To recover the missing rows, delete the file and import it again. If every row in a batch fails, the whole file fails and stays retryable, so **Retry Failed Uploads** picks it up as usual. @@ -67,8 +71,8 @@ If every row in a batch fails, the whole file fails and stays retryable, so **Re ### The file was rejected before importing Check the [file requirements](#file-requirements) above. -A common cause is a ragged row, a non-unique or blank header, or a file that isn't UTF-8 encoded. -The error message reports the row number where the problem was found. +Common causes are a ragged row, a duplicate or blank header, or an encoding OCS couldn't detect. +For a ragged row, the error message reports its row number. ## See also