Skip to content

fix: read model text files as UTF-8 regardless of platform locale - #786

Open
Aditya17-bot wants to merge 1 commit into
qdrant:mainfrom
Aditya17-bot:fix/utf8-file-reads
Open

Aditya17-bot wants to merge 1 commit into
qdrant:mainfrom
Aditya17-bot:fix/utf8-file-reads

Conversation

@Aditya17-bot

Copy link
Copy Markdown

open() without encoding= uses the locale encoding, which is cp1252 on Windows. Model files are UTF-8, so on Windows:

  • SparseTextEmbedding("Qdrant/bm25", language="russian") (also greek, arabic) fails at init with UnicodeDecodeError: 'charmap' codec can't decode byte 0x8f.
  • german and french load, but every stopword with an accent comes out garbled (über becomes über), so those stopwords are never filtered.

The same bare open() is used for miniCOIL/BM42 stopwords, the IF-SPLADE IDF file, VocabResolver vocab files, and the tokenizer/preprocessor JSON configs. Those configs can contain non-ASCII text too, e.g. SentencePiece's ▁. This PR passes encoding="utf-8" to all 15 text-mode open() calls in the package. Nothing changes on Linux/macOS, where the locale is already UTF-8.

Reproduced on Windows 11 / Python 3.12 against the files in Qdrant/bm25:

russian ERROR UnicodeDecodeError 'charmap' codec can't decode byte 0x8f in position 34
greek   ERROR UnicodeDecodeError 'charmap' codec can't decode byte 0x90 in position 136
arabic  ERROR UnicodeDecodeError 'charmap' codec can't decode byte 0x81 in position 31
german  loads, accented stopwords mismatch

With the fix, language="russian" loads 151 stopwords, and "Das ist über alles" with language="german" is all stopwords (empty embedding), as expected.

Tests: tests/test_text_file_encoding.py covers BM25 stopwords in five languages, VocabResolver txt/JSON round trips with non-ASCII words, and load_special_tokens with ▁. The test matrix runs on windows-latest, so these guard the regression there.

$ pytest tests/test_text_file_encoding.py           # with fix
8 passed
$ pytest tests/test_text_file_encoding.py           # without fix (Windows)
8 failed
$ pytest tests/test_preprocessor_utils.py tests/test_image_transform.py
50 passed

Note: #684 makes the same one-line change in Bm25._load_stopwords as part of Persian support, so whichever lands second has a trivial conflict on that line.

  • Have you followed the guidelines in our Contributing document?
  • Have you checked to ensure there aren't other open Pull Requests for the same update/change?
  • Does your submission pass the existing tests?

open() without an encoding uses the locale encoding, which is cp1252 on
Windows. Model files are UTF-8, so Bm25(language=...) raised
UnicodeDecodeError for Russian, Greek and Arabic stopwords and silently
garbled German and French ones ("über" read as "über", so it was never
filtered). The same applied to vocab files, IDF files and tokenizer and
preprocessor configs.

Pass encoding="utf-8" to every text-mode open() in the package.
@Aditya17-bot
Aditya17-bot requested a review from joein as a code owner October 6, 2026 17:53
@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 9ce226c5-c638-4a35-8d7d-5b0bdc7be642
📥 Commits

Reviewing files that changed from the base of the PR and between c0588ac and d334c6a.

📒 Files selected for processing (8)
  • fastembed/common/preprocessor_utils.py
  • fastembed/late_interaction_multimodal/colmodernvbert.py
  • fastembed/sparse/bm25.py
  • fastembed/sparse/bm42.py
  • fastembed/sparse/if_splade.py
  • fastembed/sparse/minicoil.py
  • fastembed/sparse/utils/vocab_resolver.py
  • tests/test_text_file_encoding.py

Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 6 remain after this review.


📝 Walkthrough

Priority: ➖ Normal

Change: Bug fix

Merge Risk: ⚪ Minimal · up to d334c

The affected model and vocabulary files are now read and written independently of the host locale, and the supplied tests cover non-ASCII text handling. No concrete merge-blocking risk is indicated.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 9.52% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 21 functions across 8 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main change: reading model text files as UTF-8 regardless of platform locale.
Description check ✅ Passed The description explains the Windows decoding issue, the files covered by the fix, and the tests used to validate it.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant