Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
node_modules/
dist/
public/pdfjs/
playwright-report/
test-results/
src-tauri/target/
Expand Down
1 change: 1 addition & 0 deletions README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@
- **一本地书库:** PDF、EPUB、DOCX 不再需要分别使用不同阅读器。
- **本地优先:** 书库、阅读进度、笔记、索引和备份保存在你的电脑中。
- **AI 由你主动触发:** 普通阅读和本地搜索不需要 AI;只有你明确发起 AI 操作时才会发送相关内容。
- **离线 AI 问答:** 已安装 Ollama 并下载模型后,可一键连接本地文字模型;支持识图的模型还可用于图片区域问答。LM Studio 文字通道暂为实验性,实机验证安排在 v0.2.1。
- **面向系统学习:** 本地搜索、批注、阅读历史、多语言界面、备份与恢复集中在一个桌面工作流中。
- **源码完整开放:** Rust + React + Tauri 全部源码采用 Apache-2.0 协议,开发者可以自由修改和改进。

Expand Down
2 changes: 1 addition & 1 deletion docs/ai-services.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# AI services

TextbookLens supports user-configured official OpenAI, Gemini, Anthropic, DeepSeek, and Kimi services. Requests leave from the Rust desktop layer; the browser UI does not fetch providers or configure a production base URL. Adding or replacing a key validates it before a key/profile update is committed. Keys are in Windows Credential Manager, not profile lists, SQLite/frontend state, or backups.
TextbookLens supports user-configured official OpenAI, Gemini, Anthropic, DeepSeek, and Kimi services, plus [Ollama local text/image models and experimental LM Studio text models](local-models.md). Requests leave from the Rust desktop layer; the browser UI does not fetch providers or configure a production base URL. Adding or replacing a cloud key validates it before a key/profile update is committed. Cloud keys are in Windows Credential Manager, not profile lists, SQLite/frontend state, or backups. Local profiles use a discovered loopback port and require no cloud key.

## Configured capability registry

Expand Down
32 changes: 32 additions & 0 deletions docs/local-models.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# Offline local models

In **Settings → AI services**, select **Auto-connect local AI**. The same button is available during onboarding. TextbookLens discovers an existing Ollama or LM Studio installation, starts its local service when needed, finds downloaded text models, and tests a model with a short synthetic question. It adds the discovered models and selects a tested model as the default for text answers. Repeating the action updates existing profiles instead of duplicating them.

No API key is needed for the supported unauthenticated local services. The application does not install software, download models, use cloud models, or fall back to a cloud provider when a local request fails. A service that requires authentication is reported separately; TextbookLens does not change its authentication settings.

## Supported workflow

- **Ollama:** discover its executable from the normal Windows installation or PATH; use the local `OLLAMA_HOST` port when configured, otherwise 11434. Start `ollama serve` if the service is stopped. Use `/api/tags` and `/api/show` to find installed completion models. Reject cloud/remote models before each inference. Load the model on demand through `/api/chat`.
- **LM Studio (experimental; real-machine validation deferred to v0.2.1):** discover `lms.exe` from the LM Studio home or PATH and its saved local server port (default 1234). Start the server with the installed CLI. Read `lms ls --llm --json`, exclude models on other devices, and use `lms load` when needed. LM Link-aware versions are loaded with `--local`, and the resulting local instance is verified before sending a question.
- Models are stored as separate profiles, with their own model identifier, context budget, and numeric loopback port. Use **Use for learning** to switch models. Local profiles do not create or require Windows credentials.
- Local models support text selection questions, book questions, follow-up answers, and the teaching-instruction test. Ollama models whose installed metadata advertises vision can also be selected with **Use for vision** for image-region questions. Image input is checked again before every request; text-only and remote models never receive images. LM Studio vision, structured page indexing, and cloud file extraction remain unavailable for local profiles in this version.

Vision capability is saved per local profile and survives restart. One request accepts one PNG/JPEG/WebP image, at most 2 MiB encoded, 2048 pixels on either side, and 1,048,576 decoded pixels. TextbookLens reserves additional context for that image. Discovered Ollama vision models use up to 16,384 context tokens, subject to the model's declared limit. Reconnecting retains an existing connected local text default; select a vision default independently in AI services.

The software must be functional and its inference engine and model files must already be installed. Sufficient RAM/VRAM is still required. Initial model loading can take several minutes. A missing runtime, missing text model, authentication requirement, or failed answer test is shown in the connection result.

All TextbookLens local HTTP requests use `127.0.0.1`, bypass proxies, and refuse redirects. Application-level validation checks model locality as well: an Ollama localhost endpoint can serve cloud models, and LM Studio can expose remote devices through LM Link.

## References

- [Ollama model list](https://docs.ollama.com/api/tags), [cloud/local model distinction](https://docs.ollama.com/cloud).
- [LM Studio CLI model list](https://lmstudio.ai/docs/cli/local-models/ls), [model loading](https://lmstudio.ai/docs/cli/local-models/load), [server startup](https://lmstudio.ai/docs/cli/serve/server-start).
- [LM Studio CLI implementation](https://github.com/lmstudio-ai/lms): `src/subcommands/list.ts` and `src/subcommands/load.ts` define local device metadata and local-only loading.

## Verification

Synthetic loopback tests cover local/remote filtering, credential-free persistence and deletion, repeated discovery, stream termination, and rejection of redirects. Learning preparation tests cover both selection and book questions without credential access. The database migration contract covers existing profile references and historical upgrades.

An opt-in Rust test, `ai::local::tests::installed_local_runtime_offline_smoke`, exercises discovery, persistence, the normal provider runtime, and a real streamed answer using installed software and a temporary database. It never uses a user's books or cloud keys. Run it with `cargo test --manifest-path src-tauri/Cargo.toml --lib installed_local_runtime_offline_smoke -- --ignored`, using the isolated build directories described in CONTRIBUTING.md.

Verified on 2026-09-07: 55 frontend tests, 347 Rust unit tests, and 37 binding, database, migration, and profile lifecycle tests passed. The installed Ollama 0.33.3 service with `deepseek-r1:8b` passed discovery, registration, and a complete streamed answer. A separate unused loopback port verified automatic startup and discovery of already downloaded models; that test-owned service was stopped afterward. LM Studio model filtering and stream handling have automated coverage; an actual LM Studio installation was not available for an end-to-end run on this machine.
63 changes: 63 additions & 0 deletions docs/releases/v0.2.0-validation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
# v0.2.0 发布准备验证记录

2026-09-08。本记录对应 v0.2.0 公开发布候选。安装包未进行代码签名;公开文件及校验值以 [GitHub Release](https://github.com/gh615280-maker/TextbookLens/releases/tag/v0.2.0) 为准。

## 安装包

| 包 | 字节数 | 运行时处理 |
| ------------ | ----------: | ----------------------------- |
| 默认 EXE | 10,877,755 | 复用 WebView2,缺少时联网下载 |
| 默认 MSI | 13,963,264 | 同上 |
| 完整离线 EXE | 272,870,116 | 附带 WebView2 离线运行时 |
| 完整离线 MSI | 272,834,560 | 同上 |

不包含 Ollama、LM Studio 或模型权重。微软 WebView2 原始安装文件为 258,614,480 字节;新增 PDF 字体和解码资源原始总量约 3 MB。安装文件经过压缩,原始文件大小不能直接相加。

构建目录:`D:\CodexBuild\textbooklens-v020-20260908`。最终包位于 `artifacts/standard` 和 `artifacts/offline`;校验值见 `evidence/package-checksums.json`。源码基于 `9a193045dcb3a10034393786118f88cedab144f8` 加本次受控修改,完整路径与哈希见 `evidence/source-manifest.json`。

## 已完成

- PDF 字体、CMap 和扫描页解码资源离线打包,导入/阅读/页面渲染共用配置;页面解码失败显示错误。
- 修复页底提问表单越界、书库卡片/列表布局、目录跳转后工具栏不可见、调整大小手柄与阅读滚动条重叠。前后端版本统一为 0.2.0。
- EPUB 的 xmldom 更新为 0.8.15,并加入 GHSA-6gmq-8vp8-gcm6 检查;生产 npm audit 为 0 项漏洞。
- 前端和脚本 100 个文件、411 项测试通过;浏览器端到端 65 项通过;Rust Release 490 项通过、3 项默认忽略。
- 另外执行两项真实环境测试:已安装模型的一键连接/文字回答通过;独立端口 11465 的 Ollama 自动启动通过。测试服务随后关闭,正常用户服务未停止。
- TypeScript、ESLint、Prettier、Release Clippy、生成绑定字节一致性、依赖/许可/素材检查通过。最终分发内容扫描通过:285 个文件、580,170,456 字节。

## Windows 安装与升级

环境为关闭网络的 Windows Sandbox(Windows 11,构建 22621),数据与用户书库独立。

- 干净环境原先无 WebView2:完整离线 EXE 安装、首次启动和卸载通过;完整离线 MSI 安装和启动通过。
- v0.1.0 EXE 覆盖升级:数据库 15→17,已导入 DOCX、2 条批注、2 条问答消息、3 个检索片段、设置、服务配置和保存凭据逐项保留,数据库完整性与外键检查正常。
- 升级后实际打开 DOCX 并跳转目录,工具栏保持可见。
- 已有 WebView2 时,默认小 EXE/MSI 的断网安装、启动和卸载通过;MSI 卸载后凭据保留,随后清理测试凭据。
- 默认 MSI 中导入桌面《数学分析习题课讲义下》的副本并显示扫描目录页;浏览器另验证第 1、10 页非空且没有外部资源请求。
- 原桌面 8 本教材的 SHA-256 全部与测试前一致。

中间曾有前一 Sandbox 客户端未退出;清理旧客户端后重新确认当前实例,并完成小包卸载、MSI 安装、扫描教材阅读和卸载。重复启动的中间结果不算额外独立样本。

## 本地视觉补测

使用已有 Ollama/Qwen3-VL 4B,经同一源码的 Release 应用库、真实 ProviderRuntime 和本地适配层测试,数据写入专用数据库,图片只发到本机。

| 样例 | 运行结果 | 内容核对 |
| ------------------------------ | -------------------------------- | ---------------------------------------------------------------- |
| 数学分析第 10 页,级数收敛条件 | 正常返回,约 4 秒 | 正确:通项趋于 0 是必要条件 |
| 公共经济学第 18 页,补贴图表 | 正常返回,约 5 秒 | 坐标轴和补贴对象正确,但把“明补效率损失更小”回答反了,质量不通过 |
| 1158×1700 图片 | 发送前拒绝超出本地像素限制的输入 | 当前提示较笼统 |
| 运行中取消 | 约 0.1 秒结束,无完成事件 | 适配层通过 |
| 取消后重试 | 正常返回完整正文 | 运行通过 |

证据:`evidence/textbook-vision/adapter-results.json`。其中 `passed` 表示传输/状态断言,不能理解为事实正确。初始直接请求探针用了 768 token 上限,图表没有返回正文;按应用实际的 4096 token 上限复测后返回正文,但仍存在上述事实错误。没有为单个失败样例调整产品提示词。

## 验证边界与后续项

- 未进行真实显存/内存耗尽压力测试;未覆盖所有教材的每页/手写公式,也未完整复测原生 UI 的视觉取消/重试组合。适配层验证不能替代这些界面测试。
- 未独立验证无 WebView2 时默认小包的联网下载路径;断网新装应选完整离线包。
- 大尺寸扫描 PDF 默认缩放仍需手动调整,没有自动适应页宽;测试教材的原始标题元数据还含异常字符。这两项不在此次选定修复清单内。
- 安装包未代码签名;安装验证不代表已获得 Windows 下载信誉。
- LM Studio 文字实机验证按约定留到 v0.2.1;本版不承诺 LM Studio 视觉。
- PDF 文字层、EPUB、DOCX 的文本提取不需要大模型。本地全书 OCR/视觉索引尚未实现;未来可用独立 OCR 或视觉模型,并不必然依赖本地大模型。当前全书扫描识别使用已配置且获确认的云端通道。

结论:安装包与主要阅读修复可作为候选版交付;本地视觉应明确为辅助能力,不能宣称复杂公式/图表可靠。本记录不能描述成全部场景均通过。
24 changes: 24 additions & 0 deletions docs/releases/v0.2.0.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
# TextbookLens v0.2.0

## 本次范围

- Ollama 已安装模型的一键连接与离线文字问答。
- 支持 vision 的 Ollama 模型可独立设为视觉服务,用于图片区域问答;原有文字默认模型保持独立。
- EPUB 图片问答的位置保存和重新打开修复。
- 扫描 PDF 的本地解码资源随应用提供;修复 JBIG2 页空白。
- 修复页底提问表单越界、书库卡片和列表排版、目录跳转后工具栏不可见,以及调整大小手柄遮挡阅读滚动条。

## 使用边界

- 正式目标平台为 Windows 11 x64。模型运行软件和模型文件需事先安装;应用的一键连接不会自动下载模型。
- 本地文字和视觉已在 Ollama 上实测。LM Studio 文字通道为实验性,实机验证留到 v0.2.1;本版不承诺 LM Studio 视觉输入。
- 普通 PDF 的文字层、EPUB 和 DOCX 的本地文本读取不需要大模型。
- 全书扫描识别/视觉索引仍使用已配置且支持该操作的云端服务,经用户确认后运行;本版没有独立本地 OCR 引擎或本地全书视觉索引。
- 本地图片问答当前每次接受一张图片,限制见 `docs/local-models.md`。识图成功不保证复杂公式、手写内容和所有题目都正确。
- 安装包、升级验证、签名状态和已知问题以[本次验证报告](v0.2.0-validation.md)为准。公开安装包请从 [GitHub Release](https://github.com/gh615280-maker/TextbookLens/releases/tag/v0.2.0) 下载。

## 安装包选择

- 默认安装包复用电脑已有的 WebView2;缺少运行时时需要联网下载安装。完成安装且本地模型已就绪后,阅读和本地问答可以断网使用。
- 完整离线包附带微软 WebView2 离线运行时,可用于缺少该运行时的电脑断网安装。体积约 273 MB,主要来自约 259 MB 的运行时安装文件。
- 两种安装包都不包含本地大模型文件。完整离线包使用 `npm run tauri -- build --config src-tauri/tauri.offline.conf.json` 构建;应放在单独的分发目录,避免与默认包混淆。
14 changes: 14 additions & 0 deletions docs/superpowers/plans/2026-09-08-v020-release-preparation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
# v0.2.0 release preparation

Scope approved in the conversation on 2026-09-08. Work in the current shared checkout; do not include unrelated experiments or user files. Build and test artifacts use a separate task directory. Do not tag, push, or publish as part of preparation.

- [x] Fix PDF decoding in `src/features/reader/pdf/PdfReaderAdapter.ts`, `src/features/import/parsers/pdf-parser.ts`, and `src/features/indexing/pdf-page-renderer.ts` using shared local PDF.js resources. Bundle decoder/font resources from the pinned dependency, retain script confinement, and surface page-render failures. Verify JBIG2 pages from the desktop test corpus and a synthetic browser regression.
- [x] Fix `SelectionMenu.tsx` positioning when its form grows; retain cancellation and focus behavior. Verify bottom-edge question/note forms in normal and narrow windows.
- [x] Restore library card/table layout, persistent reader navigation, and distinct resize/scroll hit areas in the library components, `ReaderLayout.tsx`, and reader styles. Verify long book titles and actual desktop interactions.
- [x] Set application, lockfile, and package versions to 0.2.0. Document Ollama text/vision scope; defer LM Studio real-machine validation to 0.2.1. Existing cloud indexing remains available; do not claim local whole-book OCR/indexing.
- [x] Run affected tests, full application regression, formatting, lint, bindings, dependency/license/fixture checks, and release artifact scans. Produce Release EXE/MSI from a controlled source snapshot, preserving existing migration bytes.
- [x] Exercise fresh Windows installation and previous-version upgrade; verify library/history/settings/profile retention and uninstall behavior. Recheck actual textbook image questions, oversized input, cancellation, stopped-service recovery, and offline startup where the environment permits. Record unexecuted gates explicitly.

Acceptance: the two observed P1 defects are fixed; the selected common UI defects are fixed; installable 0.2.0 artifacts and reproducible evidence are available. Remaining platform or model limitations are stated accurately. No automatic cloud fallback, model download, or whole-book local OCR is added.

Validation and remaining coverage limits: [v0.2.0-validation.md](../../releases/v0.2.0-validation.md). Completed preparation does not mean every platform/model scenario passed.
Loading
Loading