[RFC] Local Model Integration – Replace DeepSeek API
Background
近期 DeepSeek API 宣布即将大幅上调价格,对于 CheckPause 这种需要频繁调用 AI 进行棋局讲解的工具,长期使用成本将显著增加。
同时,部分用户对 API 调用存在隐私顾虑(棋谱数据需上传至第三方服务器)。
为了降低长期运营成本、提升数据隐私性,计划引入本地模型作为 DeepSeek 的替代方案。
DeepSeek recently announced a significant price increase, which will raise long-term costs for CheckPause due to frequent API calls for chess commentary.
Some users also have privacy concerns about uploading PGN data to third-party servers.
To reduce long-term costs and improve privacy, we plan to introduce local models as an alternative to DeepSeek API.
Goals
-
降低成本:部署本地模型后,每次分析不再按 Token 计费,运营成本趋近于零(仅需电费)。
Cost reduction: Once a local model is deployed, each analysis costs no API fees – only electricity.
-
保护隐私:棋谱数据完全本地处理,不上传任何第三方服务器。
Privacy: All PGN data stays local, never uploaded to any third-party server.
-
保持可用性:在现有硬件上,确保分析速度在可接受范围内(如 10~30 秒内完成一轮讲解)。
Usability: Ensure response time stays within an acceptable range (e.g., 10–30 seconds per round) on typical consumer hardware.
-
保留灵活性:用户仍可选择性使用 DeepSeek API(通过配置切换),不强制替换。
Flexibility: Allow users to optionally keep using DeepSeek API via a configuration toggle.
Technical Approach
Model Candidates
主要考虑以下本地模型(需在 8GB 显存内可运行):
| 模型 |
参数量 |
量化版本 |
预期效果 |
| Qwen2.5-7B-Instruct |
7B |
4-bit / 8-bit |
主流首选,推理能力强 |
| DeepSeek-R1-Distill-Qwen-7B |
7B |
4-bit |
同系模型,可能更适配棋谱分析 |
| Llama 3.1-8B-Instruct |
8B |
4-bit |
备选方案 |
| Phi-3.5-mini-instruct |
3.8B |
4-bit |
低配置硬件备选 |
All models must be runnable within 8GB VRAM (quantized).
Inference Backend
-
首选: llama.cpp(通过 llama-cpp-python)—— 通用性强,支持 CPU / GPU 混合推理,无需 CUDA 环境。
-
备选: Ollama(简化部署)—— 更易上手,适合快速原型验证。
-
备选: transformers + bitsandbytes —— 灵活性高,但需要完整 Python 环境。
-
Primary: llama.cpp (via llama-cpp-python) – versatile, supports both CPU and GPU, no CUDA required.
-
Alternative: Ollama – easier to set up, ideal for prototyping.
-
Alternative: transformers + bitsandbytes – more flexible but requires a full Python environment.
Architecture Changes
当前架构:用户 → main.py → engine.py (Stockfish) → ai.py (DeepSeek API) → 返回
新架构:用户 → main.py → engine.py (Stockfish) → local_inference.py (本地模型) → 返回
主要改动点:
- 新增
local_inference.py 模块,封装本地模型的加载与推理。
- 在
config.py 中增加 INFERENCE_BACKEND 开关("local" / "deepseek"),支持一键切换。
ai.py 根据配置选择调用本地模型或 DeepSeek API,对外接口保持一致。
Core changes:
- Add
local_inference.py to handle local model loading and inference.
- Add
INFERENCE_BACKEND switch in config.py to toggle between "local" and "deepseek".
ai.py routes requests to either local model or DeepSeek API based on config, keeping the external interface unchanged.
Open Questions
-
硬件要求:最低需要什么配置才能流畅运行?是否必须要有 GPU?
-
模型质量:本地模型(尤其是 7B 级别)能否达到与 DeepSeek V4-Flash 相近的讲解质量?
-
响应速度:在典型消费级 GPU(如 RTX 3060 12GB)上,生成一段分析需要多长时间?
-
部署复杂度:用户是否需要额外安装 CUDA 或下载数 GB 的模型文件?如何简化这一流程?
-
兼容性:现有的 compact_analysis 压缩数据格式是否适用于本地模型?
-
Hardware requirements: What is the minimum spec to run smoothly? Is GPU mandatory?
-
Quality: Can local models (especially 7B) deliver comparable analysis quality to DeepSeek V4-Flash?
-
Speed: How long does it take to generate an analysis on a typical consumer GPU (e.g., RTX 3060 12GB)?
-
Deployment complexity: Do users need to install CUDA or download multi‑GB model files? How can we streamline this?
-
Compatibility: Does the current compact_analysis data format work well with local models?
Next Steps
- 选择 1~2 个候选模型,在本地进行基准测试(质量 + 速度)。
- 确定最终采用的推理框架和量化方案。
- 实现
local_inference.py 核心模块。
- 在
config.py 中增加切换开关。
- 测试并发布测试版,收集早期用户反馈。
- Benchmark 1–2 candidate models locally for quality and speed.
- Finalize the inference backend and quantization scheme.
- Implement
local_inference.py.
- Add a configuration toggle in
config.py.
- Release a beta version for early user feedback.
References
[RFC] Local Model Integration – Replace DeepSeek API
Background
近期 DeepSeek API 宣布即将大幅上调价格,对于 CheckPause 这种需要频繁调用 AI 进行棋局讲解的工具,长期使用成本将显著增加。
同时,部分用户对 API 调用存在隐私顾虑(棋谱数据需上传至第三方服务器)。
为了降低长期运营成本、提升数据隐私性,计划引入本地模型作为 DeepSeek 的替代方案。
DeepSeek recently announced a significant price increase, which will raise long-term costs for CheckPause due to frequent API calls for chess commentary.
Some users also have privacy concerns about uploading PGN data to third-party servers.
To reduce long-term costs and improve privacy, we plan to introduce local models as an alternative to DeepSeek API.
Goals
降低成本:部署本地模型后,每次分析不再按 Token 计费,运营成本趋近于零(仅需电费)。
Cost reduction: Once a local model is deployed, each analysis costs no API fees – only electricity.
保护隐私:棋谱数据完全本地处理,不上传任何第三方服务器。
Privacy: All PGN data stays local, never uploaded to any third-party server.
保持可用性:在现有硬件上,确保分析速度在可接受范围内(如 10~30 秒内完成一轮讲解)。
Usability: Ensure response time stays within an acceptable range (e.g., 10–30 seconds per round) on typical consumer hardware.
保留灵活性:用户仍可选择性使用 DeepSeek API(通过配置切换),不强制替换。
Flexibility: Allow users to optionally keep using DeepSeek API via a configuration toggle.
Technical Approach
Model Candidates
主要考虑以下本地模型(需在 8GB 显存内可运行):
All models must be runnable within 8GB VRAM (quantized).
Inference Backend
首选:
llama.cpp(通过llama-cpp-python)—— 通用性强,支持 CPU / GPU 混合推理,无需 CUDA 环境。备选:
Ollama(简化部署)—— 更易上手,适合快速原型验证。备选:
transformers+bitsandbytes—— 灵活性高,但需要完整 Python 环境。Primary:
llama.cpp(viallama-cpp-python) – versatile, supports both CPU and GPU, no CUDA required.Alternative:
Ollama– easier to set up, ideal for prototyping.Alternative:
transformers+bitsandbytes– more flexible but requires a full Python environment.Architecture Changes
当前架构:
用户 → main.py → engine.py (Stockfish) → ai.py (DeepSeek API) → 返回新架构:
用户 → main.py → engine.py (Stockfish) → local_inference.py (本地模型) → 返回主要改动点:
local_inference.py模块,封装本地模型的加载与推理。config.py中增加INFERENCE_BACKEND开关("local"/"deepseek"),支持一键切换。ai.py根据配置选择调用本地模型或 DeepSeek API,对外接口保持一致。Core changes:
local_inference.pyto handle local model loading and inference.INFERENCE_BACKENDswitch inconfig.pyto toggle between"local"and"deepseek".ai.pyroutes requests to either local model or DeepSeek API based on config, keeping the external interface unchanged.Open Questions
硬件要求:最低需要什么配置才能流畅运行?是否必须要有 GPU?
模型质量:本地模型(尤其是 7B 级别)能否达到与 DeepSeek V4-Flash 相近的讲解质量?
响应速度:在典型消费级 GPU(如 RTX 3060 12GB)上,生成一段分析需要多长时间?
部署复杂度:用户是否需要额外安装 CUDA 或下载数 GB 的模型文件?如何简化这一流程?
兼容性:现有的
compact_analysis压缩数据格式是否适用于本地模型?Hardware requirements: What is the minimum spec to run smoothly? Is GPU mandatory?
Quality: Can local models (especially 7B) deliver comparable analysis quality to DeepSeek V4-Flash?
Speed: How long does it take to generate an analysis on a typical consumer GPU (e.g., RTX 3060 12GB)?
Deployment complexity: Do users need to install CUDA or download multi‑GB model files? How can we streamline this?
Compatibility: Does the current
compact_analysisdata format work well with local models?Next Steps
local_inference.py核心模块。config.py中增加切换开关。local_inference.py.config.py.References