Skip to content

Local Model Replacement Plan / 本地模型替代方案 #4

Description

@Comet-zzz

[RFC] Local Model Integration – Replace DeepSeek API

Background

近期 DeepSeek API 宣布即将大幅上调价格,对于 CheckPause 这种需要频繁调用 AI 进行棋局讲解的工具,长期使用成本将显著增加。

同时,部分用户对 API 调用存在隐私顾虑(棋谱数据需上传至第三方服务器)。

为了降低长期运营成本、提升数据隐私性,计划引入本地模型作为 DeepSeek 的替代方案。

DeepSeek recently announced a significant price increase, which will raise long-term costs for CheckPause due to frequent API calls for chess commentary.

Some users also have privacy concerns about uploading PGN data to third-party servers.

To reduce long-term costs and improve privacy, we plan to introduce local models as an alternative to DeepSeek API.


Goals

  1. 降低成本:部署本地模型后,每次分析不再按 Token 计费,运营成本趋近于零(仅需电费)。

    Cost reduction: Once a local model is deployed, each analysis costs no API fees – only electricity.

  2. 保护隐私:棋谱数据完全本地处理,不上传任何第三方服务器。

    Privacy: All PGN data stays local, never uploaded to any third-party server.

  3. 保持可用性:在现有硬件上,确保分析速度在可接受范围内(如 10~30 秒内完成一轮讲解)。

    Usability: Ensure response time stays within an acceptable range (e.g., 10–30 seconds per round) on typical consumer hardware.

  4. 保留灵活性:用户仍可选择性使用 DeepSeek API(通过配置切换),不强制替换。

    Flexibility: Allow users to optionally keep using DeepSeek API via a configuration toggle.


Technical Approach

Model Candidates

主要考虑以下本地模型(需在 8GB 显存内可运行):

模型 参数量 量化版本 预期效果
Qwen2.5-7B-Instruct 7B 4-bit / 8-bit 主流首选,推理能力强
DeepSeek-R1-Distill-Qwen-7B 7B 4-bit 同系模型,可能更适配棋谱分析
Llama 3.1-8B-Instruct 8B 4-bit 备选方案
Phi-3.5-mini-instruct 3.8B 4-bit 低配置硬件备选

All models must be runnable within 8GB VRAM (quantized).

Inference Backend

  • 首选: llama.cpp(通过 llama-cpp-python)—— 通用性强,支持 CPU / GPU 混合推理,无需 CUDA 环境。

  • 备选: Ollama(简化部署)—— 更易上手,适合快速原型验证。

  • 备选: transformers + bitsandbytes —— 灵活性高,但需要完整 Python 环境。

  • Primary: llama.cpp (via llama-cpp-python) – versatile, supports both CPU and GPU, no CUDA required.

  • Alternative: Ollama – easier to set up, ideal for prototyping.

  • Alternative: transformers + bitsandbytes – more flexible but requires a full Python environment.

Architecture Changes

当前架构:用户 → main.py → engine.py (Stockfish) → ai.py (DeepSeek API) → 返回

新架构:用户 → main.py → engine.py (Stockfish) → local_inference.py (本地模型) → 返回

主要改动点:

  1. 新增 local_inference.py 模块,封装本地模型的加载与推理。
  2. 在 config.py 中增加 INFERENCE_BACKEND 开关("local" / "deepseek"),支持一键切换。
  3. ai.py 根据配置选择调用本地模型或 DeepSeek API,对外接口保持一致。

Core changes:

  1. Add local_inference.py to handle local model loading and inference.
  2. Add INFERENCE_BACKEND switch in config.py to toggle between "local" and "deepseek".
  3. ai.py routes requests to either local model or DeepSeek API based on config, keeping the external interface unchanged.

Open Questions

  • 硬件要求:最低需要什么配置才能流畅运行?是否必须要有 GPU?

  • 模型质量:本地模型(尤其是 7B 级别)能否达到与 DeepSeek V4-Flash 相近的讲解质量?

  • 响应速度:在典型消费级 GPU(如 RTX 3060 12GB)上,生成一段分析需要多长时间?

  • 部署复杂度:用户是否需要额外安装 CUDA 或下载数 GB 的模型文件?如何简化这一流程?

  • 兼容性:现有的 compact_analysis 压缩数据格式是否适用于本地模型?

  • Hardware requirements: What is the minimum spec to run smoothly? Is GPU mandatory?

  • Quality: Can local models (especially 7B) deliver comparable analysis quality to DeepSeek V4-Flash?

  • Speed: How long does it take to generate an analysis on a typical consumer GPU (e.g., RTX 3060 12GB)?

  • Deployment complexity: Do users need to install CUDA or download multi‑GB model files? How can we streamline this?

  • Compatibility: Does the current compact_analysis data format work well with local models?


Next Steps

  1. 选择 1~2 个候选模型,在本地进行基准测试(质量 + 速度)。
  2. 确定最终采用的推理框架和量化方案。
  3. 实现 local_inference.py 核心模块。
  4. 在 config.py 中增加切换开关。
  5. 测试并发布测试版,收集早期用户反馈。

  1. Benchmark 1–2 candidate models locally for quality and speed.
  2. Finalize the inference backend and quantization scheme.
  3. Implement local_inference.py.
  4. Add a configuration toggle in config.py.
  5. Release a beta version for early user feedback.

References


Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions