Algorithms test engineer · Full-stack · Model evaluation · Distributed GPU scheduling · Agent loop · Infrastructure · Security enthusiast
| Project | What it is | Stack |
|---|---|---|
| yizhi-oral-history | A family-facing oral history platform: AI-guided interviews preserve an elder's voice, stories and photos | JavaScript |
| OnlineJudge | An online judge: submitted code runs in an isolated Docker container (no network, CPU/memory/process limits, timeout kill), with interactive Linux-shell support | Java · Spring · Docker |
| goodnight-storybox | Turns bedtime stories into picture-book videos, narrated in a cloned voice | Python |
| EvalCraft | A hands-on lab for learning LLM evaluation from a testing perspective: the eval loop, judge calibration, RAG layering, agent-trace reliability | TypeScript |
Two cases of taking a pattern from outside the immediate problem and landing it on a concrete need — one from distributed GPU infrastructure, one from LLM agents.
Turning idle GPUs into one "virtual big machine" — batch model preprocessing
- 🎯 Why: running the data through a model takes about 100 h on a single machine, but the GPUs on the team's machines sit idle most of the time. Pooling idle GPUs is the idea behind io.net — here scoped down to one office LAN, which beat buying hardware.
- 🛠️ How: a purpose-built scheduler plus a task queue in MySQL (reusing the team's existing stack rather than pulling in new middleware at this scale); workers poll for jobs and fetch their data over HTTP; a node that keeps failing past a threshold is marked unhealthy and removed automatically.
- ✅ Result: work shards per video, so throughput scales with node count (≈ 100 h on one machine → ≈ 25 h on four); the model's first-pass screening runs 7×24 and a human only confirms the results.
- ♻️ Reusable: the shard + queue + self-healing-node structure is independent of the business domain and transfers to other batch workloads.
Handing repetitive work to agents — speed-up for test work
- 🎯 Why: bad-case analysis is mostly mechanical and grows with the number of cases, eating time that should go into judging the results.
- 🛠️ How: scripted the mechanical parts of bad-case analysis (bulk screenshotting, frame-level localisation) and captured the know-how as skills; judged the agents' output by spot-checking and inspecting intermediate steps rather than trusting the final answer.
- ✅ Result: over 50% overall speed-up, and adopted by the team after I had used it myself for a while.
xterm.js (21k ★) — fixed two defects in @xterm/addon-serialize:
- #6205 cursor restored one column off after a full-row round trip
- #6206 the hyperlink render hint leaking into serialized underline styles
- Awards / qualifications: Provincial Second Prize, Blue Bridge Cup (C, 2022; China's national collegiate programming contest) · National Software Qualification Exam, Intermediate — Software Designer (state-run professional qualification)
- Security background: CTF — reverse engineering, unpacking, image steganography, web SQL injection, crawlers; packet capture analysis
- Mentoring / talks: mentored 10+ interns; gave an internal talk on open-source tooling for productivity
