Skip to content

Latest commit

 

History

History
174 lines (126 loc) · 14.6 KB

File metadata and controls

174 lines (126 loc) · 14.6 KB

LocalForgeLLM documentation

English · Русский · 简体中文 · Español

Back to the project · Install · Launch profiles · Measurements

Agent manual

Start with AGENTS.md and the shared LocalForgeLLM skill. The framework covers three LLM work levels: install an existing model and stack; tune or modify them; build a derived artifact from full weights. Updates and interface integration apply across the levels. Image, video and audio generation have a separate planned branch.

Read when Guide
Understand the framework and component ownership Architecture
Interpret a task, discover an implementation or resume a stack Agent workflow and stack record
Install a published model and its dependencies Level 1 — install
Change context, placement, performance or MoE behavior Level 2 — tune and adapt
Select, build, switch or compare CUDA and Vulkan Backend workflow and measured adaptation
Download a matching head; enable, disable or diagnose speculative decoding MTP operations, memory and quality
Convert/calibrate/quantize full weights Level 3 — build
Prune and rebuild full weights; place individual experts REAP → APEX and expert caching
Fit or accelerate Ternary Bonsai 2 on 8–12 GB GPUs Packing, calibration and GPU cache · Measurements
Add typed decisions and browser-use to a local agent Jev API · Hermes MCP/skills in llama UI
Choose an engine or an optional weight-processing method Engines · APEX
Connect or customize an agent/shell Interface contracts · Pi · OpenShell · llama UI · Hermes
Update, repair or roll back a working stack Operations
Establish capability, quality and resource use Validation
Work on the future generative branch Planned scope
Check versions and the evidence behind a guide Sources

The technical guides are maintained in English. The four language entry pages preserve the translated deployment examples below. Source-inspected configuration is distinguished from measured deployment results in verification scope.

How it works

  1. Give your coding agent the task, target context and memory budget. Specify a dense or MoE model or let it choose one. It inspects the CPU, GPU, RAM, storage and installed software.
  2. Using the framework, model and runtime documentation, it selects a compatible engine and weight format.
  3. It tunes CPU/GPU placement, threads, context, cache precision and batch sizes for your workload.
  4. It runs the task, checks the output, measures speed and RAM/VRAM use, and refines the settings. The result is a reusable launch profile with its runtime version and parameters.

An agent selects and tunes a local stack from your task and hardware, measures the result and saves a launch profile.

The installation and commands below are two historical examples for one PC, using Qwen and Gemma with APEX GGUF. The framework applies the same selection and tuning workflow to dense and MoE models and other quantization methods.

September 18 update

Evidence Best recorded result / outcome Procedure and data
Bonsai PTQ1_0, RTX 4060 8 GB / Ryzen 5 5600 27.14 tok/s short response; 20.22 with 31,018 input tokens; 6.24 GiB process VRAM snapshot Case study · 32K profile
Bonsai PTQ1_0, later llama UI/MCP session on the same 4060 19.58 tok/s weighted; 7,355 output tokens / 9 completed requests; 1 cancelled request excluded Session scope and CSV
Bonsai PQ2_0, community RTX 5070 12 GB / i7-10700K 59.65 tok/s over 10K thinking tokens; 5.60× the report's forced-10K Opus row ap3x0s report and CSV
Gemma APEX, RTX 4060 8 GB / Ryzen 5 5600 18.80 weighted phase mean; 19.38 best completed-response average CUDA, RAM and MTP
Gemma REAP124 rebuild, RTX 4500 Ada 24 GB / 5995WX GGUF −2.81%; no pruning decode gain; expert-cache experiment +8.07% decode Tools, commands and quality limits
Jev with a local model and browser-use Typed API routing and host tool integration; pinned browser-use/Cua review; no local Jev latency benchmark API and harness contract

These rows use different workloads. See each report for timing definitions and failures; no universal framework or model-quality ranking is implied. English graphics and regeneration.

Install

Reference hardware: Linux x86_64, a Vulkan driver, RTX 4060 8 GB, Ryzen 5 5600 and 32 GB RAM. Reserve approximately 22 GB of SSD space for Gemma with vision, or 18 GB for Qwen. These are the tested configurations, not universal minimum requirements.

  1. Download the llama.cpp b10883 Vulkan archive and extract it into runtime/.
  2. Create models/ and download one profile below. Keep the original filenames; the vision projector is mmproj.gguf.
  3. Start the selected profile, then open localhost:8080. Agent clients use http://127.0.0.1:8080/v1, with model ID gemma-apex or qwen-apex. Run one model at a time.
Profile Downloads
Gemma 4 26B-A4B Heretic · I-Balanced Model · 19.51 GB + vision · 1.19 GB
Qwen3.6 35B-A3B · I-Compact Model · 17.29 GB · text profile

Launch profiles

Run from the directory containing runtime/ and models/. The archive creates runtime/llama-b10883/. Vulkan0 is the RTX 4060 in our setup; adapt it if your device order differs.

Gemma

Requires a systemd user session. The 11 GiB service limit includes file cache; mmap allows model pages to be reclaimed and read from SSD again. This limits resident memory at the cost of possible I/O stalls. It does not reserve memory for other applications or guarantee against OOM.

mkdir -p logs
systemd-run --user --unit=gemma-apex --collect \
  --property=MemoryMax=11G --property=MemorySwapMax=0 \
  --property=OOMPolicy=kill --property=OOMScoreAdjust=1000 \
  --property=LimitCORE=0 --property=CPUQuota=600% \
  --property=Nice=5 --property=TimeoutStopSec=15 \
  --property="StandardOutput=append:$PWD/logs/gemma-server.log" \
  --property=StandardError=inherit \
  "$PWD/runtime/llama-b10883/llama-server" \
  --model "$PWD/models/gemma-4-26B-A4B-heretic-APEX-I-Balanced.gguf" \
  --mmproj "$PWD/models/mmproj.gguf" \
  --no-mmproj-offload --image-max-tokens 1120 \
  --alias gemma-apex --host 127.0.0.1 --port 8080 --cors-origins localhost \
  --device Vulkan0 --gpu-layers all --n-cpu-moe 27 --fit off \
  --load-mode mmap --threads 6 --threads-batch 6 \
  --ctx-size 32768 --parallel 1 --batch-size 1152 --ubatch-size 1152 \
  --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \
  --cache-ram 0 --ctx-checkpoints 1 \
  --temp 1.0 --top-p 0.95 --top-k 64 --min-p 0 \
  --log-colors off --log-timestamps

Stop with systemctl --user stop gemma-apex. The projector runs on the CPU. Keep the 1152 batch sizes with the 1120 image-token limit: smaller microbatches caused an image-processing assertion in this build.

For the optional MTP trial, follow the verified 0.46 GB head download and Gemma MTP settings. Add those options to a candidate copy of this command, retaining the main placement and context. Disable by restoring this original command through the same service manager. The target remains APEX I-Balanced; the separate head is Q8_0.

Qwen

Runs in the foreground; stop with Ctrl+C. This profile has no service memory limit or vision projector.

./runtime/llama-b10883/llama-server \
  --model ./models/Qwen3.6-35B-A3B-APEX-I-Compact.gguf \
  --alias qwen-apex --host 127.0.0.1 --port 8080 --cors-origins localhost \
  --device Vulkan0 --gpu-layers all --n-cpu-moe 36 --fit off \
  --load-mode none --threads 6 --threads-batch 6 \
  --ctx-size 32768 --parallel 1 --batch-size 512 --ubatch-size 128 \
  --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \
  --cache-ram 0 --ctx-checkpoints 4 \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0 \
  --log-colors off --log-timestamps

APEX example

These profiles use APEX mixed-precision GGUF weights. It assigns precision by tensor role and layer; I- profiles use importance-matrix calibration. MoE activates a subset of experts per token, while llama.cpp splits execution between CPU and GPU. The complete model spans SSD, RAM and VRAM.

Measured on an RTX 4060

Ryzen 5 5600 · 32 GB RAM · Manjaro Linux · llama.cpp b10883 / Vulkan · September 10–11, 2026.

Model / profile Decode Context window RAM / model VRAM
Qwen3.6 35B-A3B · I-Compact 21.11 tok/s 32,768 13.43 GiB peak RSS / 4.07 GiB
Gemma 4 26B-A4B Heretic · I-Balanced + vision 9.78 tok/s 32,768 11 GiB service limit / 4.88 GiB snapshot
Gemma 4, same target + Q8 MTP head (September 11) 14.83 tok/s 32,768 11 GiB service limit / startup allocations

September 10: Qwen had 28 completed requests, 19.37–22.09 tok/s, up to 29,087 reported context tokens; Gemma had one completed request with vision enabled, 599 input / 632 output tokens. September 11: Gemma MTP had two requests with final timings, 1835 logged output tokens and at most 2265 live context tokens; MTP vision quality was not tested. A configured 32K window is not a full-window stress test; decode rates exclude prompt processing.

Resource use and methodology

Qwen resource monitoring covers 314 samples over the first 10 min 44 s of the session, including pauses:

Resource Average / maximum
Model RAM, RSS 13.23 / 13.43 GiB
Model VRAM 4.07 / 4.07 GiB
Model CPU, normalized to all 12 logical threads 10.1% / 46.9%
Whole-GPU utilization, including desktop applications 89.8% / 100%
GPU temperature 50.4 / 54 °C

Available system RAM fell to 2.53 GiB; free VRAM to 811 MiB. Gemma reached its 11 GiB cgroup ceiling, which includes file cache and is not the same measure as RSS. Its 4.88 GiB VRAM value is a snapshot; CPU/GPU utilization was not recorded for that profile.

Qwen generated 7,820 tokens across 28 completed requests. Weighted decode is (7820 − 28) / 369.07315 = 21.11 tok/s, following llama.cpp's first-token accounting. New-prompt processing averaged 67.36 tok/s. Gemma spent 97.59 s processing the prompt and 64.55 s decoding: 162.13 s total. These measurements describe serving speed, not a model-quality evaluation.

Comparison with Colibri and FreeToken

Qwen3.6 comparison: LocalForgeLLM 21.11, Colibri 9.9, FreeToken 39.3 tokens per second. Local peak RSS is 13.43 GiB; Colibri reports 40 GB; FreeToken lists 32 GiB installed RAM. Different hardware and quantization.

  • Colibri: authors report 9.2 / 9.9 tok/s cold / warm, 40 GB peak RSS, Threadripper 3945WX + RTX 3070 8 GB, int4, 200-token decode.
  • FreeToken: authors report 39.3 tok/s on RTX 4060 Laptop 8 GB + i9-13900H, 32 GiB LPDDR5, NVFP4, OpenCode coding workload. Installed RAM is not measured process consumption.

Our Qwen rate is 2.13× Colibri's cited warm rate and 46.3% below FreeToken's cited rate. The rows describe different CPUs, formats, workloads and averaging methods. For memory, our figure and Colibri's are process RSS; FreeToken's 32 GiB is installed capacity. Gemma's 11 GiB is a service limit including file cache. Keep these definitions when comparing configurations.

Gemma 4 and MTP comparisons

Gemma 4 deployment results with hardware, weights and evidence scope for LocalForgeLLM, llama.cpp community profiles, Ollama and FreeToken.

Gemma MTP session change: 14.83 tok/s weighted, 15.03 over a 102.05-second decode, and a 51.1 percent observed gap that is not a controlled gain.

The September 11 no-MTP request measured 9.81 tok/s; two different MTP requests averaged 14.83 tok/s. The longer response sustained 15.03 tok/s during decode, but also took 47.29 s to process its new prompt. Main offload, CPU expert placement and the 32K allocation stayed fixed. Different requests and cache state prevent treating the +51.1% gap as causal MTP uplift. Formal quality checks remain pending.

The case study includes full-Q8 and QAT alternatives, the external measurements' methods, Colibri's missing Gemma result, FreeToken's modality limits and the sanitized timing data. Follow MTP operations for artifact discovery, verified downloads, switching and rollback.

The later CUDA, MTP and RAM adaptation adds an English infographic and 19 sanitized timing records across eight phases. CUDA + MTP with a 20 GiB cap averaged 18.80 tok/s across four responses; the initial backend switch alone did not improve the observed rate. The report separates runtime compilation time from total setup time. Use the backend workflow to apply and test these choices on another stack.