🌐 Language / Ngôn ngữ: English | Tiếng Việt
A portable native inference and deployment stack for Mage-Flow-Turbo. The repository does not train or modify model weights. Python provides configuration, model/runtime identity verification, CLI/REST orchestration, lifecycle control, telemetry and evidence collection; the actual model execution path is the native stable-diffusion.cpp sd-cli runtime.
The project name describes the execution stack, not a single quantization or serialization format. v1.0.0 supports a canonical Q8 GGUF profile and a BF16 SafeTensors profile, both through the same native runtime.
| Profile | Mage-Flow diffusion | Text encoder | VAE | Native runtime | CPU | CUDA cuda0 |
Role |
|---|---|---|---|---|---|---|---|
q8-reference |
GGUF Q8_0 |
Qwen3-VL-4B GGUF Q4_K_M |
SafeTensors | pinned stable-diffusion.cpp sd-cli |
yes | yes | canonical/default |
bf16-safetensors |
BF16 SafeTensors | Qwen3-VL-4B GGUF Q4_K_M |
SafeTensors | pinned stable-diffusion.cpp sd-cli |
yes | yes | supported alternative |
The project does not provide a Hugging Face Transformers inference backend and does not run a PyTorch/Transformers inference loop. PyTorch/Transformers wording in provenance or model-mirror paths refers to the source artifact/distribution layout of the BF16 SafeTensors weights, not the execution framework.
The project is also intentionally not described as GGUF-only: even the canonical Q8 profile uses a SafeTensors VAE.
| Role | Artifact / identity | Format |
|---|---|---|
| Q8 diffusion | Mage-Flow-Turbo-DiT-Q8_0.gguf |
GGUF Q8_0 |
| BF16 diffusion | diffusion_pytorch_model.safetensors from Mage-Flow PyTorch / default |
BF16 SafeTensors |
| Text encoder | Qwen3VL-4B-Instruct-Q4_K_M.gguf |
GGUF Q4_K_M |
| VAE | diffusion_pytorch_model.safetensors |
SafeTensors |
| Native runtime | stable-diffusion.cpp sd-cli |
pinned commit 6b3edaaf32cc19e5bb2d819c788bd557eddc8eba |
Exact SHA-256 identities are verified before real inference. The Git repository contains no model weights.
The canonical strict 2×2 measurements remain bound to the measured benchmark evidence source HEAD/TREE recorded inside the retained evidence artifacts. The final publication source contains public notebook, documentation, and contract-test corrections. A checksum-protected qualification-equivalence manifest bridges these identities only when qualification-critical Git objects are byte-identical; measured evidence provenance is not rewritten.
| Profile | Kaggle CPU | Kaggle T4/T4x2 cuda0 |
|---|---|---|
q8-reference |
retained measured evidence | retained measured evidence |
bf16-safetensors |
retained measured evidence | retained measured evidence |
Every cell uses the same frozen source HEAD/TREE and the same canonical protocol:
prompt = A small red fox sitting in a quiet green forest, natural light, detailed photography.
seed = 42
steps = 4
CFG = 1.0
threads = 4
matrix = 512 → 640 → 768 → 1024
Each resolution is recorded exactly once in its retained authority evidence. The matrix is sequential and fail-fast. If a genuine later-resolution platform/runtime limit occurs, evidence records it rather than silently changing placement or inventing a ratio.
- Kaggle
Accelerator=None; - backend exactly
cpu; - prebuilt CPU
sd-clionly; - no CUDA fallback;
- host-memory and process-RSS telemetry;
- the BF16 CPU profile retains its explicit high-memory/headroom safety gates.
- host must be T4 or T4x2;
- release qualification uses only physical GPU slot 0;
CUDA_DEVICE_ORDER=PCI_BUS_ID;CUDA_VISIBLE_DEVICES=0;- effective inference backend
cuda0; - no
cuda1, no multi-GPU split, noauto-fit, no CPU inference fallback; - prebuilt CUDA runtime only;
- successful CUDA generations require positive VRAM telemetry.
A T4x2 host is therefore allowed as a host configuration, but v1.0.0 qualification remains a strict single-T4 benchmark.
The frozen comparator verifies the measured evidence source HEAD/TREE, canonical request, model identities, runtime commit and backend-specific runtime SHA values before calculating ratios. The qualification-equivalence manifest separately proves that qualification-critical objects are byte-identical in the final publication source.
Q8/CPU retains its documented same-session full-reset recovery exception; equivalence does not relabel it as fresh-session evidence. A fresh public Q8/T4 notebook smoke is supplemental reproduction evidence, not a replacement for the retained strict 2×2 benchmark measurements. Final measured numbers remain checksum-protected GitHub Release assets/body and are not pasted into source docs.
Diffusion execution, text conditioning and VAE decoding are performed by sd-cli. Python validates identities, constructs explicit subprocess arguments with shell=False, monitors the native process, validates PNG artifacts and records structured evidence.
This architecture gives both model profiles one common runtime path, which makes the Q8/BF16 × CPU/CUDA comparison substantially easier to audit.
mageflow-native verify --manifest configs/mage-flow-turbo-q8-reference.jsonThe generic CLI may build a local runtime when deliberately developing outside release qualification:
python -m pip install -e .
mageflow-native runtime build --backend cpu
mageflow-native doctor --manifest configs/mage-flow-turbo-q8-reference.json
mageflow-native verify --manifest configs/mage-flow-turbo-q8-reference.jsonRelease qualification itself is prebuilt-runtime only.
python -m pip install -e .
mageflow-native runtime build --backend cuda
mageflow-native doctor --manifest configs/mage-flow-turbo-q8-reference.json --backend cuda0Release qualification uses deterministic cuda0 placement rather than automatic splitting.
The reference service binds to 127.0.0.1 by default.
GET /healthz
GET /readyz
GET /v1/info
POST /v1/images/generate
GET /v1/artifacts/<png>
The public notebook notebooks/kaggle-production-demo.ipynb detects supported Kaggle accelerators. For release qualification, use the dedicated matrix harness and exact prebuilt runtime/profile inputs rather than relying on notebook defaults. See docs/kaggle.md.
The prebuilt CPU/CUDA runtimes are attached as Kaggle Datasets; Mage-Flow, Qwen, and VAE inputs are Kaggle Models. Attach one runtime matching the detected accelerator and one Mage diffusion family matching the selected model profile.
For the exact UI sequence, use Session options → Accelerator, then
Input → Add Input; see the detailed Kaggle reproduction procedure.
The Dataset identities are
dangkhoa2016/stable-diffusion-cpp-6b3edaa-portable-cpu-runtime (CPU) and
dangkhoa2016/stable-diffusion-cpp-6b3edaa-cuda-t4-runtime (T4/T4x2). The
Model identities are dangkhoa2016/mage-flow-community-mage-flow-turbo and
dangkhoa2016/qwen-qwen3-vl-4b-instruct-gguf.
For a normal qualification/inference session, attach exactly one Mage-Flow-Turbo diffusion family:
q8-reference— Mage-FlowGGUF / q8-0, QwenGGUF / q4-k-m, and requiredPyTorch / vae-only;bf16-safetensors— Mage-FlowPyTorch / defaultBF16 transformer/VAE plus QwenGGUF / q4-k-m; the VAE is already included, so no separatevae-onlyattachment.
Do not attach both Mage diffusion families in an ordinary authority session. Mixed-family attachment is reserved for explicitly controlled research tooling and is not part of the retained strict release evidence.
For the recommended fresh BF16 T4x2 run, attach only the CUDA runtime Dataset,
Mage-Flow PyTorch / default, and Qwen GGUF / q4-k-m. Do not attach the CPU
runtime, Mage-Flow GGUF / q8-0, or Mage-Flow PyTorch / vae-only.
Before the strict 2×2 release redesign, a same-host paired 768×768 CPU visual study compared the Q8 and BF16 representations. That historical study remains useful as quality-oriented research, but it is not the final v1.0.0 2×2 performance authority and does not determine the default profile. Q8 remains the canonical/default profile.
See BF16 SafeTensors background and qualification policy.
CPU and CUDA outputs may legitimately differ byte-for-byte across numerical backends. Release evidence records:
- exact source HEAD and TREE;
- profile/backend identity;
- model component names/formats/SHA-256 values;
- pinned native runtime commit and binary SHA-256;
- canonical request and resolution;
- elapsed time;
- host memory / process RSS;
- CUDA peak VRAM when applicable;
- PNG filename, dimensions, byte count and SHA-256;
- explicit failure classification when a matrix stops.
Evidence archives are checksum-protected, contain internal manifests, and reject model weights and known secret patterns.
- Architecture
- Model stack
- Local Linux
- CUDA
- Kaggle
- Strict v1.0.0 benchmark contract
- BF16 SafeTensors profile
- REST API
- Testing
- Troubleshooting
- Contributing
- Security policy
MIT License. Copyright © 2026 Đăng Khoa i.am@dangkhoa.dev.