DrunkenBot LLM-IDE is an end-to-end desktop development environment engineered to build, train, fine-tune, benchmark, export, and chat with custom generative language models entirely on local hardwareβfrom raw documents to quantized .gguf binaries.
Unlike generic training scripts or opaque cloud platforms, DrunkenBot LLM-IDE provides a cohesive, hardware-adaptive desktop studio. It combines an isolated computational backend (engine/) with a responsive PySide6 desktop interface (interface/), featuring real-time training telemetry, prompt loss masking, multi-stage LoRA adaptation, process supervision, and integrated GGUF inference with reasoning controls.
- Multi-Format Ingestion: Ingest raw text (
.txt), Markdown (.md), PDFs (.pdf), and structured dialogue/code (.jsonl). - Syntax-Preserving Code Mode: Dedicated
clean_codepipeline that preserves exact indentation, whitespace, and code fences for programming languages. - Custom BPE Tokenizer Engine: Train custom Byte-Pair Encoding tokenizers with configurable vocabulary sizes (from 256 to 65,536+ tokens).
- Target Loss Masking (
IGNORE_INDEX = -100): Automatically masks system instructions and user queries during instruction, conversation, code, and tool-call preparation so the neural network trains strictly on assistant completions. - High-Performance Memory Mapping: Pre-tokenized corpora are stored as binary NumPy arrays (
train_tokens.npy,train_targets.npy) for sub-second dataset loading without RAM exhaustion. - Adaptive Diversity Filtering: Built-in repetition and diversity filters that accept large technical books and monolithic codebases while rejecting synthetic template padding.
- Transformer Block Styles: Classic GPT-style and modern LLaMA-style transformer blocks.
-
Rotary Position Embeddings (RoPE): High-precision positional encoding with configurable base
$\theta$ ($10,000$ to$500,000$ ). - Attention Variety: Multi-Head Attention (MHA), Grouped-Query Attention (GQA), Multi-Query Attention (MQA), and sliding-window attention.
- Fast Compute Backends: PyTorch Scaled Dot-Product Attention (SDPA) with automatic FlashAttention and memory-efficient kernel selection.
- Normalization & Activations: RMSNorm or standard LayerNorm with optional bias terms; GELU, SiLU, and SwiGLU feed-forward networks.
-
Specialized Workflows:
- Base Pretraining: Learn foundational language, grammar, and world knowledge from scratch.
- Instruction Fine-Tuning: Align models to follow prompts and tasks (Alpaca, Dolly, SlimOrca).
- Conversation Fine-Tuning: Multi-turn dialogue with turn-aware context pruning.
- Code Fine-Tuning: Programming syntax, docstrings, and algorithm synthesis.
- Tool-Call Fine-Tuning: Structured function calling and JSON output schemas.
-
Parameter-Efficient Fine-Tuning (PEFT / LoRA):
- Low-Rank Adaptation (
$W = W_0 + \frac{\alpha}{r} B A$ ). - Target either Attention projections (
q, k, v, out) or full Attention + MLP blocks (w1, w2, c_fc, c_proj). - Auto-configured hyperparameter presets tuned to dataset scale and hardware.
- Low-Rank Adaptation (
- Dynamic VRAM Scaling: Automatically probes host GPU capacity to calculate optimal micro-batch sizes and gradient accumulation steps.
- Detached Worker Process: The training worker runs in an independent operating system process. You can close the IDE, reopen it, or reboot the GUI without interrupting ongoing training runs.
- Live SQLite Telemetry: Batched streaming telemetry captures step-by-step training loss, validation loss, learning rate decay, tokens per second, and GPU memory utilization.
- Mixed Precision: Automatic Mixed Precision (AMP FP16 / BF16) with native gradient scaling.
- Safe Resumption: Checkpoints store optimizer states, RNG seeds, and step counters for bit-exact resumption.
- Full Model Bundling: Consolidates PyTorch weights, tokenizer configs, and training lineage into standalone distribution packages.
- FP16 & FP32 Quantization: Halves memory footprint while preserving model fidelity.
- HuggingFace Packaging: Exports model architectures into standard Transformers-compatible directories (
config.json,model.safetensors). - GGUF Conversion: Compiles models directly into GGUF format (
f16,q8_0,q4_k_m) for high-speed inference on CPU and Apple Silicon / CUDA viallama.cpp.
- Embedded Inference: Run local GGUF models directly within the desktop application using
llama-cpp-python. - Streamed Markdown Rendering: Rich formatting with live token-by-token streaming, code syntax highlighting, and copy buttons.
- Reasoning / Thinking Controls: Selectable reasoning effort profiles (
Light,Balanced,Deep) for chain-of-thought exploration. - Hyperparameter Tuning: Real-time control over temperature, top-p, top-k, repetition penalty, context window, and pinned system prompts.
- Local-First Launch: Startup checks local encrypted license metadata in ~1msβno remote network lag and zero UI freeze.
- Two-Layer Cryptography: Fernet authenticated encryption (AES-128-CBC + HMAC-SHA256 keyed via PBKDF2 with machine ID and hardware MAC node) wrapped with Windows DPAPI (
CryptProtectData). - Machine & User Bound: Ciphertext is bound to the physical device and local Windows user credentials, preventing unauthorized license transfer.
- Operating System: Windows 10/11 (64-bit), Linux (Ubuntu 22.04+ recommended), or macOS (Apple Silicon).
- Python: Version 3.12 or newer.
- GPU (Recommended): NVIDIA GPU with CUDA 12.x support (RTX 30xx/40xx/50xx or Data Center GPUs) for accelerated training. CPU training is fully supported for small models.
git clone https://github.com/drunkenbot-ai/LLM-IDE.git
cd LLM-IDE# Windows (PowerShell)
python -m venv .venv
.\.venv\Scripts\Activate.ps1
# Linux / macOS
python3 -m venv .venv
source .venv/bin/activatepip install --upgrade pip
pip install -r requirements.txtOptional (CUDA Acceleration): If running an NVIDIA GPU, ensure your PyTorch build includes CUDA:
pip install torch --index-url https://download.pytorch.org/whl/cu124
python run_app.pyOn Linux or macOS, you can also make run_app.py directly executable:
chmod +x run_app.py
./run_app.pyOn startup, the IDE executes a startup validation splash screen that verifies local cache directories, inspects core dependencies, and verifies repository integrity before presenting the Project Chooser.
The non-Qt computational core can be executed directly from the terminal or in automated CI/CD pipelines without launching the graphical desktop interface:
# Display CLI help and available commands
python -m engine.cli --help# Prepare text/code files with prompt loss masking
python -m engine.cli prepare \
--input_dir ./data/raw_corpus \
--output_dir ./runs/prepared_data \
--context_length 512 \
--vocab_size 16384# Prepare programming language corpora with indentation preservation
python -m engine.cli prepare \
--input_dir ./data/code_files \
--output_dir ./runs/code_data \
--context_length 1024 \
--code_training_modepython -m engine.cli train \
--data_dir ./runs/prepared_data \
--output_dir ./runs/my_model \
--epochs 3 \
--batch_size 16 \
--context_length 512 \
--embedding_size 512 \
--head_count 8 \
--layer_count 6 \
--device cudaThe codebase enforces a strict unidirectional dependency architecture:
engine/: Pure computational Python/PyTorch modules. Never imports GUI or Qt code. Can be imported in scripts, notebooks, or CLI workers.interface/: PySide6 desktop application, widget assemblies, screen mixins, and reactive event loops.
LLM-IDE/
βββ engine/ # Non-Qt Computational Backend
β βββ config.py # Model, Dataset, and Training dataclass configs
β βββ model.py # Transformer architecture (RoPE, GQA, SwiGLU, SDPA)
β βββ tokenizer.py # BPE tokenizer training and streaming encoding
β βββ target_masking.py # Prompt loss masking (IGNORE_INDEX = -100)
β βββ data_core.py # Dataset loading, indentation preservation, normalization
β βββ dataset_corpus.py # Corpus building, document chunking, diversity filter
β βββ training_orchestrator.py # Training loop, optimizer, loss computation, scheduler
β βββ training_runtime.py # Array chunking, dataset splits, checkpoint saving
β βββ training_worker.py # Standalone headless background worker process
β βββ lora.py # Parameter-Efficient Fine-Tuning (LoRA) layers
β βββ license_client.py # Machine-bound encrypted licensing (DPAPI + Fernet)
β βββ export.py # GGUF, SafeTensors, and HuggingFace exporters
β βββ microgpt_chat.py # Embedded PyTorch chat session with KV cache
β βββ llama_chat.py # llama.cpp GGUF streamed chat wrapper
βββ interface/ # PySide6 Desktop Application
β βββ app.py # Main application entry point and window lifecycle
β βββ license_activation_dialog.py # Activation dialog and responsive background thread
β βββ startup.py # Validation splash screen & project chooser
β βββ core/ # Window core mixin, project state, and menu logic
β βββ screens/ # Screen widgets (Dataset, Training, LoRA, Export, Chat)
β βββ tabs/ # Tab layouts and navigation adapters
β βββ widgets/ # Custom UI controls, app shell, and metric cards
βββ packaging/ # Inno Setup and PyInstaller packaging scripts
βββ ref/ # High-resolution UI screenshots and artwork
βββ tests/ # 140+ comprehensive pytest test suites
βββ README.md # Project landing page (this document)
βββ Technical_Documentation.md # Exhaustive technical specification & architecture
βββ How to Train your LLM.md # End-to-end user tutorial for building custom LLMs
βββ pyproject.toml # Project metadata, dependencies, and tooling config
To package DrunkenBot LLM-IDE into a self-contained Windows executable and Inno Setup installer that includes a private Python runtime:
# Standard CPU / Generic Build
python packaging/packager.py
# CUDA-Enabled GPU Build (auto-probes installed NVIDIA drivers)
python packaging/packager.py --gpuSee build.md for full compilation prerequisites, Inno Setup configurations, and runtime profile details.
- π How to Train your LLM.md: The practical, step-by-step handbook for creating your first language model from scratch.
- π¬ Technical_Documentation.md: In-depth architectural blueprint covering mathematical formulations, process supervisors, encrypted licensing, and algorithm implementations.
- βοΈ build.md: Build options, compiler settings, and standalone distribution instructions.
Copyright Β© DrunkenBot. All rights reserved.
See LICENSE.md for licensing terms and usage restrictions.









