Skip to content

Latest commit

Β 

History

429 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸ€– DrunkenBot LLM-IDE

The Complete Desktop Foundry for Training, Fine-Tuning, and Deploying Local Language Models

Python 3.12+ PyTorch 2.x PySide6 CUDA Accelerated License: Proprietary


Neural Forge - Model Architecture & Training Fine-Tuning Lab - LoRA & PEFT

Dataset Blueprint & Ingestion Interactive Streamed Chat with Reasoning Controls


πŸ“– Overview

DrunkenBot LLM-IDE is an end-to-end desktop development environment engineered to build, train, fine-tune, benchmark, export, and chat with custom generative language models entirely on local hardwareβ€”from raw documents to quantized .gguf binaries.

Unlike generic training scripts or opaque cloud platforms, DrunkenBot LLM-IDE provides a cohesive, hardware-adaptive desktop studio. It combines an isolated computational backend (engine/) with a responsive PySide6 desktop interface (interface/), featuring real-time training telemetry, prompt loss masking, multi-stage LoRA adaptation, process supervision, and integrated GGUF inference with reasoning controls.


🌟 Key Highlights & Capabilities

1. πŸ“Š Visual Dataset Blueprint & Ingestion

  • Multi-Format Ingestion: Ingest raw text (.txt), Markdown (.md), PDFs (.pdf), and structured dialogue/code (.jsonl).
  • Syntax-Preserving Code Mode: Dedicated clean_code pipeline that preserves exact indentation, whitespace, and code fences for programming languages.
  • Custom BPE Tokenizer Engine: Train custom Byte-Pair Encoding tokenizers with configurable vocabulary sizes (from 256 to 65,536+ tokens).
  • Target Loss Masking (IGNORE_INDEX = -100): Automatically masks system instructions and user queries during instruction, conversation, code, and tool-call preparation so the neural network trains strictly on assistant completions.
  • High-Performance Memory Mapping: Pre-tokenized corpora are stored as binary NumPy arrays (train_tokens.npy, train_targets.npy) for sub-second dataset loading without RAM exhaustion.
  • Adaptive Diversity Filtering: Built-in repetition and diversity filters that accept large technical books and monolithic codebases while rejecting synthetic template padding.

2. 🧠 Modern Neural Architecture ("Neural Forge")

  • Transformer Block Styles: Classic GPT-style and modern LLaMA-style transformer blocks.
  • Rotary Position Embeddings (RoPE): High-precision positional encoding with configurable base $\theta$ ($10,000$ to $500,000$).
  • Attention Variety: Multi-Head Attention (MHA), Grouped-Query Attention (GQA), Multi-Query Attention (MQA), and sliding-window attention.
  • Fast Compute Backends: PyTorch Scaled Dot-Product Attention (SDPA) with automatic FlashAttention and memory-efficient kernel selection.
  • Normalization & Activations: RMSNorm or standard LayerNorm with optional bias terms; GELU, SiLU, and SwiGLU feed-forward networks.

3. 🎯 Multi-Stage Fine-Tuning Lab

  • Specialized Workflows:
    • Base Pretraining: Learn foundational language, grammar, and world knowledge from scratch.
    • Instruction Fine-Tuning: Align models to follow prompts and tasks (Alpaca, Dolly, SlimOrca).
    • Conversation Fine-Tuning: Multi-turn dialogue with turn-aware context pruning.
    • Code Fine-Tuning: Programming syntax, docstrings, and algorithm synthesis.
    • Tool-Call Fine-Tuning: Structured function calling and JSON output schemas.
  • Parameter-Efficient Fine-Tuning (PEFT / LoRA):
    • Low-Rank Adaptation ($W = W_0 + \frac{\alpha}{r} B A$).
    • Target either Attention projections (q, k, v, out) or full Attention + MLP blocks (w1, w2, c_fc, c_proj).
    • Auto-configured hyperparameter presets tuned to dataset scale and hardware.

4. ⚑ Hardware-Adaptive Training & Process Supervision

  • Dynamic VRAM Scaling: Automatically probes host GPU capacity to calculate optimal micro-batch sizes and gradient accumulation steps.
  • Detached Worker Process: The training worker runs in an independent operating system process. You can close the IDE, reopen it, or reboot the GUI without interrupting ongoing training runs.
  • Live SQLite Telemetry: Batched streaming telemetry captures step-by-step training loss, validation loss, learning rate decay, tokens per second, and GPU memory utilization.
  • Mixed Precision: Automatic Mixed Precision (AMP FP16 / BF16) with native gradient scaling.
  • Safe Resumption: Checkpoints store optimizer states, RNG seeds, and step counters for bit-exact resumption.

5. πŸ“¦ Export Bay & Quantization

  • Full Model Bundling: Consolidates PyTorch weights, tokenizer configs, and training lineage into standalone distribution packages.
  • FP16 & FP32 Quantization: Halves memory footprint while preserving model fidelity.
  • HuggingFace Packaging: Exports model architectures into standard Transformers-compatible directories (config.json, model.safetensors).
  • GGUF Conversion: Compiles models directly into GGUF format (f16, q8_0, q4_k_m) for high-speed inference on CPU and Apple Silicon / CUDA via llama.cpp.

6. πŸ’¬ Interactive Streamed Chat

  • Embedded Inference: Run local GGUF models directly within the desktop application using llama-cpp-python.
  • Streamed Markdown Rendering: Rich formatting with live token-by-token streaming, code syntax highlighting, and copy buttons.
  • Reasoning / Thinking Controls: Selectable reasoning effort profiles (Light, Balanced, Deep) for chain-of-thought exploration.
  • Hyperparameter Tuning: Real-time control over temperature, top-p, top-k, repetition penalty, context window, and pinned system prompts.

7. πŸ”’ Machine-Bound Encrypted Licensing

  • Local-First Launch: Startup checks local encrypted license metadata in ~1msβ€”no remote network lag and zero UI freeze.
  • Two-Layer Cryptography: Fernet authenticated encryption (AES-128-CBC + HMAC-SHA256 keyed via PBKDF2 with machine ID and hardware MAC node) wrapped with Windows DPAPI (CryptProtectData).
  • Machine & User Bound: Ciphertext is bound to the physical device and local Windows user credentials, preventing unauthorized license transfer.

πŸ–ΌοΈ Visual Tour of the IDE

Tab Screen Preview Core Purpose
Dataset Blueprint Blueprint Inspect, categorize, and verify source text, code, and PDF documents with estimated token counts.
Data Ingestion Ingestion Train custom BPE tokenizers, apply target loss masks, and compile fast binary .npy token streams.
Neural Forge Training Configure model parameters (layers, heads, dimensions, RoPE, SDPA) and launch pretraining.
Fine-Tuning Lab Fine-Tuning Select instruction, code, conversation, or tool-call adaptation with full fine-tune or LoRA.
Live Telemetry Live Monitor real-time training and validation loss curves, hardware utilization, and step throughput.
Job Manager Jobs Supervise background worker processes, view run manifests, inspect heartbeats, and reattach.
Benchmarks Benchmarks Run standardized benchmark prompts across checkpoints to quantify perplexity and latency.
Export Bay Export Export to SafeTensors, HuggingFace format, or quantize directly into GGUF (Q4_K_M, Q8_0).
Chat Studio Chat Test trained models in an interactive, streaming chat interface with reasoning and sampler controls.
Activation License Machine-bound cryptographic activation dialog for seamless offline and online licensing.

πŸš€ Installation & Quickstart

Prerequisites

  • Operating System: Windows 10/11 (64-bit), Linux (Ubuntu 22.04+ recommended), or macOS (Apple Silicon).
  • Python: Version 3.12 or newer.
  • GPU (Recommended): NVIDIA GPU with CUDA 12.x support (RTX 30xx/40xx/50xx or Data Center GPUs) for accelerated training. CPU training is fully supported for small models.

Step 1: Clone the Repository

git clone https://github.com/drunkenbot-ai/LLM-IDE.git
cd LLM-IDE

Step 2: Create & Activate Virtual Environment

# Windows (PowerShell)
python -m venv .venv
.\.venv\Scripts\Activate.ps1

# Linux / macOS
python3 -m venv .venv
source .venv/bin/activate

Step 3: Install Dependencies

pip install --upgrade pip
pip install -r requirements.txt

Optional (CUDA Acceleration): If running an NVIDIA GPU, ensure your PyTorch build includes CUDA:

pip install torch --index-url https://download.pytorch.org/whl/cu124

Step 4: Launch DrunkenBot LLM-IDE

python run_app.py

On Linux or macOS, you can also make run_app.py directly executable:

chmod +x run_app.py
./run_app.py

On startup, the IDE executes a startup validation splash screen that verifies local cache directories, inspects core dependencies, and verifies repository integrity before presenting the Project Chooser.


πŸ’» Headless CLI Interface

The non-Qt computational core can be executed directly from the terminal or in automated CI/CD pipelines without launching the graphical desktop interface:

# Display CLI help and available commands
python -m engine.cli --help

1. Ingest & Tokenize Data

# Prepare text/code files with prompt loss masking
python -m engine.cli prepare \
  --input_dir ./data/raw_corpus \
  --output_dir ./runs/prepared_data \
  --context_length 512 \
  --vocab_size 16384

2. Ingest Source Code

# Prepare programming language corpora with indentation preservation
python -m engine.cli prepare \
  --input_dir ./data/code_files \
  --output_dir ./runs/code_data \
  --context_length 1024 \
  --code_training_mode

3. Train a Model Headless

python -m engine.cli train \
  --data_dir ./runs/prepared_data \
  --output_dir ./runs/my_model \
  --epochs 3 \
  --batch_size 16 \
  --context_length 512 \
  --embedding_size 512 \
  --head_count 8 \
  --layer_count 6 \
  --device cuda

πŸ—οΈ Architecture & Project Structure

The codebase enforces a strict unidirectional dependency architecture:

  • engine/: Pure computational Python/PyTorch modules. Never imports GUI or Qt code. Can be imported in scripts, notebooks, or CLI workers.
  • interface/: PySide6 desktop application, widget assemblies, screen mixins, and reactive event loops.
LLM-IDE/
β”œβ”€β”€ engine/                          # Non-Qt Computational Backend
β”‚   β”œβ”€β”€ config.py                    # Model, Dataset, and Training dataclass configs
β”‚   β”œβ”€β”€ model.py                     # Transformer architecture (RoPE, GQA, SwiGLU, SDPA)
β”‚   β”œβ”€β”€ tokenizer.py                 # BPE tokenizer training and streaming encoding
β”‚   β”œβ”€β”€ target_masking.py            # Prompt loss masking (IGNORE_INDEX = -100)
β”‚   β”œβ”€β”€ data_core.py                 # Dataset loading, indentation preservation, normalization
β”‚   β”œβ”€β”€ dataset_corpus.py            # Corpus building, document chunking, diversity filter
β”‚   β”œβ”€β”€ training_orchestrator.py     # Training loop, optimizer, loss computation, scheduler
β”‚   β”œβ”€β”€ training_runtime.py          # Array chunking, dataset splits, checkpoint saving
β”‚   β”œβ”€β”€ training_worker.py           # Standalone headless background worker process
β”‚   β”œβ”€β”€ lora.py                      # Parameter-Efficient Fine-Tuning (LoRA) layers
β”‚   β”œβ”€β”€ license_client.py            # Machine-bound encrypted licensing (DPAPI + Fernet)
β”‚   β”œβ”€β”€ export.py                    # GGUF, SafeTensors, and HuggingFace exporters
β”‚   β”œβ”€β”€ microgpt_chat.py             # Embedded PyTorch chat session with KV cache
β”‚   └── llama_chat.py                # llama.cpp GGUF streamed chat wrapper
β”œβ”€β”€ interface/                       # PySide6 Desktop Application
β”‚   β”œβ”€β”€ app.py                       # Main application entry point and window lifecycle
β”‚   β”œβ”€β”€ license_activation_dialog.py # Activation dialog and responsive background thread
β”‚   β”œβ”€β”€ startup.py                   # Validation splash screen & project chooser
β”‚   β”œβ”€β”€ core/                        # Window core mixin, project state, and menu logic
β”‚   β”œβ”€β”€ screens/                     # Screen widgets (Dataset, Training, LoRA, Export, Chat)
β”‚   β”œβ”€β”€ tabs/                        # Tab layouts and navigation adapters
β”‚   └── widgets/                     # Custom UI controls, app shell, and metric cards
β”œβ”€β”€ packaging/                       # Inno Setup and PyInstaller packaging scripts
β”œβ”€β”€ ref/                             # High-resolution UI screenshots and artwork
β”œβ”€β”€ tests/                           # 140+ comprehensive pytest test suites
β”œβ”€β”€ README.md                        # Project landing page (this document)
β”œβ”€β”€ Technical_Documentation.md       # Exhaustive technical specification & architecture
β”œβ”€β”€ How to Train your LLM.md         # End-to-end user tutorial for building custom LLMs
└── pyproject.toml                   # Project metadata, dependencies, and tooling config

πŸ“¦ Building Standalone Installers

To package DrunkenBot LLM-IDE into a self-contained Windows executable and Inno Setup installer that includes a private Python runtime:

# Standard CPU / Generic Build
python packaging/packager.py

# CUDA-Enabled GPU Build (auto-probes installed NVIDIA drivers)
python packaging/packager.py --gpu

See build.md for full compilation prerequisites, Inno Setup configurations, and runtime profile details.


πŸ“š Further Reading

  • πŸ“˜ How to Train your LLM.md: The practical, step-by-step handbook for creating your first language model from scratch.
  • πŸ”¬ Technical_Documentation.md: In-depth architectural blueprint covering mathematical formulations, process supervisors, encrypted licensing, and algorithm implementations.
  • βš™οΈ build.md: Build options, compiler settings, and standalone distribution instructions.

πŸ“„ License & Intellectual Property

Copyright Β© DrunkenBot. All rights reserved.
See LICENSE.md for licensing terms and usage restrictions.

About

A modern desktop IDE for building, training, fine-tuning, benchmarking, and deploying local language models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages