Skip to content

Latest commit

Β 

History

99 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

AI Engineering Cockpit

A production-grade AIDevOps platform: a modular framework for testing, evaluating, red-teaming, securing, monitoring, and governing AI systems β€” plus a set of 15 runnable example projects that demonstrate each framework in real-world scenarios.

Platform: Windows, Mac (Intel & Apple Silicon), Linux (Ubuntu/Debian/Fedora/Arch/Alpine)
Setup time: ~5 minutes per machine
Status: All three tiers complete. Every framework ships with real, working logic. Nothing here is a placeholder.
Secrets: Production-grade secrets management (encrypted .env.gpg, external vaults, env vars).


πŸš€ Quick start (5 minutes)

All Platforms (Unified Setup)

git clone https://github.com/AshraHossain/AI_Engineering_Cockpit.git
cd AI_Engineering_Cockpit

# Windows (PowerShell):
powershell -ExecutionPolicy Bypass -File scripts\setup-windows.ps1

# Mac or Linux (Bash):
bash scripts/setup-mac.sh    # or setup-linux.sh

Set Up Secrets (Interactive)

# Initialize API keys interactively
python -m cockpit.config.secrets_cli init

# Verify secrets loaded
python -m cockpit.config.secrets_cli verify

# (Optional) Encrypt with GPG for production
python -m cockpit.config.secrets_cli encrypt

Run Your First Project

cd projects/01-hello-world
uv sync
uv run python src/main.py

Each setup script:

  1. Installs uv if missing
  2. Installs Python 3.11+ toolchain via uv
  3. Optionally installs Ollama (local model inference)
  4. Syncs all dependencies
  5. Creates .env from .env.example
  6. Runs the test suite (mocked, no API keys needed)

Secrets management (new):

  • Interactive setup: python -m cockpit.config.secrets_cli init
  • Load from: env vars β†’ encrypted .env.gpg β†’ plaintext .env β†’ external vaults
  • All projects inherit secure secrets automatically

See SETUP.md for detailed platform-specific guidance and docs/SECRETS.md for production secrets workflows.


πŸ“‹ What makes this different

Secrets management is production-grade

cockpit/security/secrets_manager.py handles API keys securely across all environments:

  • Development: Plaintext .env (git-ignored)
  • Team: Encrypted .env.gpg (GPG-encrypted, shareable)
  • Production: External vaults (AWS Secrets Manager, HashiCorp Vault, 1Password)

All projects access secrets the same way β€” no hardcoding, no plaintext in logs. See docs/SECRETS.md for setup and integration.

Security is first-class, not bolted on

cockpit/security/ ships real regex-based prompt-injection detection and PII masking with Luhn validation (cuts credit-card false positives). Not a TODO.

The red-team harness grades your own defenses

It runs 32-payload corpus through the injection detector and reports which attacks slip past, so security blind spots are measured, not assumed. See docs/SECURITY_COVERAGE.md.

Injection detection sees through obfuscation

Scans match against the raw input, Unicode-normalized form (zero-width characters, letter-spacing, accented homoglyphs), and base64/hex/rot13 decodings β€” because an encoded instruction is still an instruction.

Every framework is feature-flagged

cockpit/config/feature_flags.py gates testing/evaluation/red-teaming/security/monitoring/governance independently, with pre-built profiles for startup, enterprise, research, and safety-focused use cases.

Examples are real, isolated, and tested

Each projects/0X-*/ is its own uv-managed project with mocked tests β€” no live API key needed to run pytest, but the code talks to real provider SDKs.

Cross-platform from the start

  • Windows: Cloud-only (Gemini, OpenAI, Anthropic APIs)
  • Mac (Intel): Hybrid (cloud + CPU-only Ollama)
  • Mac (Apple Silicon): Hybrid (cloud + GPU-accelerated Ollama)
  • Linux (NVIDIA/AMD): Hybrid (cloud + GPU-accelerated Ollama)

Setup is identical across all platforms (setup-windows.ps1, setup-mac.sh, setup-linux.sh). See docs/WINDOWS_VS_MAC.md and docs/LINUX_SETUP.md.


πŸ—οΈ Repository layout

cockpit/              Shared framework: testing, evaluation, red_teaming,
                      security, monitoring, governance, config, utils.
                      See docs/ARCHITECTURE.md.

projects/             15 standalone example projects (01-hello-world through
                      16-production-agent; 15 is ORION, its own repository).
                      Each has its own pyproject.toml.

models/               Model registry + Ollama/Hugging Face reference material.

docs/                 Architecture, getting started, deployment, troubleshooting,
                      API comparison, recipes, platform-specific guides.

scripts/              Setup scripts (Windows PowerShell, Mac Bash, Linux Bash),
                      test runners, formatters, project scaffolding.

config/               pytest, ruff, mypy, pre-commit, and coverage configuration.

tests/                Root-level tests covering cockpit/config/* and repo structure.

.github/              CI workflows (test, security, docs) and issue/PR templates.

πŸ“š The 15 Example Projects

Projects 01-05 are standalone. Projects 06-14 and 16 import cockpit/ frameworks.

# Name Demonstrates
01 hello-world Minimal Gemini API call
02 rag-chatbot Retrieval-augmented generation with task-asymmetric embeddings
03 multi-model-orchestrator Comparing Gemini vs OpenAI in one app
04 streaming-responses Streaming model output to client
05 hybrid-orchestrator Local Ollama + cloud Gemini routing (Mac/Linux only)
06 eval-harness Scoring a pipeline against a reference set (quality, safety, cost)
07 red-team-runner Attacking a target with injection corpus; reporting defense leaks
08 cost-dashboard Instrumenting model calls for spend and latency observability
09 agent-tool-use Function calling with automatic and manual authorization
10 batch-pipeline Bulk processing with rate limiting, retries, partial-failure tolerance
11 secure-gateway RBAC, input/output security, tamper-evident audit chain
12 governed-deployment Model promotion gated by N-of-M approvals
13 compliance-report GDPR/HIPAA/SOX findings with evidence
14 threat-monitor Brute force, exfiltration, privilege-escalation detection
16 production-agent Claude agent traced in Phoenix, loop/failure alerts, canary with approval-gated promotion and automatic rollback

Every project runs offline tests with no API key. Projects 06-14 and 16 support --dry-run to see them work before adding credentials.


πŸ”§ The Six Core Frameworks

1. Testing Framework

Unit, integration, and benchmark tests. All projects include mocked test suites.

2. Evaluation Framework

Quality metrics (BLEU, ROUGE, custom scoring), safety evaluation (toxicity, bias detection), and cost evaluation (per-token pricing, spend tracking).

3. Red Teaming Framework

Adversarial testing with 32-payload injection corpus. Tests your own defenses and reports which attacks slip through.

4. Security Framework ⭐ First-Class

  • Input security: Prompt injection detection, input validation, rate limiting
  • Output security: PII masking (email, phone, SSN, credit card, passport, IP)
  • Data security: Encryption at-rest/in-transit, secrets management
  • Access control: RBAC (admin, developer, viewer, external)
  • Audit & logging: Tamper-evident audit trail, real-time monitoring
  • Compliance: GDPR, HIPAA, SOX findings with evidence
  • Threat detection: Brute force, exfiltration, privilege escalation

See docs/SECURITY_COVERAGE.md for detailed coverage.

5. Monitoring & Governance

  • Cost tracking per model/API call
  • Performance metrics (latency, throughput)
  • Model versioning and promotion workflows
  • Approval workflows (N-of-M gating)
  • Audit logging with hash chains

6. Configuration & Feature Flags

  • Enable/disable frameworks independently
  • Pre-built use-case profiles: Startup, Enterprise, Research, Safety-Focused
  • Per-environment configuration (dev, staging, prod)

πŸ“Š Cross-Platform Comparison

Windows Mac (Intel) Mac (Apple Silicon) Linux
Setup time 5 min 5 min 5 min 5 min
Cloud APIs βœ… Gemini/OpenAI/Anthropic βœ… βœ… βœ…
Local Ollama ❌ Not supported βœ… CPU-only (1-5 tok/s) βœ… GPU-accel (10-50+ tok/s) βœ… GPU-accel (NVIDIA/AMD)
Recommended use Development, cloud-only apps Hybrid (cost-effective) Hybrid (fast inference) Hybrid (GPU clusters)
Setup script setup-windows.ps1 setup-mac.sh setup-mac.sh setup-linux.sh

All platforms run identical code. Platform-specific behavior is opt-in via configuration (cloud-only vs hybrid mode).


🚦 Installation & Setup

Prerequisites (All Platforms)

  • Git
  • No Python install needed β€” uv handles the Python toolchain
  • Optional: Homebrew (Mac), apt/dnf/pacman/apk (Linux), or curl (for manual installs)

Install

See SETUP.md for:

  • Detailed per-platform instructions
  • Troubleshooting (PATH issues, dependency conflicts, Ollama setup)
  • Re-running the setup safely multiple times

Verify

uv run python -c "import sys; print(sys.version)"   # should print 3.11.x
uv run pytest -c config/pytest.ini --rootdir=. -v    # root suite, no API keys needed

πŸ“– Documentation


πŸ” Secrets Management

Quick Setup (Interactive)

# Initialize API keys
python -m cockpit.config.secrets_cli init

# Verify secrets loaded
python -m cockpit.config.secrets_cli verify

# Display loaded secrets (masked)
python -m cockpit.config.secrets_cli show

# Encrypt with GPG (production)
python -m cockpit.config.secrets_cli encrypt

How Projects Access Secrets

Automatic (most projects):

from cockpit.config.settings import get_settings
settings = get_settings()
api_key = settings.gemini_api_key  # Loaded securely

Direct access:

from cockpit.security.secrets_manager import get_secret
api_key = get_secret("GEMINI_API_KEY")

With verification:

from cockpit.security.secrets_manager import get_secrets_manager
manager = get_secrets_manager()
if not manager.verify(["GEMINI_API_KEY"]):
    raise RuntimeError("Missing required API key")

Secrets load from (first match wins):

  1. Environment variables (highest priority)
  2. Encrypted .env.gpg (GPG-encrypted, recommended for teams)
  3. Plaintext .env (development only)
  4. External vaults (AWS Secrets Manager, Vault, 1Password, etc.)

See docs/SECRETS.md for production workflows and integration examples.


πŸ§ͺ Running Tests

Quick validation (no API keys)

cd ~/AI_Engineering_Cockpit
uv run pytest -c config/pytest.ini --rootdir=. -q

Full test suite (all projects, with coverage)

bash scripts/run-tests.sh

Lint and typecheck

make lint      # ruff check (no fixes)
make typecheck # mypy against cockpit/
make security  # bandit against cockpit/

πŸ” Security & Compliance

This repo is intended for research, development, and learning. If you're building production systems, please read:

Key highlights:

  • Injection detector catches 24 of 32 corpus payloads (known gap documented)
  • Audit log is tamper-evident (hash chain reveals tampering) not tamper-proof
  • Compliance findings are evidence-based; no verdicts (compliance is a legal determination)
  • Secrets storage example in code; integrate with real KMS for production

πŸ“¦ Contributing

See CONTRIBUTING.md and CODE_OF_CONDUCT.md.


πŸ“„ License

MIT


🎯 Next Steps

  1. Clone and setup using the quick-start above (5 minutes)

    git clone https://github.com/AshraHossain/AI_Engineering_Cockpit.git
    cd AI_Engineering_Cockpit
    bash scripts/setup-mac.sh  # or setup-linux.sh or setup-windows.ps1
  2. Set up secrets (interactive, 2 minutes)

    python -m cockpit.config.secrets_cli init
    python -m cockpit.config.secrets_cli verify
  3. Run the first example projects/01-hello-world

    cd projects/01-hello-world
    uv sync
    uv run python src/main.py
  4. Try a framework that interests you:

  5. Explore the frameworks in cockpit/ β€” each module is standalone and well-commented

  6. Read the architecture docs/ARCHITECTURE.md to understand how pieces fit

  7. Deploy to production following docs/DEPLOYMENT.md and docs/SECRETS.md


Questions?

About

Production-grade AIDevOps platform: testing, evaluation, red-teaming, security, monitoring, and governance for AI systems, with 14 runnable example projects.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages