A production-grade AIDevOps platform: a modular framework for testing, evaluating, red-teaming, securing, monitoring, and governing AI systems β plus a set of 15 runnable example projects that demonstrate each framework in real-world scenarios.
Platform: Windows, Mac (Intel & Apple Silicon), Linux (Ubuntu/Debian/Fedora/Arch/Alpine)
Setup time: ~5 minutes per machine
Status: All three tiers complete. Every framework ships with real, working logic. Nothing here is a placeholder.
Secrets: Production-grade secrets management (encrypted .env.gpg, external vaults, env vars).
git clone https://github.com/AshraHossain/AI_Engineering_Cockpit.git
cd AI_Engineering_Cockpit
# Windows (PowerShell):
powershell -ExecutionPolicy Bypass -File scripts\setup-windows.ps1
# Mac or Linux (Bash):
bash scripts/setup-mac.sh # or setup-linux.sh# Initialize API keys interactively
python -m cockpit.config.secrets_cli init
# Verify secrets loaded
python -m cockpit.config.secrets_cli verify
# (Optional) Encrypt with GPG for production
python -m cockpit.config.secrets_cli encryptcd projects/01-hello-world
uv sync
uv run python src/main.pyEach setup script:
- Installs
uvif missing - Installs Python 3.11+ toolchain via uv
- Optionally installs Ollama (local model inference)
- Syncs all dependencies
- Creates
.envfrom.env.example - Runs the test suite (mocked, no API keys needed)
Secrets management (new):
- Interactive setup:
python -m cockpit.config.secrets_cli init - Load from: env vars β encrypted .env.gpg β plaintext .env β external vaults
- All projects inherit secure secrets automatically
See SETUP.md for detailed platform-specific guidance and docs/SECRETS.md for production secrets workflows.
cockpit/security/secrets_manager.py handles API keys securely across all
environments:
- Development: Plaintext
.env(git-ignored) - Team: Encrypted
.env.gpg(GPG-encrypted, shareable) - Production: External vaults (AWS Secrets Manager, HashiCorp Vault, 1Password)
All projects access secrets the same way β no hardcoding, no plaintext in logs. See docs/SECRETS.md for setup and integration.
cockpit/security/ ships real regex-based prompt-injection detection and PII
masking with Luhn validation (cuts credit-card false positives). Not a TODO.
It runs 32-payload corpus through the injection detector and reports which attacks slip past, so security blind spots are measured, not assumed. See docs/SECURITY_COVERAGE.md.
Scans match against the raw input, Unicode-normalized form (zero-width characters, letter-spacing, accented homoglyphs), and base64/hex/rot13 decodings β because an encoded instruction is still an instruction.
cockpit/config/feature_flags.py gates testing/evaluation/red-teaming/security/monitoring/governance
independently, with pre-built profiles for startup, enterprise, research,
and safety-focused use cases.
Each projects/0X-*/ is its own uv-managed project with mocked tests β
no live API key needed to run pytest, but the code talks to real provider SDKs.
- Windows: Cloud-only (Gemini, OpenAI, Anthropic APIs)
- Mac (Intel): Hybrid (cloud + CPU-only Ollama)
- Mac (Apple Silicon): Hybrid (cloud + GPU-accelerated Ollama)
- Linux (NVIDIA/AMD): Hybrid (cloud + GPU-accelerated Ollama)
Setup is identical across all platforms (setup-windows.ps1, setup-mac.sh, setup-linux.sh).
See docs/WINDOWS_VS_MAC.md and docs/LINUX_SETUP.md.
cockpit/ Shared framework: testing, evaluation, red_teaming,
security, monitoring, governance, config, utils.
See docs/ARCHITECTURE.md.
projects/ 15 standalone example projects (01-hello-world through
16-production-agent; 15 is ORION, its own repository).
Each has its own pyproject.toml.
models/ Model registry + Ollama/Hugging Face reference material.
docs/ Architecture, getting started, deployment, troubleshooting,
API comparison, recipes, platform-specific guides.
scripts/ Setup scripts (Windows PowerShell, Mac Bash, Linux Bash),
test runners, formatters, project scaffolding.
config/ pytest, ruff, mypy, pre-commit, and coverage configuration.
tests/ Root-level tests covering cockpit/config/* and repo structure.
.github/ CI workflows (test, security, docs) and issue/PR templates.
Projects 01-05 are standalone. Projects 06-14 and 16 import cockpit/ frameworks.
| # | Name | Demonstrates |
|---|---|---|
| 01 | hello-world | Minimal Gemini API call |
| 02 | rag-chatbot | Retrieval-augmented generation with task-asymmetric embeddings |
| 03 | multi-model-orchestrator | Comparing Gemini vs OpenAI in one app |
| 04 | streaming-responses | Streaming model output to client |
| 05 | hybrid-orchestrator | Local Ollama + cloud Gemini routing (Mac/Linux only) |
| 06 | eval-harness | Scoring a pipeline against a reference set (quality, safety, cost) |
| 07 | red-team-runner | Attacking a target with injection corpus; reporting defense leaks |
| 08 | cost-dashboard | Instrumenting model calls for spend and latency observability |
| 09 | agent-tool-use | Function calling with automatic and manual authorization |
| 10 | batch-pipeline | Bulk processing with rate limiting, retries, partial-failure tolerance |
| 11 | secure-gateway | RBAC, input/output security, tamper-evident audit chain |
| 12 | governed-deployment | Model promotion gated by N-of-M approvals |
| 13 | compliance-report | GDPR/HIPAA/SOX findings with evidence |
| 14 | threat-monitor | Brute force, exfiltration, privilege-escalation detection |
| 16 | production-agent | Claude agent traced in Phoenix, loop/failure alerts, canary with approval-gated promotion and automatic rollback |
Every project runs offline tests with no API key. Projects 06-14 and 16 support --dry-run to see them work before adding credentials.
Unit, integration, and benchmark tests. All projects include mocked test suites.
Quality metrics (BLEU, ROUGE, custom scoring), safety evaluation (toxicity, bias detection), and cost evaluation (per-token pricing, spend tracking).
Adversarial testing with 32-payload injection corpus. Tests your own defenses and reports which attacks slip through.
- Input security: Prompt injection detection, input validation, rate limiting
- Output security: PII masking (email, phone, SSN, credit card, passport, IP)
- Data security: Encryption at-rest/in-transit, secrets management
- Access control: RBAC (admin, developer, viewer, external)
- Audit & logging: Tamper-evident audit trail, real-time monitoring
- Compliance: GDPR, HIPAA, SOX findings with evidence
- Threat detection: Brute force, exfiltration, privilege escalation
See docs/SECURITY_COVERAGE.md for detailed coverage.
- Cost tracking per model/API call
- Performance metrics (latency, throughput)
- Model versioning and promotion workflows
- Approval workflows (N-of-M gating)
- Audit logging with hash chains
- Enable/disable frameworks independently
- Pre-built use-case profiles: Startup, Enterprise, Research, Safety-Focused
- Per-environment configuration (dev, staging, prod)
| Windows | Mac (Intel) | Mac (Apple Silicon) | Linux | |
|---|---|---|---|---|
| Setup time | 5 min | 5 min | 5 min | 5 min |
| Cloud APIs | β Gemini/OpenAI/Anthropic | β | β | β |
| Local Ollama | β Not supported | β CPU-only (1-5 tok/s) | β GPU-accel (10-50+ tok/s) | β GPU-accel (NVIDIA/AMD) |
| Recommended use | Development, cloud-only apps | Hybrid (cost-effective) | Hybrid (fast inference) | Hybrid (GPU clusters) |
| Setup script | setup-windows.ps1 |
setup-mac.sh |
setup-mac.sh |
setup-linux.sh |
All platforms run identical code. Platform-specific behavior is opt-in via configuration (cloud-only vs hybrid mode).
- Git
- No Python install needed β
uvhandles the Python toolchain - Optional: Homebrew (Mac), apt/dnf/pacman/apk (Linux), or curl (for manual installs)
See SETUP.md for:
- Detailed per-platform instructions
- Troubleshooting (PATH issues, dependency conflicts, Ollama setup)
- Re-running the setup safely multiple times
uv run python -c "import sys; print(sys.version)" # should print 3.11.x
uv run pytest -c config/pytest.ini --rootdir=. -v # root suite, no API keys needed- SETUP.md β Detailed per-platform setup guide with troubleshooting
- docs/SECRETS.md β NEW: Secrets management (encrypted .env.gpg, external vaults, CLI)
- docs/ARCHITECTURE.md β System design, module overview
- docs/WINDOWS_VS_MAC.md β Why platforms differ
- docs/LINUX_SETUP.md β Linux-specific guidance (GPU setup, package managers)
- docs/SECURITY_COVERAGE.md β Prompt injection detection limits (24/32 payloads caught)
- docs/DEPLOYMENT.md β Containerization, secrets, production checklist
- docs/API_COMPARISON.md β Gemini vs OpenAI vs Anthropic
- docs/GETTING_STARTED.md β First project walkthrough
- docs/TROUBLESHOOTING.md β Common issues and fixes
# Initialize API keys
python -m cockpit.config.secrets_cli init
# Verify secrets loaded
python -m cockpit.config.secrets_cli verify
# Display loaded secrets (masked)
python -m cockpit.config.secrets_cli show
# Encrypt with GPG (production)
python -m cockpit.config.secrets_cli encryptAutomatic (most projects):
from cockpit.config.settings import get_settings
settings = get_settings()
api_key = settings.gemini_api_key # Loaded securelyDirect access:
from cockpit.security.secrets_manager import get_secret
api_key = get_secret("GEMINI_API_KEY")With verification:
from cockpit.security.secrets_manager import get_secrets_manager
manager = get_secrets_manager()
if not manager.verify(["GEMINI_API_KEY"]):
raise RuntimeError("Missing required API key")Secrets load from (first match wins):
- Environment variables (highest priority)
- Encrypted
.env.gpg(GPG-encrypted, recommended for teams) - Plaintext
.env(development only) - External vaults (AWS Secrets Manager, Vault, 1Password, etc.)
See docs/SECRETS.md for production workflows and integration examples.
cd ~/AI_Engineering_Cockpit
uv run pytest -c config/pytest.ini --rootdir=. -qbash scripts/run-tests.shmake lint # ruff check (no fixes)
make typecheck # mypy against cockpit/
make security # bandit against cockpit/This repo is intended for research, development, and learning. If you're building production systems, please read:
- docs/SECURITY_COVERAGE.md β Honest gaps in injection detection
- docs/DEPLOYMENT.md β Pre-launch checklist, secrets management, audit trail limitations
- SECURITY.md β How to report security issues
Key highlights:
- Injection detector catches 24 of 32 corpus payloads (known gap documented)
- Audit log is tamper-evident (hash chain reveals tampering) not tamper-proof
- Compliance findings are evidence-based; no verdicts (compliance is a legal determination)
- Secrets storage example in code; integrate with real KMS for production
See CONTRIBUTING.md and CODE_OF_CONDUCT.md.
-
Clone and setup using the quick-start above (5 minutes)
git clone https://github.com/AshraHossain/AI_Engineering_Cockpit.git cd AI_Engineering_Cockpit bash scripts/setup-mac.sh # or setup-linux.sh or setup-windows.ps1
-
Set up secrets (interactive, 2 minutes)
python -m cockpit.config.secrets_cli init python -m cockpit.config.secrets_cli verify
-
Run the first example
projects/01-hello-worldcd projects/01-hello-world uv sync uv run python src/main.py -
Try a framework that interests you:
- 06-eval-harness β Evaluation framework
- 07-red-team-runner β Security testing
- 11-secure-gateway β Access control
- 12-governed-deployment β Approval workflows
-
Explore the frameworks in
cockpit/β each module is standalone and well-commented -
Read the architecture docs/ARCHITECTURE.md to understand how pieces fit
-
Deploy to production following docs/DEPLOYMENT.md and docs/SECRETS.md
Questions?
- Setup issues? See SETUP.md and docs/TROUBLESHOOTING.md
- Secrets/security? See docs/SECRETS.md
- Linux/platform-specific? See docs/LINUX_SETUP.md
- Contributing? See CONTRIBUTING.md
- Security issues? See SECURITY.md