A knowledge base for building cloud data & AI solutions with GitHub Copilot β organised into specialised "brains", one per technology domain, plus cross-cutting meta-tooling.
Contents: What this is Β· Watch it Β· What it produces Β· Take only what you need Β· Start here Β· Brains Β· Layout Β· Setup Β· Use it from another repo Β· Umbrella knowledge Β· Testing Β· Add a brain Β· Contributing
This repo contains no application. There is nothing to build, nothing to run, no entry point.
What it contains is agent instruction files β the accumulated knowledge of what actually works on Microsoft Fabric, Foundry, Azure databases and the apps built on top of them. An AI coding agent reads the relevant file on demand and follows it.
Two things worth internalising before you use it:
- The instructions exist because of real failures. Rules that look arbitrary usually encode a production incident. Don't improvise around them.
- Every claim carries its evidence. A statement marked observed was seen in a tenant; one marked doc came from Microsoft Learn and may not survive contact with reality. Nothing is labelled "verified" without a trace or a test output behind it β that discipline is the whole value of the repo, and it degrades the moment someone writes down something merely plausible.
teaser-azure-brain-en.mp4
Full quality:
marketing/teaser-azure-brain-en.mp4Β· 40 s, no audio
Forty seconds, no narration. Two projects that share nothing β a retail customer 360 and a
telco network operations console β built by the same agents reading the same instruction files.
The teaser states the case; the rest of this README is the evidence for it. The cut is owned by
marketing/build_teaser.py, so it can be remade when the copy changes.
These are examples, not the product. They come from one build β a customer 360 on retail data β but nothing in it is bespoke. The same agents, reading the same instruction files, produce the equivalent for manufacturing, energy, supply chain, healthcare or finance: the industry is an axis, the tables change, the guidelines do not.
This particular build is B1 + M-ONTO + M-AGENT β the retail starter kit (B1 + M-AGENT) with the
ontology module bolted on. That is the whole idea: a named recipe when one fits, a composition
when none does. The six starter kits and the ten modules are in
Meta-Brain/SCENARIOS.md.
A question in plain language, and the DAX it actually ran. The Data Agent answers "how much customer lifetime value is exposed to churn?", shows the query it generated and names its source β next to the report that number comes from.
What an agent reads before it touches anything. This is the part that transfers between domains. Rule 2 is the three-call OneLake DFS protocol; rule 3 is the polling loop that exists because the SQL endpoint is not ready when creation returns. Both are there because they failed first, and both apply whatever you are modelling.
The semantic layer underneath. Eight entity types bound to lakehouse tables, nine relations β the model answers the numbers, the ontology answers the links, and both are queryable. Swap the industry and this becomes sites, equipment and sensors instead of customers and campaigns.
Served as an application. One entrance per persona, over the same Fabric artifacts.
You are not expected to adopt all of it. Pick the granularity that matches what you're doing:
| Level | You take | When |
|---|---|---|
| One agent | a single instructions.md plus the companions it names |
one task β "build the semantic model" |
| One brain | e.g. Fabric-Brain/ |
you work in one technology all day |
| One scenario | a preset from SCENARIOS.md β base + modules |
you're building something end to end |
| The whole brain, from your own repo | reference it by path, pinned to a tag | it's your team's standing knowledge base |
This is a measured property, not an intention: 38 of the 42 agents depend on no umbrella file at all, and each technology brain resolves the large majority of its links internally β Fabric 83 %, Database 84 %, Foundry 89 %. Lifting one out is a copy, not a surgery.
The exception is Apps-Brain β 16 % internal, deliberately: it is the
layer that consumes the others, so it points at them constantly. Take it together with the
brains it references, not on its own.
The real entry point is AGENTS.md β the routing table plus the index of all 42
agents. It is auto-loaded by the GitHub Copilot CLI and the Copilot app;
.github/copilot-instructions.md is its VS Code counterpart.
Both point at the same tree, so nothing is duplicated.
If you'd rather be pointed straight at a starting file:
| I want to⦠| Brain | Open first |
|---|---|---|
| Land data, model it, ship a report on Fabric | Fabric | lakehouse-agent β then semantic-model-agent β report-builder-agent |
| Build AI agents that orchestrate other agents | Foundry | generation_map.md first, then foundry-orchestration-agent |
| Let an app or a portal consume the platform | Apps | Apps-Brain/README.md β the runtime is a decision inside that brain |
| Deploy or migrate a database | Database | postgres-deploy-agent Β· Oracle β PG track under 03-oracle-to-postgres/ |
| Ask a question over data in natural language | Fabric β Foundry | ai-skills-agent creates the Data Agent; foundry-fabric-bridge-agent consumes it |
| Find out why it did that | Foundry | foundry-observability-agent β a trace is the only place a multi-agent system is legible |
| Build a whole project end to end | Meta | SCENARIOS.md β pick a preset (base + modules), or compose your own; project-orchestrator-agent drives the 12 steps |
| Write tests, a deck, a diagram, a README | Meta | Meta-Brain/README.md |
Something broke? known_issues.md, then
ERROR_RECOVERY.md β decision trees by HTTP status. Most errors you will hit
are already written down.
| Brain | Scope | Agents | Status |
|---|---|---|---|
| Fabric-Brain | Microsoft Fabric β Lakehouse, Warehouse, semantic models, reports, Real-Time Intelligence, Data Agents, Ontology, migrations | 24 | β Active |
| Foundry-Brain | Microsoft Foundry β agent service, tools, knowledge (Foundry IQ), orchestration, observability, governance, the Fabric bridge | 7 active / 11 catalogued | π‘ Bootstrap |
| Apps-Brain | Applications β the layer that consumes the platform brains: runtime, identity, embedding, in-app intelligence, frontend, operations | 3 active / 9 catalogued | π‘ Bootstrap |
| Database-Brain | Azure databases β Azure SQL, PostgreSQL, Cosmos DB, MySQL, cross-engine migration (Oracle β PostgreSQL track live) | 4 active / 22 catalogued | β Active |
| Meta-Brain | Cross-cutting β testing, PowerPoint, HTML diagrams, README authoring, project orchestration | 5 | β Active |
| Databricks-Brain | Databricks on Azure | β | π Planned |
| Synapse-Brain | Azure Synapse legacy | β | π Planned |
One owner per domain. Any agent may read any artifact; only its owner modifies it. Crossing a
boundary is an explicit handoff β state what was produced, name the next agent, list the affected
files and IDs. Boundary notes between confusable agents live in each brain's
agents/_catalog.yaml and in AGENTS.md.
Azure-Brain/ β umbrella (this repo)
βββ AGENTS.md β entry point: routing table + index of all 42 agents
βββ Fabric-Brain/ β Microsoft Fabric (24 agents, flat)
βββ Foundry-Brain/ β Microsoft Foundry (7 active / 11 catalogued, flat)
βββ Apps-Brain/ β Applications (3 active / 9 catalogued, flat)
βββ Database-Brain/ β Azure databases (4 active / 22 catalogued, nested)
βββ Meta-Brain/ β cross-cutting + tests (5 agents, flat)
βββ (future brains) β Databricks-Brain, Synapse-Brain, β¦
Layout note. Fabric-Brain, Foundry-Brain, Apps-Brain and Meta-Brain keep agents flat (
agents/<agent>/). Database-Brain nests them by domain (agents/<NN-domain>/<agent>/). Any tooling that walks agents must handle both depths β see the depth-awareagent_dirs()inMeta-Brain/tests/conftest.py.
Every agent folder holds instructions.md (the agent itself), a README.md (human summary),
usually a known_issues.md, and whatever domain files that agent declares in its own load order.
Read instructions.md in full before acting β it names the companion files it needs, and
guessing which ones matter is how the documented failures come back.
git clone https://github.com/Statyx/Azure-Brain.git
cd Azure-BrainNothing to install for reading. Copy the config templates only for the brains you actually use β
each pair is gitignored and ships with a committed .example twin:
# Fabric work
cp Fabric-Brain/resource_ids.example.md Fabric-Brain/resource_ids.md
cp Fabric-Brain/environment.example.md Fabric-Brain/environment.md
# Foundry work
cp Foundry-Brain/resource_ids.example.md Foundry-Brain/resource_ids.md
cp Foundry-Brain/environment.example.md Foundry-Brain/environment.md
# Database work
cp Database-Brain/resource_ids.example.md Database-Brain/resource_ids.md
cp Database-Brain/environment.example.md Database-Brain/environment.mdApps-Brain and Meta-Brain need no local config. You don't need every ID up front β fill them in
as you deploy. Full walkthrough: GETTING_STARTED.md.
Then just work. The Copilot CLI and the Copilot app load AGENTS.md; VS Code loads
.github/copilot-instructions.md; both discover the agents from there. Ask for what you want to
build β "create a Fabric workspace and lakehouse for a finance demo", "stand up a Foundry
supervisor over two Fabric data agents" β and the routing table does the rest.
MCP servers available to this workspace are catalogued in
Meta-Brain/mcp_registry.md.
Azure-Brain holds the knowledge; your project repo holds the project. Keep the two checked out
side by side and point the agent at the relevant instructions.md by path. Nothing here needs to
be installed or copied.
Do not fork or duplicate
instructions.mdinto the consuming repo. The brain is the single source of truth; a copy goes stale silently and reintroduces exactly the failures these files prevent.
Pin a version. This brain drives agent behaviour, so an unpinned reference means your agents
change the day this repo does. Check out a tag, not main:
git clone --branch v1.0.0 https://github.com/Statyx/Azure-Brain.git ../Azure-BrainThen paste this into your own repo's AGENTS.md (or .github/copilot-instructions.md),
keeping only the rows you actually use:
## Knowledge base
This project uses [Azure-Brain](https://github.com/Statyx/Azure-Brain), checked out at
`../Azure-Brain`, pinned to `v1.0.0`. It is the single source of truth β never copy an
`instructions.md` into this repo.
Before acting on one of these topics, read the matching file **in full**:
| Task | Read first |
| --- | --- |
| Lakehouse, Delta tables, OneLake | `../Azure-Brain/Fabric-Brain/agents/lakehouse-agent/instructions.md` |
| Semantic model, DAX, Direct Lake | `../Azure-Brain/Fabric-Brain/agents/semantic-model-agent/instructions.md` |
| Power BI report | `../Azure-Brain/Fabric-Brain/agents/report-builder-agent/instructions.md` |
Always: read before write Β· idempotent re-runs Β· async-first (HTTP 202, then poll).
Something failed β `../Azure-Brain/known_issues.md`, then `../Azure-Brain/ERROR_RECOVERY.md`.
Lessons learned go back into the brain, never into this repo.Corrections and lessons learned go back into the brain (the relevant agent's
known_issues.md) β see CONTRIBUTING.md. Details:
Using this brain from another working directory.
Applies to every brain. A per-agent instructions.md may be stricter, and wins on its own
domain.
| File | Purpose |
|---|---|
AGENTS.md |
Entry point β routing table + index of all 42 agents |
agent_principles.md |
Mandatory β plan first, verify before done, capture the lesson after any correction |
shared_constraints.md |
The 9 hard rules β read before write, config-driven, idempotent, async-first |
PUBLIC_SAFETY.md |
Write as if already public β the company is always Zava, GUIDs are visibly fake, no account name in a path, secrets read at runtime |
known_issues.md |
Cross-cutting gotchas and workarounds |
ERROR_RECOVERY.md |
Decision trees by HTTP status, retry patterns |
GETTING_STARTED.md |
15-minute setup walkthrough |
Conventions: Python 3.12+ with pathlib and type hints Β· UTF-8 everywhere, no BOM
([System.IO.File]::WriteAllText() in PowerShell, never Out-File for JSON) Β· conventional
commits (feat(foundry):, fix(core):, docs(brain):).
Any change to agent instructions, catalogs or shared docs must keep both gates green:
cd Meta-Brain
pip install -r requirements.txt # pytest + PyYAML, first run only
python -m pytest tests/ -v --tb=short
cd ..
python Meta-Brain/tools/scan_public_safety.py .The suite validates, for every brain listed in BRAINS (Meta-Brain/tests/conftest.py β the
single source of truth for coverage): catalogs parse and match what is on disk, every agent folder
has a non-trivial instructions.md, internal markdown links resolve, Python compiles, JSON parses,
root markdown is non-empty.
The scanner is the second gate, and the one that matters most in practice: it is what stops a
customer name, a real endpoint or a tenant GUID reaching a public repo. It only works if
something runs it. CI (.github/workflows/no-client-leak.yml) runs both on every push and pull
request β run them locally first anyway.
The suite is green on a fresh clone, and it must stay that way. A permanently-red test gives no signal: it becomes indistinguishable from a real regression. If you see a failure here now, it is real. (Narrow exemption: a link may resolve to a missing file only when that filename is gitignored and a committed
<name>.example.mdsits beside it β both conditions required.)
For reference, on main at 2026-08-26 both gates give 1777 passed, 6 skipped and clean.
The count grows as agents and documents are added β that is expected. A count that drops is not:
it means assertions stopped being collected. The badge at the top of this page reports the same two
gates as CI last ran them on main.
- Create the folder plus
README.mdandagents/_catalog.yaml. - Add the brain to
BRAINSinMeta-Brain/tests/conftest.pyβ single source of truth, covers every test module at once. - Update three places by hand: the brain table and layout tree in this README, the brain list
in
.github/copilot-instructions.md, and the layout + routing table + agent index inAGENTS.md. - Re-run both gates.
Agent counts appear in several files. If you change one, change them all β the tests check catalogs against disk, but nothing checks prose.
The most valuable contribution here is a lesson learned the hard way β an error message and its real cause, an API that behaves differently from its documentation, a rule that turned out to be wrong. That is what this repo is made of. The second most valuable is telling us a rule is wrong: being contradicted by a tenant is the point.
Read CONTRIBUTING.md first β it names the five rules that actually fail a PR
(the 20 KB cap, evidence labels, public-safety, catalog sync, expiry clocks) and the two gate
commands to run locally before you open one.
Version history and upgrade notes: CHANGELOG.md.
MIT



