Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

context-engineer

You describe what you want. It writes the prompt that gets it done — properly.

Most prompts fail for a boring reason: the AI session doing the work knows nothing you didn't write down. It hasn't seen your files, your constraints, or your standards — and a rambled request leaves it guessing at all three. Guesswork produces generic work.

This skill treats writing-it-down as a discipline. It reads your ramble and your project, asks you only the questions it genuinely cannot answer itself, and compresses everything the working session needs into one prompt you paste into a fresh session: the goal, the exact facts, the quality bar, the decisions it's trusted to make, and how to prove its own work before calling it done.

Install

npx skills add kasoob/context-engineer

Works with Claude Code and any agent that reads SKILL.md-style skills — or copy the context-engineer/ folder into your skills directory.

How much it asks you

As little as the task allows, and it tells you its read up front so you can override it:

  • A clear request gets zero questions — gaps are filled with sensible defaults and listed as assumptions you can correct in one line.
  • A few genuine gaps get one batch of questions, each arriving with a recommended answer, so you can mostly just confirm.
  • A fuzzy or high-stakes ask gets a real interview — questions in dependency order until nothing important is left open, and a restatement of your intent in its own words before a single line of prompt is written.

Three rules keep the interview honest. It never asks what your files already answer — it goes and reads them. It never asks you to make a call it can responsibly make itself; it decides, tells you, and lets you veto. And every question it does ask arrives contextualized, with its own recommended answer and the reason behind it, so you're reacting to a proposal rather than filling in a blank.

The standards every prompt meets

  • No invented facts, ever. A missing name, number, or path becomes a visible [PLACEHOLDER] — with instructions for what the session does when it reaches one.
  • Failure is named, not just success. The prompt says what the generic version of the work would look like, so "safe and mediocre" reads as failure rather than as done.
  • Real decision authority. Whatever you didn't pin down is explicitly handed to the session with criteria to decide by and an instruction to record what it chose — not left as a gap it discovers halfway through.
  • Proof before done. Every prompt defines what evidence of completion looks like — counts that reconcile, pages that load, checks that pass — never a feeling of being finished.
  • Unattended runs are braked. Overnight prompts carry a stop condition and honest-reporting rules: only evidence-backed progress gets claimed, and failures get reported with their output.
  • Nothing sensitive travels. Prompts are portable, so they carry the location of secrets ("the key lives in .env"), never the values.
  • The final check is cold. Could a session with no memory of your conversation act on this prompt without asking a single question? If not, it isn't finished.

What it deliberately won't do

It won't build the thing in your session — the prompt is the deliverable. It won't debate whether or what to build. It won't run the prompt for you. Those boundaries are what keep it excellent at the one job.

Measured, not vibes

Three test rambles run with and without the skill on the same model: 96.7% vs 80.0% across 30 graded checks — the gap concentrated exactly where prompts usually rot: missing self-checks, micromanaged methods, no stop conditions, buried goals. Twenty trigger queries against a menu of nine competing skills: 80/80 correct across four independent judges. Version 1.1 absorbed two independent reviews — 38 findings, every accepted fix re-verified in output. The eval sets ship in evals/; the full history is in CHANGELOG.md.


A complete run

An illustrative session, start to finish: a researcher screening abstracts for a systematic review.

What you say

"ok so i've got my systematic review going, i exported like 400 abstracts from pubmed into a csv and screening them by hand is killing me. i want a fresh claude session to do a first-pass screen against my criteria: adults with type 2 diabetes, RCTs only, any GLP-1 agonist vs standard care, has to report HbA1c, english, 2015 or later. i still wanna double-check the maybes myself obviously. make me the prompt"

What it checks before it asks you anything

It finds screening/pubmed_export_2026-08-21.csv, opens it, and reads the header row and record count for itself: 412 records, columns title, abstract, authors, year, doi. That's four questions you never get asked. Your six criteria came through the ramble intact, so it doesn't play them back to you either.

What it asks

Three questions, because three things in that ramble genuinely cannot be settled without you — each one changes what comes out the other end.

1. "GLP-1 agonist vs standard care" — how wide is the comparator? Most GLP-1 trials are placebo-controlled or run against another glucose-lowering drug, so a strict usual-care reading would throw away most of your evidence base.

  • ✅ Standard care, placebo, and active comparators — each one flagged per record (recommended: you can subset later without re-screening 412 abstracts)
  • Strictly usual care / treat-as-usual
  • Standard care and placebo, but not active comparators

2. Tirzepatide — in or out? It's a dual GIP/GLP-1 agonist, so "any GLP-1 agonist" is genuinely ambiguous here, and this is a protocol decision rather than a screening one.

  • ✅ In, flagged as dual-agonist (recommended: dropping a flagged subset later costs you nothing; adding it back costs a full re-screen)
  • Out — GLP-1 mono-agonists only

3. You said you'd double-check the maybes. Is that your whole second screen? PRISMA reviewers expect two independent screeners across all records, not just the uncertain ones — worth deciding now, because it changes what the output file has to support.

  • ✅ You screen all 412 independently too; the file is built so both sets can be compared and an agreement rate reported (recommended: it's what a published review is held to)
  • Maybes plus a random 10% audit of the excludes, with the limitation declared in your methods
  • Maybes only

What you get

I'm running a systematic review and you're doing the first-pass screen. Work through
`screening/pubmed_export_2026-08-21.csv` — 412 records, columns title, abstract, authors, year,
doi — and decide every one against my protocol criteria. This review is heading for
publication under PRISMA, and I'm screening the same records independently, so your decisions
get compared against mine and an agreement rate gets reported. A silently dropped eligible
study is the one unforgivable outcome here: it never resurfaces, and it is invisible in the
finished review.

So screen for sensitivity, not tidiness. At title and abstract, a wrong include costs me
ninety seconds at full text; a wrong exclude costs a missing study in a published paper. When
a record is genuinely unclear, it is a maybe — never a quiet exclude.

The criteria, exactly, and in this order — the first one a record fails is the reason it is
excluded, so every exclusion carries exactly one countable reason:

1. Population — adults with type 2 diabetes. Mixed type 1 and type 2 populations are maybes
   unless type 2 results are reported separately.
2. Design — randomized controlled trials. Protocols without results, secondary or post-hoc
   analyses of a trial, and open-label extensions are maybes, not excludes; I settle those at
   full text.
3. Intervention — any GLP-1 receptor agonist, injectable or oral. Tirzepatide counts and gets
   flagged as a dual agonist, since I may drop that subset later.
4. Comparator — standard care, placebo, or an active glucose-lowering comparator. Record which,
   so I can subset without re-screening.
5. Outcome — HbA1c reported, primary or secondary. An abstract citing "glycaemic control"
   without naming HbA1c is a maybe.
6. English language, published 2015 or later.

Judge only what the title and abstract actually say. Never infer what a full text probably
contains, and never lean on anything you know about a trial that is not in the record in front
of you.

Deliver `screening/first_pass_decisions.csv`: my original columns untouched, plus decision
(include / exclude / maybe), failed_criterion (the number above; blank for includes), reason
(one plain sentence I can audit a year from now — "no HbA1c outcome named", never "does not
fit"), comparator_type, dual_agonist_flag, and for every maybe, the_question — the single thing
I need to check at full text. How you work through 412 records is yours to decide, as is how
you phrase the reasons; consistency across records is what matters.

Then prove the screen held before you report it. Every record decided exactly once, 412 in and
412 out. Re-screen 20 random records from scratch without looking at your first decision, and
report any that flipped along with what caused it. Confirm that each exclusion's stated reason
really is the first criterion that record fails. Report the tally by decision and by failed
criterion — that is my PRISMA flow diagram — and say plainly if anything was ambiguous enough
that you would want a second opinion.

A fully screened file, every record decided once, every decision auditable, and nothing
eligible silently lost, is your goal.

Assumptions — correct any with one line:

  • Title-and-abstract screening only; no full-text retrieval and no data extraction yet.
  • Your ramble's "2015 or later" reads as publication year, matching the year column in the export.
  • Deduplication already happened in your reference manager, so all 412 records get screened.

What to notice

Three questions, not thirteen. The file, its columns, its size, and your criteria were all knowable without you. The comparator scope, the tirzepatide boundary, and the screening protocol were not — and each one silently changes the result. The recommended answer on each carries the actual reason, and two of the three recommend the reversible option, which is a judgment worth having in a research assistant.

The exclusion hierarchy is the load-bearing invention. Nothing in the ramble asked for ordered criteria. But PRISMA needs exclusion counts that don't overlap, so "first criterion that fails" turns a pile of judgments into a flow diagram — a research need the prompt anticipated and solved.

The domain shows up where it matters. Mixed populations, protocols, post-hoc analyses, extension studies, "glycaemic control" without a named measure — these are the specific ways an abstract screen goes wrong, and each is given a decision rule instead of being left to improvised judgment on record 300.

Rules come with their reasoning. "Ninety seconds at full text versus a missing study in a published paper" is why the screen leans inclusive. A session that understands the asymmetry applies it correctly to a case nobody anticipated; a session told only "be inclusive" doesn't.

Method is delegated; judgment is not. How to process 412 records is explicitly the session's call. What counts as an include, what a reason has to look like, and what proof is required are not.

Verification is real work, not a checkbox. Counts that reconcile, a blind re-screen of 20 records that would catch drift, and an audit that the stated reason is the true first failure. And the closing line invites the session to flag what it wasn't sure about — because a screener that never says "I wasn't certain here" is not being careful, it's being quiet.

What's inside

context-engineer/
├── SKILL.md                    the skill — how it reads, asks, writes, and checks
└── references/
    ├── interview.md            how it decides what to ask, and how
    ├── craft.md                how the prompts themselves are engineered
    ├── template.md             the structured format, for machine-read outputs
    └── example.md              a second worked run, in a different domain

License

MIT — see LICENSE.

About

You describe what you want; it writes the prompt another AI session needs — asking only what it can't figure out itself.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors