diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..b5544a0 --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,71 @@ +# Writing Ark apps in this repository + +This repository teaches how to write apps for the Ark. Read `README.md` first, +then `docs/01-app-model.md` through `docs/06-reports.md` in order. They are +short and they are the rules. When a real Ark is involved, `ark help agents` +comes before anything else. + +## Commands + +```sh +make build APP= # one module into build/.wasm +make run APP= # build, then both passes against fixtures/ +make run APP= FIXTURES=fixtures/no-call +make run APP= FIXTURES=fixtures/unanswered +tools/check.sh # what CI runs: links, builds, runs, ports +``` + +`make run` mounts only the paths an app's manifest declares, read-only, and +runs it with `/` as the data directory, the way an Ark does. The three fixture +roots are described in `fixtures/README.md`. An app whose grants a root lacks +stops before its run pass there, which is expected. + +## Adding an app + +- Copy the nearest example. One source file, in `apps/--/` for + a mini and `apps/-/` for a full app, with a short `README.md` that + says what it shows and how to run it. +- The manifest names the app, its version and the narrowest grants that answer + the question. Spell every path exactly as `ark data paths` shows it, with + `v1/` kept and no trailing `/`. `docs/02-manifest.md` has the rules. +- Read grants as plain files. Absent means "not found" and is a result. Any + other error goes to standard error with a non-zero exit. The shapes a + genotype can take are in `docs/04-reading-data.md`. +- Nothing in the sandbox is random or timed. Sort anything you list. +- Keep the module small. Rust apps carry the release profile from any sibling's + `Cargo.toml`. +- Run `tools/check.sh` before proposing a change. New fixtures follow + `fixtures/README.md`, with no trailing newline in any value file. + +## The report + +Follow `docs/06-reports.md`. In short: + +- One `#` title in Title Case naming the subject, then the finding straight + under it with no heading, within the first ten lines. No app name, version + or grant list; the Ark reports those beside the report. +- Then sections in this order, `Evidence`, `Method`, `Limitations`, + `Sources`, each when it has content, headings in sentence case. +- The finding is a sentence and a value. A verdict label may lead it when + the value follows at once. Say "carries" and "is associated with", never + "you have" or "you will". An absent answer is stated as a finding, never + read as reference. +- Every value beside its coordinate. Genotypes exactly as the file holds + them, in prose or code when they contain a pipe. Coordinates labelled with + the assembly only when the app granted `v1/genome/reference` and read + `build`. +- Studies named inline as author and year, listed in Sources with a stable + URL. Links nowhere else. +- Pick one voice and keep it. Light or serious, never both in one report. + Genomics words, never clinic words. +- Plain Markdown only. No images, HTML, footnotes, encoded content, dates or + run ids. Tables with as few columns as carry the evidence, rows about 80 + characters. + +## Never + +- Never read absence as homozygous reference. +- Never assume the ALT allele is the risk allele. Derive it from the study. +- Never grant more than the question needs. +- Never depend on randomness, time, the network or writable storage. +- Never put anything in a report the owner can't read. diff --git a/README.md b/README.md index 1df132a..5190c07 100644 --- a/README.md +++ b/README.md @@ -23,7 +23,7 @@ An app is a [WebAssembly](https://webassembly.org/) module built for - **With no arguments**, the app prints a short TOML manifest naming itself and the data it wants. The Ark checks it and shows it to the owner for approval. - **With a data directory as its first argument**, the app reads its granted - files and prints its report. + files and prints its report, a Markdown page the owner reads first. ```rust fn main() { @@ -88,6 +88,8 @@ and a full app turns the same read into a report. - [04-reading-data.md](docs/04-reading-data.md) covers absence, errors, sequences and genotypes, with a grant for every need. - [05-running.md](docs/05-running.md) covers running apps locally and on an Ark. +- [06-reports.md](docs/06-reports.md) covers the report, its shape and its + voice. ## A note on scope diff --git a/apps/03-cilantro-soapiness/Cargo.lock b/apps/03-cilantro-soapiness/Cargo.lock index 0bdcb31..8a26069 100644 --- a/apps/03-cilantro-soapiness/Cargo.lock +++ b/apps/03-cilantro-soapiness/Cargo.lock @@ -4,4 +4,4 @@ version = 4 [[package]] name = "cilantro-soapiness" -version = "0.3.0" +version = "0.4.0" diff --git a/apps/03-cilantro-soapiness/Cargo.toml b/apps/03-cilantro-soapiness/Cargo.toml index 08a235e..529e035 100644 --- a/apps/03-cilantro-soapiness/Cargo.toml +++ b/apps/03-cilantro-soapiness/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "cilantro-soapiness" -version = "0.3.0" +version = "0.4.0" edition = "2024" # A smaller module uploads and starts faster on an Ark. diff --git a/apps/03-cilantro-soapiness/README.md b/apps/03-cilantro-soapiness/README.md index fbacc2b..2b6a313 100644 --- a/apps/03-cilantro-soapiness/README.md +++ b/apps/03-cilantro-soapiness/README.md @@ -1,10 +1,12 @@ # 03 - cilantro-soapiness The single variant from [03-cilantro-mini-rust](../03-cilantro-mini-rust), -turned into a full report with a verdict, the copies behind it, caveats and -further reading. A no-call and an absent genotype each get their own -explanation instead of a verdict, and a call without exactly two copies shows -its copies without a homozygous or heterozygous label. +turned into a full report in the shape [06-reports.md](../../docs/06-reports.md) +describes, with a finding, the evidence behind it, the method, its limitations +and sources. A no-call and an absent genotype each get their +own finding instead of a verdict, and a call without exactly two copies shows +its copies without a homozygous or heterozygous label. The `reference` grant is +read for one file, the assembly build the coordinate is reported on. ## Build and run diff --git a/apps/03-cilantro-soapiness/src/main.rs b/apps/03-cilantro-soapiness/src/main.rs index 40d9105..58792bf 100644 --- a/apps/03-cilantro-soapiness/src/main.rs +++ b/apps/03-cilantro-soapiness/src/main.rs @@ -1,11 +1,12 @@ -//! Cilantro Soapiness: the cilantro taste test as a polished Ark report. +//! Cilantro Soapiness: the cilantro taste test as a full Ark report. //! //! This is the full version of the cilantro example. The mini version //! (../03-cilantro-mini-rust) shows the bare lens read in a handful of lines; -//! this one wraps the same single genotype in a complete Markdown report with a -//! verdict, a result table, caveats, references, and a technical-details section. -//! The data access is identical. Everything extra here is presentation, which is -//! what separates a demo from an app someone would actually want to run. +//! this one wraps the same single genotype in a complete report in the shape +//! docs/06-reports.md describes: a finding, the evidence behind it, the method, +//! its limitations and sources. The data access is identical. +//! Everything extra here is presentation, which is what separates a demo from +//! an app someone would actually want to run. //! //! Some people taste cilantro as soap. The trait tracks a SNP near the olfactory //! receptor gene OR6A2: the more copies of the C allele at rs72921001 you carry, @@ -18,58 +19,78 @@ use std::{fs, io}; const RSID: &str = "rs72921001"; /// Allele associated with tasting cilantro as soapy; the app counts copies of it. const RISK_BASE: &str = "C"; -/// dbSNP page for the SNP, linked from the report's further-reading section. -const NIH_SNP_URL: &str = "https://www.ncbi.nlm.nih.gov/snp/rs72921001"; fn main() { - // With no data directory, print the manifest the host reads to grant access. + // With no data directory, print the manifest the Ark reads to grant access. + // The reference grant is there for one file, the assembly build the + // coordinate is reported on. let Some(dir) = std::env::args().nth(1) else { print!( "[package]\n\ name = \"cilantro-soapiness\"\n\ - version = \"0.3.0\"\n\ - datasets = [\"v1/genome/rsids/rs72921001\"]\n" + version = \"0.4.0\"\n\ + datasets = [\"v1/genome/rsids/rs72921001\", \"v1/genome/reference\"]\n" ); return; }; // Everything comes from one lens directory: the user's genotype and the // variant's coordinate. The app reads plain files, never a genomic format. - let base = Path::new(&dir).join("v1/genome/rsids").join(RSID); + let genome = Path::new(&dir).join("v1/genome"); + let base = genome.join("rsids").join(RSID); + let build = leaf(&genome.join("reference"), "build"); - print_header(); - match read_call(&base) { + println!("# 🌿 Cilantro Taste Test"); + println!(); + + let reference = leaf(&base, "reference"); + let call = read_call(&base); + match &call { Call::Called(genotype) => { - let copies = genotype - .split(['/', '|']) - .filter(|a| *a == RISK_BASE) - .count(); - let ploidy = genotype + let alleles: Vec<&str> = genotype .strip_prefix(['/', '|']) - .unwrap_or(&genotype) + .unwrap_or(genotype) .split(['/', '|']) - .count(); - if ploidy == 2 { - print_result(&genotype, copies); + .collect(); + let copies = alleles.iter().filter(|a| **a == RISK_BASE).count(); + if alleles.len() == 2 { + print!("{}", finding(copies)); + if reference.as_deref() == Some(RISK_BASE) && copies == 2 { + print!( + " C is also the reference base here, so a call file that lists only \ + variants would have stayed quiet about it. Yours didn't." + ); + } + println!(); } else { - println!("### Your result: copy count\n"); println!( - "Genotype `{genotype}` has {copies} copies of `{RISK_BASE}` across {ploidy} alleles.\n" + "**No verdict.** Genotype `{genotype}` holds {copies} copies of C across \ + {} alleles, and the taste comparison is built for two-allele calls.", + alleles.len() ); - println!("The taste comparison uses two-copy SNP calls.\n"); } - print_footer(); - print_technical_details(&base); } Call::Uncertain(genotype) => { - print_uncertain(&genotype); - print_footer(); + println!( + "**Inconclusive** ⚠️. Your genotype at {RSID} is `{genotype}`, with at least \ + one allele missing, so the copies of C can't be counted." + ); } Call::Unanswered => { - print_unanswered(); - print_footer(); + println!( + "**No answer.** No genotype is available at {RSID}, so this app has no \ + result. An absent genotype is not a reference call. A call file that \ + records only variants has no record at a reference site, and none at a \ + site it didn't cover, and the two can't be told apart here." + ); } } + println!(); + + print_evidence(&base, &call, reference.as_deref(), build.as_deref()); + print_method(); + print_limitations(); + print_sources(); } /// What the lens can tell us about the user at this site. @@ -98,141 +119,99 @@ fn leaf(base: &Path, name: &str) -> Option { Ok(value) => Some(value), Err(err) if err.kind() == io::ErrorKind::NotFound => None, Err(err) => { - eprintln!("Could not read {RSID} {name}: {err}"); + eprintln!("Could not read {}: {err}", base.join(name).display()); std::process::exit(1); } } } -/// How a given risk-allele copy count is presented in the report. -struct Verdict { - emoji: &'static str, - heading: &'static str, - detail: &'static str, -} - -/// Maps the risk-allele copy count (0, 1, or 2) to its report presentation. -fn interpret(copies: usize) -> Verdict { +/// The finding for a two-allele call, by the number of C copies it holds. The +/// verdict leads, and the count it rests on follows in the same breath. +fn finding(copies: usize) -> &'static str { match copies { - 0 => Verdict { - emoji: "🌿", - heading: "Cilantro lover", - detail: "You carry **zero** copies of the risk allele at this locus. \ - Most people with this genotype perceive cilantro as herby or citrusy.", - }, - 1 => Verdict { - emoji: "🫧", - heading: "On the fence", - detail: "You carry **one** copy of the risk allele (heterozygous). \ - The effect is partial. You may notice a mild soapy note \ - or none at all, depending on other genetic and environmental factors.", - }, - _ => Verdict { - emoji: "🧼", - heading: "Soap detector", - detail: "You carry **two** copies of the risk allele (homozygous). \ - This is the genotype most strongly associated with perceiving \ - cilantro as soapy or unpleasant in published GWAS data.", - }, + 0 => { + "**Cilantro lover** 🌿. You carry no copy of C at rs72921001, the genotype with \ + no soapy association. Most people built this way taste cilantro as herby or \ + citrusy." + } + 1 => { + "**On the fence** 🫧. You carry one copy of C at rs72921001, so the soapy \ + association is partial. You may notice a mild soapy note, or none at all, \ + depending on the rest of your genome and what you grew up eating." + } + _ => { + "**Soap detector** 🧼. You carry two copies of C at rs72921001, the genotype most \ + strongly tied to tasting cilantro as dish soap. Blame *OR6A2*, the olfactory \ + receptor next door." + } } } -/// Intro: what the test measures and the gene behind it. -fn print_header() { - println!("## 🌿 Cilantro Taste Test"); +/// The values the finding rests on: the variant, its coordinate on the assembly +/// this Ark holds, and the genotype exactly as the call file records it. The +/// lens exposes chromosome and position separately; the app composes the +/// `chr:pos` form itself, and labels it only when the build could be read. +fn print_evidence(base: &Path, call: &Call, reference: Option<&str>, build: Option<&str>) { + println!("## Evidence"); println!(); - println!( - "Some people love cilantro. Others think it tastes like dish soap. \ - Blame *OR6A2*, an olfactory receptor gene on chromosome 11. \ - A 2012 GWAS (Eriksson *et al.*) pinpointed `{RSID}` as the SNP \ - behind it. More copies of the **{RISK_BASE}** allele, more soap." - ); - println!(); -} - -/// The result for a confident genotype: the copy count and its interpretation. -fn print_result(genotype: &str, copies: usize) { - let verdict = interpret(copies); - println!("### Your result: {} {}", verdict.heading, verdict.emoji); - println!(); - println!("| Genotype | Risk-allele copies |"); - println!("| :---: | :---: |"); - println!( - "| `{}` | **{copies}**Γ—{RISK_BASE} |", - genotype.replace('|', "\\|") - ); + println!("| Variant | Nearest gene | Position |"); + println!("| :-- | :-- | :-- |"); + let position = match (leaf(base, "chromosome"), leaf(base, "position"), build) { + (Some(chrom), Some(pos), Some(build)) => format!("{chrom}:{pos}, {build}"), + (Some(chrom), Some(pos), None) => format!("{chrom}:{pos}"), + _ => "unknown".to_string(), + }; + println!("| {RSID} | *OR6A2* | {position} |"); println!(); - println!("{}", verdict.detail); + let recorded = match call { + Call::Called(genotype) | Call::Uncertain(genotype) => { + format!("Your call file says `{genotype}` here") + } + Call::Unanswered => "Your call file has no record here".to_string(), + }; + match reference { + Some(reference) => println!("{recorded}, and the reference base is {reference}."), + None => println!("{recorded}."), + } println!(); } -/// The result when an allele is missing, so no copy count can be reported. -fn print_uncertain(genotype: &str) { - println!("### Your result: inconclusive ⚠️"); +/// How the finding follows from the evidence, and where the association comes from. +fn print_method() { + println!("## Method"); println!(); println!( - "Your genotype at `{RSID}` is `{genotype}` - at least one allele is \ - missing, so the risk-allele count can't be determined." + "Some people love cilantro. Others think it tastes like dish soap. Eriksson et al. \ + (2012) went looking for why, in a genome-wide study of self-reported preference \ + among European-ancestry participants, and {RSID} is what they found. The app \ + counts your copies of C. Two is the strongest association, one is somewhere in \ + between, and zero is none. More copies, more soap." ); println!(); } -/// The result when the lens has no genotype at this site. -fn print_unanswered() { - println!("### No genotype answer"); +/// What the app didn't read and what the finding doesn't establish. +fn print_limitations() { + println!("## Limitations"); println!(); println!( - "No genotype answer is available for `{RSID}`. Reference sites in a \ - variants-only file can be absent, as can missing calls or ambiguous \ - records or placements. Absence does not imply two reference alleles." + "One variant, one modest effect. The study measured what people said they \ + preferred, not what they tasted, and taste is polygenic and shaped by diet, \ + culture and exposure besides. Two copies doesn't mean you hate cilantro, and \ + zero doesn't mean you love it. The study was mostly people of European ancestry, \ + so elsewhere the effect is less well known, and this app doesn't know how common \ + your genotype is." ); println!(); } -/// Caveats and references, shown after every result. -fn print_footer() { - println!("### Fine print"); - println!(); - println!( - "One SNP doesn't tell the whole story. Taste is polygenic and shaped by \ - diet, culture, and exposure. Two copies doesn't mean you hate cilantro, \ - zero copies doesn't mean you love it." - ); +/// The studies Method names, in that order, with a stable address for each. +fn print_sources() { + println!("## Sources"); println!(); - println!("### Further reading"); - println!(); - println!( - "- Eriksson N *et al.* (2012). \"A genetic variant near olfactory receptor genes influences cilantro preference.\" *Flavour* 1:22." - ); - println!( - "- NCBI dbSNP [*{RSID}*]({NIH_SNP_URL}): population frequencies, genomic context, and submission history." - ); println!( - "- NCBI Gene [*OR6A2*](https://www.ncbi.nlm.nih.gov/gene/8590): olfactory receptor family 6 subfamily A member 2." + "1. Eriksson N, et al. A genetic variant near olfactory receptor genes influences \ + cilantro preference. Flavour. 2012;1:22. https://doi.org/10.1186/2044-7248-1-22" ); - println!(); -} - -/// The resolved locus and reference allele, read from the lens's scalar leaves. -/// The lens exposes chromosome and position separately; the app composes the -/// `chr:pos` form itself. -fn print_technical_details(base: &Path) { - let (chrom, pos, reference) = ( - leaf(base, "chromosome"), - leaf(base, "position"), - leaf(base, "reference"), - ); - - println!("### Technical details"); - println!(); - println!("| Field | Value |"); - println!("| :--- | :--- |"); - println!("| **rsID** | `{RSID}` |"); - if let (Some(chrom), Some(pos)) = (chrom, pos) { - println!("| **Locus** | `{chrom}:{pos}` |"); - } - if let Some(reference) = reference { - println!("| **Reference allele** | `{reference}` |"); - } - println!(); + println!("2. NCBI dbSNP, {RSID}. https://www.ncbi.nlm.nih.gov/snp/{RSID}"); } diff --git a/apps/04-drunk-o-type/Cargo.lock b/apps/04-drunk-o-type/Cargo.lock index 647dad6..65164e0 100644 --- a/apps/04-drunk-o-type/Cargo.lock +++ b/apps/04-drunk-o-type/Cargo.lock @@ -4,4 +4,4 @@ version = 4 [[package]] name = "drunk-o-type" -version = "0.2.0" +version = "0.3.0" diff --git a/apps/04-drunk-o-type/Cargo.toml b/apps/04-drunk-o-type/Cargo.toml index 793d079..15d10d9 100644 --- a/apps/04-drunk-o-type/Cargo.toml +++ b/apps/04-drunk-o-type/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "drunk-o-type" -version = "0.2.0" +version = "0.3.0" edition = "2024" # A smaller module uploads and starts faster on an Ark. diff --git a/apps/04-drunk-o-type/src/main.rs b/apps/04-drunk-o-type/src/main.rs index 8d97fc5..41f8f9e 100644 --- a/apps/04-drunk-o-type/src/main.rs +++ b/apps/04-drunk-o-type/src/main.rs @@ -141,7 +141,6 @@ const TARGETS: &[SnpTarget] = &[ #[derive(Debug, Clone)] struct GenotypeCall { - genotype: String, alleles: Vec, // A missing allele makes the count inconclusive, never zero. risk_copies: Option, @@ -375,14 +374,7 @@ fn combo_callout(flush_idx: usize, like_idx: usize) -> &'static str { // ── Output sections ──────────────────────────────────────────────────────── fn print_app_header() { - println!("## 🍻 Drunk-o-type"); - println!(); - println!( - "Eleven SNPs across three axes: **Flush** (does alcohol make you sick?), \ - **Like** (does your brain enjoy it?), and **Junk** (do drink additives - \ - histamine, tannins, the fermented-stuff load - wreck you?). Each axis \ - bins into 5 tiers." - ); + println!("# 🍻 How Your Body Handles a Drink"); println!(); } @@ -400,9 +392,34 @@ fn print_sample_report(results: &[(&'static SnpTarget, Option)]) { ("Inconclusive", "❓") }; + // The labels are the fun part; the counts beside them are the finding. println!( - "### Your drunk-o-type: {} {} + {} {} + {} {}", - flush_emoji, flush_label, like_emoji, like_label, junk_emoji, junk_label + "**Your drunk-o-type:** {flush_emoji} {flush_label} + {like_emoji} {like_label} + \ + {junk_emoji} {junk_label}. Flush from {} ALDH2 risk alleles and {} fast ADH1B \ + alleles, Like from {} reward alleles, Junk from {} histamine alleles. Eleven \ + variants across three axes: **Flush** (does alcohol make you sick?), **Like** \ + (does your brain enjoy it?), and **Junk** (does the fermented-stuff load wreck \ + you?), each binned into five tiers.", + copy_count( + flush_score.aldh2_risk, + flush_score.aldh2_copies, + flush_score.aldh2_observed + ), + copy_count( + flush_score.adh1b_risk, + flush_score.adh1b_copies, + flush_score.adh1b_observed == flush_score.adh1b_expected + ), + copy_count( + reward_score.risk, + reward_score.copies, + reward_score.fully_called() + ), + copy_count( + histamine_score.risk, + histamine_score.copies, + histamine_score.fully_called() + ), ); println!(); @@ -421,7 +438,7 @@ fn print_sample_report(results: &[(&'static SnpTarget, Option)]) { } // Flush axis breakdown - println!("#### 🍻 Flush axis - {} {}", flush_label, flush_emoji); + println!("## 🍻 Flush: {flush_label} {flush_emoji}"); println!(); let aldh2_marker = if flush_score.aldh2_observed { "βœ“ found" @@ -447,13 +464,10 @@ fn print_sample_report(results: &[(&'static SnpTarget, Option)]) { println!(); println!("{}", flush_blurb(&flush_score)); println!(); - print_axis_table(results, Axis::Flush); // Like axis breakdown println!( - "#### 🧠 Like axis - {} {} ({} risk Β· {}/{} markers{})", - like_label, - like_emoji, + "## 🧠 Like: {like_label} {like_emoji} ({} risk Β· {}/{} markers{})", copy_count( reward_score.risk, reward_score.copies, @@ -468,25 +482,20 @@ fn print_sample_report(results: &[(&'static SnpTarget, Option)]) { }, ); println!(); - if !reward_score.fully_called() { + if reward_score.fully_called() { + println!("{}", like_blurb(reward_score.risk)); + } else { println!( "*Incomplete calls: {} of {} reward markers have every allele \ called. The Like score is inconclusive.*", reward_score.observed, reward_score.expected ); - println!(); - } - if reward_score.fully_called() { - println!("{}", like_blurb(reward_score.risk)); } println!(); - print_axis_table(results, Axis::Reward); // Junk axis breakdown println!( - "#### 🀧 Junk axis - fermented-drink sensitivity - {} {} ({} risk Β· {}/{} markers{})", - junk_label, - junk_emoji, + "## 🀧 Junk: {junk_label} {junk_emoji} ({} risk Β· {}/{} markers{})", copy_count( histamine_score.risk, histamine_score.copies, @@ -501,19 +510,16 @@ fn print_sample_report(results: &[(&'static SnpTarget, Option)]) { }, ); println!(); - if !histamine_score.fully_called() { + if histamine_score.fully_called() { + println!("{}", junk_blurb(histamine_score.risk)); + } else { println!( "*Incomplete calls: {} of {} histamine markers have every allele \ called. The Junk score is inconclusive.*", histamine_score.observed, histamine_score.expected ); - println!(); - } - if histamine_score.fully_called() { - println!("{}", junk_blurb(histamine_score.risk)); } println!(); - print_axis_table(results, Axis::Histamine); } #[derive(Debug, Default, Clone, Copy)] @@ -605,42 +611,84 @@ fn score_sample( (flush, reward, histamine) } -fn print_axis_table(results: &[(&'static SnpTarget, Option)], axis: Axis) { - println!("| Variant | Gene | Genotype | Risk copies | Effect |"); - println!("| :--- | :--- | :--- | :---: | :--- |"); - for (target, call) in results.iter().filter(|(t, _)| t.axis == axis) { - let (gt, copies) = match call { - Some(c) => ( - format!("`{}`", c.genotype.replace('|', "\\|")), - c.risk_copies - .map_or_else(|| "inconclusive".to_string(), |n| n.to_string()), - ), - None => ("no answer".to_string(), "inconclusive".to_string()), - }; - println!( - "| `{}` | *{}* | {} | {} | {} |", - target.rsid, target.gene, gt, copies, target.short - ); - } +// One table per axis. Alleles get a column each, so a phased call prints +// exactly as the file holds it without a pipe fighting the table. +fn print_evidence(results: &[(&'static SnpTarget, Option)], build: Option<&str>) { + println!("## Evidence"); println!(); + println!( + "Resolved through the `rsids/` lens{}, without the app reading the call \ + file. A missing genotype means there is no answer; it does not imply \ + homozygous reference.", + match build { + Some(build) => format!(", with positions on {build}"), + None => String::new(), + } + ); + println!(); + for (heading, axis) in [ + ("Flush", Axis::Flush), + ("Like", Axis::Reward), + ("Junk", Axis::Histamine), + ] { + println!("### {heading}"); + println!(); + println!("| Variant | Gene | Position | Reference | Allele 1 | Allele 2 | Risk copies |"); + println!("| :-- | :-- | :-- | :-- | :-- | :-- | --: |"); + for (target, call) in results.iter().filter(|(t, _)| t.axis == axis) { + let (position, reference, first, second, copies) = match call { + Some(c) => ( + if c.chromosome.is_empty() || c.position.is_empty() { + "no answer".to_string() + } else { + format!("{}:{}", c.chromosome, c.position) + }, + if c.reference.is_empty() { + "no answer".to_string() + } else { + c.reference.clone() + }, + c.alleles.first().cloned().unwrap_or_default(), + c.alleles.get(1).cloned().unwrap_or_default(), + c.risk_copies + .map_or_else(|| "inconclusive".to_string(), |n| n.to_string()), + ), + None => ( + "no answer".to_string(), + "no answer".to_string(), + "no answer".to_string(), + String::new(), + "inconclusive".to_string(), + ), + }; + println!( + "| {} | *{}* | {} | {} | {} | {} | {} |", + target.rsid, target.gene, position, reference, first, second, copies + ); + } + println!(); + } } fn print_science_section() { - println!("### How this works"); + println!("## Method"); println!(); println!( "Alcohol metabolism is a two-step pipeline. **ADH** (alcohol dehydrogenase) \ turns ethanol into **acetaldehyde** - a toxic intermediate responsible for \ flushing, headaches, nausea, and rapid heart rate. **ALDH2** then converts \ acetaldehyde into harmless acetate. The flush axis measures how much \ - acetaldehyde you accumulate (faster ADH1B in, slower ALDH2 out β†’ more flush)." + acetaldehyde you accumulate (faster ADH1B in, slower ALDH2 out β†’ more flush). \ + ALDH2 is dominant-negative, one broken subunit poisons the whole tetramer, so \ + the Flush tier is set by the ALDH2 count first and the ADH1B count after." ); println!(); println!( "Whether you *like* drinking is a different question. The **mu-opioid \ receptor** (OPRM1) and **dopamine system** (DRD2/ANKK1) determine how \ rewarding alcohol feels. **GABA-A** receptors govern its anxiolytic \ - effect - the chill. The like axis stacks these three reward pathways." + effect - the chill. The like axis adds up the risk alleles across these \ + three reward pathways." ); println!(); println!( @@ -649,114 +697,81 @@ fn print_science_section() { and your body clears it via two enzymes: **DAO** (gut/blood, encoded \ by AOC1) and **HNMT** (CNS/airways). Variants in either pathway slow \ clearance, and histamine then drives the wine headache, the stuffy \ - nose, the racing heart, the post-drink itch. The junk axis stacks DAO \ - and HNMT loss-of-function variants. *Sulfites and tyramine don't have \ + nose, the racing heart, the post-drink itch. The junk axis adds up DAO \ + and HNMT loss-of-function alleles. *Sulfites and tyramine don't have \ clean common SNPs* - those phenotypes exist but the panel-level \ genetics isn't there yet." ); println!(); + println!( + "Each variant counts the allele the studies tied to the effect, which is not \ + always the ALT allele. What each one does:" + ); + println!(); + for target in TARGETS { + println!( + "- **{}** ({}, *{}*), {}. {}", + target.rsid, + match target.axis { + Axis::Flush => "Flush", + Axis::Reward => "Like", + Axis::Histamine => "Junk", + }, + target.gene, + target.short, + target.blurb + ); + } + println!(); } fn print_fine_print() { - println!("### Fine print"); + println!("## Limitations"); println!(); println!( "Eleven SNPs is a long way from the whole story. Body weight, gut \ microbiome, sleep, food in your stomach, history of drinking, sulfite \ and tyramine and tannin loads, hormones, and dozens of other genes \ - all shape your response. This is a toy demo, not medical advice. \ + all shape your response. The GABRA2 direction is debated in the \ + literature, and the DAO effects are modest with mixed replication. \ Don't drink your way around your genotype." ); println!(); } fn print_further_reading() { - println!("### Further reading"); + println!("## Sources"); println!(); println!( - "- Brooks PJ *et al.* (2009). \"The Alcohol Flushing Response: An \ - Unrecognized Risk Factor for Esophageal Cancer.\" *PLoS Med* 6(3):e50." + "1. Brooks PJ, et al. The alcohol flushing response: an unrecognized risk factor \ + for esophageal cancer from alcohol consumption. PLoS Medicine. 2009;6(3):e1000050. \ + https://doi.org/10.1371/journal.pmed.1000050" ); println!( - "- Edenberg HJ (2007). \"The Genetics of Alcohol Metabolism.\" \ - *Alcohol Research & Health* 30(1):5-13." + "2. Edenberg HJ. The genetics of alcohol metabolism: role of alcohol \ + dehydrogenase and aldehyde dehydrogenase variants. Alcohol Research & Health. \ + 2007;30(1):5-13. https://pubmed.ncbi.nlm.nih.gov/17718394/" ); println!( - "- Ray LA, Barr CS, Blendy JA, Oslin D, Goldman D, Anton RF (2012). \ - \"The role of the OPRM1 gene in alcohol use disorder and treatment \ - response.\" *Addiction Biology* 17(3):525-540." + "3. Ray LA, et al. The role of the OPRM1 gene in alcohol use disorder and \ + treatment response. Addiction Biology. 2012;17(3):525-540." ); println!( - "- Edenberg HJ *et al.* (2004). \"Variations in GABRA2, encoding the Ξ±2 \ - subunit of the GABA-A receptor, are associated with alcohol \ - dependence and with brain oscillations.\" *Am J Hum Genet* 74(4):705-714." + "4. Edenberg HJ, et al. Variations in GABRA2, encoding the Ξ±2 subunit of the \ + GABA-A receptor, are associated with alcohol dependence and with brain \ + oscillations. American Journal of Human Genetics. 2004;74(4):705-714. \ + https://doi.org/10.1086/383283" ); println!( - "- Maintz L, Yu CF, RodrΓ­guez E, *et al.* (2011). \"Association of \ - single nucleotide polymorphisms in the diamine oxidase gene with \ - diamine oxidase serum activities.\" *Allergy* 66(7):893-902." + "5. Maintz L, et al. Association of single nucleotide polymorphisms in the \ + diamine oxidase gene with diamine oxidase serum activities. Allergy. \ + 2011;66(7):893-902." ); println!( - "- Hrubisko M *et al.* (2021). \"Histamine intolerance - the more we know, \ - the less we know. A review.\" *Nutrients* 13(7):2228." - ); - println!(); -} - -fn print_technical_details(results: &[(&'static SnpTarget, Option)]) { - println!("### Technical details"); - println!(); - println!( - "Resolved through the `rsids/` lens on the Ark's reference assembly, \ - without the app reading the variant file. A missing genotype means \ - there is no answer; it does not imply homozygous reference." + "6. Hrubisko M, et al. Histamine intolerance, the more we know the less we know. \ + A review. Nutrients. 2021;13(7):2228. https://doi.org/10.3390/nu13072228" ); println!(); - println!("| rsID | Gene | Locus | Reference | Status |"); - println!("| :--- | :--- | :--- | :--- | :--- |"); - for (target, call) in results { - let (locus, reference, status) = match call { - Some(c) => ( - if c.chromosome.is_empty() || c.position.is_empty() { - "no answer".to_string() - } else { - format!("`{}:{}`", c.chromosome, c.position) - }, - if c.reference.is_empty() { - "no answer".to_string() - } else { - format!("`{}`", c.reference) - }, - if c.risk_copies.is_some() { - "βœ“ found" - } else { - "⚠️ missing allele" - }, - ), - None => ("-".to_string(), "-".to_string(), "⚠️ no genotype answer"), - }; - println!( - "| `{}` | *{}* | {} | {} | {} |", - target.rsid, target.gene, locus, reference, status - ); - } - println!(); - println!("#### What each SNP does"); - println!(); - for target in TARGETS { - println!( - "- **`{}`** ({}, *{}*) - {}", - target.rsid, - match target.axis { - Axis::Flush => "Flush", - Axis::Reward => "Like", - Axis::Histamine => "Junk", - }, - target.gene, - target.blurb - ); - } - println!(); } // ── Entry point ──────────────────────────────────────────────────────────── @@ -764,7 +779,7 @@ fn print_technical_details(results: &[(&'static SnpTarget, Option) const METADATA: &str = "\ [package] name = \"drunk-o-type\" -version = \"0.2.0\" +version = \"0.3.0\" datasets = [ \"v1/genome/rsids/rs671\", \"v1/genome/rsids/rs1229984\", @@ -777,6 +792,7 @@ datasets = [ \"v1/genome/rsids/rs1049793\", \"v1/genome/rsids/rs2052129\", \"v1/genome/rsids/rs11558538\", + \"v1/genome/reference\", ] "; @@ -812,7 +828,6 @@ fn resolve_target(base: &Path, target: &SnpTarget) -> Result Result<(), Box> { let dir = std::env::args().nth(1).unwrap(); let rsids = Path::new(&dir).join("v1/genome/rsids"); + // The reference grant is read for one file, the assembly the positions are on. + let build = read_leaf(&Path::new(&dir).join("v1/genome/reference"), "build")?; let results: Vec<(&'static SnpTarget, Option)> = TARGETS .iter() @@ -837,10 +854,10 @@ fn run() -> Result<(), Box> { print_app_header(); print_sample_report(&results); + print_evidence(&results, build.as_deref()); print_science_section(); print_fine_print(); print_further_reading(); - print_technical_details(&results); Ok(()) } diff --git a/apps/05-bitter-meter/Cargo.lock b/apps/05-bitter-meter/Cargo.lock index 6338f59..d028aad 100644 --- a/apps/05-bitter-meter/Cargo.lock +++ b/apps/05-bitter-meter/Cargo.lock @@ -4,4 +4,4 @@ version = 4 [[package]] name = "bitter-meter" -version = "0.1.0" +version = "0.2.0" diff --git a/apps/05-bitter-meter/Cargo.toml b/apps/05-bitter-meter/Cargo.toml index a224f66..ccf1b71 100644 --- a/apps/05-bitter-meter/Cargo.toml +++ b/apps/05-bitter-meter/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "bitter-meter" -version = "0.1.0" +version = "0.2.0" edition = "2024" # A smaller module uploads and starts faster on an Ark. diff --git a/apps/05-bitter-meter/src/main.rs b/apps/05-bitter-meter/src/main.rs index 638b5dc..0429ffb 100644 --- a/apps/05-bitter-meter/src/main.rs +++ b/apps/05-bitter-meter/src/main.rs @@ -1,14 +1,14 @@ -//! Bitter Meter: a small Markdown report built from the genome `genes/` lens. +//! Bitter Meter: a report built from the genome `genes/` lens. //! //! TAS2R38 is a bitter-taste receptor made famous by PTC/PROP tasting demos. This //! app deliberately keeps the claim narrower: it does not predict whether someone //! likes bitter greens or tastes a lab strip. It shows what a gene-scoped lens can //! reveal: gene metadata, reference sequence summary, and the non-reference changes -//! carried inside that gene. +//! carried inside that gene, in the report shape docs/06-reports.md describes. //! -//! Tutorial note for app authors: the app asks Ark for exactly one dataset path -//! (`v1/genome/genes/TAS2R38`). The user-facing report below avoids teaching -//! path details; those details belong here and in the README. +//! Tutorial note for app authors: the app asks Ark for the gene's path +//! (`v1/genome/genes/TAS2R38`) and for `v1/genome/reference`, which it reads only +//! for the assembly build its coordinates are reported on. use std::collections::BTreeSet; use std::fs::{self, File}; @@ -25,18 +25,31 @@ fn main() { print!( "[package]\n\ name = \"bitter-meter\"\n\ - version = \"0.1.0\"\n\ - datasets = [\"v1/genome/genes/TAS2R38\"]\n" + version = \"0.2.0\"\n\ + datasets = [\"v1/genome/genes/TAS2R38\", \"v1/genome/reference\"]\n" ); return; }; let base = Path::new(&dir).join(LENS_PATH); + let build = match read_leaf(&Path::new(&dir).join("v1/genome/reference"), "build") { + Ok(build) => build, + Err(err) => { + eprintln!("{err}"); + std::process::exit(1); + } + }; print_header(); match GeneReport::load(&base) { - Ok(Some(report)) => report.print(), - Ok(None) => println!("No metadata answer for {GENE}."), + Ok(Some(report)) => report.print(build.as_deref()), + Ok(None) => { + println!( + "No answer. The annotations on this Ark carry no coordinates for *{GENE}*, \ + so the app has nothing to describe." + ); + println!(); + } Err(err) => { eprintln!("{err}"); std::process::exit(1); @@ -98,89 +111,110 @@ impl GeneReport { })) } - fn print(&self) { - let Some(changes) = &self.changes else { - println!("No changes answer is available for this span."); - println!( - "Reference length: {} bp; GC: {:.1}%.", - self.sequence.len, - self.sequence.gc_percent() - ); - return; - }; + fn print(&self, build: Option<&str>) { let span = self.end.saturating_sub(self.start).saturating_add(1); let gc = self.sequence.gc_percent(); - let density = per_kb(changes.len(), span); - let heterozygous = changes.iter().filter(|c| c.call == "heterozygous").count(); - let homozygous = changes - .iter() - .filter(|c| c.call == "homozygous change") - .count(); - let uncertain = changes.iter().filter(|c| c.call == "uncertain").count(); - let indels = changes.iter().filter(|c| c.kind == "indel").count(); - - println!("## Your Result"); - println!(); - println!("| Signal | Value |"); - println!("| :--- | ---: |"); - println!("| Gene span | {} bp |", format_int(span)); - println!("| Reference GC | {:.1}% |", gc); - println!("| Observed changes | {} |", changes.len()); - println!("| Change density | {:.2} / kb |", density); - println!("| Heterozygous calls | {heterozygous} |"); - println!("| Homozygous change calls | {homozygous} |"); - println!("| Indels or complex alleles | {indels} |"); - println!("| Uncertain calls | {uncertain} |"); - println!( - "| Positions without an answer | {} |", - changes.iter().filter(|c| c.call == "no answer").count() - ); + + // The finding states the counts in words; the tables beneath carry + // every value they rest on. + match &self.changes { + Some(changes) if changes.is_empty() => { + println!( + "Your call file lists no changes inside *{GENE}* across {} bp on {}. A \ + postcard of the gene as recorded, not a verdict on how you taste.", + format_int(span), + self.chromosome + ); + } + Some(changes) => { + let density = per_kb(changes.len(), span); + let heterozygous = changes.iter().filter(|c| c.call == "heterozygous").count(); + let homozygous = changes + .iter() + .filter(|c| c.call == "homozygous change") + .count(); + let uncertain = changes.iter().filter(|c| c.call == "uncertain").count(); + let unanswered = changes.iter().filter(|c| c.call == "no answer").count(); + let indels = changes.iter().filter(|c| c.kind == "indel").count(); + print!( + "Your bitter receptor gene *{GENE}* carries {} {} in your call file, \ + {heterozygous} heterozygous and {homozygous} homozygous, across {} bp on {}", + changes.len(), + if changes.len() == 1 { + "change" + } else { + "changes" + }, + format_int(span), + self.chromosome + ); + if indels > 0 { + print!( + ", {indels} of them {}", + plural(indels, "an indel", "indels") + ); + } + if uncertain + unanswered > 0 { + print!(", with {uncertain} uncertain and {unanswered} without an answer"); + } + println!( + ". That's {density:.2} per kb, {} for a gene this small. A postcard of \ + the gene as your call file records it, not a verdict on how you taste.", + density_label(density) + ); + } + None => { + println!( + "No changes answer is available for *{GENE}*, so the app can't say what \ + your call file holds there. The reference sequence is {} bp with {gc:.1}% GC.", + format_int(self.sequence.len) + ); + } + } println!(); - println!("## Gene Postcard"); + println!("## Evidence"); + println!(); + println!("### Gene postcard"); println!(); println!("| Field | Value |"); - println!("| :--- | :--- |"); - println!("| Gene | `{GENE}` |"); - println!( - "| Location | `{}:{}-{}` |", - self.chromosome, self.start, self.end - ); - println!("| Strand | `{}` |", self.strand); - println!("| Biotype | `{}` |", self.biotype); + println!("| :-- | :-- |"); + println!("| Gene | *{GENE}* |"); + let location = format!("{}:{}-{}", self.chromosome, self.start, self.end); + match build { + Some(build) => println!("| Location | {location}, {build} |"), + None => println!("| Location | {location} |"), + } + println!("| Strand | {} |", self.strand); + println!("| Biotype | {} |", self.biotype); println!( "| Reference length | {} bp |", format_int(self.sequence.len) ); - println!("| Local label | {} |", density_label(density)); + println!("| Reference GC | {gc:.1}% |"); println!(); - println!("## Change Ruler"); - println!(); - println!( - "The ruler marks observed non-reference changes across the gene. `+` means more than one change landed in the same text column." - ); - println!(); - print_ruler(self.start, self.end, changes); - println!(); + if let Some(changes) = &self.changes { + println!("### Change ruler"); + println!(); + println!( + "Each `*` marks a non-reference change along the gene, and `+` more than one \ + in the same column." + ); + println!(); + print_ruler(self.start, self.end, changes); + println!(); + print_change_table(changes); + } - print_change_table(changes); - print_data_receipt(); - print_meaning(); - print_fine_print(); - print_references(); + print_method(); + print_limitations(); + print_sources(); } } fn print_header() { - println!("# Bitter Meter: {GENE}"); - println!(); - println!( - "Some bitter flavors are sensed by tiny receptor proteins on the tongue. \ - `{GENE}` is the famous bitter-taste receptor from PTC and PROP classroom \ - genetics. This report keeps the claim modest: it summarizes the visible \ - changes in this gene without trying to predict your taste preferences." - ); + println!("# TAS2R38, the Bitter Taste Receptor Gene"); println!(); } @@ -404,13 +438,11 @@ fn print_ruler(start: u64, end: u64, changes: &[Change]) { } fn print_change_table(changes: &[Change]) { - println!("## Change Receipt"); - println!(); if changes.is_empty() { - println!("No non-reference changes were listed under `changes/` for this gene."); - println!(); return; } + println!("### Change receipt"); + println!(); println!("| Position | Reference | Genotype | Call | Kind |"); println!("| ---: | :---: | :---: | :--- | :--- |"); @@ -433,62 +465,56 @@ fn print_change_table(changes: &[Change]) { println!(); } -fn print_data_receipt() { - println!("## What Was Checked"); - println!(); - println!("This report checked:"); - println!(); - println!("- `{GENE}` chromosome, start, end, strand, and biotype"); - println!("- The reference sequence for this gene, summarized as length and GC content"); - println!("- Observed non-reference changes inside this gene, with reference and genotype"); - println!(); -} - -fn print_meaning() { - println!("## What It Means"); +fn print_method() { + println!("## Method"); println!(); println!( - "This is a gene postcard. It tells you where `{GENE}` sits, how large \ - its reference sequence is, and which observed non-reference calls appear \ - inside that span. The change count is useful for seeing how much personal \ - variation is visible in this small taste-receptor gene." - ); - println!(); - println!( - "It is intentionally not a PTC or PROP taster prediction. The classic \ - bitter-taste story involves specific coding variants and haplotypes, and \ - this report does not assume phase or infer named rsIDs from the gene view." + "Some bitter flavors are sensed by tiny receptor proteins on the tongue, and \ + *{GENE}* is the famous one from PTC and PROP classroom genetics. The app reads the \ + gene's coordinates, strand and biotype from the annotations, \ + streams its reference sequence for length and GC content, and lists every position \ + under the gene's changes where a record starting inside the gene holds an ALT \ + allele. Each genotype is classified against the reference base as heterozygous, \ + homozygous or uncertain, and as a substitution or an indel. A position listed \ + without leaves, where several records start, counts as no answer. Kim et al. \ + (2003) mapped taste sensitivity to phenylthiocarbamide to this gene; that story \ + rests on specific coding variants and haplotypes the app doesn't call." ); println!(); } -fn print_fine_print() { - println!("## Fine Print"); +fn print_limitations() { + println!("## Limitations"); println!(); println!( - "The change list contains record starts within this span whose genotype \ - holds an ALT allele. Records starting before the span are excluded. A position not shown here should not be \ - treated as a standalone trait result. Taste is also shaped by other \ - genes, age, exposure, diet, and plain preference." + "Not a PTC or PROP taster test. The classic bitter-taste story rests on specific \ + coding variants and haplotypes, and this app doesn't phase calls or name rsIDs \ + from the gene view. The change list holds records starting inside the span, so \ + one starting just before it is left out, and a position not listed isn't a \ + result on its own. Taste is also shaped by other genes, age, exposure, diet and \ + plain preference." ); println!(); - println!("This report is for demo and education only."); - println!(); } -fn print_references() { - println!("## Further Reading"); +fn print_sources() { + println!("## Sources"); println!(); - println!("- NCBI Gene search for `TAS2R38`: https://www.ncbi.nlm.nih.gov/gene/?term=TAS2R38"); - println!( - "- Ensembl gene summary for `TAS2R38`: https://www.ensembl.org/Homo_sapiens/Gene/Summary?g=TAS2R38" - ); println!( - "- Kim U et al. (2003). Positional cloning of the human PTC taste-sensitivity locus. PubMed: https://pubmed.ncbi.nlm.nih.gov/12595690/" + "1. Kim UK, et al. Positional cloning of the human quantitative trait locus \ + underlying taste sensitivity to phenylthiocarbamide. Science. 2003;299:1221-1225. \ + https://pubmed.ncbi.nlm.nih.gov/12595690/" ); + println!("2. NCBI Gene, {GENE}. https://www.ncbi.nlm.nih.gov/gene/?term={GENE}"); + println!("3. Ensembl, {GENE}. https://www.ensembl.org/Homo_sapiens/Gene/Summary?g={GENE}"); println!(); } +/// Picks the singular or plural form for a count. +fn plural(n: usize, one: &'static str, many: &'static str) -> &'static str { + if n == 1 { one } else { many } +} + fn format_int(n: u64) -> String { let s = n.to_string(); let mut out = String::new(); diff --git a/apps/06-powerhouse-of-the-cell/Cargo.lock b/apps/06-powerhouse-of-the-cell/Cargo.lock index b3c676f..26347a1 100644 --- a/apps/06-powerhouse-of-the-cell/Cargo.lock +++ b/apps/06-powerhouse-of-the-cell/Cargo.lock @@ -4,4 +4,4 @@ version = 4 [[package]] name = "powerhouse-of-the-cell" -version = "0.1.0" +version = "0.2.0" diff --git a/apps/06-powerhouse-of-the-cell/Cargo.toml b/apps/06-powerhouse-of-the-cell/Cargo.toml index 916464f..437d3bf 100644 --- a/apps/06-powerhouse-of-the-cell/Cargo.toml +++ b/apps/06-powerhouse-of-the-cell/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "powerhouse-of-the-cell" -version = "0.1.0" +version = "0.2.0" edition = "2024" # A smaller module uploads and starts faster on an Ark. diff --git a/apps/06-powerhouse-of-the-cell/src/main.rs b/apps/06-powerhouse-of-the-cell/src/main.rs index 5e1b452..990e694 100644 --- a/apps/06-powerhouse-of-the-cell/src/main.rs +++ b/apps/06-powerhouse-of-the-cell/src/main.rs @@ -1,12 +1,13 @@ -//! Powerhouse of the Cell: a Markdown report from one genome `regions/` lens. +//! Powerhouse of the Cell: a report from one genome `regions/` lens. //! //! The mitochondrial chromosome is small enough to read as one range, which makes //! it a useful demo of the region lens: one granted interval exposes a streamed -//! reference sequence plus all non-reference changes inside that span. +//! reference sequence plus all non-reference changes inside that span, printed in +//! the report shape docs/06-reports.md describes. //! -//! Tutorial note for app authors: the app asks Ark for exactly one dataset path -//! (`v1/genome/regions/chrM/1-16569`). The user-facing report below avoids -//! teaching path details; those details belong here and in the README. +//! Tutorial note for app authors: the app asks Ark for the interval's path +//! (`v1/genome/regions/chrM/1-16569`) and for `v1/genome/reference`, which it +//! reads only for the assembly build its coordinates are reported on. use std::collections::BTreeSet; use std::fs::{self, File}; @@ -25,17 +26,24 @@ fn main() { print!( "[package]\n\ name = \"powerhouse-of-the-cell\"\n\ - version = \"0.1.0\"\n\ - datasets = [\"v1/genome/regions/chrM/1-16569\"]\n" + version = \"0.2.0\"\n\ + datasets = [\"v1/genome/regions/chrM/1-16569\", \"v1/genome/reference\"]\n" ); return; }; let base = Path::new(&dir).join(LENS_PATH); + let build = match read_leaf(&Path::new(&dir).join("v1/genome/reference"), "build") { + Ok(build) => build, + Err(err) => { + eprintln!("{err}"); + std::process::exit(1); + } + }; print_header(); match RegionReport::load(&base) { - Ok(report) => report.print(), + Ok(report) => report.print(build.as_deref()), Err(err) => { eprintln!("{err}"); std::process::exit(1); @@ -69,85 +77,110 @@ impl RegionReport { Ok(Self { sequence, changes }) } - fn print(&self) { - let Some(changes) = &self.changes else { - println!("No changes answer is available for this span."); - println!( - "Reference length: {} bp; GC: {:.1}%.", - self.sequence.len, - self.sequence.gc_percent() - ); - return; - }; + fn print(&self, build: Option<&str>) { let span = END - START + 1; let gc = self.sequence.gc_percent(); - let density = per_kb(changes.len(), span); - let mixed = changes.iter().filter(|c| c.call == "mixed alleles").count(); - let fixed = changes - .iter() - .filter(|c| c.call == "single change" || c.call == "all copies changed") - .count(); - let uncertain = changes.iter().filter(|c| c.call == "uncertain").count(); - let indels = changes.iter().filter(|c| c.kind == "indel").count(); - let transitions = changes - .iter() - .filter(|c| c.substitution == "transition") - .count(); - let transversions = changes - .iter() - .filter(|c| c.substitution == "transversion") - .count(); - println!("## Your Result"); + // The finding states the counts in words; the map and tables beneath + // carry every value they rest on. + match &self.changes { + Some(changes) if changes.is_empty() => { + println!( + "Your call file lists no changes across the {} bp mitochondrial \ + chromosome. A map of where your calls differ from the reference, not a \ + haplogroup.", + format_int(span) + ); + } + Some(changes) => { + let density = per_kb(changes.len(), span); + let mixed = changes.iter().filter(|c| c.call == "mixed alleles").count(); + let fixed = changes + .iter() + .filter(|c| c.call == "single change" || c.call == "all copies changed") + .count(); + let uncertain = changes.iter().filter(|c| c.call == "uncertain").count(); + let unanswered = changes.iter().filter(|c| c.call == "no answer").count(); + let indels = changes.iter().filter(|c| c.kind == "indel").count(); + let transitions = changes + .iter() + .filter(|c| c.substitution == "transition") + .count(); + let transversions = changes + .iter() + .filter(|c| c.substitution == "transversion") + .count(); + print!( + "Your mitochondrial chromosome, all {2} bp of it, carries {0} {1} in your \ + call file, {transitions} {3}, {transversions} {4} and {indels} {5}", + changes.len(), + plural(changes.len(), "change", "changes"), + format_int(span), + plural(transitions, "transition", "transitions"), + plural(transversions, "transversion", "transversions"), + plural(indels, "indel", "indels") + ); + if mixed > 0 { + print!(", {mixed} with mixed alleles and {fixed} in every copy"); + } + if uncertain + unanswered > 0 { + print!(", with {uncertain} uncertain and {unanswered} without an answer"); + } + println!( + ". That's {density:.2} per kb. A map of where your calls differ from the \ + reference, not a haplogroup." + ); + } + None => { + println!( + "No changes answer is available for this interval, so the app can't say \ + what your call file holds there. The reference sequence is {} bp with \ + {gc:.1}% GC.", + format_int(self.sequence.len) + ); + } + } println!(); - println!("| Signal | Value |"); - println!("| :--- | ---: |"); - println!("| Interval | {}:{}-{} |", CHROM, START, END); + + println!("## Evidence"); + println!(); + println!("### Interval"); + println!(); + println!("| Field | Value |"); + println!("| :-- | :-- |"); + match build { + Some(build) => println!("| Interval | {CHROM}:{START}-{END}, {build} |"), + None => println!("| Interval | {CHROM}:{START}-{END} |"), + } println!( "| Reference bases read | {} bp |", format_int(self.sequence.len) ); - println!("| Reference GC | {:.1}% |", gc); - println!("| Observed changes | {} |", changes.len()); - println!("| Change density | {:.2} / kb |", density); - println!("| Mixed-allele calls | {mixed} |"); - println!("| Fixed or single-copy changes | {fixed} |"); - println!("| Indels or complex alleles | {indels} |"); - println!("| Transitions | {transitions} |"); - println!("| Transversions | {transversions} |"); - println!("| Uncertain calls | {uncertain} |"); - println!( - "| Positions without an answer | {} |", - changes.iter().filter(|c| c.call == "no answer").count() - ); + println!("| Reference GC | {gc:.1}% |"); println!(); - println!("## Variant Compass"); - println!(); - println!( - "Each `*` marks at least one non-reference change. `+` means multiple changes share a text column." - ); - println!(); - print_star_map(changes); - println!(); + if let Some(changes) = &self.changes { + println!("### Variant compass"); + println!(); + println!( + "Each `*` marks a non-reference change along the chromosome, and `+` more \ + than one in the same column." + ); + println!(); + print_star_map(changes); + println!(); + print_windows(changes); + print_change_table(changes); + } - print_windows(changes); - print_change_table(changes); - print_data_receipt(); - print_meaning(); - print_fine_print(); - print_references(); + print_method(); + print_limitations(); + print_sources(); } } fn print_header() { - println!("# Powerhouse of the Cell"); - println!(); - println!( - "Your mitochondrial chromosome is compact enough to fit on a single \ - report page. This map summarizes observed non-reference changes across \ - `chrM` and shows where they fall along the 16.6 kb sequence." - ); + println!("# Your Mitochondrial Genome"); println!(); } @@ -401,7 +434,7 @@ fn print_star_map(changes: &[Change]) { } fn print_windows(changes: &[Change]) { - println!("## Density Windows"); + println!("### Density windows"); println!(); println!("| Window | Changes | Label | Bar |"); println!("| :--- | ---: | :--- | :--- |"); @@ -446,13 +479,11 @@ fn bar(count: usize) -> String { } fn print_change_table(changes: &[Change]) { - println!("## Change Receipt"); - println!(); if changes.is_empty() { - println!("No non-reference changes were listed under this mitochondrial interval."); - println!(); return; } + println!("### Change receipt"); + println!(); println!("| Position | Reference | Genotype | Call | Kind | Substitution |"); println!("| ---: | :---: | :---: | :--- | :--- | :--- |"); @@ -476,62 +507,56 @@ fn print_change_table(changes: &[Change]) { println!(); } -fn print_data_receipt() { - println!("## What Was Checked"); - println!(); - println!("This report checked:"); - println!(); - println!("- The mitochondrial interval `{CHROM}:{START}-{END}`"); - println!("- The reference sequence for that interval, summarized as length and GC content"); - println!("- Observed non-reference changes inside the interval, with reference and genotype"); - println!(); -} - -fn print_meaning() { - println!("## What It Means"); +fn print_method() { + println!("## Method"); println!(); println!( - "This is a mitochondrial shape report. It shows how many observed \ - non-reference calls appear across `chrM`, where they sit, and whether \ - the calls look like substitutions or indels." - ); - println!(); - println!( - "It is not a haplogroup caller. Haplogroups need curated marker trees, \ - careful handling of build and strand conventions, and usually more \ - interpretation than this demo should do." + "Your mitochondrial chromosome is compact enough to fit on a single report page, so \ + the app reads it as one interval. It streams the reference sequence for length and \ + GC content, and lists every position under its changes where a record starting \ + inside the \ + interval holds an ALT allele. Each genotype is classified against the reference \ + base by how many copies carry the change, as a substitution or an indel, and as \ + a transition or a transversion. A position listed without leaves, where several \ + records start, counts as no answer. The reference is the revised Cambridge \ + sequence of Andrews et al. (1999), which GRCh38 carries as chrM." ); println!(); } -fn print_fine_print() { - println!("## Fine Print"); +fn print_limitations() { + println!("## Limitations"); println!(); println!( - "The change list contains record starts within this span whose genotype \ - holds an ALT allele. Records starting before the span are excluded. The absence of a listed change is not a \ - medical or ancestry result. Mitochondrial data can also involve \ - heteroplasmy and platform-specific calling choices that a simple genotype \ - string does not fully describe." + "A map, not a haplogroup call. Haplogroups need curated marker trees and care with \ + build and strand conventions, none of which this app has. The change list holds \ + records starting inside the interval, and a position not listed isn't a result on \ + its own. Mitochondria can also carry several versions at once, heteroplasmy, which \ + a single genotype string doesn't capture." ); println!(); - println!("This report is for demo and education only."); - println!(); } -fn print_references() { - println!("## Further Reading"); +fn print_sources() { + println!("## Sources"); println!(); - println!("- NCBI Nucleotide `NC_012920.1`: https://www.ncbi.nlm.nih.gov/nuccore/NC_012920.1"); println!( - "- MITOMAP human mitochondrial sequence resources: https://www.mitomap.org/MITOMAP/HumanMitoSeq" + "1. Andrews RM, et al. Reanalysis and revision of the Cambridge reference sequence \ + for human mitochondrial DNA. Nature Genetics. 1999;23:147. \ + https://pubmed.ncbi.nlm.nih.gov/10508508/" ); + println!("2. NCBI Nucleotide, NC_012920.1. https://www.ncbi.nlm.nih.gov/nuccore/NC_012920.1"); println!( - "- Andrews RM et al. (1999). Reanalysis and revision of the Cambridge reference sequence. PubMed: https://pubmed.ncbi.nlm.nih.gov/10508508/" + "3. MITOMAP, the human mitochondrial genome. https://www.mitomap.org/MITOMAP/HumanMitoSeq" ); println!(); } +/// Picks the singular or plural form for a count. +fn plural(n: usize, one: &'static str, many: &'static str) -> &'static str { + if n == 1 { one } else { many } +} + fn format_int(n: u64) -> String { let s = n.to_string(); let mut out = String::new(); diff --git a/apps/07-vcf-roll-call/Cargo.lock b/apps/07-vcf-roll-call/Cargo.lock index 1a7dfd5..b061b75 100644 --- a/apps/07-vcf-roll-call/Cargo.lock +++ b/apps/07-vcf-roll-call/Cargo.lock @@ -204,7 +204,7 @@ checksum = "e6e4313cd5fcd3dad5cafa179702e2b244f760991f45397d14d4ebf38247da75" [[package]] name = "vcf-roll-call" -version = "0.1.0" +version = "0.2.0" dependencies = [ "noodles-vcf", ] diff --git a/apps/07-vcf-roll-call/Cargo.toml b/apps/07-vcf-roll-call/Cargo.toml index b50827c..8c45fee 100644 --- a/apps/07-vcf-roll-call/Cargo.toml +++ b/apps/07-vcf-roll-call/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "vcf-roll-call" -version = "0.1.0" +version = "0.2.0" edition = "2024" [dependencies] diff --git a/apps/07-vcf-roll-call/src/main.rs b/apps/07-vcf-roll-call/src/main.rs index a74ee8d..445da98 100644 --- a/apps/07-vcf-roll-call/src/main.rs +++ b/apps/07-vcf-roll-call/src/main.rs @@ -1,4 +1,5 @@ -//! VCF Roll Call: a full-VCF Markdown sanity report. +//! VCF Roll Call: a full-VCF sanity report, in the shape docs/06-reports.md +//! describes. //! //! Tutorial note for app authors: this is deliberately a broad-permission demo. //! It asks Ark for `v1/genome/snp-indel` and streams `vcf` with `noodles-vcf`. @@ -28,7 +29,7 @@ fn main() -> Result<(), Box> { print!( "[package]\n\ name = \"vcf-roll-call\"\n\ - version = \"0.1.0\"\n\ + version = \"0.2.0\"\n\ datasets = [\"v1/genome/snp-indel\"]\n" ); return Ok(()); @@ -177,41 +178,32 @@ enum GenotypeClass { impl Report { fn print(&self) { - println!("## Your Result"); - println!(); - println!("| Signal | Value |"); - println!("| :--- | ---: |"); - println!("| VCF version | `{}` |", self.file_format); - println!("| Samples | {} |", self.samples); - println!("| Records scanned | {} |", format_int(self.records)); + // The finding says what the file is and how it reads, in words, and + // the check table beneath it grades the same facts. let (primary_count, other_count) = contig_counts(&self.chromosomes); - println!("| Primary chromosomes seen | {primary_count} |"); - println!("| Other contigs seen | {other_count} |"); - println!( - "| Variant records | {} |", - format_int(self.variant_records()) - ); - println!( - "| All-reference records | {} |", - format_int(self.genotype.hom_ref) - ); - println!( - "| Reference-block span | {} bp |", - format_int(self.reference_blocks.bases) - ); - println!("| Median DP | {} |", optional_metric(self.depth.median())); - println!("| Median GQ | {} |", optional_metric(self.quality.median())); - println!( - "| Missing GT calls | {} |", + print!( + "Your call file holds {} records for {} {} in {} format, across {primary_count} primary \ + chromosomes and {other_count} other contigs. {} carry a non-reference \ + genotype, {} are all-reference and {} have a missing call", + format_int(self.records), + self.samples, + if self.samples == 1 { + "sample" + } else { + "samples" + }, + self.file_format, + format_int(self.variant_records()), + format_int(self.genotype.hom_ref), format_int(self.genotype.missing) ); - println!( - "| Parse errors skipped | {} |", - format_int(self.parse_errors) - ); - println!(); - - println!("## VCF Roll Call"); + if self.parse_errors > 0 { + print!( + ", and {} could not be parsed and were skipped", + format_int(self.parse_errors) + ); + } + println!(". {}", self.shape_sentence()); println!(); println!("| Check | Status | Detail |"); println!("| :--- | :--- | :--- |"); @@ -242,12 +234,29 @@ impl Report { ); println!(); + println!("## Evidence"); + println!(); print_genotype_table(&self.genotype); print_variant_table(&self.variants); print_quality_table(self); print_chromosome_table(&self.chromosomes); - print_meaning(self); - print_fine_print(); + print_method(); + print_limitations(self); + print_sources(); + } + + /// One sentence on what the file's shape lets a reader conclude. + fn shape_sentence(&self) -> &'static str { + if self.reference_blocks.records > 0 { + "The file carries reference blocks, so it says where calls were possible, \ + not only where they differ." + } else if self.genotype.hom_ref > 0 { + "The file carries all-reference calls at some sites, but no reference \ + blocks, so it doesn't say how much of the genome was callable." + } else { + "The file records variants only, so a sparse stretch is low variant density, \ + not evidence of low coverage." + } } fn variant_records(&self) -> u64 { @@ -610,19 +619,12 @@ fn sample_integer(header: &vcf::Header, record: &vcf::Record, key: &str) -> Opti } fn print_header() { - println!("# VCF Roll Call"); - println!(); - println!( - "This report scans the loaded variant calls for signs that the file is \ - healthy: chromosomes are present, genotype calls are readable, depth and \ - quality fields appear when available, and reference blocks are counted \ - when the VCF contains them." - ); + println!("# Call File Health Check"); println!(); } fn print_genotype_table(stats: &GenotypeStats) { - println!("## Genotype Calls"); + println!("### Genotype calls"); println!(); println!("| Class | Records |"); println!("| :--- | ---: |"); @@ -643,7 +645,7 @@ fn print_genotype_table(stats: &GenotypeStats) { } fn print_variant_table(stats: &VariantStats) { - println!("## Variant Shape"); + println!("### Variant shape"); println!(); println!("| Signal | Records |"); println!("| :--- | ---: |"); @@ -660,7 +662,7 @@ fn print_variant_table(stats: &VariantStats) { } fn print_quality_table(report: &Report) { - println!("## Depth And Quality"); + println!("### Depth and quality"); println!(); println!("| Signal | Value |"); println!("| :--- | ---: |"); @@ -686,7 +688,7 @@ fn print_quality_table(report: &Report) { } fn print_chromosome_table(chromosomes: &BTreeMap) { - println!("## Chromosome Scoreboard"); + println!("### Chromosome scoreboard"); println!(); println!( "Primary chromosomes are listed separately from alternate, random, unplaced, and patch contigs." @@ -713,7 +715,7 @@ fn print_chromosome_table(chromosomes: &BTreeMap) { .map(|(_, stats)| stats.reference_block_bases) .sum(); println!(); - println!("## Other Contigs"); + println!("### Other contigs"); println!(); println!("| Signal | Value |"); println!("| :--- | ---: |"); @@ -752,48 +754,46 @@ fn print_contig_row(name: &str, stats: &ChromStats) { ); } -fn print_meaning(report: &Report) { - println!("## What It Means"); - println!(); - if report.reference_blocks.records > 0 { - println!( - "This VCF includes reference-block style records, so the report can \ - make a rough callable-span estimate from those blocks. That is a \ - stronger coverage hint than variant density alone." - ); - } else if report.genotype.hom_ref > 0 { - println!( - "This VCF includes homozygous-reference calls, which proves some \ - reference sites were emitted. It does not provide a complete callable \ - span estimate unless those calls cover intervals." - ); - } else { - println!( - "This looks like a variant-only VCF. Sparse windows in this report \ - should be read as low variant density, not as proof of low sequencing \ - coverage." - ); - } +fn print_method() { + println!("## Method"); println!(); println!( - "The most useful red flags are missing genotype calls, absent DP/GQ fields \ - when you expected sequencing data, unexpectedly missing chromosomes, or \ - parse errors while scanning records." + "The app streams the whole call file one record at a time and classifies the \ + first sample's genotype, each record's allele shape, its filters, and its DP \ + and GQ fields when present. A record whose END reaches past its reference \ + allele, or whose alternate is a non-reference placeholder, counts as a \ + reference block, and their spans add up to a rough callable span. A record \ + that can't be parsed is counted and skipped, while a failed read stops the \ + app. The signs it looks for are missing genotype calls, absent DP or GQ fields \ + where sequencing data was expected, missing chromosomes and parse errors." ); println!(); } -fn print_fine_print() { - println!("## Fine Print"); +fn print_limitations(report: &Report) { + println!("## Limitations"); println!(); println!( - "This is a VCF sanity check, not a formal coverage calculator. True base-by-base \ - depth normally comes from alignment data or a gVCF with reference blocks. \ - Genotyping arrays, imputed VCFs, and variant-only exports can all look \ - sparse while still being valid for their intended purpose." + "This is a check of the file, not of how much of a genome was sequenced. \ + Base-by-base depth comes from alignment data or a gVCF with reference blocks, \ + and this file {}. Genotyping arrays, imputed files and variant-only exports \ + all look sparse while being valid for their purpose.", + if report.reference_blocks.records > 0 { + "carries such blocks, so its callable span is an estimate from them" + } else { + "carries none, so nothing here measures coverage" + } ); println!(); - println!("This report is for demo and data-quality triage only."); +} + +fn print_sources() { + println!("## Sources"); + println!(); + println!( + "1. The Variant Call Format specification, VCFv4.5. \ + https://samtools.github.io/hts-specs/VCFv4.5.pdf" + ); println!(); } diff --git a/apps/08-motif-finder/Cargo.lock b/apps/08-motif-finder/Cargo.lock index 50b63a4..86a5e19 100644 --- a/apps/08-motif-finder/Cargo.lock +++ b/apps/08-motif-finder/Cargo.lock @@ -4,4 +4,4 @@ version = 4 [[package]] name = "motif-finder" -version = "0.1.0" +version = "0.2.0" diff --git a/apps/08-motif-finder/Cargo.toml b/apps/08-motif-finder/Cargo.toml index 1b8dbd2..298d57c 100644 --- a/apps/08-motif-finder/Cargo.toml +++ b/apps/08-motif-finder/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "motif-finder" -version = "0.1.0" +version = "0.2.0" edition = "2024" # A smaller module uploads and starts faster on an Ark. diff --git a/apps/08-motif-finder/src/main.rs b/apps/08-motif-finder/src/main.rs index 97ae244..accc5f2 100644 --- a/apps/08-motif-finder/src/main.rs +++ b/apps/08-motif-finder/src/main.rs @@ -24,7 +24,7 @@ fn main() { print!( "[package]\n\ name = \"motif-finder\"\n\ - version = \"0.1.0\"\n\ + version = \"0.2.0\"\n\ datasets = [\"v1/genome/genes/TAS2R38\"]\n" ); return; @@ -57,12 +57,66 @@ fn run(root: &Path) -> Result<(), String> { return Err("The gene sequence is empty".into()); } - println!("## Restriction map of {GENE}\n"); - println!("Scanning {length} bp of reference sequence.\n"); + // The finding names the totals; the table beneath it carries every count. + let total: u64 = cuts.iter().sum(); + let (most, most_cuts) = ENZYMES + .iter() + .zip(cuts) + .max_by_key(|(_, count)| *count) + .map(|((name, _), count)| (*name, count)) + .expect("the enzyme panel is not empty"); + let uncut: Vec<&str> = ENZYMES + .iter() + .zip(cuts) + .filter(|(_, count)| *count == 0) + .map(|((name, _), _)| *name) + .collect(); + + println!("# Restriction Map of {GENE}\n"); + print!( + "{} enzymes cut the {} bp reference sequence of *{GENE}* {total} times between \ + them. {most} cuts most often, {most_cuts} times", + ENZYMES.len(), + commas(length) + ); + match uncut.as_slice() { + [] => println!("."), + [one] => println!(", and {one} not at all."), + many => println!(", and {} never do.", many.join(", ")), + } + println!(); + println!("## Evidence\n"); println!("| Enzyme | Site | Cuts |"); - println!("| :--- | :--- | ---: |"); + println!("| :-- | :-- | --: |"); for ((name, site), count) in ENZYMES.iter().zip(cuts) { println!("| {name} | `{site}` | {count} |"); } + println!(); + println!("## Method\n"); + println!( + "The app streams the gene's reference sequence through a four-base window and \ + counts every position where an enzyme's recognition site appears, overlaps \ + included. Repeats are soft-masked in lowercase, so bases are uppercased before \ + matching.\n" + ); + println!("## Limitations\n"); + println!( + "This is the reference sequence, not yours. The grant covers your changes in the \ + gene, but the app never reads them, so a variant that creates or destroys a site \ + isn't counted." + ); Ok(()) } + +/// Formats a count with thousands separators, so 1143 reads as 1,143. +fn commas(n: u64) -> String { + let digits = n.to_string(); + let mut out = String::with_capacity(digits.len() + digits.len() / 3); + for (i, c) in digits.chars().enumerate() { + if i > 0 && (digits.len() - i) % 3 == 0 { + out.push(','); + } + out.push(c); + } + out +} diff --git a/apps/09-fortune-cookie/Cargo.lock b/apps/09-fortune-cookie/Cargo.lock index 00c0904..2554822 100644 --- a/apps/09-fortune-cookie/Cargo.lock +++ b/apps/09-fortune-cookie/Cargo.lock @@ -4,4 +4,4 @@ version = 4 [[package]] name = "fortune-cookie" -version = "0.1.0" +version = "0.2.0" diff --git a/apps/09-fortune-cookie/Cargo.toml b/apps/09-fortune-cookie/Cargo.toml index 06d1077..c76cfcf 100644 --- a/apps/09-fortune-cookie/Cargo.toml +++ b/apps/09-fortune-cookie/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "fortune-cookie" -version = "0.1.0" +version = "0.2.0" edition = "2024" # A smaller module uploads and starts faster on an Ark. diff --git a/apps/09-fortune-cookie/src/main.rs b/apps/09-fortune-cookie/src/main.rs index 80ea9b4..5bf4715 100644 --- a/apps/09-fortune-cookie/src/main.rs +++ b/apps/09-fortune-cookie/src/main.rs @@ -28,7 +28,7 @@ fn main() { print!( "[package]\n\ name = \"fortune-cookie\"\n\ - version = \"0.1.0\"\n\ + version = \"0.2.0\"\n\ datasets = [\ \"v1/genome/rsids/rs72921001\", \ \"v1/genome/rsids/rs671\", \ @@ -61,11 +61,18 @@ fn main() { } let pick = (acc % FORTUNES.len() as u64) as usize; - println!("## Your genomic fortune\n"); + println!("# Your Genomic Fortune\n"); println!("> {}\n", FORTUNES[pick]); - println!("Genotypes without an answer: {unanswered}.\n"); + println!("## Method\n"); println!( - "The sandbox has no randomness, so this is derived entirely from your \ - genotypes. Run it again on the same data and the fortune is the same." + "The sandbox has no randomness, so the fortune is derived entirely from your \ + genotypes at three variants, folded into one number that picks a line. Run it \ + again on the same data and the fortune is the same. A genotype without an answer \ + adds nothing to the number, and this run had {unanswered} of those.\n" + ); + println!("## Limitations\n"); + println!( + "A fortune is a fortune. Nothing here reads your genes for meaning, only for \ + bytes." ); } diff --git a/docs/01-app-model.md b/docs/01-app-model.md index 2dec026..a7ccf82 100644 --- a/docs/01-app-model.md +++ b/docs/01-app-model.md @@ -29,8 +29,8 @@ An app picks its pass by counting arguments: | C | `argc < 2` | otherwise | | Python | `len(sys.argv) < 2` | otherwise | -Write the report as Markdown, since that is how it is shown. Headings, tables -and emphasis all help. +The report is Markdown, since that is how it is shown, and +[06-reports.md](06-reports.md) covers its shape. ## Before the owner is asked diff --git a/docs/06-reports.md b/docs/06-reports.md new file mode 100644 index 0000000..29142f3 --- /dev/null +++ b/docs/06-reports.md @@ -0,0 +1,208 @@ +# Reports + +A report is what an app prints in its run pass, and it is the only thing an app +can send out of the sandbox. The owner reads it first, on their phone, and +decides whether it goes any further. Whoever ran the app then gets its exact +bytes, in a terminal from `ark app run`, or rendered as Markdown. So a report is +read three ways, on a phone, as raw text and as a rendered page, and it has to +work in all of them. This page is a guideline for that, not a format the Ark +checks. The Ark returns whatever the app printed, up to 1 MiB. + +## What a report has to do + +The owner approved the app before seeing a word of it, so the title and the +finding fit in the first ten lines, and the owner sees the answer before +scrolling. Everything the finding rests on is in the report, each value beside +the coordinate it came from, so a reader can check it and a second app could +reproduce it. The Ark reports the app's name and version and the paths it +granted beside the report, so the report carries none of them. And the report +is readable to the last byte. A report is only as trustworthy as it is +readable. + +Two things follow from the sandbox. The run is deterministic, so a report +carries no date and no run id. The same app over the same data prints the same +report, which is what lets last year's be diffed against this year's, and an +app that lists a directory sorts the listing before printing it, since listing +order isn't promised. And the output is capped, so a report summarises. It +lists the evidence for its finding and counts the rest, since output past the +cap is dropped without a mark. + +An app that exits non-zero returns nothing. Without `develop = true` the Ark +withholds standard output on failure and never returns standard error, so the +owner is left with a blank screen after approving the run. A failure exit is +for data that couldn't be read. Everything an app can state, including that +it has no answer, is a report. + +## The shape + +A report follows one order of ideas. The answer, what it rests on, how it +follows, what it doesn't establish, and where it came from. The answer sits +straight under the title with no heading of its own, and the rest are sections +with these default names. + +``` +# what it answers, in Title Case, and then at once the finding, + the answer in plain language, a paragraph or one small table + +## Evidence the values the finding rests on, each with its coordinate +## Method how the finding follows, and from which study or guideline +## Limitations what the app doesn't establish, as facts +## Sources the studies named in Method, in that order +``` + +The title and the finding are the minimum. The four sections appear when they +have something to say. A finding with several parts, the axes of a panel, gives +each a `##` of its own before Evidence. A report that isn't a trait finding, a +map, a scan or a quality check, keeps the order and names its sections for what +they hold. + +The title is the subject, not the app's name, which the Ark reports beside the +report, so an app called `drunk-o-type` prints a report titled "How Your Body +Handles a Drink". The same shape holds one variant, a panel, a gene, a whole +call file, and the summary a study asked for. + +[03-cilantro-soapiness](../apps/03-cilantro-soapiness) is a light app, and +its report in this shape: + +```markdown +# 🌿 Cilantro Taste Test + +**Soap detector** 🧼. You carry two copies of C at rs72921001, the genotype +most strongly tied to tasting cilantro as dish soap. Blame *OR6A2*, the +olfactory receptor next door. C is also the reference base here, so a call +file that lists only variants would have stayed quiet about it. Yours didn't. + +## Evidence + +| Variant | Nearest gene | Position | +| :-- | :-- | :-- | +| rs72921001 | *OR6A2* | chr11:6868417, GRCh38.p14 | + +Your call file says `|C|C` here, and the reference base is C. + +## Method + +Some people love cilantro. Others think it tastes like dish soap. Eriksson et +al. (2012) went looking for why, in a genome-wide study of self-reported +preference among European-ancestry participants, and rs72921001 is what they +found. The app counts your copies of C. Two is the strongest association, one +is somewhere in between, and zero is none. More copies, more soap. + +## Limitations + +One variant, one modest effect. The study measured what people said they +preferred, not what they tasted, and taste is polygenic and shaped by diet, +culture and exposure besides. Two copies doesn't mean you hate cilantro, and +zero doesn't mean you love it. The study was mostly people of European +ancestry, so elsewhere the effect is less well known, and this app doesn't +know how common your genotype is. + +## Sources + +1. Eriksson N, et al. A genetic variant near olfactory receptor genes + influences cilantro preference. Flavour. 2012;1:22. + https://doi.org/10.1186/2044-7248-1-22 +2. NCBI dbSNP, rs72921001. https://www.ncbi.nlm.nih.gov/snp/rs72921001 +``` + +## The sections + +**The finding** says what was found, in the owner's own terms, and never more +than the evidence supports. It says "carries" and "is associated with", never +"you have" or "you will". A finding is a sentence and a value. A verdict label +can lead it, "Soap detector" or "Tomato Mode", as long as the value it rests on +follows in the same breath, since a reader shouldn't have to scroll to learn +what a label means. An absent answer is a finding too, stated as one, and +never read as two reference alleles: + +```markdown +# 🌿 Cilantro Taste Test + +**No answer.** No genotype is available at rs72921001, so this app has no +result. An absent genotype is not a reference call. A call file that records +only variants has no record at a reference site, and none at a site it didn't +cover, and the two can't be told apart here. +``` + +**Evidence** holds the values the finding rests on, one row per variant, gene +or interval, with its coordinate. A genotype is printed exactly as the file +holds it, `|C|C` and all. A pipe can't survive a table cell, where it has to be +escaped and reads as `\|` raw, so a genotype holding one goes in prose or in +code, or a panel table carries one column per allele. Never swap a slash for a +pipe, that is a different genotype. Absent values stay in the table, marked +absent. A long list is cut, with a count of what was left out. + +Coordinates name their assembly when the app can read it. It sits in +`v1/genome/reference/build`, and reading it means granting +`v1/genome/reference`, which puts the stored reference genome on the approval +screen. The demos take that trade. An app that doesn't leaves the coordinate +unlabelled rather than labelling it wrongly, since an Ark may hold a GRCh37 +genome. + +**Method** explains how the finding follows from the evidence, in plain words, +and names each study inline, as author and year, with the population it was +established in. An app may carry allele frequencies or effect sizes from the +literature, and cites where they came from like any other claim. + +**Limitations** says what the app didn't read and what the finding doesn't +establish, as facts rather than reassurance. One variant rarely tells a whole +story, and this section says how much this one tells. + +**Sources** lists the studies Method named, in that order, in full, author, +year, title and journal, with a URL where the work has a stable one. Links +appear only here, and each is the fixed public address of the work it cites. +Nothing in a URL comes from the data. + +## Voice + +Pick a voice and keep it from the first line to the last. A serious subject +gets a professional report throughout. A light subject can be light +throughout. What a report never does is switch, a joke verdict over a clinical +table, or a punchline wrapped around a finding that touches real risk. A light +report can carry a heavy fact, and when it does, that sentence is plain and +says so, the way a friend drops the joke for a moment. + +The examples are toys, and they read that way, charming and friendly whatever +they count. The serious voice is for heavier subjects, and none of the examples +has one yet. So that its register is on record, the cilantro finding above in +the serious voice: + +```markdown +You carry two copies of the C allele at rs72921001, the genotype most strongly +associated with perceiving cilantro as soapy. C is the reference base at this +position, so a call file that lists only variants may not record it. +``` + +Whichever voice, the evidence, the limitations and the sources are the same, +and the finding states the value in words. Emoji belong in prose and headings +in a light report, never in a table, where they break alignment in a terminal, +and never in place of the value. + +Genomics vocabulary, not clinic vocabulary. An app reports an rsID, a gene, a +genotype and a coordinate on a named assembly. It has no patient, no specimen, +no test, no result that is positive or negative, no clinical significance and +no recommendation. It names what a study found and how strongly, and stops. + +Second person for the owner's own data. Genes in italics, alleles, genotypes +and paths in code. Every number with its unit, every percentile with its +reference population. Title Case for the title, sentence case for headings. + +## Form + +Plain Markdown, and a small subset of it, because the subset is what keeps a +report readable everywhere. One `#`, the `##` sections above, and `###` to +group the rows of an Evidence section. Paragraphs, lists and +tables with as few columns as carry the evidence, in rows of about 80 +characters, the values the owner needs in the leftmost columns since a phone +is narrower still. A paragraph is one line of output, since every renderer +wraps it. A text figure in a code fence when it earns its place. Nothing else. + +No images, no HTML, no footnotes and no encoded content, since the owner can't +read them. No dates and no run ids, which the sandbox can't give and the +journal already has. Standard error is for the developer and never part of the +report. + +The minis print a bare line on purpose. They demonstrate a read, and the full +app beside each one demonstrates the report. `ark help apps` covers the +manifest and the sandbox limits, and `ark help output` how the CLI returns a +report. diff --git a/fixtures/README.md b/fixtures/README.md index d8827a3..7a9b87d 100644 --- a/fixtures/README.md +++ b/fixtures/README.md @@ -20,8 +20,9 @@ The rsID apps are `02-permissions`, the cilantro apps, `04-drunk-o-type` and make run APP=03-cilantro-mini-rust FIXTURES=fixtures/no-call ``` -The two smaller roots hold only the variant data those apps need, so any other -app stops at its missing grants there. +The two smaller roots hold only the variant data those apps need, plus the +reference build they name coordinates on, so any other app stops at its missing +grants there. ## Editing them diff --git a/fixtures/no-call/v1/genome/reference/build b/fixtures/no-call/v1/genome/reference/build new file mode 100644 index 0000000..292aab6 --- /dev/null +++ b/fixtures/no-call/v1/genome/reference/build @@ -0,0 +1 @@ +GRCh38.p14 \ No newline at end of file diff --git a/fixtures/unanswered/v1/genome/reference/build b/fixtures/unanswered/v1/genome/reference/build new file mode 100644 index 0000000..292aab6 --- /dev/null +++ b/fixtures/unanswered/v1/genome/reference/build @@ -0,0 +1 @@ +GRCh38.p14 \ No newline at end of file