Skip to content

Repository files navigation

English | 日本語

carbonroute

Which of these two ways to make the same product has the lower carbon footprint — and how sure are we?

carbonroute answers that question for one comparison at a time, or for an entire reaction database at once. It never tells you the absolute carbon footprint of a synthesis — that needs background data for every single input, most of which nobody can get for free. It only needs data for what's different between two routes, because everything they share cancels out. That is a much smaller, much cheaper problem, and it is one public data can actually solve — at a scale most LCA tooling never attempts.

This is v0, and it is a screening tool, not a certification: its output is not an ISO 14067-conformant carbon footprint. See "What this tool does not do" and docs/limitations.md before relying on it for anything.

pip install -e .
carbonroute compare route.yaml --a routeA --b routeB

Background: why nobody has answered "are enzymes actually greener?" at scale

Environmental awareness is rising worldwide, and countries are pursuing carbon neutrality across every industry. In chemical manufacturing, that attention has landed on biomanufacturing — using enzymatic reactions in place of organic-chemical ones. Enzymes generally work in aqueous solvent at ambient temperature and pressure, and they are stereo-, regio- and enantioselective, so the route is expected to waste less energy and generate less waste.

The number that quantifies environmental burden is LCA (Life Cycle Assessment): the method (ISO 14040/14044) that inventories the resources, energy and emissions of every life-cycle stage of a product — raw-material extraction, manufacture, transport, use, disposal — and converts them into environmental impact. Restricted to greenhouse gases and expressed per unit of product in kg CO₂-equivalent, it is the carbon footprint (ISO 14067).

Yet there is very little research that compares enzymatic against conventional organic-chemical routes quantitatively and comprehensively on an LCA basis. The cause is structural: LCA is designed to accumulate absolute values. Computing one product's footprint requires a background emission factor for every single input — each of which is itself the output of another LCA. Most of those live behind paid commercial databases, and for the substances that dominate the enzymatic side — the cofactors UDP-glucose, NADPH, S-adenosyl-L-methionine, acetyl-CoA — there is effectively no openly licensed cradle-to-gate factor at all. One comparison therefore costs days to weeks of primary-literature work, which puts database-scale coverage out of reach in principle. The bottleneck is not compute, it is the cost of acquiring data.

This work starts from the observation that you do not need absolute values to answer "which one is lower." The approach is to accumulate only the delta set between two routes, and to carry missing factors not as point estimates but as bounded intervals. The tool built on it, carbonroute, computes the relative carbon burden of an enzymatic versus an organic-chemical route, and the critical condition at which that verdict flips. When two routes make the same product, everything common to both cancels out of the difference. What cancels is precisely the expensive part — the substrate and product, different in every reaction. What survives is the cofactor on the enzymatic side and the protecting groups, activator, base and solvents on the chemical side: two closed vocabularies of a few dozen substances each. So the primary-research effort is bounded by the size of the vocabulary, not by the number of reactions.

carbonroute has four defining properties. (1) It invents no numbers — missing data enters as an interval, and if interval arithmetic does not settle the comparison it prints "indeterminate" and stops. (1b) And it says which number the answer is made of — provenance alone cannot, and this repository has withdrawn three published results to learn it: each time one term dominated the delta while its value came from an assumption. In the shipped class, 451 of 451 verdicts rest at least half on a single material, the same one every time. (2) Its output is a critical value, not a winner — not "the enzyme wins" but "at what solvent recovery rate does that conclusion stop holding." (3) It scales per reaction class — template one real published chemical procedure for a class, and every reaction in that class costs one arithmetic evaluation. That makes it useful for choosing biomanufacturing targets — deciding which enzymatic reactions deserve scarce research budget — and for auditing claims of environmental advantage.

Applying it, we set out to compare enzymatic and organic-chemical routes comprehensively: which enzymatic reactions contribute most to emissions reduction; where the advantage lies once enzymatic yield and solvent recycling rate are accounted for; and where the emission-reduction effect of today's commercialised biomanufacturing ranks among all enzymatic reactions.

Where this actually stands — and what is still aspirational

The no-invented-numbers rule applies to our own progress too. The paragraph above states the goal of this repository; only part of it is done. Here is each question, what it needs, and where it stands.

Question What it needs Status
Q1. Which enzymatic reactions contribute most A metric comparable across reaction classes, and a ranking Metric done, coverage growing. Reactions rank on kg CO₂e saved per kg of product, which means the same thing in any class. Twenty-six classes now built — 8,041 of 18,558 reactions matched (43.3%), 3,906 decided (21.0%). See "How far coverage can actually go" for the honest ceiling on this number
Q2. The advantage once yield and solvent recycling are accounted for A 2-D break-even curve over (enzymatic yield × solvent recovery) Done. Both axes are modelled; the frontier is below. The answer is not the one the enzymatic route wanted
Q3. Where commercialised biomanufacturing ranks A mapping from commercial processes to Rhea reactions, and percentiles Not started

Q2, answered: the break-even frontier

The bias this section used to describe has been removed, and removing it changed the conclusion.

The asymmetry was this. On the chemical side, the shipped template's quantities are carried through the source paper's own real yields — 62% glycosylation × 85% deprotection = 52.7% overall — so the chemical route pays the penalty of charging nearly twice the material to obtain 1 kg of product. The enzymatic side was billed at pure stoichiometry: an implicit 100% conversion, and no equivalent penalty. Enzymatic conversion is now a declared variable that divides the cofactor demand, because a reaction converting half its acceptor consumes twice the cofactor per kg of product.

Sweeping it gives the verdict's real boundary — a curve, not a number. Every row below re-screens all 451 decided reactions at a different conversion:

enzymatic conversion min threshold median max
100% 84.56% 86.42% 91.54%
90% 82.80% 84.87% 90.54%
80% 80.61% 82.94% 89.30%
70% 77.79% 80.44% 87.70%
60% 74.02% 77.12% 85.57%
50% 68.76% 72.47% 82.59%
40% 60.86% 65.50% 78.11%
30% 47.69% 53.87% 70.65%

A coin-flip enzyme costs the class about 14 points of median threshold. But the sharper finding is on the other axis. Industrial distillation recovers about 90%, and at that recovery only 25 of the 451 reactions are still decided at all — the other 426 lose their verdict however well the enzyme performs. The 25 that survive need conversions of 85.3% at minimum, 93.3% at the median, and 100% at the worst. The calibration case itself (RHEA:12560, β-arbutin) is one of the 426: its threshold is 85.87%, below the plant's 90%.

So the honest answer to Q2 is that in this class, at a realistic solvent loop, the enzymatic advantage mostly is not there — and where it is, it demands a near-quantitative enzyme. Run it yourself:

carbonroute screen --template ... --bounds ... --assumptions-from ... \
  --enzymatic-yield 0.5 --frontier

Q1's metric: what can be compared between classes, and what cannot

A recovery threshold is measured against whatever solvent load a template happens to carry, so 86% in a glycosylation class and 86% in a methylation class are not the same claim. An absolute saving is: kg CO₂e per kg of product, read at the same operating point everywhere. Every row now carries that as an interval, evaluated at 90% recovery rather than the bench's zero.

Ranking on an interval cannot be a total order, so rank_by_advantage returns a rank range: a reaction is outranked only by reactions whose worst case still beats its best case. Sorting by midpoint and printing 1, 2, 3 would manufacture the precision this project refuses to manufacture anywhere else.

Run on the shipped class, that reports something about the bounds rather than the chemistry. Every rank range comes back 1–451, because four chemical-side materials are deliberately asserted with no upper ceiling — an honest refusal to invent one — and an unbounded chemical side makes every enzymatic advantage unbounded above. The report names those four, so the gap is actionable: put defensible ceilings on them and the ranking bites.

What is available meanwhile is the guaranteed floor, the saving that holds everywhere in the asserted bounds. At 90% recovery only 25 of 451 reactions have a floor above zero. And it reproduces the mechanism the screen exists to measure — the top ten carry 27 to 34 protectable groups, led by an N-glycan at +4.15 kg CO₂e per kg — which is the check that it measures the same thing the threshold did.

The road from here

  1. Add classes, now that they are cheap Done, and continuing. A template may declare a process_model — reagents at stated equivalents, solvent from a stated reaction concentration — instead of quoting one paper's charged amounts. The parameters are chemistry-independent, so a new class supplies its reagents and inherits the process. That is what unblocked coverage: three rounds of literature search produced zero new paper-sourced templates, and the paper-per-class route does not scale. The shipped udp-glucosyltransferase.yaml class itself grew for free the same way galactose originally did: six more hexose-nucleotide donors (GDP-mannose, ADP-glucose, GDP-glucose, UDP-galactofuranose, dTDP-glucose, CDP-glucose — each the same 162.14 g/mol hexose transfer, just carried by a different nucleotide) added 72 more matched reactions and 63 more decided, taking it from 406/388 to 478/451 with no new research — every downstream Q1/Q2 finding below was re-derived against the larger set, not assumed to carry over. Six more classes are built the same way, no paper cited in any: sam-methyltransferase.yaml (EC 2.1.1, SAM-dependent O/N/S-methylation — 449 matched, 351 decided), nad-oxidoreductase.yaml (EC 1.1.1, NAD(P)+-dependent oxidation, the largest EC-3 group — 515 matched, but 0 decided, honestly: its process model is too small relative to the cofactor's wide, unevidenced bound to guarantee a verdict either way, and that is reported rather than papered over with inflated reagent amounts), atp-kinase.yaml (EC 2.7.1, ATP-dependent phosphorylation — 209 matched, 202 decided, all of them decisively favouring the enzyme; also where ChEBI's phosphate-dianion charge-state convention was pinned down, correcting the class's mass delta from the textbook 79.980 to the observed 77.963), dmapp-prenyltransferase.yaml (EC 2.5.1, DMAPP-dependent prenylation — 47 matched, 38 decided, again all decisive; two more real confounds — chain elongation and DMAPP homodimerisation, sharing the same cofactor but a different transformation — correctly excluded rather than mis-decided), and acetyl-coa-acyltransferase.yaml (EC 2.3.1, acetyl-CoA-dependent O/N-acetylation — 138 matched, a second charge-state unification (this time the acceptor carries the extra proton, not the product), and a second honest non-result: all 124 resolved reactions are indeterminate, for a reason related to but distinct from the NAD(P)+ class's — small acceptors mean a comparatively large cofactor cost per kilogram of product even at acetyl-CoA's cheapest plausible value), and udp-glucuronosyltransferase.yaml (UDP-glucuronate-dependent glucuronidation — the real UGT drug-metabolism family — 102 matched, 95 decided, all decisive; a third charge-state split, one proton apart, unified the same way as the other two), and three more siblings of the DMAPP class — gpp-prenyltransferase.yaml, fpp-prenyltransferase.yaml and ggpp-prenyltransferase.yaml (the other three allylic-diphosphate prenyl donors, transferring two/three/four isoprene units instead of DMAPP's one — 11/13/9 matched, 10/2/4 decided). FPP is the honest finding there: it decides only 2 of 13 (15.4%, the opposite of DMAPP's 80.9%) because most of its real chemistry is chain elongation or homodimerisation, not transfer onto a foreign nucleophile — a smaller class, correctly reported as one rather than padded. Building these three also surfaced a real bug shared by every class: the acceptor-identifying step excluded a bare proton from the product side of an equation but not the reactant side, silently failing every reaction that genuinely needs one as a co-reactant (2,673 of Rhea's 18,558, 14.4%). Fixed symmetrically, and checked to be purely additive (it can only add matches, never remove one) against all ten shipped classes before landing — three gained a reaction each. And an eleventh, o2-monooxygenase.yaml (EC 1.14.13, aromatic/aliphatic hydroxylation, Baeyer-Villiger oxidation, sulfoxidation and N-oxidation — 198 matched, 134 decided, all decisive; a fourth independent charge-state split), is the first class needing genuinely two cofactors: O2 supplies the atom the product gains, and NAD(P)H drives the mechanism in the same step. Nothing in this architecture could express that before, so ClassTemplate gained unpriced_co_cofactor_chebi — NAD(P)H is excluded from the acceptor search but deliberately never priced, a stated gap (this class's enzymatic side is understated by whatever NAD(P)H's real regeneration costs) rather than an invented zero, flagged on every report this class produces. A twelfth, p450-monooxygenase.yaml (EC 1.14.14, cytochrome P450s and related heme-thiolate monooxygenases — Rhea's single largest O2-consuming EC group, bigger even than EC 1.14.13's — 256 matched, 176 decided, all decisive), reuses the same unpriced_co_cofactor_chebi mechanism with a different electron-donor identity (CHEBI:57618, the same ChEBI entity Rhea's equation text variously labels "reduced [NADPH--hemoprotein reductase]", "FMNH2" or "reduced [flavodoxin]" depending on biological context, plus FADH2). A thirteenth, o2-desaturase.yaml (EC 1.14.19, fatty acyl desaturation — 115 matched, 92 decided, all decisive), shares the O2 cofactor with the two monooxygenase classes but targets the opposite mass signature: this EC group removes two hydrogens to form a C=C bond (−2.016 g/mol, the same 2H-loss signature the NAD(P)+-oxidoreductase class targets), not insert an oxygen atom, so it reuses that class's process model rather than the monooxygenases' mCPBA one. A fourteenth, o2-dioxygenase.yaml (EC 1.13.11, dioxygenation incorporating both O2 atoms with no separate electron donor at all — 102 matched, 59 decided, all decisive), needs no co-cofactor and shows a graded three-way version of the charge-state pattern found four times before (32.0 / 30.99 / 29.98 g/mol, unified with a widened tolerance the same way the UDP-glucuronosyltransferase class's two-way split was). A fifteenth, 2og-dioxygenase.yaml (EC 1.14.11, Fe(II)/2-oxoglutarate- dependent dioxygenases — prolyl/lysyl hydroxylases, gibberellin and clavulanic acid biosynthesis — 101 matched, 60 decided, all decisive), inserts one oxygen atom the same way the monooxygenase classes do, but the second oxygen atom oxidatively decarboxylates the co-substrate 2-oxoglutarate to succinate + CO2 rather than reducing NAD(P)H — declared unpriced the same way, a real co-cofactor identity this mechanism now covers five times over. A sixteenth, ferredoxin-monooxygenase.yaml (EC 1.14.15, steroid/bile acid hydroxylases, camphor monooxygenase, alkane hydroxylases — 72 matched, 50 decided, all decisive), reuses the same mCPBA process model with a third electron-donor family (ferredoxin, adrenodoxin, putidaredoxin and rubredoxin, three ChEBI ids covering all 72 reactions). Combined: 2,815 of 18,558 reactions matched (15.2%), 1,724 decided (9.3%), up from one class's original 406 / 388. More classes are being added the same way — see the coverage ceiling below for what that number can and cannot reach.

    UPDATE, after re-deriving the ceiling below (see the corrected section): ec_prefix on every class above was an unnecessary extra restriction, not a required one. Rhea's own EC annotation covers only 41.1% of its reactions, and a real share of the rest is genuine class membership that expected_mass_delta/transferred_bond_smarts already verify structurally without needing an EC number at all. Two real bugs were found and fixed in the process — _identify was treating water as a candidate acceptor the same way a proton once was (fixed the same way, closing three false positives already live in the shipped glycosylation/glucuronidation classes and recovering one true positive), and all four prenyltransferase classes were treating a single isopentenyl-diphosphate chain-elongation hop as a genuine transfer, arithmetically indistinguishable from one by mass delta alone (fixed with a new excluded_co_cofactor_chebi field). With both fixed, ec_prefix was removed from ten classes (SAM, ATP-kinase, all four prenyltransferases, and the O2-monooxygenase/desaturase/2-oxoglutarate/ ferredoxin classes), each disambiguated instead by which co-cofactor it actually requires or excludes — a new required_co_cofactor_chebi field, the mirror image of excluded_co_cofactor_chebi. Verified against all ten with zero reactions lost from any class's previous decided set, and every newly decided reaction spot-checked against its real equation. Corrected total: 6,505 of 18,558 reactions matched (35.1%), 3,305 decided (17.8%) — see docs/screening.md for the full per-class breakdown and data/reaction-classes/*.yaml file headers for each class's own verified numbers.

    A seventeenth class, paps-sulfotransferase.yaml (3'-phosphoadenylyl sulfate-dependent sulfation of a phenol, alcohol, amine or glycosaminoglycan/steroid/bile-acid hydroxyl — 130 matched, 128 decided, all decisive), was added the same EC-free way from the start: the mass-delta check alone (net +79.056 g/mol, the sulfonate group replacing a hydrogen) separates 128 of 130 PAPS-consuming reactions with zero confounds found — the tightest signal-to-noise ratio of any class this project has built. Its process model is SO3-pyridine complex in pyridine solvent, the textbook mild sulfating agent. Updated total: 6,635 of 18,558 reactions matched (35.8%), 3,433 decided (18.5%).

    An eighteenth class, coa-ligase.yaml (ATP/GTP-dependent CoA thioester formation, the acid-thiol ligase mechanism — 156 matched, 0 decided), is a second honest non-result, more extreme than the NAD(P)+-oxidoreductase and acetyl-CoA-acyltransferase classes' own: CoA's transferred mass (746.502 g/mol, forming the whole acyl-CoA thioester) is the largest of any class this project has built, and applied to the same wide, unevidenced [0.5, 100] cofactor bound every unpriced cofactor here uses, the enzymatic side's own uncertainty span ([0.37, 74.6] kg CO2e per functional unit) is wide enough alone to straddle any plausible chemical-route footprint — not close, structurally undecidable at these bounds. All 156 matches resolve cleanly (zero excluded by mass delta), so the class is real, verified Q1 coverage even though it decides nothing; its marginal contribution to unique matched coverage is also small (+4, not +156), because most of its 156 already fall inside atp-kinase's own widened matched set once ec_prefix was dropped from that class (both require ATP; they never double-decide, since atp-kinase's own mass-delta check correctly rejects a CoA thioester's ~746 g/mol addition as nowhere near its 77.963 target). Final total: 6,639 of 18,558 reactions matched (35.8%), 3,433 decided (18.5%) — unchanged from the seventeen-class figure at this precision, an honest reflection of how little marginal coverage this particular class adds once its overlap and its own undecidability are both counted.

    A systematic recovery, not a new class: AH2. Rhea uses CHEBI:17499 ("AH2") as its own generic placeholder for "some unspecified reduced donor" wherever curators knew a monooxygenase-type reaction needed one but not its specific identity. 169 reactions across every class this project has built were blocked from resolving to a single acceptor by AH2 sitting unexcluded in the acceptor search — the same shape as the water and proton exclusions before it, but per-class rather than universal, because AH2 is genuinely a different transformation on different classes (radical-SAM chemistry on sam-methyltransferase, tRNA sulfur-relay and cobalamin biosynthesis on atp-kinase — both correctly left excluded). Verified as genuine only for two classes: o2-monooxygenase (+64 decided, zero lost, zero double-decided with any other O2 class) and 2og-dioxygenase (+1). o2-dioxygenase's own exclusion list gained AH2 too, to keep its "matched" honest now that those 64 reactions correctly belong elsewhere. One bug caught before landing: adding AH2 to required_co_cofactor_chebi on both classes independently briefly made o2-monooxygenase and 2og-dioxygenase double-decide the same 64 reactions, because required_co_cofactor_chebi uses OR semantics — fixed by declaring AH2 required only where it is a genuine alternative donor (o2-monooxygenase) and unpriced-only, not required, where a different cofactor already does the real gating (2og-dioxygenase, still gated on 2-oxoglutarate alone). Total after this fix: 6,639 of 18,558 reactions matched (35.8%), 3,498 decided (18.8%).

    A nineteenth class, gdp-fucosyltransferase.yaml (GDP-fucose- dependent fucosylation — 77 matched, 76 decided, all decisive), found by surveying every not-yet-covered cofactor for mass-delta homogeneity directly (a script checking the real acceptor/product delta for each candidate, not EC-group inspection) rather than picked one at a time. The cleanest new class this session: 76 of 77 land within one proton of 146.14 g/mol (the fucosyl group), the other 1 being GDP-fucose SYNTHASE (the biosynthetic route that makes GDP-fucose, correctly excluded by mass delta as a different transformation). Includes RHEA:14257, the human-milk-oligosaccharide fucosylation this project's own Q3 target list already names. Process model is Koenigs-Knorr fucosylation with a real, commercially catalogued glycosyl bromide donor (CAS 16741-27-8, verified via web search), the same mechanism udp-glucuronosyltransferase's own donor uses. The same survey surfaced a cleaner sialyltransferase candidate too (CMP-sialic acid, 122 reactions, 97%+ purity) — not built yet, because sialylation donors are custom-synthesized per lab rather than sold with a simple, verifiable commercial CAS number the way the fucosyl bromide is; left as a documented candidate rather than guessed at. Total after this class: 6,715 of 18,558 reactions matched (36.2%), 3,574 decided (19.3%).

    A twentieth class, udp-acetylhexosaminyltransferase.yaml (UDP- GlcNAc/UDP-GalNAc-dependent glycosylation — 152 matched, 130 decided, all decisive), merges the two donors the same way udp-glucosyltransferase merges UDP-glucose/UDP-galactose: C4 epimers, identical transferred mass (203.194 g/mol). The least clean of the survey's top candidates, reported at its real number rather than rounded up — 85.5%, not the 95%+ the flagship glycosylation classes reach — because this cofactor pair has real competing chemistry a simple mass-delta check has to separate out: GlcNAc-1-phosphate transfer onto dolichyl phosphate (RHEA:13289, EC 2.7.8.15, a different enzyme family transferring the whole phosphosugar, not just GlcNAc) and oxidation of the donor itself to the uronate (RHEA:13325, NAD+- dependent), both correctly excluded rather than folded in. Process model uses a glycosyl CHLORIDE donor (CAS 3068-34-6, verified via web search) rather than the bromide the other sugar-donor classes use, because the C2 acetamido group's neighbouring-group participation favours the more stable chloride for this specific sugar. Total after this class: 6,863 of 18,558 reactions matched (37.0%), 3,704 decided (20.0%) — crossing 20% decided for the first time this session.

    gdp-fucosyltransferase.yaml extended, not a new class: UDP- rhamnose. Fucose and rhamnose are both L-configured 6-deoxyhexoses — different stereochemistry, identical formula and identical transferred mass (146.14 g/mol, verified against RHEA:61160, quercetin -> quercitrin) — so UDP-beta-L-rhamnose was added to the fucosylation class's own cofactor_chebi list rather than built as a twenty-first class, the same merge already used three times over (UDP-glucose/ UDP-galactose; UDP-GlcNAc/UDP-GalNAc; now GDP-fucose/UDP-rhamnose). 14 of 15 rhamnose-consuming reactions land on the identical delta (rhamnosylation of flavonoids and saponins — myricetin, quercetin, kaempferol among them); the other 1 is chain elongation of a growing rhamnogalacturonan, correctly excluded. Renamed to "6-deoxyhexose-nucleotide-dependent glycosylation" to reflect the merge. Total after this merge: 6,878 of 18,558 reactions matched (37.1%), 3,718 decided (20.0%).

    A twenty-first class, udp-xylosyltransferase.yaml (UDP-xylose- dependent xylosylation — 24 matched, 22 decided, all decisive), the third and final class from the mass-delta survey's clean-candidate tier. Xylose is a pentose, not a hexose, so it cannot merge with the sugar classes above despite the identical Koenigs-Knorr mechanism — 132.12 g/mol transferred, verified against RHEA:22244 (kaempferol -> kaempferol 3-O-beta-D-xyloside), one carbon short of the hexose donors' mass. Covers protein O-xylosylation on EGF-like/Notch domains as well as flavonoid and saponin xylosylation. The 2 excluded: a phosphoxylosyl transfer at a different net mass, and chain elongation of a growing proteoglycan linkage region. Total after this class: 6,902 of 18,558 reactions matched (37.2%), 3,740 decided (20.2%).

    A twenty-second class, cmp-sialyltransferase.yaml (CMP-sialic acid-dependent sialylation — 122 matched, 119 decided, all decisive), the largest candidate left unbuilt from the mass-delta survey, deferred only because a simple, verifiable commercial sialylation donor took longer to find than the sugar-nucleotide classes' peracetylated halides did. Net mass 290.25 g/mol (the sialyl group), verified against RHEA:11836 (673.598 − 383.350 = 290.248 exactly) and confirmed identical across ganglioside (RHEA:18021, GM1 → GD1a) and glycolipid (RHEA:18417) acceptors. A genuine doubled-transfer cluster — 4 reactions installing two sialyl residues in one step onto a branched N-glycan, landing at exactly 2× 290.25 = 580.50 — is correctly decided via the same cofactor_coeff mechanism the SAM class's di/tri- methylations use. The chemical route is chemically harder than the other sugar-donor classes' Koenigs-Knorr chemistry (a ketose anomeric centre with no neighbouring-group participation), so the process model uses an NIS/TfOH-activated thioglycoside donor (CAS 155155-64-9) instead of a halide, at a lower stage yield (65%, not the other sugar-donor classes' 80%) reflecting that real difficulty. The 3 excluded: 2 (RHEA:79555, RHEA:81827) are O-acetylation of the sialic acid's own side chain via acetyl-CoA, a different transformation the same cofactor also participates in; 1 (RHEA:16145) is hydroxylation to the N-glycoloyl form via O2/cytochrome b5, a three-reactant shape this model does not attempt to handle. Total after this class: 7,024 of 18,558 reactions matched (37.8%), 3,859 decided (20.8%).

    A twenty-third class, cdp-cholinetransferase.yaml (CDP-choline- dependent phosphocholine transfer — 18 matched, 17 decided, all decisive), the smaller sibling candidate found alongside the sialic acid class in the same survey, deferred behind it earlier this session. Net mass 165.13 g/mol (the phosphocholine group), verified against RHEA:32939 (diacylglycerol → phosphatidylcholine) and confirmed identical across a ceramide acceptor (RHEA:16273, → sphingomyelin) and a protein-serine acceptor (RHEA:56080, Rab1 phosphocholination by the Legionella effector AnkX) — genuinely different EC numbers unified by the same mass-delta mechanism, so no ec_prefix is declared. The chemical route is the Aneja method, the established synthesis for synthetic phosphatidylcholines: phosphorylation with a cyclic chlorophosphate (2-chloro-2-oxo-1,3,2-dioxaphospholane, CAS 6609-64-9) then ring-opening with trimethylamine (CAS 75-50-3), both verified via web search. The 1 excluded, RHEA:32487, is plain hydrolysis of CDP-choline itself — no organic acceptor remains once the cofactor, H+ and H2O are excluded. Total after this class: 7,042 of 18,558 reactions matched (37.9%), 3,876 decided (20.9%).

    A twenty-fourth class, nadph-ketoreductase.yaml (NADPH-dependent carbonyl reduction — 687 matched, 0 decided, another honest non-result), is the mirror image of nad-oxidoreductase: it matches the reduced cofactor form (NADPH) instead of the oxidised one, the reverse direction (ketone/aldehyde → alcohol). NADPH's dominant mass-delta cluster (+2.016 g/mol) turns out to mix two different reductions that weigh the same — carbonyl reduction and alkene reduction — the same confound sam-methyltransferase solved with a bond check, so this class declares transferred_bond_smarts: "[CX4][OX2H]" (a new C-OH) to separate them, verified directly against real Rhea structures. 148 of 687 (21.5%) are structurally clean carbonyl reductions — genuine coverage — but every one is indeterminate, not decided: NADPH is as heavy as NAD(P)+ (~745 g/mol), and its own wide, unevidenced cofactor bound produces the same honest non-result already documented for nad-oxidoreductase, acetyl-coa-acyltransferase and coa-ligase. Chemical route: sodium borohydride (CAS 16940-66-2) in methanol — the standard hydride source for a carbonyl reduction, not claimed to cover the alkene reductions the bond check excludes. Total after this class: 7,729 of 18,558 reactions matched (41.6%), 3,876 decided (20.9%).

    A twenty-fifth class, nadh-ketoreductase.yaml, is the direct sibling of the NADPH class above — same session, same bond-check reasoning, same sodium borohydride process model, matching NADH instead. 278 matched (smaller than NADPH's 687), 18 structurally clean carbonyl reductions, all 18 indeterminate — the same honest non-result, same heavy-cofactor reason (NADH ~665 g/mol). Total after this class: 8,007 of 18,558 reactions matched (43.1%), 3,876 decided (20.9%).

    A twenty-sixth class, hcn-cyanohydrin.yaml (HCN-dependent cyanohydrin formation — 34 matched, 30 decided, all decisive), an unusually small, unusually clean candidate caught on a second, wider discovery pass. Hydroxynitrile lyase chemistry: HCN adds across an aldehyde or ketone to form the cyanohydrin, net mass 27.026 g/mol (HCN itself), verified against RHEA:77427 (benzaldehyde → (S)-mandelonitrile). No ec_prefix — every real member carries no EC annotation at all in this Rhea release. Chemical route: potassium cyanide (CAS 151-50-8) buffered with acetic acid, the standard textbook synthesis — a bench-stable surrogate, not neat HCN gas. Final total for this session: 8,041 of 18,558 reactions matched (43.3%), 3,906 decided (21.0%) — up from 2,815/1,724 (15.2%/9.3%) at the start.

  2. Build a solvent-lean chemical template (the fair fight). Both routes now have an effort dial — the chemical route's solvent recovery, and the enzymatic route's cofactor regeneration — and --fair-fight moves them together. Recycling is not free, and a template declares how it is paid for rather than the code assuming one shape: per_turnover measures (a co-substrate, charged every cycle, never bought down by recycling) and amortised ones (an immobilised enzyme and its carrier, divided by the batches one purchase serves — the number immobilisation exists to raise). The shipped class declares sucrose as per_turnover, at the amount a published cascade actually charges (Liu et al. 2021, CC BY: 4.196 mol sucrose per mole of product, not the theoretical 1:1, which understated it 4.2x in the enzyme's favour) and no immobilisation measure, because the bridge from enzyme-production GWP to enzyme mass per mole of product is not held from a document this repository has read. Screened at Liu's published 240 cofactor turnovers and industrial distillation's 90% recovery, all 451 reactions have a guaranteed saving — but one number still dominates: the template's 159 kg/mol ethyl-acetate isolation. Recovery divides that; it cannot un-choose it. Solvent recovery and solvent avoidance are different levers, and comparing a solvent-lean enzymatic route against a chemical route merely tidying up after a wasteful one is not the fair fight it looks like. The process_model mechanism built for coverage is also the fix here — a process-model isolation already charges 5 reaction volumes instead of the paper's ~30 — and running the fair fight against it rather than the paper is next.

  3. Add paper-sourced classes where a source exists (optional, and secondary to process_model coverage). Each needs one real published procedure, chosen by the same rule in every class, or the cross-class ranking measures how hard each paper's authors tried rather than the chemistry.

  4. Locate the commercial processes (Q3). Map already-commercialised enzymatic reactions — human-milk-oligosaccharide fucosylation (RHEA:14257 and others), anthocyanin glucosylation (RHEA:20093), β-arbutin (RHEA:12560) — onto Rhea ids and report their percentile in the guaranteed-floor ranking above.

How far coverage can actually go, and why not further

Say this plainly, because the numbers above invite the question directly: a class-template screen of Rhea cannot reach anywhere near comprehensive coverage, and that is a fact about the data, not about how much work gets done.

CORRECTION: the paragraph below originally claimed a 43.5% ceiling and called 30–40% the honest target. Both were wrong — an artefact of counting only sixteen cofactors chosen by convenience (CoA, acetyl-CoA, SAM, NAD(P)(H), ATP, ADP, the UDP-sugars, FAD, FMN, O₂, H₂O₂) rather than measuring true structural reachability. Re-derived properly: count every distinct species that appears on the LEFT side of any Rhea reaction, weighted by how many reactions it co-occurs with, excluding only the truly ubiquitous spectators (H₂O, H⁺, and the handful of universal metal ions) that carry no class-defining information. By that measure, 11.4% of Rhea (2,110 reactions) shares literally zero left-side participants with any other reaction — those are the only ones structurally unreachable by any class-template method, no matter how many classes get built. The real ceiling is 88.6%, not 43.5%.

The second error the original paragraph made was treating O₂ as untemplatable because it is "consumed by oxidations, oxygenations and radical chemistry with no shared chemical counterpart." That is true of O₂ alone, but O₂ is never used alone — every real O₂-consuming reaction also needs a specific electron donor (NAD(P)H, a reduced flavin, a reduced iron-sulfur cluster, 2-oxoglutarate, or none at all), and that co-cofactor's identity is a clean, mechanistic discriminator. This project has since built six classes on exactly that basis (o2-monooxygenase, p450-monooxygenase, o2-desaturase, o2-dioxygenase, 2og-dioxygenase, ferredoxin-monooxygenase), covering 2,931 O₂-consuming Rhea reactions between them and correctly separating one mechanism from another via required_co_cofactor_chebi / excluded_co_cofactor_chebi (see the UPDATE two sections above) rather than by EC prefix.

None of this means 80–100% is reachable soon, or that the remaining gap between today's 43.3% matched / 21.0% decided and the 88.6% structural ceiling is small or easy — it is a large, multi-year undertaking (roughly 800 more class templates would be needed to reach the tier where a cofactor is common enough, at 5+ reactions, to be worth templating at all). But it is not capped at 30–40% by the data itself, and reporting that cap as a hard limit was itself a fabricated number by this project's own standard, corrected here rather than left standing. What remains true from the original argument: reporting 80% coverage today by loosening the cofactor-plus-structural-check methodology would be exactly the kind of shortcut this project exists to refuse. The number that matters is still decided, not matched — an honestly verified 21.0% is worth more than a fabricated 80%, and closing the gap to 88.6% is a matter of building more classes the same rigorous way, not lowering the bar.


18,558 real enzymatic reactions. About one second. Zero invented numbers.

Most comparative-LCA tooling answers one question about one pair of routes, because building the background data for even one comparison is a day of work. carbonroute screen answers the same question — enzyme or chemistry? — across an entire curated reaction database, and does it fast enough that the database stops being the bottleneck.

Run against all 478 UDP-hexose-dependent glycosylation reactions in Rhea — every reaction that consumes UDP-glucose, its diastereomer UDP-galactose, or one of six sibling hexose-nucleotide donors — screened in about one second against a real published chemical procedure:

statistic solvent recovery threshold
minimum 84.56%
median 86.42%
maximum 91.54%

These three figures are upper bounds, computed with the enzyme held at 100% conversion. Pricing a real conversion pushes them lower, and at the 90% solvent recovery a plant achieves, 426 of the 451 lose their verdict outright. See "Q2, answered: the break-even frontier".

That threshold is the honest headline, not a win count. It is the chemical-route solvent recovery rate at which each reaction's verdict stops holding — and none of the 451 decided reactions survives 99% recovery, against the 90–95% real industrial distillation achieves. The finding is not "enzymes win 451 times." It is: in this class, the enzymatic advantage is real but bounded, and does not survive industrial solvent recycling — a falsifiable claim, not a marketing number.

What "about one second" does and doesn't cover. That's the runtime of scoring 478 reactions against a class template that already exists — pure arithmetic, no lookup. It is not the cost of producing that template. A template starts from one real, fully-quantified published chemical procedure, found and verified by actually reading the paper — the same discipline every row of data/factors/ is held to, and no faster than it. Screening scales for free with the number of reactions; it does not make the next class free to add — except when it does: UDP-galactose is 143 of the original 406 reactions added to this same class for zero new research, because it is glucose's C4 epimer (identical formula, identical mass delta), verified by RDKit rather than assumed from the name — and the same trick added 72 more (55 mannose, carried by GDP; the rest the same glucose, carried by ADP, GDP, dTDP or CDP) later. Picking which cofactor deserves that check next isn't guesswork either — see "Picking the next class by EC number" below for the method and the reactions across Rhea it already screens as tractable.

The mechanism the screen exists to measure shows up directly in the spread: the threshold rises with the number of groups on the substrate a chemical route would have to mask — 85.58% where there is one or none, 91.54% for a 34-group oligosaccharide. That is an enzyme's regioselectivity, quantified from molecular structure alone, at database scale.

Why 18,558 reactions is a bounded problem, not an unbounded one

Not a shortcut — the same cancellation that makes a single compare work, applied to a whole database. When two routes make the same product, everything they share cancels out of the difference. For enzyme-versus- chemistry, what cancels is exactly the expensive part to look up: the substrate and product, different in every reaction. What survives is small and repetitive:

side what survives the diff
enzymatic the cofactor — UDP-glucose, NADPH, SAM, acetyl-CoA
chemical the protecting groups, activator, base and solvents

That is a claim about data, so it was measured, not asserted. Rhea's 18,558 curated reactions involve 14,251 distinct chemical participants — but only 63 appear in 100 reactions or more, and 12 in over a thousand, and the top 30 alone cover 47.8% of every participant slot in the database. They are the cofactor list a biochemist would recite from memory. The long tail those 30 miss is precisely the per-reaction substrate and product — the part that cancels.

So the emission-factor work needed is bounded by that vocabulary — tens of substances — not by the reaction count. Every reaction after that costs one arithmetic evaluation.

Picking the next class by EC number, not by which cofactor is most common

The obvious next move — pick whichever cofactor has the most reactions — is wrong. By raw frequency, CoA (1,649 reactions) and SAM (954) both beat UDP-glucose. Both are also chemically mixed: CoA covers acetylations and Claisen condensations and redox steps that share nothing structurally, and a template built for one would be silently misapplied to the rest. Frequency measures population, not homogeneity.

The EC (Enzyme Commission) number is the fix, because it is already the field's own 60-year-old classification for "what transformation does this enzyme perform." Grouping Rhea's reactions at the third EC level (2.4.1, not the full 2.4.1.218) gives 252 tractable groups, and the four largest were checked directly — for each, does the mass a reaction adds to its acceptor actually cluster, the way UDP-glucose's does?

EC group reactions dominant cofactor mass-delta clusters coverage
1.1.1 (oxidoreductases) 526 NAD(+)/NADP(+) −2.0 (an oxidation) 100%
2.1.1 (methyltransferases) 500 SAM +14, +15, +28, +42 (mono/di/tri-methylation) 99.5%
2.3.1 (acyltransferases) 382 acetyl-CoA +41, +42 (acetylation) 99.3%
2.4.1 (glycosyltransferases) 400 UDP-glucose +162, +163, +324 (hexosylation) 100%

Every group collapses into 2–4 tight bins — the same signature-mass pattern the shipped class's expected_mass_delta check already polices, reproduced independently three more times. Restricting to an EC group is what turns "shares a cofactor" into "shares a cofactor and a transformation": CoA reactions in general are mixed, but CoA reactions restricted to EC 2.3.1 and its dominant participant land in two bins covering 99.3%. This is what made UDP-galactose a checkable, zero-research addition to the shipped class rather than a guess — and it is a concrete, reproducible answer for which class to build next, not just this one. It does not replace finding a real published procedure for a new class — mass-delta homogeneity confirms the check will work, not that a template exists yet — but it does mean that step no longer starts from a guess. Full data and method in docs/screening.md.

Proof this isn't a fantasy: calibration against a hand-built case

A screen pairs each curated enzymatic reaction against a class template: one real published chemical procedure, applied to every substrate in the class. That extrapolation is the method's central assumption, so a screen's output is a ranked shortlist of reactions worth a real compare — never a verdict about any one of them on its own.

One reaction in the screened class, RHEA:12560 (hydroquinone + UDP-α-D-glucose → β-arbutin), is also the subject of a fully hand-sourced, independently built ledger in examples/case-studies/beta-arbutin-chemical-vs-enzymatic/. The screen reproduces that ledger's product mass, acceptor, cofactor charge and verdict direction exactly, and a test asserts it. Without that agreement, the other 387 rows would only be reporting on their own template — with it, the template has been checked against real, independently sourced chemistry.

Full method in docs/screening.md; the measured cofactor vocabulary in data/rhea/README.md.

The building block: one comparison, decided despite the worst data of any case tried

Everything above rests on being able to decide a single comparison correctly even when the public data is thin. Here is that ability, on its own, on the case with the worst factor coverage this project has tried.

Grimaldi et al., ACS Sustainable Chem. Eng. 2021 compares two ibuprofen syntheses: a flow-chemistry route (bogdan) and a variant with one step replaced by an enzyme (enzymatic). Public factors resolve only 52.9% of the differing mass — nine materials stay unresolved, which would normally force an automatic indeterminate.

Ranking two routes, though, is an easier problem than measuring either one: you don't need to know what a missing factor is, only whether it's large enough to possibly flip which route is lower. carbonroute compare --bounds bounds.yaml supplies each unresolved material with a defensible interval instead of a value — a mass-balance argument, an already-held factor for a close relative, two disagreeing published estimates used as a floor and a ceiling — and checks whether the sign of GWP_A − GWP_B holds across every combination those intervals allow:

Decided: bogdan is lower than enzymatic everywhere in the asserted bounds.

material delta_mass kg/FU needs to be asserted bound clears it
1-butyl-3-methylimidazolium hexafluorophosphate -8.412 above 1.715 kgCO2e/kg [3.5, unbounded] yes
trimethyl orthoformate -1.666 any value — cannot flip it [0.27, unbounded] yes
phosphate buffer solution, 0.05 M -1.616 any value — cannot flip it [0.55, 2] yes
…5 more, every one of them any value — cannot flip it yes

Seven of the nine unresolved materials cannot change the outcome at any admissible value. The whole comparison reduces to one inequality about one ionic liquid — is its factor above 1.715 kgCO2e/kg? Two published estimates for the closest studied analogue disagree with each other by a factor of eight, and both still clear that threshold, by margins of 2.0× and 15.9×. The estimates don't agree on a value. They agree on the verdict, which is all a ranking needs.

That ionic liquid is a recoverable reaction solvent, so the result is conditional on recycling: it holds up to 51.0% recovery. The source paper reports its own results at 50% and 100% recycling scenarios and reaches the same qualitative conclusion — this project's number falls where the paper's does, from a fraction of the data and none of the commercial database the paper relies on.

A bound is never treated as a factor: it never enters the Monte Carlo, never contributes to a reported total, and never changes the coverage percentage. Full method in docs/bounds.md; every bound and its justification in examples/case-studies/ibuprofen-bogdan-vs-enzymatic/.

And when there really isn't enough data, it says so

The other half of being trustworthy is refusing to answer when the evidence doesn't support one — on the case above, carbonroute could have guessed and been wrong. Here it is doing exactly that, correctly, on a different real paper.

Sorgenfrei et al., J. Am. Chem. Soc. 2025, 147, 40944 (open access) computed the cradle-to-gate footprint of two real routes to the antiviral letermovir — Merck's industrial route and a novel route the authors designed — using ecoinvent, a commercial database this project cannot redistribute. They reported Merck at 382 kgCO₂e/kg and the new route at 369, a 3% gap in the new route's favor.

carbonroute can't see ecoinvent — only what's openly licensed, which is the situation almost anyone doing this work actually starts from. Watch what happens on the exact same two routes as public sources are added to the factor table:

Bar chart showing the share of the two routes' differing mass resolved to an emission factor rising from 8.9% to 75.5% across four stages of adding public data sources

At the first stage the tool would have been wrong — with only 8.9% of the differing mass resolved, the visible numbers leaned toward Merck being lower, the opposite of the published result. carbonroute refused to say so:

## Conclusion

**The comparison is undecided**, because only 8.9% of the differing mass
(2 of 43 materials) resolved to a factor, below the declared minimum of 80%.
No ranking is reported.

As real, citable sources were added — and as the tool learned to derive factors for chemicals no database had, from published production recipes (see docs/bootstrap.md) — coverage climbed to 75.5%, and the evidence flipped to agree with the paper. The verdict stays indeterminate, because 75.5% is still short of the 80% the tool requires before committing to a ranking:

Resolved part of the difference: 50.28 kgCO2e/FU.
Unresolved differing mass: -4.205 kg/FU (signed).

The ranking reverses if the 4.205 kg/FU of unresolved material averages
more than 11.96 kgCO2e/kg. Compare that against the factors you do have
before treating the ranking as settled.

That's the tool stating exactly how wrong the missing 24.5% would have to be to change the answer — a number you can check your intuition against, instead of a false sense of certainty. Full story in benchmarks/README.md.

A third real paper — a ZIF-8 metal-organic framework route (6.8% coverage) — hit the same wall for the same reason: public factor data for specialty solvents is thin, so the tool declines to rank rather than guess. See examples/case-studies/, including one candidate paper investigated and rejected because its own underlying data was AI/ML-modeled rather than measured.

Try it yourself in 30 seconds

Two invented routes to the same invented product — nothing here but what's already in this repository:

carbonroute compare examples/route.yaml \
  --a legacy --b denovo \
  --factors examples/factors_illustrative.csv
## Conclusion

**New route is very likely lower** than Published route (P > 0.9999).

Delta (Published route - New route): median 126.4 kgCO2e/FU,
90% interval [70.94, 237.3], 10000 draws, seed 20240101.

Every one of those 10,000 draws is a full Monte Carlo simulation of "if the real emission factors are anywhere in their plausible range, which route wins this time?" Here, all 10,000 agree:

Histogram of 10,000 Monte Carlo draws of the emissions difference between the two example routes, entirely on the side favoring the new route

Nothing here is production data — every factor in examples/factors_illustrative.csv is an obviously fake round number, which the tool itself flags loudly in the full report. It exists so you can see the whole pipeline run before you've sourced a single real number of your own.

How it works

flowchart LR
    L["route.yaml<br/>(the ledger)"] --> V["validate<br/>schema check"]
    V --> R["resolve<br/>factor tables + synonyms"]
    R --> D["diff<br/>shared materials cancel"]
    D --> M["Monte Carlo<br/>10,000 draws"]
    M --> G{"coverage of the<br/>differing mass ≥ 80%?"}
    G -->|no| I["indeterminate<br/>or decided from --bounds"]
    G -->|yes| O["ranking + P(A &lt; B)"]
Loading

Every result above is built from the same diff step. Two routes to the same product usually share a lot — the same solvent, the same reagent, sometimes the same catalyst — and none of it needs a citable emission factor, because it's identical on both sides of the subtraction:

flowchart LR
    subgraph A["Route A"]
        a1["toluene — 10 kg"]
        a2["water — 20 kg"]
        a3["catalyst X — 0.01 kg"]
    end
    subgraph B["Route B"]
        b1["toluene — 10 kg"]
        b2["water — 15 kg"]
        b3["catalyst Y — 0.02 kg"]
    end
    a1 --> cancel["same mass in both routes<br/>→ cancels exactly, contributes zero"]
    b1 --> cancel
    a2 --> delta1["Δ water = +5 kg<br/>→ this is what needs a factor"]
    b2 --> delta1
    a3 --> delta2["different catalysts<br/>→ both need a factor"]
    b3 --> delta2
Loading

(This is the idea behind DeltaLCA, applied here to synthetic routes instead of electronics hardware — and, via screen, to a whole reaction class at a time instead of one pair of routes.)

Everything a human has to decide — the electricity grid to assume, how much solvent gets recovered, the GWP time horizon, what counts as a statistical tie — lives in one place, the ledger's assumptions: block. Every step after that is deterministic: the same ledger and the same factor tables always produce the same numbers, down to the last decimal.

Install

Python 3.11 or later.

pip install -e .

Dependencies are limited to pydantic, numpy, PyYAML, and click. RDKit is an optional extra (pip install -e ".[chem]"), needed only for material identification and for carbonroute screen's structure-derived quantities (molecular weight, protectable-group count) — it is never required for validate, resolve, coverage, compare, bootstrap or lock.

The ledger

A route ledger is one YAML file. Its canonical shape is defined by src/carbonroute/schema.py (pydantic models) and mirrored in schemas/route-ledger.schema.json for external validation.

schema_version: "0.1"

assumptions:
  functional_unit: {mass_kg: 1.0, basis: product}
  boundary: cradle-to-gate
  grid_factor:
    id: JP-2024
    value_kgCO2e_per_kWh: 0.43
    source: "Analyst-declared placeholder; replace with the grid factor you can cite."
    uncertainty_class: assumption
  gwp_method: {name: IPCC-AR6, horizon_years: 100, feedbacks: false}
  solvent_recovery_default: 0.0
  waste_treatment: excluded
  monte_carlo: {iterations: 10000, seed: 20240101}
  indeterminate_band: {low: 0.4, high: 0.6}

routes:
  legacy:
    label: "Published route"
    steps:
      - id: 1
        yield: 0.82
        inputs:
          - {name: toluene, cas: "108-88-3", mass_kg: 12.0, role: solvent}
          - {name: "substrate A", cas: null, mass_kg: 1.0, role: reactant}
        electricity_kWh: 30.0
  denovo:
    label: "New route"
    steps: [...]

Points worth knowing:

  • mass_kg on an input is what was actually charged in that step. The tool scales it to the functional unit by dividing by the cumulative yield of every downstream step (spec section 7.1); you never do that arithmetic yourself.
  • role is one of solvent, reactant, reagent, catalyst, auxiliary, used for the contribution breakdown by role.
  • cas should be filled in whenever known — it is the primary join key against the factor table and against the same material appearing in a different route or step. A material without a CAS falls back to a normalized-name key, which is a weaker match and reported as such.
  • assumptions.solvent_recovery_default (and the per-material solvent_recovery override) sets how much of a solvent's charged mass is treated as make-up rather than fresh input (spec section 7.2). The default is 0 — no recovery is assumed unless you say so.
  • routes must be linear step lists. v0 has no way to express a route where two branches converge.

The complete route ledger is the only place assumptions are allowed to live. There is no other configuration surface for them.

The commands

Command What it does
carbonroute validate route.yaml Schema check only. No factor lookup, no computation.
carbonroute resolve route.yaml [--show-missing] Looks every material up in the factor table(s); reports what matched and what didn't. No emissions math.
carbonroute coverage route.yaml --a A --b B How much of the A-vs-B differing mass the loaded tables can actually reach, by count and by mass. Exits 3 if anything is unresolved.
carbonroute compare route.yaml --a A --b B [--bounds B.yaml] -o report.md The full comparison: diff, Monte Carlo ranking, reversal thresholds, and (with --bounds) a bounded verdict when factors alone don't reach 80% coverage.
carbonroute bootstrap --processes data/processes -o out.csv Derives factors for substances no open database covers, from production recipes — see docs/bootstrap.md.
carbonroute screen --template CLASS.yaml --bounds B.yaml Screens a whole reaction database against one chemical-route template, reporting the solvent recovery threshold per reaction — see docs/screening.md.
carbonroute lock route.yaml -o route.lock.json Pins the factor table versions, every resolved value and its provenance, and the RNG seed, so someone else can reproduce the exact numbers later.

resolve, coverage, compare, lock and bootstrap accept --factors PATH (repeatable; defaults to every CSV under data/factors/) and --synonyms PATH (defaults to every CSV under data/synonyms/, which maps the names a ledger uses onto identifiers — see docs/data.md). compare and lock accept --uncertainty PATH (defaults to the bundled config/uncertainty.yaml); resolve does not, because it never touches the uncertainty model. compare additionally takes --iterations, --seed and --no-thresholds. validate takes no options. A --fetch flag exists on resolve and compare for a future network-backed factor lookup; in v0 it exits with an error, because network access is off by default and there is no code path in this tool that opens a socket. The only side effect any command has is writing the file named by -o; without -o, output goes to stdout.

The full worked example

# 1. Structural check only.
carbonroute validate examples/route.yaml

# 2. See what resolves against the illustrative factor table, and what
#    would be missing if it were the only table available.
carbonroute resolve examples/route.yaml \
  --factors examples/factors_illustrative.csv \
  --show-missing

# 3. Compare the two routes and write a Markdown report.
carbonroute compare examples/route.yaml \
  --a legacy --b denovo \
  --factors examples/factors_illustrative.csv \
  -o report.md

# 4. Pin the exact factor values, versions, and RNG seed used, for
#    someone else to reproduce.
carbonroute lock examples/route.yaml \
  --factors examples/factors_illustrative.csv \
  -o route.lock.json

report.md opens with a statement that the result is not an ISO 14067-conformant calculation, the full text of the assumptions applied, the factor table versions and their SHA-256 hashes, and the provenance breakdown of the resolution — in that order — before it states any conclusion. Because every row in examples/factors_illustrative.csv is marked ILLUSTRATIVE, the report also carries a prominent warning that its conclusion is not usable for anything beyond exercising the pipeline. The conclusion itself is always a ranking and a probability (P(GWP_legacy < GWP_denovo), a verdict of "A<B" / "B<A" / "indeterminate", and the median and 90% interval of the difference) — never a single absolute footprint presented as the headline result.

What this tool does not do

This list is deliberate, not an oversight, and it is unlikely to shrink quickly (see docs/spec-ja.md section 5.2 and docs/limitations.md). v0 does not:

  • Handle convergent routes. Only linear step sequences are supported. A route where two synthesis branches merge into one cannot be expressed in the ledger schema at all.
  • Extrapolate from lab scale to plant scale on its own. Inputs are taken at face value from the ledger; there is no automatic model for how solvent use, heating efficiency, or yield change between a bench reaction and an industrial process. --bounds and solvent_recovery let you test that sensitivity explicitly, but nothing is assumed for you.
  • Model the use phase or end-of-life. The boundary is fixed at cradle-to-gate. Nothing downstream of the product leaving the gate is in scope.
  • Cover impact categories beyond climate change. The only output quantity is GWP (kg CO2e). Water use, toxicity, land use, and every other ISO 14044 impact category are out of scope.
  • Use a language model anywhere in the calculation path. No factor value, no resolution decision, and no number in a report is ever produced or adjusted by a language model. This is a deliberate response to measured unreliability of general-purpose LLMs on LCA-adjacent tasks (arXiv:2510.19886 found 37% of answers across 11 models and 22 LCA tasks contained inaccurate or misleading content, with fabricated-citation rates up to 40% for some models).
  • Generate routes by retrosynthesis. carbonroute compares routes you give it, or reactions a curated database already documents; it does not propose or search for candidate syntheses of its own invention.

And separately from the list above: the output of this tool is not an ISO 14067-conformant product carbon footprint. It is a screening comparison meant to help decide which route deserves a full assessment, not a substitute for one. Every report says so explicitly, as required by spec section 9.

What is in data/factors/

data/factors/ ships real emission factors, every one of them fetched from an openly licensed source by a script in scripts/ that you can re-run to regenerate the table. Each row names the dataset and the record it came from, the licence it is distributed under, the date it was retrieved, and — where the source published one — its own uncertainty.

At the time of writing that means 27 substances from five sources: ADEME's Base Carbone (Licence Ouverte), the US LCI Database (US government work), ProBas/GEMIS (German environment agency, free for all users and uses), figures published directly by producers and industry associations (PlasticsEurope eco-profiles, a Nobian EPD), and carbonroute bootstrap itself, deriving factors for chemicals none of those databases cover from cited production recipes (2-MeTHF, ethyl acetate, isopropyl acetate, acetone, DMF, MTBE, triethylamine and more — see docs/bootstrap.md).

Several of those substances carry two or more independent public values, and they don't always agree — hydrochloric acid, for instance, is 1.199 kgCO2e/kg by one source and 1.700 by another. Reports print every value in play. How far openly available data spreads for the same material is one of the things worth knowing here, not something to average away.

27 factors is deliberately not the ceiling on what this tool can decide. --bounds and screen's cofactor-vocabulary approach exist precisely because coverage this small still resolves real comparisons, as shown above. carbonroute coverage tells you exactly how far your own factors-only comparison is from 80%, so the gap is a number in front of you rather than a silent omission. Adding a source means writing another ingestion script — see docs/data.md and docs/sources-investigated.md, which records what was already checked and why it was or wasn't used. Who else holds this kind of data and why most of it sits in commercial databases this project may not redistribute is covered in docs/what-others-do.md, including how to point carbonroute at a licensed table if you have one.

Nothing here is estimated, interpolated, or recalled from memory. That is a consequence of the "public data only, nothing invented" rule (spec sections 2 and 13): a table of plausible-looking numbers nobody can check would defeat the entire purpose of the tool.

examples/factors_illustrative.csv exists purely so the pipeline can be run end to end. Every value in it is an obviously-fake round number, every row's source column starts with ILLUSTRATIVE, and any report built from it says so prominently. Do not cite it, and do not use it for anything but exercising the commands above.

Reproducing everything without a live API

carbonroute itself never touches a network — enforced by a test that parses its import graph, not just claimed in prose. The scripts that built data/factors/ and data/rhea/ used to require one, though, and every one of ADEME's, PubChem's, ProBas's, the Federal LCA Commons' and Rhea's APIs is outside this project's control. Each ingestion script now has a --offline flag that replays from a durable, committed snapshot under data/raw/ instead of touching the network — see docs/reproducibility.md for which snapshot covers which source, and its one known gap.

The letermovir benchmark's entire empirical basis — a small, CC BY licensed Excel workbook — is committed at benchmarks/letermovir/source-material/ for the same reason: scripts/extract_letermovir_ledger.py --offline, with no arguments at all, reproduces benchmarks/letermovir/ledger.yaml byte-for-byte using only files already in this repository.

Benchmarks

Two test sets, both with their acceptance conditions written before the assertions (see benchmarks/README.md for the full account of both).

B1, the analytic case, is small enough to check by hand. It pins the functional-unit conversion, solvent make-up, the exact cancellation of materials common to both routes, and bit-for-bit reproducibility from a seed.

B2, the letermovir comparison is the one demonstrated above. It exists because a benchmark written before the calculation code catches things a benchmark written after cannot: before it existed, an unresolved material was silently treated as worth zero, and on 8.9% of the differing mass the tool reported P > 0.9999 for the ranking opposite the one the paper published. The coverage floor and the break-even calculation shown above both exist because this benchmark ran first and failed.

Further reading

  • docs/screening.md — screening a whole reaction database, and why that costs less than it sounds.
  • docs/research-brief.md — the four numbers this repository does not yet hold from a document it has read, written as a ready-to-run brief for anyone with library access.
  • docs/bounds.md — deciding a ranking from bounds when the factors themselves are missing.
  • docs/data.md — the factor-table format and how to build a table you can cite.
  • docs/bootstrap.md — deriving factors from production recipes when no database has them.
  • docs/uncertainty.md — how the Monte Carlo model works and the status of its parameters.
  • docs/convergence.md — how many Monte Carlo iterations are enough, and what has not yet been checked.
  • docs/limitations.md — what this tool can and cannot be expected to get right.
  • docs/reproducibility.md — the non-API route: reproducing every factor table without live network access.
  • docs/what-others-do.md — what industry LCA tools do instead, and how to plug a licensed database into this one.
  • docs/spec-ja.md — the full design specification (Japanese), the authority on intent for everything above.
  • docs/internal-api.md — the module-level contract for contributors.

License

Apache License 2.0. See LICENSE. Code and data have separate licensing: the code in src/ is Apache-2.0; any factor table you add to data/factors/ carries whatever license its own source imposes, tracked per row (see docs/data.md).

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages