Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 10 additions & 10 deletions CITATION.cff
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@ cff-version: 1.2.0
message: "If you use this software, please cite it as below."
type: software
title: "cybedtools: A Package for Reproducible Analysis of Cybersecurity Workforce and Learning Frameworks"
version: 0.4.0
version: 0.4.1
# Zenodo concept DOI: the parent identifier that resolves to the latest
# released version (currently 10.5281/zenodo.20076116, verified against
# the Zenodo project's "Cite all versions" link). This DOI does not
Expand Down Expand Up @@ -34,15 +34,15 @@ keywords:
- cybersecurity education
- competency mapping
abstract: >-
cybedtools provides a reproducible R pipeline for ingesting, semantically
representing (via a two-tier JSON-LD schema), and querying (via SPARQL)
cybersecurity-related competency and learning-standards frameworks,
including NICE, DCWF, SFIA, ENISA ECSF, Cyber.org K-12, CSTA K-12 CS,
ACM/IEEE CSEC2017, DigComp 3.0, CyQUAL, CCSSF, and OTCCF. See
framework_summary() in the package for the current list and counts. The
cybed: vocabulary uses
cybedtools is an R package and a Python package (on PyPI), sharing one
API and a shared conformance-test suite, for building and querying a
cross-framework RDF/JSON-LD graph of cybersecurity workforce and
learning frameworks. The two-tier cybed: vocabulary uses
cybed:OrganizingUnit as the cross-framework abstract for parent units
and cybed:Role as a workforce-restricted subtype, with cybed:Subpoint
and cybed:Example separating framework-as-specified enumerations from
pedagogical-scaffolding examples. Includes a layered data-integrity
verification rig and per-framework provenance manifests.
pedagogical-scaffolding examples. Per-framework data releases are
fetched and hash-verified with cybed_fetch(); see framework_summary()
in the package for the current list of frameworks and counts. Includes
a layered data-integrity verification rig and per-framework provenance
manifests.
17 changes: 9 additions & 8 deletions DESCRIPTION
Original file line number Diff line number Diff line change
Expand Up @@ -2,20 +2,21 @@ Package: cybedtools
Type: Package
Title: A Package for Reproducible Analysis of Cybersecurity Workforce
and Learning Frameworks
Version: 0.4.0
Version: 0.4.1
Authors@R: person(
"Ryan", "Straight",
email = "ryanstraight@arizona.edu",
role = c("aut", "cre", "cph"),
comment = c(ORCID = "0000-0002-6251-5662")
)
Description: A reproducible pipeline for ingesting, representing, and
analyzing cybersecurity workforce competency frameworks ('NICE',
'DCWF', 'SFIA', 'ENISA ECSF', 'CyQUAL', 'CCSSF', 'OTCCF') and
cybersecurity-related pedagogical frameworks ('Cyber.org K-12',
'CSTA K-12 CS', 'ACM/IEEE CSEC2017',
'DigComp 3.0'). Provides a framework-agnostic 'JSON-LD' schema
('cybed:') and 'SPARQL' query patterns for cross-framework analysis.
Description: Builds and queries a cross-framework 'RDF'/'JSON-LD' graph
of cybersecurity workforce and learning frameworks, with a shared
'API' and conformance tests against a companion 'Python' package
(on 'PyPI'). Per-framework data releases are fetched and
hash-verified with cybed_fetch(). Provides a framework-agnostic
'JSON-LD' schema ('cybed:') and 'SPARQL' query patterns for
cross-framework analysis; see framework_summary() for the current
list of frameworks and counts.
License: MIT + file LICENSE
Encoding: UTF-8
Language: en-US
Expand Down
13 changes: 13 additions & 0 deletions NEWS.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,16 @@
# cybedtools 0.4.1

First-time-user defect fixes found against the published v0.4.0 / PyPI 0.1.0, fixed identically in R and Python.

- **CSEC2017 held from the 2026.09.2 data release** (owner correction, on this branch before the fixes below): `docs/data-release.yml` moves `csec2017` from `shipped` to `held`, pending steward confirmation.
- **Slug alias resolution.** `cybed_license()` and `framework_similarity()` now accept either the versioned framework slug (e.g. `"nice-v2"`) or the short release-file slug (e.g. `"nice"`), matching `cybed_fetch()`'s existing behavior. The alias map is derived once, in `R/slug-alias.R` (`cybedtools/_slug_alias.py` in Python), from the shipped `framework_summary`/`framework_licenses` tables, reusing the short-slug rule `scripts/030-export-release.R`'s `release_license_row()` applies. `cybed_fetch()`'s result now carries both `framework_slug` (canonical, versioned) and `release_slug` (short) columns explicitly, rather than silently substituting one for the other.
- **`framework_similarity()` errors on an unknown `from`/`to` slug** with class `cybedtools_framework_not_found` (a specific exception in Python), listing the graph's known slugs, instead of silently returning an empty result.
- **Python `cybed_fetch()`/`load_graph()` accept a bare string** for `frameworks` as one slug, not an iterable of its characters.
- **`version` accepts a `data-v`-prefixed value** (e.g. `"data-v2026.09.2"`, matching a GitHub release tag) in both `cybed_fetch()` implementations; the leading `data-v` is stripped.
- **README and quick-start docs lead with the user path**: install, then `cybed_fetch()`/`load_graph()`, then query examples. The ingestion pipeline (`scripts/000-build.R`) moves to a separate "Rebuilding the graph from source (maintainers)" section, consistent across `README.qmd`, `concordance/start/install.qmd`, and the `getting-started` vignette. The Zenodo data-release DOI (`10.5281/zenodo.22884320`) is now named alongside the data release.
- **Concordance framework pages' footer `LICENSE`/asset links** are root-relative, fixing a 404 from any page under a subdirectory (e.g. `frameworks/`).
- `CITATION.cff` and `DESCRIPTION`'s `Description` no longer describe cybedtools as an R-only pipeline or hand-list frameworks; both now point to `framework_summary()` for the current list and note the shared R/Python API.

# cybedtools 0.4.0

## New framework
Expand Down
37 changes: 28 additions & 9 deletions R/cybed-fetch.R
Original file line number Diff line number Diff line change
Expand Up @@ -129,15 +129,32 @@ cybed_download <- function(url, destfile) {
#' Already-cached files whose hash still matches the manifest are not
#' re-downloaded.
#'
#' @param frameworks Character vector of framework slugs to fetch (as
#' carried by [framework_summary]`$framework_slug`, e.g. `"nice-v2"`, or the release file slug, e.g. `"nice"`), or
#' `NULL` (the default) for every framework the release manifest ships.
#' @param version Character scalar release version (e.g. `"1.0.0"`), or
#' `NULL` (the default) for the data release this package version was built against, `cybed_data_release`.
#' @section Two slug vocabularies:
#' Two framework-slug vocabularies exist in this package. [framework_summary]
#' and [framework_licenses] use **versioned** slugs (`"nice-v2"`,
#' `"otccf-v1.1"`), minted into every framework node's IRI. The public data
#' release's files use **short** slugs (`"nice"`, `"otccf"`), assigned by
#' `docs/data-release.yml` and recorded in the release manifest. This
#' function, [cybed_license()], and [framework_similarity()] accept either
#' form for a `frameworks`/`slug`/`from`/`to` argument and resolve it to the
#' canonical versioned slug. The returned tibble names both forms explicitly
#' (`framework_slug` and `release_slug`) rather than silently substituting
#' one for the other.
#'
#' @param frameworks Character vector of framework slugs to fetch, either the
#' versioned form carried by [framework_summary]`$framework_slug` (e.g.
#' `"nice-v2"`) or the short release-file slug (e.g. `"nice"`; see "Two
#' slug vocabularies" above), or `NULL` (the default) for every framework
#' the release manifest ships.
#' @param version Character scalar release version, either `"2026.09.2"` or
#' `"data-v2026.09.2"` (a leading `"data-v"`, matching a GitHub release
#' tag, is stripped), or `NULL` (the default) for the data release this
#' package version was built against, `cybed_data_release`.
#' @return Invisibly, a tibble with one row per fetched framework: columns
#' `framework_slug`, `path` (the cached file's local path),
#' `sha256_verified` (logical, always `TRUE` on return -- a mismatch
#' aborts instead of returning `FALSE`).
#' `framework_slug` (the canonical versioned slug, e.g. `"nice-v2"`),
#' `release_slug` (the short release-file slug, e.g. `"nice"`), `path`
#' (the cached file's local path), `sha256_verified` (logical, always
#' `TRUE` on return -- a mismatch aborts instead of returning `FALSE`).
#' @family data loading
#' @export
#' @examples
Expand All @@ -152,6 +169,7 @@ cybed_download <- function(url, destfile) {
cybed_fetch <- function(frameworks = NULL, version = NULL) {
base_url <- cybed_release_base_url()
version <- if (is.null(version)) cybed_data_release else version
version <- sub("^data-v", "", version)
manifest <- cybed_read_manifest(base_url, version)

files <- manifest$files
Expand Down Expand Up @@ -219,7 +237,8 @@ cybed_fetch <- function(frameworks = NULL, version = NULL) {
}

tibble::tibble(
framework_slug = slug,
framework_slug = if (is.null(entry$license_slug)) slug else entry$license_slug,
release_slug = slug,
path = dest,
sha256_verified = TRUE
)
Expand Down
5 changes: 5 additions & 0 deletions R/data.R
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,11 @@
#' (workforce vs pedagogy), and license are hand-curated because they
#' originate outside the JSON-LD graph.
#'
#' `framework_slug` is the **versioned** slug vocabulary. A second, **short**
#' vocabulary is used by the public data release's files (e.g. `"nice"` for
#' `"nice-v2"`); see the "Two slug vocabularies" section of [cybed_fetch()]
#' for the full mapping and which functions accept either form.
#'
#' Several frameworks were added in v0.3.0 on steward terms: CyQUAL
#' (Czech Republic, open data, attribution to CyQUAL and Masaryk
#' University), CCSSF (Canada, Government of Canada copyright, used with
Expand Down
15 changes: 12 additions & 3 deletions R/licenses.R
Original file line number Diff line number Diff line change
Expand Up @@ -35,9 +35,11 @@ license_table <- function() {
#' says which of the two a given row carries.
#'
#' @param slug Character scalar, or `NULL`. Either `"cybedtools"` for the
#' package's own code, or a framework slug as carried by
#' package's own code, a framework slug as carried by
#' [framework_summary]`$framework_slug` (for example `"nice-v2"`,
#' `"otccf-v1.1"`). `NULL`, the default, returns every row.
#' `"otccf-v1.1"`), or the short release-file slug (e.g. `"nice"`,
#' `"otccf"`) documented on [cybed_fetch()]. Either form resolves to the
#' same row. `NULL`, the default, returns every row.
#' @return A tibble. One row per licence layer when `slug` is `NULL`, and a
#' single-row tibble otherwise.
#' @family licensing
Expand Down Expand Up @@ -68,7 +70,14 @@ cybed_license <- function(slug = NULL) {
)
}

row <- licenses[licenses$slug == slug, , drop = FALSE]
resolved <- if (slug %in% licenses$slug) {
slug
} else {
alias_map <- framework_slug_alias_map()
if (slug %in% names(alias_map)) unname(alias_map[[slug]]) else slug
}

row <- licenses[licenses$slug == resolved, , drop = FALSE]

if (nrow(row) == 0L) {
rlang::abort(
Expand Down
12 changes: 10 additions & 2 deletions R/similarity-helpers.R
Original file line number Diff line number Diff line change
Expand Up @@ -214,8 +214,12 @@ similarity_strength <- function(x) {
#' @param rdf An rdf object.
#' @param from,to Character scalars, the `framework_slug` values (from
#' [organizing_unit_framework_bindings()]) whose organizing units are
#' compared. May be identical, to find near-duplicate units within one
#' framework.
#' compared, either as the versioned slug (`"nice-v2"`) or the short
#' release-file slug (`"nice"`; see [cybed_fetch()]). May be identical, to
#' find near-duplicate units within one framework. An unknown slug (in
#' either vocabulary, and not present in `rdf`) errors with class
#' `cybedtools_framework_not_found` rather than silently returning an
#' empty result.
#' @param n Integer, matches to keep per from-unit (default `5`).
#' @return A tibble with one row per (from unit, match): columns
#' `from_unit`, `to_unit` (both full IRIs), `score` (numeric in `[0, 1]`),
Expand All @@ -242,6 +246,10 @@ framework_similarity <- function(rdf, from, to, n = 5) {
stopifnot(is.character(to), length(to) == 1L)

units <- organizing_unit_framework_bindings(rdf)
known_slugs <- unique(units$framework_slug)
from <- resolve_framework_slug(from, known_slugs, arg = "from")
to <- resolve_framework_slug(to, known_slugs, arg = "to")

texts <- element_text(rdf)

child_text <- unit_element_bindings(rdf) |>
Expand Down
69 changes: 69 additions & 0 deletions R/slug-alias.R
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
# R/slug-alias.R
#
# Two framework-slug vocabularies exist in this package (documented in full
# on cybed_fetch() and framework_summary()):
#
# - versioned slugs, e.g. "nice-v2", "otccf-v1.1" -- carried by
# framework_summary()$framework_slug and framework_licenses()$slug, and
# minted into every framework node's IRI in the graph.
# - release (short) slugs, e.g. "nice", "otccf" -- the names of the public
# per-framework release files, assigned by docs/data-release.yml and
# recorded in the release manifest's `slug` field (with the versioned
# form alongside it as `license_slug`).
#
# cybed_fetch() already resolves either form because it reads the release
# manifest at call time. Every OTHER public function that takes a framework
# slug (cybed_license(), framework_similarity()) has no manifest to consult,
# so this file derives the same short-slug rule
# `scripts/030-export-release.R`'s `release_license_row()` applies (strip a
# trailing version suffix; a framework that is its own edition, not a
# version of a shorter-named sibling, keeps its full slug) from the shipped
# `framework_summary`/`framework_licenses` tibbles alone, with no network or
# build-time input. This is the ONE place that map is built; both
# `cybed_license()` and `framework_similarity()` call it.

#' Alias map from release (short) slug to canonical versioned slug
#'
#' @return A named character vector: `names()` are release/short slugs,
#' values are the canonical versioned slugs.
#' @noRd
framework_slug_alias_map <- function() {
versioned <- cybedtools::framework_summary$framework_slug
short <- sub("-v?[0-9][0-9.]*$", "", versioned)

# csta-2026 is a framework in its own right, not an edition of csta --
# csta-2017 already claims the "csta" short slug (mirrors the `reserved`
# handling in scripts/030-export-release.R's release_license_row()).
short[versioned == "csta-2026"] <- "csta-2026"

stats::setNames(versioned, short)
}

#' Resolve a framework slug (either vocabulary) against a set of known
#' (already-versioned) slugs
#'
#' @param slug Character scalar to resolve.
#' @param known Character vector of valid versioned slugs to resolve against.
#' @param arg Character scalar, the argument name to use in error messages.
#' @return The canonical versioned slug, if resolvable.
#' @noRd
resolve_framework_slug <- function(slug, known, arg = "slug") {
if (slug %in% known) {
return(slug)
}

alias_map <- framework_slug_alias_map()
if (slug %in% names(alias_map) && unname(alias_map[[slug]]) %in% known) {
return(unname(alias_map[[slug]]))
}

rlang::abort(
c(
"Unknown framework slug.",
"x" = paste0("`", arg, "`: '", slug, "'."),
"i" = paste0("Known slugs: ", paste(sort(known), collapse = ", "), ".")
),
class = "cybedtools_framework_not_found",
framework_slug = slug
)
}
102 changes: 92 additions & 10 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -892,15 +892,97 @@ vignette. Extending the schema to a new framework is covered in

## Getting started

Install the package, then fetch the public data release: hash-verified
per-framework files, downloaded once and cached locally. No source
framework text needs to be staged for this path.

### R

``` r
# install.packages("pak")
pak::pak("ryanstraight/cybedtools")
```

### Python

``` sh
pip install cybedtools
```

### Fetch data and load a graph

R:

``` r
library(cybedtools)

# Downloads and hash-verifies every shipped framework's release file into
# the R user cache directory, then parses them into one rdf object.
rdf <- load_graph()

framework_summary
```

Python:

``` python
from cybedtools import load_graph, framework_summary

graph = load_graph()
framework_summary()
```

`load_graph()` calls `cybed_fetch()` internally (a no-op for files
already cached with a verified hash). Both accept either the versioned
framework slug (`"nice-v2"`) or the short release-file slug (`"nice"`);
see `cybed_fetch()`’s documentation for the two-vocabulary mapping. The
data release backing this package version is archived on Zenodo:
[10.5281/zenodo.22884320](https://doi.org/10.5281/zenodo.22884320).

### Three queries

R:

``` r
# install.packages("remotes")
remotes::install_github("ryanstraight/cybedtools")
# 1. What's in the corpus?
framework_summary |>
dplyr::select(framework_slug, framework_name, jurisdiction, organizing_unit_count)

# 2. Which elements does an organizing unit carry?
unit_element_bindings(rdf) |>
head(10)

# 3. Where do two frameworks say the same thing?
framework_similarity(rdf, from = "nice", to = "ecsf", n = 3)
```

Python:

``` python
# 1. What's in the corpus?
framework_summary()[["framework_slug", "framework_name", "jurisdiction", "organizing_unit_count"]]

# 2. Which elements does an organizing unit carry?
from cybedtools import unit_element_bindings
unit_element_bindings(graph).head(10)

# 3. Where do two frameworks say the same thing?
from cybedtools import framework_similarity
framework_similarity(graph, from_="nice", to="ecsf", n=3)
```

The package **does not** redistribute source framework text. To run the
pipeline end-to-end, clone the repository and stage each framework’s
source file at `data/raw/<framework>/` per
[`docs/framework-data-sources.md`](docs/framework-data-sources.md):
The [function
reference](https://ryanstraight.github.io/cybedtools/reference/) indexes
the full public API for both languages.

## Rebuilding the graph from source (maintainers)

Fetching the public release (above) is the path for using the corpus.
Rebuilding it from primary sources is a separate, maintainer-only path:
the package does not redistribute source framework text, so this
requires cloning the repository and staging each framework’s source file
at `data/raw/<framework>/` per
[`docs/framework-data-sources.md`](docs/framework-data-sources.md).

``` sh
git clone https://github.com/ryanstraight/cybedtools
Expand All @@ -910,11 +992,11 @@ cd cybedtools
Rscript scripts/000-build.R # ingestion + verification + assembly + export
```

The
`scripts/000-build.R` is the pipeline’s entry point: it runs ingestion,
verification, JSON-LD assembly, and N-Triples export in order (see
`scripts/README.md` for the full stage list). The
[`getting-started`](https://ryanstraight.github.io/cybedtools/articles/getting-started.html)
vignette walks through each stage, and the [function
reference](https://ryanstraight.github.io/cybedtools/reference/) indexes
the public API.
vignette walks through each stage.

## Citing

Expand Down
Loading
Loading