This guide describes the runnable GS-mark configuration. Architecture sizes, loss weights, learning rate and attack levels follow the paper. Dataset releases, splits, subgraph sampling, optimization schedules and the concrete forgery attack are specified here to complete details left open in the paper.
| Configuration | Public release | Nodes | Features | Classes |
|---|---|---|---|---|
| computers.yaml | gnn-benchmark Amazon Computers | 13,752 | 767 | 10 |
| cs.yaml | gnn-benchmark Coauthor CS | 18,333 | 6,805 | 15 |
| flickr.yaml | GraphSAINT Flickr | 89,250 | 500 | 7 |
prepare downloads these releases under data/raw/ and writes processed data
under data/processed/. The manifest records source URLs, SHA-256 checksums,
normalization and split settings. Existing caches are checked before reuse.
Dataset files are obtained from their providers; see data attribution.
Computers and CS use binary bag-of-words features. All three datasets use row L1 normalization, dividing each feature row by its sum of absolute values; zero rows remain zero. Adjacency is symmetrized and binarized, with explicit self-loops removed. Each GCN adds its own self-loops during normalization.
Computers and CS use a class-stratified 60/20/20 train/validation/test node split
with seed 42. Flickr uses its released role.json. The three node sets are
disjoint. No edge crossing a partition is used to construct its examples.
The validation portion is divided into two fixed, disjoint pools: half for
checkpoint selection and half for null calibration. Thus calibration nodes are
excluded from both gradient updates and checkpoint selection.
Each example is an induced subgraph of up to 128 nodes. Sampling starts from a random node and visits neighbors in randomized breadth-first order, restarting when a connected component is exhausted. Nodes are unique within an example. Smaller partitions are padded and masked. Features stay sparse on disk where supported and are expanded only for the sampled batch.
Training, checkpoint validation, calibration and testing draw from their respective node partitions. Examples within a partition can overlap. The evaluation unit is a sampled graph signal; the resulting measurements describe this subgraph protocol.
The source datasets provide node labels, so the downstream task is node classification. A two-layer GCN is trained on clean training examples and then frozen. Its parameters stay fixed during watermark training while its input gradients supply the functional preservation loss.
| Setting | Default |
|---|---|
| GCN hidden width | 128 |
| Latent, watermark and key dimensions | 64 |
| Optimizer | Adam, learning rate 0.0002 |
| Batch size | 8 graph signals |
| Classifier pretraining | 50 epochs, 16 batches per epoch |
| Watermark training | 200 epochs, 16 batches per epoch |
| Validation | Every 5 epochs, 8 fixed batches |
| Gradient clipping | Norm 5 for generator and extractor |
| Loss weights: adversarial / functional / perturbation / KL / decoding | 1 / 0.5 / 20 / 0.1 / 4 |
Fresh uniformly random watermark bits are sampled for each training example.
The watermark projection maps (2*w-1)/sqrt(64) through a fixed orthogonal
linear map. Each run uses one private Gaussian key, stored in owner-key.npy.
Pass train --key path/to/key.npy to use a previously generated key. Retain the
key alongside the run when repeating or resuming it.
Training alternates one discriminator update with one joint generator/extractor update. The discriminator sees detached watermarked signals. For the joint update, its parameters are frozen and gradients flow through it to the signal. Training samples the variational posterior; evaluation uses its mean.
Loss reductions are explicit:
- Perturbation: sum squared changes over valid nodes and features in each graph, then average graphs, matching the squared Frobenius form.
- KL: sum latent dimensions, average valid nodes within each graph, then average graphs.
- Functional: cross-entropy averaged over valid labeled nodes.
- Decoding: mean
1 - cosine(extracted, target)over graphs. - Adversarial: mean non-saturating generator loss, with real graphs labeled 1.
The best checkpoint minimizes the total validation objective. Test samples do
not select the checkpoint. last.pt stores model, optimizer and random-generator
states for --resume; checkpoint.pt stores the best validation checkpoint.
Alternative decoder and loss reductions are explicit configuration options.
Evaluation samples one target watermark and reuses it across conditions. Each condition uses 256 null examples from the calibration pool and 128 separate test examples. Positives are watermarked test signals; negatives are their unwatermarked counterparts. Calibration and test examples receive the same type and level of corruption. The default evaluation seed is 2026.
| Condition | Implementation |
|---|---|
| Clean | Original topology and unmodified watermarked signals |
| Node deletion | Uniformly delete floor(rate * valid_nodes) nodes and their incident edges, at rates 10%, 30%, or 50% |
| Edge perturbation | Uniformly delete floor(rate * undirected_edges) edges at rates 10%, 30%, or 50%; both symmetric entries change together |
| Signal noise | Add Gaussian noise with absolute standard deviation 0.1, 0.3, or 0.5 to the normalized signal |
| Isomorphism | Apply the same node permutation to signal rows, adjacency rows/columns and masks |
| PGD invalidation | Minimize the claimed watermark's cosine score with 20 sign-gradient steps, step size 0.005 and an L-infinity budget of 0.05 |
Edge deletion, absolute noise scale and the PGD norm/step size make the paper's partially specified attacks concrete. PGD starts at the observed signal and projects back into its perturbation ball after each step. It leaves padding and adjacency unchanged. The training objective uses the paper's discriminator game; attacks are applied during evaluation.
Detection reports tie-aware ROC AUC and TPR/FPR at the calibrated significance
level. For a score s, the upper-tail Monte Carlo p-value is
(1 + count(null >= s)) / (M + 1), with the decision p < alpha and default
alpha = 0.01. A threshold is calibrated separately for each attack condition.
When the null sample count cannot resolve the requested alpha, no detection is
accepted and alpha_resolvable is false. Calibration applies to the specified
watermark, extractor, graph distribution and corruption condition.
The forgery evaluation starts from an unwatermarked test signal and ascends the public extractor's cosine score for the claimed watermark under the same PGD budget. It reports the fraction accepted against the clean calibration null, along with the signal MSE. This defines forgery success directly; the paper does not specify the labels or attack construction behind its forgery AUC.
Transparency uses clean and watermarked node accuracy from the same frozen task model on the same sampled test nodes. Accuracy drop is their difference; MSE is the mean squared change per valid node-feature element. Accuracies and their difference are fractions; multiply the difference by 100 for percentage points.
Run all three configurations on separate devices:
python scripts/run_main.py --devices cuda:0 cuda:1 cuda:2 --seeds 0Use --seeds 0 1 2 for multiple training runs or --resume to continue saved
runs. A single device processes the selected jobs sequentially. Individual
configurations can also be run with the commands in the README.
Each run writes its configuration, dataset manifest, private key, checkpoints
and training log under results/<dataset>/seed-<seed>/. Evaluation writes
metrics.json, raw scores.npz, the target watermark and clean null scores in
that run's evaluation/ directory. These generated files are ignored by Git.
For a short synthetic workflow check without dataset downloads:
python -m gsmark train --config configs/smoke.yaml --seed 0
python -m gsmark evaluate --config configs/smoke.yaml \
--checkpoint results/smoke/seed-0/checkpoint.pt