-
Notifications
You must be signed in to change notification settings - Fork 2
Expand file tree
/
Copy pathCITATION.cff
More file actions
90 lines (83 loc) · 3.22 KB
/
Copy pathCITATION.cff
File metadata and controls
90 lines (83 loc) · 3.22 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
cff-version: 1.2.0
title: "Medical Document Anonymizer (deidentifier)"
message: >-
If you use this software in your research, please cite it using the metadata
from this file. This is a follow-up implementation and not the code evaluated
in the LLM-Anonymizer paper (Wiest et al., NEJM AI 2025, listed under
"references"); please cite that paper as well when referring to the
underlying approach.
type: software
authors:
- family-names: "Wolf"
given-names: "Fabian"
affiliation: >-
Else Kroener Fresenius Center for Digital Health, Faculty of Medicine
and University Hospital Carl Gustav Carus, TUD Dresden University of
Technology, Dresden, Germany
email: "fabian.wolf2@tu-dresden.de"
- name: "KatherLab"
website: "https://github.com/KatherLab"
repository-code: "https://github.com/KatherLab/deidentifier"
url: "https://github.com/KatherLab/deidentifier"
abstract: >-
A locally deployable web application that anonymizes German clinical
documents. Rule-based recognizers and a prompted large language model behind
any OpenAI-compatible endpoint propose character spans; deterministic code
applies every replacement on the immutable source text, and an independent
leakage-validation pass re-scans the output. Documents are processed in
memory only and nothing is persisted. The repository also contains a
standalone evaluation harness that scores the pipeline against annotated
ground truth, reporting document-level leakage alongside character- and
span-level metrics.
keywords:
- "de-identification"
- "anonymization"
- "large language models"
- "named entity recognition"
- "clinical text"
- "German clinical documents"
- "medical natural language processing"
- "privacy-preserving artificial intelligence"
license: "AGPL-3.0-or-later"
version: "0.4.0"
date-released: "2026-08-24"
references:
- type: article
title: >-
Deidentifying Medical Documents with Local, Privacy-Preserving Large
Language Models: The LLM-Anonymizer
authors:
- family-names: "Wiest"
given-names: "Isabella C."
- family-names: "Leßmann"
given-names: "Marie-Elisabeth"
- family-names: "Wolf"
given-names: "Fabian"
- family-names: "Ferber"
given-names: "Dyke"
- family-names: "Van Treeck"
given-names: "Marko"
- family-names: "Zhu"
given-names: "Jiefu"
- family-names: "Ebert"
given-names: "Matthias P."
- family-names: "Westphalen"
given-names: "Christoph Benedikt"
- family-names: "Wermke"
given-names: "Martin"
- family-names: "Kather"
given-names: "Jakob Nikolas"
journal: "NEJM AI"
year: 2025
volume: 2
issue: 4
number: "AIdbp2400537"
doi: "10.1056/AIdbp2400537"
url: "https://ai.nejm.org/doi/full/10.1056/AIdbp2400537"
notes: >-
The predecessor work this software builds on. It established the
approach — local, privacy-preserving LLMs proposing identifiers in
clinical documents — and reported the accuracy figures for that
implementation. This repository is an independent reimplementation with a
different architecture, and its behavior and accuracy are not those
measured in the paper.