Skip to content

trans: Add Panjabi (pa-IN) translation of AISVS 1.0 - #1128

Open
GeeksikhSecurity wants to merge 18 commits into
OWASP:mainfrom
GeeksikhSecurity:panjabi-translation-v1
Open

GeeksikhSecurity wants to merge 18 commits into
OWASP:mainfrom
GeeksikhSecurity:panjabi-translation-v1

Conversation

@GeeksikhSecurity

Copy link
Copy Markdown

Pull Request: Panjabi Translation of AISVS 1.0

Summary

This PR introduces the first Panjabi (ਪੰਜਾਬੀ) translation of OWASP AISVS 1.0, reaching 130+ million Panjabi speakers worldwide. Every section is bilingual — English first, Panjabi (Gurmukhi script) immediately below — so readers can cross-reference for technical precision. The translation follows the same bilingual methodology as the OWASP ASVS 5.0 Panjabi translation (PR #3254).

Status

Complete draft — all 18 source files translated (3 front matter + 12 control-family chapters + 3 appendices). This has not yet had a Panjabi-speaking sangat (community) review pass — see "How to Review" below. Submitting now, in the open, so that review can happen against a real PR rather than a private draft.

Files Added

1.0/pa-IN/
├── 0x01-Frontispiece.md
├── 0x02-Preface.md
├── 0x03-Using-AISVS.md
├── 0x10-C01-Training-Data-Integrity-and-Traceability.md
├── 0x10-C02-Input-Validation.md
├── 0x10-C03-Model-Lifecycle-Management.md
├── 0x10-C04-Infrastructure.md
├── 0x10-C05-Access-Control-and-Identity.md
├── 0x10-C06-Supply-Chain.md
├── 0x10-C07-Model-Behavior.md
├── 0x10-C08-Memory-Embeddings-and-Vector-Database.md
├── 0x10-C09-Orchestration-and-Agentic-Action.md
├── 0x10-C10-MCP-Security.md
├── 0x10-C11-Adversarial-Robustness.md
├── 0x10-C12-Monitoring-and-Logging.md
├── 0x90-Appendix-A_Glossary.md
├── 0x91-Appendix-B_AI_Security_Controls_Inventory.md
├── 0x92-Appendix-C_AI_for_Code_Generation.md
├── CLAUDE.md              # translation rules (dictionary sources, Gurmat-safety constraints)
├── GLOSSARY.md            # shared terminology reference
├── OPEN-QUESTIONS.md      # terminology judgment-call log, cites the English source term
└── TRANSLATION-RULES.md

Terminology Approach

Each security term is classified as:

Type When to Use Example
Translated (T) Natural Panjabi equivalent exists Authentication → ਪ੍ਰਮਾਣੀਕਰਨ
Loan Word (L) Term is universal in English API → ਏ.ਪੀ.ਆਈ.
Retained (R) Acronym or proper noun OWASP, LLM, MCP
Hybrid (H) Part translates, part stays Prompt Injection → ਪ੍ਰੌਮਪਟ ਇੰਜੈਕਸ਼ਨ

Terminology judgment calls (new AI-security vocabulary not covered by the ASVS glossary) are logged in OPEN-QUESTIONS.md with the English source term, the pick, and the reasoning — same format as the ASVS translation's terminology log.

Mechanical Verification

Every file passes zero-dependency lint checks before being marked complete: requirement-ID completeness against the English source, orthography (no Devanagari contamination, Western numerals), modal-strength consistency (must/must-not), and spelling-consistency against a 177k-word Panjabi Wikipedia frequency list. These are hypothesis-generating checks, not a substitute for human sangat review.

How to Review

  • No GitHub experience needed: Email gurvinder@securityleader.ai with subject "AISVS Panjabi Review"
  • GitHub users: Leave inline comments on any file in 1.0/pa-IN/

Related

  • Sibling translation: OWASP ASVS 5.0 Panjabi (PR #3254), same bilingual methodology and terminology framework
  • Follows the pattern of existing OWASP translations into other languages

Maintainer

GeeksikhSecurity — committed to ongoing maintenance and community engagement

GeeksikhSecurity and others added 13 commits August 26, 2026 17:16
- CLAUDE.md, TRANSLATION-RULES.md forked from ASVS panjabi-translation-v5
- GLOSSARY.md seeds locked/normalised terms from the ASVS corpus
- OPEN-QUESTIONS.md skeleton for AISVS-specific terminology log

Phase 0 of the ASVS+AISVS print-edition handoff.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Front matter (0x01-0x03) + first 8 requirement chapters (training data
integrity, input validation, model lifecycle, infrastructure, access
control, supply chain, model behavior, memory/embeddings/vector DB).

Each file: AI-assisted draft, then independent adversarial Gurmat-safety +
fidelity review (fresh context), then a cross-file terminology-consistency
pass. Consistency pass found and fixed 15 cross-file terminology drifts
across 6 files (unsafe/threshold/-based/control/f-nukta spelling), and
corrected 2 false self-references in OPEN-QUESTIONS.md. New AI-specific
terms logged as Q1-Q86 for Sangat/community review.

Known follow-up: two locked terms (unsafe, threshold) had already drifted
despite being documented in OPEN-QUESTIONS.md before this pass caught it —
prose documentation didn't prevent the drift. A mechanical lint reading
pinned terms from OPEN-QUESTIONS.md is the right fix before batch B, so the
next chapters fail fast instead of needing a full audit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Encodes the 5 pinned-terminology decisions the 2026-08-26 cross-file audit
found already drifting despite being documented in OPEN-QUESTIONS.md
(unsafe, threshold, -based, control(s), English /f/ nukta) as an
executable check instead of prose alone.

Verified: passes clean on current corpus; catches injected violations
with file:line + pinned-pick + source citation (tested against a scratch
copy, not committed).

Run before every pa-IN commit: python3 1.0/tools/lint-terminology.py

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…l 18/18 AISVS translation complete

C10 (MCP Security), plus the 3 appendices (Glossary, AI Security Controls
Inventory, AI for Code Generation) — the largest single files in the corpus
(1,759 lines together). Independent review of C09/C11/C12 (translated in
the prior interrupted run, never reviewed) confirmed clean.

Full-corpus consistency pass (18 files, 145 open terminology questions):
- 3 real cross-file terminology drifts found and fixed (confidential
  computing, immutable, defense-in-depth) — same requirement ID rendered
  two ways between a chapter and the Controls Inventory appendix that
  restates every requirement.
- 2 false 'harmonised with' claims in OPEN-QUESTIONS.md corrected against
  the actual files (matches the pattern already found twice in batch A).
- 6 new rules added to tools/lint-terminology.py, specifically covering a
  finding worth flagging: every Gurmat-vocabulary rejection in the log was
  prose-only enforcement, which is exactly how a rejected term (ਸੱਚ/ਸਤਿ,
  ਮੁਦਰਾ, bare ਕਰਤਾ, noun ਗ਼ਲਤੀ) could silently re-enter through the new
  appendix content. Verified each new rule matches 0 false positives
  against its legitimate near-miss before committing.
- Q145 records the audit method, findings, and two cases deliberately left
  unguarded (unmatchable without false positives) for reviewer visibility.

Verified independently: python3 1.0/tools/lint-terminology.py -> clean,
all 18 files present and non-truncated (spot-checked).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ge PDF)

Print-ready binder edition: consolidated markdown + per-chapter files
(print/chapters/) + PDF, mirroring the sibling ASVS 5.0 print pipeline
(same margins, same font, same build script).

- 207 terminology footnotes cited to OPEN-QUESTIONS.md, zero orphans
  (independently verified, not just self-reported).
- Zero images, zero raw LaTeX-breaking escape sequences, zero missing
  Gurmukhi-font glyphs except one deliberate exception: the homoglyph
  glossary entry keeps its literal Cyrillic 'а' example (the whole point
  of that definition) with a U+0430 codepoint citation alongside, since no
  installed font covers both Gurmukhi and Cyrillic.
- Full 163-page PDF rendered and visually spot-checked (title/TOC, a
  requirement chapter, the homoglyph entry, the final page) — binder
  margins correct, footnotes render properly at page-bottom via pandoc's
  native footnote handling, dual-block EN/PA structure intact throughout.
- Two real bugs caught and fixed during my own independent verification
  (not the workflow's self-report): one chapter's print copy still had 8
  unescaped arrow/hyphen characters after the batch run; my first rebuild
  attempt of the consolidated file accidentally dropped all page breaks
  (caught before rendering, fixed by rebuilding from the per-chapter
  source files with explicit \newpage insertion).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Same logic as the sibling ASVS check (>=2 occurrences per EN id, not just
presence, since the bilingual file always contains the English block too).
Result: 0 missing, all 12 control-family chapters.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ported from the sibling ASVS corpus's identically-purposed, already-verified
script (same logic, same acronym-chain false-positive fix). Clean on the
AISVS corpus: 0 violations across all 18 files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Same systemic issue found and fixed in the sibling ASVS corpus (commit
eeda4eb2 there): 'must not' rendered with the weak ਨਹੀਂ...ਚਾਹੀਦਾ form
instead of the strong ਲਾਜ਼ਮੀ ਤੌਰ 'ਤੇ ਨਹੀਂ...ਚਾਹੀਦਾ pattern. A corpus-wide
sweep found only these 2 sites in AISVS (much smaller footprint than
ASVS's 12) -- both in 0x92-Appendix-C_AI_for_Code_Generation.md.

Also ports verify-modal-strength.py from the ASVS corpus as the mechanical
gate against regression, adapted for AISVS's AC.N.N appendix requirement
ID format. Verified clean: modal-strength, orthography, requirement-ID,
and terminology lints all pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Rebuilds the consolidated print file and re-renders the 163-page PDF with
the AC.4.1/AC.12.5 fix from commit a1e7fed. Verified: footnote integrity
(207/207), no images, no glyph-gap characters, zero LaTeX errors. Fix
visually confirmed rendering correctly on page 156.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ency lint

Same rule and same script as the sibling ASVS corpus (commit 6c096c79
there). Found 18 candidates; a fresh-context review is confirming/fixing
each (separate commit to follow).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
All 18 candidates from the 2026-08-27 fresh-context review turned out to
be legitimate distinct words (idioms, loanwords, inflections) -- zero
genuine typos found in this corpus. Allowlist suppresses them so re-runs
stay quiet instead of re-flagging forever. Same mechanism as the sibling
ASVS corpus (commit b529d4a7 there), already verified there to still
catch genuinely new candidates.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Commit b11d52d fixed the title in the source markdown but the PDF itself
was never re-rendered after. Fixed now: 163 pages, zero unexpected
rendering warnings (only the documented Cyrillic homoglyph exception),
visually confirmed the title and ਮਿਆਰ render correctly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…hunspell attempt)

New spelling check backed by a 177k-word Panjabi Wikipedia frequency list
(CC BY, Leipzig Corpora Collection -- see
pa-wikipedia-wordlist/LICENSE_AND_PROVENANCE.md). An earlier attempt using
the LibreOffice pa_IN Hunspell dictionary (~37k entries) was built, tested,
and dropped entirely after proving too imprecise for this technical corpus
(didn't recognize ਮਿਆਰ 'standard' or ਮਨੁੱਖ 'human'; fuzzy-matching produced
100-900 candidates depending on tuning, mostly false positives). This
Wikipedia-derived list is ~5x larger and, being usage-derived rather than
curated, correctly recognizes ordinary vocabulary the dictionary missed.

Fresh-context review of all 106 flagged candidates (once-only word not in
the Wikipedia list): 0 typos found, 106/106 legitimate -- specialized
AI-security loanwords, inflected forms, and coined compounds that simply
don't appear in general Wikipedia prose. Several non-obvious ones
(ਦੁਸ਼ਮਣਾਨਾ, ਅਸਰਦਾਰੀ, ਵਸਤੂਪਰਕ, ਸੀਮਾਬੰਦੀ, ਪੁੱਗਣਾ) independently confirmed via
shabdkosh.com since Gurbani-focused dictionary sources don't carry modern
AI/security vocabulary. All added to the script's ALLOWLIST.

Verified clean: spelling-wikipedia, orthography, requirement-ids,
modal-strength, spelling-consistency, and lint-terminology all pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
GeeksikhSecurity added a commit to SayvaInc/securityleaderai-blog that referenced this pull request Aug 27, 2026
…pages

- New hidden /blog/aisvs-panjabi-review-* series: hub + all 18 documents
  (frontispiece, preface, using-aisvs, C1-C12, appendices A-C), mirroring
  the ASVS review-page pattern (hidden: true frontmatter, email-feedback
  CTA, prev/next nav chain).
- scripts/aisvs/build-review-posts.py generates the 18 chapter posts from
  the OWASP-AISVS-Panjabi fork (1.0/pa-IN/); hub is hand-authored (no prior
  page existed to base an incremental updater on -- see scripts/aisvs/README.md).
- Cross-linked from/to the ASVS hub as sibling translations.
- Source: OWASP/AISVS#1128 (GeeksikhSecurity/AISVS, branch
  panjabi-translation-v1) -- forked and opened this session.
- Verified: tsc --noEmit clean, lint-content.mjs 0 errors, npm run build
  succeeds (81 pages total), <h1> count = 1 on sampled pages, hidden posts
  confirmed absent from /blog index.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@jmanico

jmanico commented Sep 22, 2026

Copy link
Copy Markdown
Member

Thanks @GeeksikhSecurity for this. Translating all 18 files, with the English kept alongside and the terminology choices logged, is a lot of careful work, and a Panjabi edition could reach a large new audience. I'm sorry it has taken us so long to reply.

To be upfront, we haven't yet decided how AISVS will publish translations. The main open questions are:

  • which release translations should follow, now that 1.01 is in progress
  • where translations should live, and which files belong in this repository
  • what review a translation should have before we publish it

We'll keep this PR open and follow up here once we've decided. Thank you for your patience, and for offering to maintain the translation.

@jmanico

jmanico commented Sep 28, 2026

Copy link
Copy Markdown
Member

In the world of AI with auto-translation, we are thinking of only publishing this in US_English and leave translations outside the repo. cc @ottosulin @RicoKomenda

@GeeksikhSecurity

GeeksikhSecurity commented Sep 30, 2026 •

Copy link
Copy Markdown
Author

Thanks @jmanico, and no worries about the wait. I appreciate you laying out the open questions.

Rigor: this is not auto-translation. Rigor: this is not automatic translation. A person translates the 18 files with framework support using Claude, keeping English alongside, while terminology choices are recorded. A glossary and style guide manage recurring terms. Machine translation tends to drift on security terms (e.g., "compromise" in the breach sense mapped to a word meaning "agreement"), which is why I don't think it's a safe substitute for a standard whose controls people audit against.

Proposal on your three questions:

  1. Which release: pin each translation to a specific AISVS version (currently 1.0). I'll rebase and update to 1.01 once it's stable, with the version stated at the top of every file.
  2. Where it lives: either a translations/pa-IN/ directory marked non-normative, or hosted externally and linked from the AISVS site. English (US) remains the only authoritative text. I'm fine with either.
  3. Review before publishing: a staged status on each translation: Draft, then Community-reviewed (Sikh/Panjabi-speaking practitioners and community), then Academic-reviewed (Panjabi linguistics and information-security academics). Only files with completed review would carry the higher labels, and reviewer names/affiliations would be recorded in the PR history.

If the project prefers to keep translations outside the repo, I'd still like a link from AISVS to the reviewed version. Happy to take on maintenance either way. Let me know what would make this easiest for you and @ottosulin and @RicoKomenda to decide.

G.S. and others added 5 commits September 30, 2026 01:05
The Gurmukhi-numerals patch accidentally committed a compiled
.pyc alongside the source tool; untrack it and gitignore future
__pycache__/*.pyc so it can't return.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@ottosulin

ottosulin commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

I think that is actually a good point @jmanico, it is very easy for everyone to get their own translation now.

Also keeping possibly a dozen translations up to date is more overhead for already busy maintainers.

We can always also revisit this discussion later. I'd say this is our policy at least for 1.xx releases.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants