trans: Add Panjabi (pa-IN) translation of AISVS 1.0 - #1128
GeeksikhSecurity wants to merge 18 commits into
Conversation
- CLAUDE.md, TRANSLATION-RULES.md forked from ASVS panjabi-translation-v5 - GLOSSARY.md seeds locked/normalised terms from the ASVS corpus - OPEN-QUESTIONS.md skeleton for AISVS-specific terminology log Phase 0 of the ASVS+AISVS print-edition handoff. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Front matter (0x01-0x03) + first 8 requirement chapters (training data integrity, input validation, model lifecycle, infrastructure, access control, supply chain, model behavior, memory/embeddings/vector DB). Each file: AI-assisted draft, then independent adversarial Gurmat-safety + fidelity review (fresh context), then a cross-file terminology-consistency pass. Consistency pass found and fixed 15 cross-file terminology drifts across 6 files (unsafe/threshold/-based/control/f-nukta spelling), and corrected 2 false self-references in OPEN-QUESTIONS.md. New AI-specific terms logged as Q1-Q86 for Sangat/community review. Known follow-up: two locked terms (unsafe, threshold) had already drifted despite being documented in OPEN-QUESTIONS.md before this pass caught it — prose documentation didn't prevent the drift. A mechanical lint reading pinned terms from OPEN-QUESTIONS.md is the right fix before batch B, so the next chapters fail fast instead of needing a full audit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Encodes the 5 pinned-terminology decisions the 2026-08-26 cross-file audit found already drifting despite being documented in OPEN-QUESTIONS.md (unsafe, threshold, -based, control(s), English /f/ nukta) as an executable check instead of prose alone. Verified: passes clean on current corpus; catches injected violations with file:line + pinned-pick + source citation (tested against a scratch copy, not committed). Run before every pa-IN commit: python3 1.0/tools/lint-terminology.py Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…l 18/18 AISVS translation complete C10 (MCP Security), plus the 3 appendices (Glossary, AI Security Controls Inventory, AI for Code Generation) — the largest single files in the corpus (1,759 lines together). Independent review of C09/C11/C12 (translated in the prior interrupted run, never reviewed) confirmed clean. Full-corpus consistency pass (18 files, 145 open terminology questions): - 3 real cross-file terminology drifts found and fixed (confidential computing, immutable, defense-in-depth) — same requirement ID rendered two ways between a chapter and the Controls Inventory appendix that restates every requirement. - 2 false 'harmonised with' claims in OPEN-QUESTIONS.md corrected against the actual files (matches the pattern already found twice in batch A). - 6 new rules added to tools/lint-terminology.py, specifically covering a finding worth flagging: every Gurmat-vocabulary rejection in the log was prose-only enforcement, which is exactly how a rejected term (ਸੱਚ/ਸਤਿ, ਮੁਦਰਾ, bare ਕਰਤਾ, noun ਗ਼ਲਤੀ) could silently re-enter through the new appendix content. Verified each new rule matches 0 false positives against its legitimate near-miss before committing. - Q145 records the audit method, findings, and two cases deliberately left unguarded (unmatchable without false positives) for reviewer visibility. Verified independently: python3 1.0/tools/lint-terminology.py -> clean, all 18 files present and non-truncated (spot-checked). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ge PDF) Print-ready binder edition: consolidated markdown + per-chapter files (print/chapters/) + PDF, mirroring the sibling ASVS 5.0 print pipeline (same margins, same font, same build script). - 207 terminology footnotes cited to OPEN-QUESTIONS.md, zero orphans (independently verified, not just self-reported). - Zero images, zero raw LaTeX-breaking escape sequences, zero missing Gurmukhi-font glyphs except one deliberate exception: the homoglyph glossary entry keeps its literal Cyrillic 'а' example (the whole point of that definition) with a U+0430 codepoint citation alongside, since no installed font covers both Gurmukhi and Cyrillic. - Full 163-page PDF rendered and visually spot-checked (title/TOC, a requirement chapter, the homoglyph entry, the final page) — binder margins correct, footnotes render properly at page-bottom via pandoc's native footnote handling, dual-block EN/PA structure intact throughout. - Two real bugs caught and fixed during my own independent verification (not the workflow's self-report): one chapter's print copy still had 8 unescaped arrow/hyphen characters after the batch run; my first rebuild attempt of the consolidated file accidentally dropped all page breaks (caught before rendering, fixed by rebuilding from the per-chapter source files with explicit \newpage insertion). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Same logic as the sibling ASVS check (>=2 occurrences per EN id, not just presence, since the bilingual file always contains the English block too). Result: 0 missing, all 12 control-family chapters. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ported from the sibling ASVS corpus's identically-purposed, already-verified script (same logic, same acronym-chain false-positive fix). Clean on the AISVS corpus: 0 violations across all 18 files. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Same systemic issue found and fixed in the sibling ASVS corpus (commit eeda4eb2 there): 'must not' rendered with the weak ਨਹੀਂ...ਚਾਹੀਦਾ form instead of the strong ਲਾਜ਼ਮੀ ਤੌਰ 'ਤੇ ਨਹੀਂ...ਚਾਹੀਦਾ pattern. A corpus-wide sweep found only these 2 sites in AISVS (much smaller footprint than ASVS's 12) -- both in 0x92-Appendix-C_AI_for_Code_Generation.md. Also ports verify-modal-strength.py from the ASVS corpus as the mechanical gate against regression, adapted for AISVS's AC.N.N appendix requirement ID format. Verified clean: modal-strength, orthography, requirement-ID, and terminology lints all pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Rebuilds the consolidated print file and re-renders the 163-page PDF with the AC.4.1/AC.12.5 fix from commit a1e7fed. Verified: footnote integrity (207/207), no images, no glyph-gap characters, zero LaTeX errors. Fix visually confirmed rendering correctly on page 156. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ency lint Same rule and same script as the sibling ASVS corpus (commit 6c096c79 there). Found 18 candidates; a fresh-context review is confirming/fixing each (separate commit to follow). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
All 18 candidates from the 2026-08-27 fresh-context review turned out to be legitimate distinct words (idioms, loanwords, inflections) -- zero genuine typos found in this corpus. Allowlist suppresses them so re-runs stay quiet instead of re-flagging forever. Same mechanism as the sibling ASVS corpus (commit b529d4a7 there), already verified there to still catch genuinely new candidates. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Commit b11d52d fixed the title in the source markdown but the PDF itself was never re-rendered after. Fixed now: 163 pages, zero unexpected rendering warnings (only the documented Cyrillic homoglyph exception), visually confirmed the title and ਮਿਆਰ render correctly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…hunspell attempt) New spelling check backed by a 177k-word Panjabi Wikipedia frequency list (CC BY, Leipzig Corpora Collection -- see pa-wikipedia-wordlist/LICENSE_AND_PROVENANCE.md). An earlier attempt using the LibreOffice pa_IN Hunspell dictionary (~37k entries) was built, tested, and dropped entirely after proving too imprecise for this technical corpus (didn't recognize ਮਿਆਰ 'standard' or ਮਨੁੱਖ 'human'; fuzzy-matching produced 100-900 candidates depending on tuning, mostly false positives). This Wikipedia-derived list is ~5x larger and, being usage-derived rather than curated, correctly recognizes ordinary vocabulary the dictionary missed. Fresh-context review of all 106 flagged candidates (once-only word not in the Wikipedia list): 0 typos found, 106/106 legitimate -- specialized AI-security loanwords, inflected forms, and coined compounds that simply don't appear in general Wikipedia prose. Several non-obvious ones (ਦੁਸ਼ਮਣਾਨਾ, ਅਸਰਦਾਰੀ, ਵਸਤੂਪਰਕ, ਸੀਮਾਬੰਦੀ, ਪੁੱਗਣਾ) independently confirmed via shabdkosh.com since Gurbani-focused dictionary sources don't carry modern AI/security vocabulary. All added to the script's ALLOWLIST. Verified clean: spelling-wikipedia, orthography, requirement-ids, modal-strength, spelling-consistency, and lint-terminology all pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…pages - New hidden /blog/aisvs-panjabi-review-* series: hub + all 18 documents (frontispiece, preface, using-aisvs, C1-C12, appendices A-C), mirroring the ASVS review-page pattern (hidden: true frontmatter, email-feedback CTA, prev/next nav chain). - scripts/aisvs/build-review-posts.py generates the 18 chapter posts from the OWASP-AISVS-Panjabi fork (1.0/pa-IN/); hub is hand-authored (no prior page existed to base an incremental updater on -- see scripts/aisvs/README.md). - Cross-linked from/to the ASVS hub as sibling translations. - Source: OWASP/AISVS#1128 (GeeksikhSecurity/AISVS, branch panjabi-translation-v1) -- forked and opened this session. - Verified: tsc --noEmit clean, lint-content.mjs 0 errors, npm run build succeeds (81 pages total), <h1> count = 1 on sampled pages, hidden posts confirmed absent from /blog index. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
Thanks @GeeksikhSecurity for this. Translating all 18 files, with the English kept alongside and the terminology choices logged, is a lot of careful work, and a Panjabi edition could reach a large new audience. I'm sorry it has taken us so long to reply. To be upfront, we haven't yet decided how AISVS will publish translations. The main open questions are:
We'll keep this PR open and follow up here once we've decided. Thank you for your patience, and for offering to maintain the translation. |
|
In the world of AI with auto-translation, we are thinking of only publishing this in US_English and leave translations outside the repo. cc @ottosulin @RicoKomenda |
|
Thanks @jmanico, and no worries about the wait. I appreciate you laying out the open questions. Rigor: this is not auto-translation. Rigor: this is not automatic translation. A person translates the 18 files with framework support using Claude, keeping English alongside, while terminology choices are recorded. A glossary and style guide manage recurring terms. Machine translation tends to drift on security terms (e.g., "compromise" in the breach sense mapped to a word meaning "agreement"), which is why I don't think it's a safe substitute for a standard whose controls people audit against. Proposal on your three questions:
If the project prefers to keep translations outside the repo, I'd still like a link from AISVS to the reviewed version. Happy to take on maintenance either way. Let me know what would make this easiest for you and @ottosulin and @RicoKomenda to decide. |
… locked styleguide)
…ties); tool, lint and rules updated
The Gurmukhi-numerals patch accidentally committed a compiled .pyc alongside the source tool; untrack it and gitignore future __pycache__/*.pyc so it can't return. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
I think that is actually a good point @jmanico, it is very easy for everyone to get their own translation now. Also keeping possibly a dozen translations up to date is more overhead for already busy maintainers. We can always also revisit this discussion later. I'd say this is our policy at least for 1.xx releases. |
Pull Request: Panjabi Translation of AISVS 1.0
Summary
This PR introduces the first Panjabi (ਪੰਜਾਬੀ) translation of OWASP AISVS 1.0, reaching 130+ million Panjabi speakers worldwide. Every section is bilingual — English first, Panjabi (Gurmukhi script) immediately below — so readers can cross-reference for technical precision. The translation follows the same bilingual methodology as the OWASP ASVS 5.0 Panjabi translation (PR #3254).
Status
Complete draft — all 18 source files translated (3 front matter + 12 control-family chapters + 3 appendices). This has not yet had a Panjabi-speaking sangat (community) review pass — see "How to Review" below. Submitting now, in the open, so that review can happen against a real PR rather than a private draft.
Files Added
Terminology Approach
Each security term is classified as:
Terminology judgment calls (new AI-security vocabulary not covered by the ASVS glossary) are logged in
OPEN-QUESTIONS.mdwith the English source term, the pick, and the reasoning — same format as the ASVS translation's terminology log.Mechanical Verification
Every file passes zero-dependency lint checks before being marked complete: requirement-ID completeness against the English source, orthography (no Devanagari contamination, Western numerals), modal-strength consistency (must/must-not), and spelling-consistency against a 177k-word Panjabi Wikipedia frequency list. These are hypothesis-generating checks, not a substitute for human sangat review.
How to Review
1.0/pa-IN/Related
Maintainer
GeeksikhSecurity — committed to ongoing maintenance and community engagement