Skip to content

Transliterate Pali and Tibetan texts in Events - #798

Merged
tentamdin merged 10 commits into
developfrom
feat/reader-transliteration
Sep 22, 2026
Merged

tentamdin merged 10 commits into
developfrom
feat/reader-transliteration

Conversation

@harshal-2304

@harshal-2304 harshal-2304 commented Sep 21, 2026 •

Copy link
Copy Markdown
Member

Readers can now see a text's original verses in another script. Everything runs on the phone; the backend is untouched.

What's in it

  • Languages sheet. The Original section gets a switch and a script picker. Rows show each script's own name in its own script. One layer is always on: turning off the last one switches the other on.
  • Pali texts (24 on dev and prod) convert into 13 scripts through pali_script_convertor: Devanagari, Sinhala, Thai, Myanmar, Khmer, Bengali, Gurmukhi, Gujarati, Telugu, Kannada, Malayalam, Tibetan and Cyrillic, plus Roman. A post-pass repairs the package's Thai private-use glyphs and Myanmar/Khmer dangling stack marks.
  • Tibetan texts get a Roman row (THL Simplified Phonetics, how the verse is chanted) and the same scripts derived from it. Wylie is an internal step only, via a Dart port of BDRC's EWTS converter checked against their 210-case corpus.
  • Tibetan marks (༄༅, །) can be kept, dropped, turned into line breaks or a separator; default is a line break per shad, set in TransliterationService.standard().
  • The choice is saved per language and applies wherever the reader is embedded: reader, events, chants and plans.
  • Interlinear spacing now groups a verse with its translation (6 px) and separates verses (28 px).

Not in it

  • Sanskrit and Chinese sources: converters researched, not built.
  • Word joining in phonetics ("gyagar" rather than "gya gar"): needs a word list; ewts-js port proposed as a follow-up.
  • Tibetan font size in the reader.
image image image image

- Introduced sample text files for Tibetan (Unicode and Wylie) transliteration.
- Implemented PaliScriptConverter for detecting and converting Pali scripts.
- Added tests for PaliScriptConverter to ensure correct script detection and conversion.
- Created ReaderScriptPreferenceProvider to manage user script preferences with local storage.
- Developed TibetanPhonetics for phonetic transcription of Tibetan names and texts.
- Implemented TibetanScriptConverter for handling Tibetan script conversion and phonetic output.
- Added tests for TibetanScriptConverter to validate script detection and conversion functionality.
- Enhanced TransliterationService to support multiple scripts and caching for performance.
- Included tests for TransliterationService to verify conversion and caching behavior.
- Call the Roman row "Roman", its own name like every other script, and
  drop the reader_roman_transliteration l10n key from all six locales
- Map the Thai private-use glyphs the Pali package emits back to ญ and ฐ
- Close Myanmar syllables with asat and drop dangling Khmer coeng, so
  Tibetan phonetics such as "chak" render without dotted circles
- Give ཤ and ཞ each script's ś letter instead of an s + h cluster
- Drop Brahmi from the script list: no phone ships a font for it
- Put the reader's gap between verses (28 px) instead of between the
  original and its translation (6 px)
@greptile-apps

greptile-apps Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

The PR is not yet safe to merge because a valid HTML5 void footnote marker can prevent all following text in its segment from being transliterated.

Findings

  1. P1 Void Tag Swallows Verse ▶
  2. P1 Original text can disappear ▶
Fix with agent prompt
### Issue 1
lib/features/reader/domain/transliteration/transliteration_service.dart:100-104
A valid HTML void element such as `<img class="footnote-marker">` does not need a `/>` suffix. This code enters footnote-skip mode for that tag, but can leave the mode only after a matching `</img>`, which a void element never has. As a result, all following verse text in the segment is left untranslated. Detect HTML void elements independently of self-closing syntax before starting a skipped region.

### Issue 2
lib/features/reader/presentation/widgets/reader_content/reader_content_part.dart:701-707
In translation-only mode, a selected secondary version is treated as available before its content has loaded successfully. If the secondary request fails or a segment has no aligned translation, the original remains hidden and the reader displays only loading text or em dashes. Keep the original visible until usable secondary content exists, or restore it when secondary content fails or is missing.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

This PR adds persistent, client-side script selection and transliteration for Pali and Tibetan reader content, including an EWTS converter, Tibetan phonetics, HTML-aware conversion, translation-only display behavior, and focused tests.

  • Adds Pali conversion across supported scripts and Tibetan phonetic/script conversion.
  • Persists per-language script choices and original-text visibility.
  • Integrates transliteration into reader rendering, metadata, copying, and language settings.
  • Preserves editor footnotes and handles Tibetan punctuation and generated line breaks.
  • The previously reported original-visibility, startup preference, lockfile, and copy-selection issues are fixed or resolved.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Loaded source segments] --> B[Detect source script]
  B --> C[Languages sheet]
  C --> D[Persist per-language script choice]
  D --> E[Reader widgets and copy action]
  E --> F[TransliterationService]
  F --> G{Source language}
  G -->|Pali| H[Pali script converter]
  G -->|Tibetan| I[EWTS and phonetics converter]
  H --> J[Rendered HTML]
  I --> J
Loading

Reviews (6) · Last reviewed commit: "Fix transliteration review findings"

Comment on lines 701 to +706
final secondaryVersionId = dualSettings.secondary.versionId;
final secondaryActive =
dualSettings.secondaryEnabled && secondaryVersionId != null;
// "Translation only": the original can hide behind an active translation,
// never on its own.
final showOriginal = dualSettings.originalVisible || !secondaryActive;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Original text can disappear

In translation-only mode, a selected secondary version is treated as available before its content has loaded successfully. If the secondary request fails or a segment has no aligned translation, the original remains hidden and the reader displays only loading text or em dashes. Keep the original visible until usable secondary content exists, or restore it when secondary content fails or is missing.

Prompt To Fix With AI
This is a comment left during a code review.
Path: lib/features/reader/presentation/widgets/reader_content/reader_content_part.dart
Line: 701-706

Comment:
**Original text can disappear**

In translation-only mode, a selected secondary version is treated as available before its content has loaded successfully. If the secondary request fails or a segment has no aligned translation, the original remains hidden and the reader displays only loading text or em dashes. Keep the original visible until usable secondary content exists, or restore it when secondary content fails or is missing.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Comment thread lib/features/reader/presentation/providers/reader_script_preference_provider.dart Outdated
Comment thread pubspec.lock
@harshal-2304 harshal-2304 changed the title Transliterate Pali and Tibetan texts on the phone Transliterate Pali and Tibetan texts in Events Sep 21, 2026
@harshal-2304
harshal-2304 marked this pull request as draft September 21, 2026 07:08
@harshal-2304 harshal-2304 self-assigned this Sep 21, 2026
@harshal-2304
harshal-2304 marked this pull request as ready for review September 21, 2026 07:08
@harshal-2304 harshal-2304 mentioned this pull request Sep 21, 2026
5 tasks done
…anslation line and update copy action logic; add tests for translation retrieval scenarios.
Detect the original's script from the verses, not the title. Catalogue
titles are romanised as a matter of course, so a Tibetan text called
"Derge Kangyur" was read as Roman: the Original field claimed Roman
while the body stayed in Uchen, and the real Roman row - THL phonetics,
the feature worth having - was filtered out as "the script already on
screen". The reader now hands the sheet a sample of the loaded segments,
and detectScript lets a Latin script win only when no other script's
letters appear, so a siglum or a loanword inside a verse cannot outvote
it. A pick naming the source script now ticks the "as written" row
instead of leaving the list with nothing ticked.

Leave footnotes as the editor wrote them. convertHtml transliterated
every text node, so "So PTS; see Burmese ed. page 12." came back as
Sinhala letters. It now tracks element nesting and skips anything
classed footnote or footnote-marker.

Convert only the line breaks the converter introduced. The decision
keyed off the source chunk, so one raw newline - HTML whitespace, not a
break - suppressed every <br> the mark style produced in it, and a
pretty-printed document ran its verses together. Each source line is
now converted on its own and rejoined as it came.

Accept a hex escape in either case in EwtsConverter, as BDRC's Java
converter and pyewts do. Sloppy normalisation lower-cases a stray b but
not an F, so ་ was rejected as invalid and the character dropped.
Comment on lines +100 to +104
} else if (!closing &&
!raw.endsWith('/>') &&
_footnoteClass.hasMatch(raw)) {
skippedTag = name;
skipDepth = 1;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Void Tag Swallows Verse

A valid HTML void element such as <img class="footnote-marker"> does not need a /> suffix. This code enters footnote-skip mode for that tag, but can leave the mode only after a matching </img>, which a void element never has. As a result, all following verse text in the segment is left untranslated. Detect HTML void elements independently of self-closing syntax before starting a skipped region.

Prompt To Fix With AI
This is a comment left during a code review.
Path: lib/features/reader/domain/transliteration/transliteration_service.dart
Line: 100-104

Comment:
**Void Tag Swallows Verse**

A valid HTML void element such as `<img class="footnote-marker">` does not need a `/>` suffix. This code enters footnote-skip mode for that tag, but can leave the mode only after a matching `</img>`, which a void element never has. As a result, all following verse text in the segment is left untranslated. Detect HTML void elements independently of self-closing syntax before starting a skipped region.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

@tentamdin
tentamdin changed the base branch from main to develop September 22, 2026 20:46
@tentamdin
tentamdin merged commit 00dc969 into develop Sep 22, 2026
1 of 2 checks passed
@TenzDelek
TenzDelek deleted the feat/reader-transliteration branch September 23, 2026 04:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants