Skip to content

Support non-canonical/peripheral books #369

Description

@imnasnainaec

Summary

The interlinearizer assumes every book has real chapters and verses.

Paratext's 15 non-canonical/peripheral book codes don't: XXA–XXG, FRT, BAK, OTH, INT, CNC, GLO, TDX, NDX. They're paragraph-structured, not verse-structured.

Today, opening one of these books doesn't error. It silently tokenizes to zero segments and renders a blank view — no message explains why.

paranext-core support

  • Canon.nonCanonicalIds / Canon.isExtraMaterial model these 15 codes explicitly.
  • Storage, project settings, and VerseRef validity treat them uniformly with canonical books. VerseRef.internalValid skips chapter/verse bounds checking for them entirely.
  • Some platform UI (Find, the book/chapter quick-nav control, Manage Books) deliberately excludes or special-cases them, since they have no real chapter/verse structure to browse or search by.

So paranext-core has the vocabulary and partial precedent. This extension has none of it.

What core would need to add

  • A way to open/view/edit extra-material content at all. The platform's own Find feature excludes these books today for exactly this reason (tracked as PT-4414: "drop this exclusion once extra material can be opened and addressed"). Until that lands, there's no host capability for this extension to build on.
  • A paragraph-structured content API. platformScripture.USJ_Book is the only USJ access point this extension uses, and it's requested through a chapter/verse-shaped reference. Peripheral books need a way to fetch and navigate their content that doesn't route through chapter/verse at all.
  • Chapter/verse quick-nav support, or an equivalent. The platform's book-chapter control currently drops all 15 peripheral ids because "it cannot browse to them." Some paragraph-level navigation primitive would need to take that spot.
  • Optionality for verse/chapter in the shared scripture-reference types, so a peripheral book can be the active reference without a meaningless chapter/verse pair attached.

Current gaps

No awareness of the distinction

  • Zero references anywhere in the codebase to Canon.isCanonical, Canon.isExtraMaterial, Canon.nonCanonicalIds, or any of the 15 codes.
  • Canon is imported in exactly two places, both display-only (bookIdToEnglishName) — never for filtering.

Tokenizer/segmentation pipeline assumes verse structure

  • usjBookExtractor.ts only accumulates text into state.currentVerse, which is set only by a \c/\v node handler. A peripheral book has neither, so every character of body text is silently discarded.
  • bookTokenizer.ts then maps zero verses to zero segments. No error is raised.
  • useInterlinearizerBookData.ts only checks for a missing USJ book, not one that loaded but tokenized to nothing. Result: a blank view, indistinguishable from "this book has no analyzable text."
  • ScriptureRef declares chapter/verse as non-optional. There's no way to represent "this paragraph belongs to no chapter/verse."
  • The alignment model's documented default — one segment per verse — inherits the same assumption.

No book-selection UI to gate or allow this

  • The extension has no book picker of its own. Project creation only collects name, description, and analysis languages.
  • Book/chapter/verse is entirely driven by the platform's shared scroll-group reference. Whatever restriction (or lack of one) the platform's navigation control applies is the only gate that exists today.

Test-fixture coverage: zero

  • No test, fixture, or comment anywhere in the suite references any of the 15 codes.
  • Every shared book-fixture builder (GEN_1_1_BOOK, defaultScrRef, makeRawBook, makeVerseBook) is hardcoded to, or defaults to, canonical verse-structured books.
  • makeSegment actively throws unless its sid parses as "<BOOK> <chapter>:<verse>". Building a segment fixture for a peripheral book is structurally impossible with this helper today.
  • Two tests exercise "no verse markers" generically (usjBookExtractor.test.ts, bookTokenizer.test.ts), both using book code GEN. Neither is framed around, or asserts anything about, the peripheral-book case.

Non-goals

  • Deuterocanon/apocrypha books (TOB, WIS, 1MA, etc.) are full canon members in paranext-core's model already. Out of scope here.

Open questions

  • What should the UI show when a peripheral book is the active scripture reference: a blank interlinearizer, or an explicit "not supported" message (mirroring the platform's own Find-feature pattern)?
  • Does paragraph-structured content need its own segment-identity scheme, or can it reuse a synthetic verse-like key?
  • Should this wait on the platform's own extra-material viewing/editing support? The interlinearizer can't display these books meaningfully until the platform can.

Related: #129, #230 (alignment format's one-segment-per-verse default shares this gap).

Size: L — touches parsing, tokenization, segmentation, and every verse-keyed data structure.

Activity

  1. changed the title [-]Support non-canonical/peripheral books (XXA–XXG, FRT, GLO, etc.)[/-] [+]Support non-canonical/peripheral books (XXA, BAK, etc.)[/+] on Sep 25, 2026
  2. changed the title [-]Support non-canonical/peripheral books (XXA, BAK, etc.)[/-] [+]Support non-canonical/peripheral books[/+] on Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions