Skip to content

ru: addition of speech rules improvements and fixing conflicts with main branch - #822

Open
Kostenkov-2021 wants to merge 188 commits into
daisy:rufrom
Kostenkov-2021:ru
Open

Kostenkov-2021 wants to merge 188 commits into
daisy:rufrom
Kostenkov-2021:ru

Conversation

@Kostenkov-2021

@Kostenkov-2021 Kostenkov-2021 commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Clarify Russian wording for membership of standard number sets (ℕ, ℤ, ℚ, ℝ). Add a dedicated set-member-number-set rule and update ClearSpeak/SimpleSpeak shared rules to speak e.g. "принадлежащих множеству целых чисел" instead of ambiguous/calque forms. Fix related translations in unicode-full (remove English OT markers, replace «член» → «элемент»), simplify interval terminology in general rules, add FunctionApplicationWord connector, and refine several UI phrases (table/diacritical labels). Update and add unit tests in tests/Languages/ru to assert the new phrasing.
Also conflicts with the Main branch has been resolved.
Resolves #798

Dependency

This PR depends on #840 for the intent inference required by the new norm,
limit, and gathered-expression speech rules. Merge #840 first, then update
this branch before rerunning CI.

NSoiffer and others added 30 commits September 27, 2024 15:28
24 test passes, 3 test failures
… atomic numbers in prescripts to bump up their liklihood number. This allows single character elements with pre subscripts that match their atomic number to be considered chemistry.
Add French to the braille testing file.
New French braille tests.
Updated French test files and added the chemistry file to the tests
Incorporate the braille.rs trait-based refactor from main. The French
braille code is now wired through the BrailleCode trait/registry:
- register "French" in get_braille_code()
- French struct impls cleanup (french_cleanup), get_braille_chars
  (get_braille_ueb_chars), and needs_grouping (needs_grouping_for_french)
- keep French-specific helpers (FRENCH_INDICATOR_REPLACEMENTS,
  french_cleanup, needs_grouping_for_french)

French now uses SPEECH_DEFINITIONS (like the other codes) instead of
BRAILLE_DEFINITIONS: braille.rs is taken from main wholesale and only the
French-specific items are re-added, with needs_grouping_for_french
converted to SPEECH_DEFINITIONS.

No new test failures: tests/braille results are identical to the
pre-merge French baseline (1394 passed, 57 failed).

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…was overly broad, so I removed scripted letters and numbers. That rule made its way into other braille codes, so I removed them everywhere. All tests pass.
…(for accented letters) so that the pattern for them looks like:

```
 - "Ñ": [t: "CA⠘⠻L⠝"]             # 0x00D1
```

Added a few tests.
…or!() message. May want to change the promoted name...
…bers don't conflict with letters.

Add code to handle `;` special case if there are "blocks" in the expr
Now at 185 passed; 4 failed
Added several more chem arrows.
…ween different arrows that go in the same direction. That seems wrong, but the spec is silent on the alternatives
…or a more restricted version for nested munder/mover. They add a "," between lines.

This was done for limit_x_tends_1_14_2_01_corrected and limit_x_tends_1_14_2_02_corrected
moritz-gross and others added 6 commits September 24, 2026 23:07
Clarify Russian wording for membership of standard number sets (ℕ, ℤ, ℚ, ℝ). Add a dedicated set-member-number-set rule and update ClearSpeak/SimpleSpeak shared rules to speak e.g. "принадлежащих множеству целых чисел" instead of ambiguous/calque forms. Fix related translations in unicode-full (remove English OT markers, replace «член» → «элемент»), simplify interval terminology in general rules, add FunctionApplicationWord connector, and refine several UI phrases (table/diacritical labels). Update and add unit tests in tests/Languages/ru to assert the new phrasing.
Kostenkov-2021 and others added 5 commits September 25, 2026 11:49
This update improves Russian mathematical speech and navigation by fixing set-membership wording, compact Unicode operator and relation names, and table navigation boundary phrasing. It also adds targeted regression tests covering the corrected column navigation message and compact Unicode output.
Update Russian Unicode names in `Rules/Languages/ru/unicode-full.yaml` to use the orthographically correct `ё` variants for overline and underline descriptions. This keeps the terminology consistent across related symbols and aligns the translations with the language's spelling conventions.
@moritz-gross

moritz-gross commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

a job fails. Have you looked into this? does this look caused by changes in this PR or somewhere else?

@Kostenkov-2021

Kostenkov-2021 commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor Author

Hi @moritz-gross
As far as I understood, this appears to be the existing intermittent CI failure tracked in #781, rather than a failure caused by this PR.

The only failing job was Test (no-unsafe, Rules-minimized.zip). It passed 5,150 tests and then failed in Languages::ru::alphabets::script with:

RefCell already borrowed at src/speech.rs:2811

At that line, replace_single_char attempts to clear and reload unicode_full while an outer recursive replacement still holds a shared borrow of the same RefCell. Whether that reload is attempted depends on the extracted rule files' path/timestamp state, which explains why this has appeared intermittently with either Rules.zip or Rules-minimized.zip and with different languages/tests, as described in #781.

The corresponding Test (no-unsafe, Rules.zip) job passed, as did both normal Rust test jobs, Clippy, Python tests, packaging, and fuzzing. This PR does not change the failing script test or its mathematical-script mappings; its unicode-full.yaml changes are unrelated entries.

So I don't see evidence that the failure is caused by this PR. A rerun should confirm the intermittent nature of the issue.

moritz-gross and others added 3 commits September 27, 2026 13:48
…aisy#820)

* replace buggy Display implementation for Pronounce with Debug trait

* remove unused `eloquence` field and related code from `Pronounce` struct

* fix outdated syntax
Add Russian rules to recognize norms (double bars and magnitude) and nested double overline; treat paired absolute-value as "норма" and ordinary as "модуль". Fix multiline handling to skip duplicated '=' and final '.' in gathered/continued lines and announce visual lines. Add translations for named operators (max/min/sup/lim) including munder lower-limit phrasing and context-sensitive wording for overline and infinity in unicode. Adjust munder/mover default wording. Update and add unit tests (mtable, shared, linear_algebra) to cover these cases.
NSoiffer and others added 8 commits September 29, 2026 12:45
…his looks much better. All language `definitions.yaml` are updated.

Fix parsing of `IntentMappings`. Parsing now lives in one place, so ratio-style "name : glue" pairs, arity templates, and bracketing speech stay consistent across languages.

This all started as a fix to daisy#709, but then I discovered that the format used in `IntentMappings` was parsed incorrectly in some cases (such as "ratio").
* v1 of rule coverage report

* Add Python script to generate Markdown and HTML rule coverage reports

* clarify name of logfiles

* Switch rule coverage events to JSONL format

* switch to jinja2 templates

* Exclude Unicode mapping files from rule coverage report and improve report generation

* Exclude definition files from rule coverage report and update related documentation

* Add hit counts and test details to rule coverage report

Enhance the rule coverage report to include hit counts, test details, and hoverable tooltips for matched rules. Improve report formatting and coverage fraction calculation. Update related tests and documentation to reflect these changes.

* fix title

* Simplify rule coverage event recording

* Stream rule coverage events and trim unused report fields

* Suppress unused variable warning when `rule-coverage` feature is disabled

* Migrate `rule_coverage` implementation to a dedicated `rulecoverage` package, separating logic, templates, and CLI.

* uv run ruff format

* simplify dataclasses in Python

* remove "loaded" as statistic for rule coverage

* remove MD report and only keep HTML.
simplify rest accordingly.

* uv ruff format
Adds Russian-specific speech handling for arrays, gathered derivations, function aliases, typographic quotes, and linear algebra constructs such as vector products, projections, and norms. It also introduces a RussianFunctionAliases definition table so common names like det, arctan, and Cauchy are spoken correctly in ClearSpeak and SimpleSpeak.
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Linux library size: 0.58 MiB (+0.32%)
Revision Release liblibmathcat.so
Base (2b530d1) 0.58 MiB (603,896 bytes)
PR (cbf0d94) 0.58 MiB (605,832 bytes)
Change +1,936 bytes (+0.32%)

Built with default features, Rust 1.96.0, and Ubuntu 24.04. Workflow run.

@Kostenkov-2021 Kostenkov-2021 changed the title ru: use explicit "множество" for standard number sets ru: addition of speech rules improvements and fixing conflicts with main branch Oct 5, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Triage

Development

Successfully merging this pull request may close these issues.