Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,9 @@ The release run heads these entries with the version and opens a fresh

## Unreleased

- Legacy Word validates FIB signatures and extension lengths. It stores the
shared header fields by value, removing unsafe deletion through base pointers.

- Legacy PowerPoint validates nested picture bounds, record lengths, and formatting
counts, limits text-container nesting, and avoids overflow in frame dimensions.

Expand Down
20 changes: 9 additions & 11 deletions src/odr/internal/oldms/text/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ text.

| File (`oldms/text/`) | Role |
|---|---|
| `doc_structs.hpp` | `#pragma pack(1)` PODs (`FibBase`, the `FibRgFcLcb97` to `FibRgFcLcb2007` chain, `Sprm`, `FcCompressed`, `Pcd`, `PnFkpChpx`, `FfnFixed`), `PlcMap` (`PlcPcdMap`, `PlcBteChpxMap`), `ParsedFib`, the character SPRM opcodes |
| `doc_structs.hpp` | `#pragma pack(1)` PODs (`FibBase`, `FibRgFcLcb97`, `Sprm`, `FcCompressed`, `Pcd`, `PnFkpChpx`, `FfnFixed`), `PlcMap` (`PlcPcdMap`, `PlcBteChpxMap`), `ParsedFib`, the character SPRM opcodes |
| `doc_io.{hpp,cpp}` | `read(...)` over `std::istream`: variable-length FIB, Clx walk, string decoding, `uncompress_char` |
| `doc_helper.{hpp,cpp}` | `CharacterIndex` (decoded piece table) and `read_character_index`; `CharacterRuns` (fc-keyed style-index runs) and `read_character_runs` |
| `doc_style.{hpp,cpp}` | `StyleRegistry` (resolved `TextStyle`s by index, font-name store), `apply_character_sprms`, `read_font_names` (SttbfFfn) |
Expand All @@ -35,14 +35,12 @@ body with headers, footnotes and annotations, and the FIB's `ccp*` counts
partition the CP space. The parser takes the first `ccpText` CPs by clamping
each piece to the remaining budget.

**Self-describing FIB read.** `read(ParsedFib&)` trusts the on-disk counts,
reads what the module models and ignores the surplus. Version dispatch picks
the `FibRgFcLcb*` layout by `nFib` (`nFib97` 0x00C1, `nFib2000` 0x00D9,
`nFib2002` 0x0101, `nFib2003` 0x010C, `nFib2007` 0x0112). A newer `nFib` uses
the 2007 layout. The `FcLcb` block is copied clamped to
`min(sizeof(layout), cbRgFcLcb·8)`, so extra entries are ignored and a shorter
block leaves the rest zero. `clx` sits in the `FibRgFcLcb97` base, so it is
always covered.
**Self-describing FIB read.** Counts bound each FIB array. The parser keeps the
`FibRgFcLcb97` prefix by value; later versions append entries that it does not
use. A shorter block zero-fills missing entries. `cswNew` bounds the version
extension, whose unused fields are skipped. Known versions are validated;
a future base `nFib` may still use the common prefix. No versioned object is
allocated or deleted through a non-virtual base.

**Piece table.** The `PlcPcd` is `n+1` ascending CP boundaries followed by `n`
`Pcd`s. `PlcPcdMap` is a zero-copy view over the raw bytes. Each `Pcd`'s
Expand Down Expand Up @@ -101,8 +99,8 @@ are dropped.
- `html_output_test` compares the real `.doc` fixtures against the reference
output. There is no assertion-based test over a real fixture.

Not tested: FIB robustness (negative `ccpText`, newer-than-2007 fallback) and
`page_break` emission.
FIB versions, counted extensions, and negative `ccpText` have synthetic coverage.
`page_break` emission still lacks a focused test.

## Binary format reference (FIB)

Expand Down
115 changes: 38 additions & 77 deletions src/odr/internal/oldms/text/doc_io.cpp
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
#include <odr/internal/oldms/text/doc_io.hpp>

#include "odr/internal/util/string_util.hpp"

#include <odr/internal/util/byte_stream_util.hpp>
#include <odr/internal/util/byte_string.hpp>
#include <odr/internal/util/string_util.hpp>

#include <algorithm>
#include <cstring>
Expand All @@ -12,34 +12,20 @@ namespace odr::internal::oldms::text {

namespace {

template <typename T> struct TypeTag {
using type = T;
};

template <typename F>
auto type_dispatch_FibRgFcLcb(const std::uint16_t nFib, const F &f) {
switch (nFib) {
case nFib97: {
return f(TypeTag<FibRgFcLcb97>{});
}
case nFib2000: {
return f(TypeTag<FibRgFcLcb2000>{});
}
case nFib2002: {
return f(TypeTag<FibRgFcLcb2002>{});
}
case nFib2003: {
return f(TypeTag<FibRgFcLcb2003>{});
}
case nFib2007: {
return f(TypeTag<FibRgFcLcb2007>{});
}
void validate_version(const std::uint16_t version, const bool allow_future) {
switch (version) {
case nFib97:
case nFib2000:
case nFib2002:
case nFib2003:
case nFib2007:
return;
default:
// Newer FIBs only append entries, so the newest modelled layout fits.
if (nFib > nFib2007) {
return f(TypeTag<FibRgFcLcb2007>{});
if (allow_future && version > nFib2007) {
return;
}
throw std::runtime_error("Unknown nFib value: " + std::to_string(nFib));
throw std::runtime_error("doc: unsupported FIB version " +
std::to_string(version));
}
}

Expand All @@ -55,24 +41,26 @@ void text::read(std::istream &in, FibBase &out) {

void text::read(std::istream &in, ParsedFib &out) {
read(in, out.base);
if (out.base.wIdent != fib_wIdent) {
throw std::runtime_error("doc: invalid FIB signature");
}

util::byte_stream::read(in, out.csw);
if (static_cast<std::size_t>(out.csw) * 2 < sizeof(out.fibRgW)) {
throw std::runtime_error("Unexpected Fib.csw value: " +
std::to_string(out.csw));
}
util::byte_stream::read(in, out.fibRgW);
in.ignore(static_cast<std::streamsize>(static_cast<std::size_t>(out.csw) * 2 -
sizeof(out.fibRgW)));
util::byte_stream::skip(in, std::uint64_t{out.csw} * 2 - sizeof(out.fibRgW));

util::byte_stream::read(in, out.cslw);
if (static_cast<std::size_t>(out.cslw) * 4 < sizeof(out.fibRgLw)) {
throw std::runtime_error("Unexpected Fib.cslw value: " +
std::to_string(out.cslw));
}
util::byte_stream::read(in, out.fibRgLw);
in.ignore(static_cast<std::streamsize>(
static_cast<std::size_t>(out.cslw) * 4 - sizeof(out.fibRgLw)));
util::byte_stream::skip(in,
std::uint64_t{out.cslw} * 4 - sizeof(out.fibRgLw));

// ccpText MUST be >= 0 ([MS-DOC] 2.5.5).
if (out.ccpText() < 0) {
Expand All @@ -81,53 +69,26 @@ void text::read(std::istream &in, ParsedFib &out) {
}

util::byte_stream::read(in, out.cbRgFcLcb);
const auto fibRgFcLcb =
std::make_unique<char[]>(static_cast<std::size_t>(out.cbRgFcLcb) * 8);
in.read(fibRgFcLcb.get(), static_cast<std::streamsize>(out.cbRgFcLcb) * 8);
const std::string offsets = util::byte_stream::read_u8s(
in, std::uint32_t{out.cbRgFcLcb} * sizeof(FcLcb));
out.fibRgFcLcb = {};
std::memcpy(&out.fibRgFcLcb, offsets.data(),
std::min(sizeof(out.fibRgFcLcb), offsets.size()));

util::byte_stream::read(in, out.cswNew);
out.nFibNew.reset();
if (out.cswNew > 0) {
out.fibRgCswNew.emplace();
read(in, *out.fibRgCswNew);
}
const std::uint16_t nFib =
out.fibRgCswNew.has_value() ? out.fibRgCswNew->nFibNew : out.base.nFib;

out.fibRgFcLcb = type_dispatch_FibRgFcLcb(
nFib, [&]<typename T>(const T) -> std::unique_ptr<FibRgFcLcb97> {
using FibRgFcLcbType = T::type;
auto result = std::make_unique<FibRgFcLcbType>();
// Clamped: surplus entries are dropped, a short block stays zeroed.
const std::size_t copy =
std::min<std::size_t>(sizeof(FibRgFcLcbType),
static_cast<std::size_t>(out.cbRgFcLcb) * 8);
std::memcpy(result.get(), fibRgFcLcb.get(), copy);
return result;
});
}

void text::read(std::istream &in, ParsedFibRgCswNew &out) {
util::byte_stream::read(in, out.nFibNew);

switch (out.nFibNew) {
case nFib97:
break;
case nFib2000:
case nFib2002:
case nFib2003: {
auto rgCswNewData = std::make_unique<FibRgCswNewData2000>();
util::byte_stream::read(in, *rgCswNewData);
out.rgCswNewData = std::move(rgCswNewData);
} break;
case nFib2007: {
auto rgCswNewData = std::make_unique<FibRgCswNewData2007>();
util::byte_stream::read(in, *rgCswNewData);
out.rgCswNewData = std::move(rgCswNewData);
} break;
default:
throw std::runtime_error("Unsupported nFibNew value: " +
std::to_string(out.nFibNew));
}
const std::string extension = util::byte_stream::read_u8s(
in, std::uint32_t{out.cswNew} * sizeof(std::uint16_t));
util::byte_string::Reader cursor(extension);
out.nFibNew = cursor.read<std::uint16_t>();
validate_version(*out.nFibNew, false);
// [MS-DOC] 2.5.11–13: only the version affects the fields we use.
cursor.skip(*out.nFibNew == nFib97 ? 0
: (*out.nFibNew == nFib2007 ? 8 : 2));
}
validate_version(out.nFibNew.value_or(out.base.nFib),
!out.nFibNew.has_value());
}

void text::read_Clx(std::istream &in, const HandlePrc &handle_Prc,
Expand All @@ -151,7 +112,7 @@ void text::skip_Prc(std::istream &in) {
}

const auto cbGrpprl = util::byte_stream::read<std::uint16_t>(in);
in.ignore(cbGrpprl);
util::byte_stream::skip(in, cbGrpprl);
}

std::string text::read_string(std::istream &in, const std::size_t length_cp,
Expand Down
1 change: 0 additions & 1 deletion src/odr/internal/oldms/text/doc_io.hpp
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,6 @@
namespace odr::internal::oldms::text {

void read(std::istream &in, FibBase &out);
void read(std::istream &in, ParsedFibRgCswNew &out);
void read(std::istream &in, ParsedFib &out);

using HandlePrc = std::function<void(std::istream &in)>;
Expand Down
6 changes: 3 additions & 3 deletions src/odr/internal/oldms/text/doc_parser.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -171,16 +171,16 @@ ElementIdentifier text::parse_tree(ElementRegistry &registry,
const auto table_stream = table_file->stream();

const std::vector<std::string> font_names =
read_font_names(*table_stream, fib.fibRgFcLcb->sttbfFfn);
read_font_names(*table_stream, fib.fibRgFcLcb.sttbfFfn);

// Direct character formatting only ([MS-DOC] 2.4.6.2).
std::vector<TextStyle> styles{default_character_style()}; // index 0
const CharacterRuns character_runs =
read_character_runs(*document_stream, *table_stream,
fib.fibRgFcLcb->plcfBteChpx, styles, font_names);
fib.fibRgFcLcb.plcfBteChpx, styles, font_names);
style_registry = StyleRegistry(std::move(styles));

table_stream->seekg(fib.fibRgFcLcb->clx.fc);
table_stream->seekg(fib.fibRgFcLcb.clx.fc);
const CharacterIndex character_index = read_character_index(*table_stream);

const auto ccp_text = static_cast<std::size_t>(fib.ccpText());
Expand Down
Loading
Loading