Skip to content

perf(serde): Compile per-shape JSON and XML serde plans with caching - #3363

Open
rohangavankar wants to merge 26 commits into
aws:masterfrom
rohangavankar:serde-cache
Open

rohangavankar wants to merge 26 commits into
aws:masterfrom
rohangavankar:serde-cache

Conversation

@rohangavankar

@rohangavankar rohangavankar commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Compile per-shape JSON and XML serialize/deserialize instructions once, cache them on
the shape, and reuse them on later calls. No public API change; wire output stays
byte-identical. Serialization runs in the build step; parsing runs as the handler
result resolves.

This is the JSON + XML half of the serde-plan work (shared ShapePlanCache foundation

  • JSON encode/decode + XML encode/decode). HTTP bindings and Query serialization are a
    separate workstream and are not in this PR.

What's included

  • Foundation: ShapePlanCache slot registry; plan cache on AbstractModel / ShapeMap
    keyed off a monotonic generation, with invalidation on shape mutation.
  • JSON: JsonEncodePlanProvider / JsonDecodePlanProvider, wired into JsonBody and
    JsonParser.
  • XML: XmlEncodePlanProvider / XmlDecodePlanProvider, wired into XmlBody and
    XmlParser. Also benefits AWS Query response parsing (shared XmlParser).

Correctness

  • Full unit suite green: 1,866 tests (Serde + Serializer + Parser), plus 1,024 protocol
    compliance tests. Output byte-identical (encode) / result-array-identical (decode) in
    all cases.
  • Merged onto latest aws/aws-sdk-php master with no conflicts.

Performance (x86 m7i.xlarge, PHP 8.1.34, OPcache on, JIT tracing)

Isolated body-codec, warm p50, median of 2 runs (SmallNoList / NestedLarge / MapHeavy):

Direction SmallNoList NestedLarge MapHeavy
JSON encode -48% -35% -33%
JSON decode -16% -33% -42%
XML encode -40% -46% -47%
XML decode -12% -22% -18%

All output byte-identical (encode) / result-array-identical (decode). JSON encode
SmallNoList shows a first-use cost (+63%, plan compiled on first call) that the warm
path recovers many times over.

Whole-operation REST-XML gains are diluted until HTTP-binding plans land (separate
workstream), as anticipated in the design.

Note on the XML decode COERCE scalar tag

The XML decode plan precomputes a scalar coercion kind (COERCE_STRING/INT/FLOAT) rather
than re-reading $shape['type'] per leaf element. This was added after benchmarking
showed the per-element model read regressed the map-heavy case. Re-verified via an
ablation on x86 (two runs each, same baseline):

MapHeavy decode p50 vs legacy Without COERCE With COERCE
Run 1 +7.2% slower -18.0% faster
Run 2 +11.7% slower -17.3% faster

The tag is the same "precompute metadata into the plan" pattern used throughout; it
turns a regression into a solid gain on leaf-heavy payloads.

Design-doc conformance

Follows the JSON and XML serde-plan design. The XML decode descriptor carries a few
fields beyond the original proposal (M_ATTRKEY, M_TSFORMAT, M_COERCE) to remove
per-element model reads from the hot path, consistent with the design's acceptance
criterion ("removes repeated metadata interpretation from the recursive hot path"). The
COERCE ablation table above is the evidence those fields earn their keep.

rohangavankar and others added 22 commits September 17, 2026 10:52
Add the shared cache foundation for serde plans (PR 0). Serializers and
parsers can compile protocol instructions once per shape or operation and
reuse them, replacing repeated model-array interpretation on every call.

- ShapePlanCache: central slot registry (protocol + direction).
- AbstractModel::getCachedPlan/setCachedPlan: per-slot plan storage.
- ShapeMap: monotonic graph generation for cross-object invalidation.
- clearResolvedModelCache hook on StructureShape, ListShape, MapShape,
  Operation, and Service; invoked on offsetSet/offsetUnset before the
  generation bump so plans derived from mutated shapes rebuild.
- Service::setDefinition clears the operation cache and getOperation now
  returns stable instances so operation-level plans can warm.

No protocol providers and no public API changes. Plan accessors are internal.
Add the json-serde-plans spec (requirements, design, tasks) grounded in the
committed plan-cache foundation and the docs/serde design docs. Add a focused
benchmark/json-serde.php that compares the legacy JsonBody::format() path
against the compiled-plan path in one process, reporting first-use, p50, and
p90 on small, nested-large, and map-heavy payloads.

Local benchmark numbers are dev-grade. Merge evidence requires the x86
m7i.xlarge runbook flow.
Replace JsonBody's per-request model reads with a compiled JsonEncodePlan
cached per shape via ShapePlanCache::JSON_ENCODE. JsonShapeType maps model
strings to integer tags, and JsonEncodePlanProvider compiles one shape level
lazily, fetching child plans when traversal reaches composite members. build()
runs the plan path; buildLegacy() retains the old format() path so the
benchmark can compare both in one process.

Wire output is byte-identical: all tests/Api/Serializer + ComplianceTest pass
(1252 tests). Local dev-grade encode numbers (200k iters, Xdebug off): warm p50
improves 27-47% (NestedLarge -26.8%, MapHeavy -36.8%, SmallNoList -46.7%).
Small scalar-only payloads pay a ~2us first-use plan-compile cost. The
PutItemRequest_Nested_L regression from the design doc did not reproduce here.

Numbers in benchmark/results/json-encode-2026-09-21.md. Merge evidence still
requires the x86 m7i.xlarge runbook flow.
Replace JsonParser's per-response model reads with a compiled JsonDecodePlan
cached per shape via ShapePlanCache::JSON_DECODE. Structure members are stored
as an ordered list to preserve V3 result order; the plan carries a union flag
for the Unknown fallback. JsonDecodePlanProvider compiles one shape level
lazily, fetching child plans when traversal reaches composite members. parse()
runs the plan path; parseLegacy() retains the old path for the benchmark.

Decode timestamp default stays null (DateTimeResult), null map values are
skipped, and blob base64_decode is preserved. Output is identical: all
tests/Api/Parser pass (677 tests). Extend benchmark/json-serde.php with
--direction=decode.

Local dev-grade decode numbers (200k iters, Xdebug off): warm p50 improves
21-51% (MapHeavy -50.7%, NestedLarge -31.5%, SmallNoList -21.0%), no
small-payload first-use regression. Numbers in
benchmark/results/json-decode-2026-09-21.md. Merge evidence still requires the
x86 m7i.xlarge runbook flow.
Add --items option to the focused harness to match the runbook JSON command
(--items=N sizes collections; 0 exercises the small scalar-only path). Capture
x86 before/after encode numbers on benchmark-v4 (m7i.xlarge, PHP 8.1.34, JIT
tracing, OPcache on).

Warm p50 improves 27-52% across SmallNoList, NestedLarge, and MapHeavy at both
items=50 and items=0. Small/empty payloads pay a few-microsecond first-use
plan-compile cost that repays within a few warm calls. Wire output identical.
The PutItemRequest_Nested_L regression did not reproduce on x86 (nested-large
-37.6% p50). Full data in benchmark/results/json-encode-x86-2026-09-21.md.
Capture x86 before/after decode numbers on benchmark-v4 (m7i.xlarge, PHP
8.1.34, JIT tracing, OPcache on), comparing JsonParser::parseLegacy against the
plan path.

Warm p50 improves 13-44% across SmallNoList, NestedLarge, and MapHeavy at both
items=50 and items=0. Empty-collection first-use pays a few-microsecond
plan-compile cost. A one-off MapHeavy first-use spike was confirmed a
measurement artifact across three isolated re-runs. Decoded output identical.
Full data in benchmark/results/json-decode-x86-2026-09-21.md.
Add --mode=memory to the focused harness (retained plan heap, both directions
on one graph) and capture the remaining evidence the design docs require.

Retained memory: bounded by the model, not payload size (identical at items=50
and items=2000). NestedLarge ~11.7 KB, MapHeavy ~3.9 KB per graph.

Large payload (items=2000, ~3 ms): encode warm p50 -37.6%, decode -40.7%,
output identical.

Full compliance corpus (scripts/benchmarks/serde_benchmark.php, 70 cases, all 5
protocols, foundation baseline vs serde-cache candidate): JSON RPC weighted p50
-35.9%; REST-JSON flat (HTTP-binding bound); REST-XML/Query/CBOR controls within
+-0.3% (no regression). Raw result JSON preserved under x86-compliance-corpus/.

Docs: json-memory-and-large-x86-2026-09-22.md,
compliance-corpus-x86-2026-09-22.md.
Remove the JSON serde benchmark result docs and raw result JSON from the SDK
tree and gitignore benchmark/results/. These are internal evidence and live in
php-fork-dev, not the public aws-sdk-php package. Only benchmark/json-serde.php
(the harness code, which needs the SDK autoloader) stays in the SDK.
Add the focused unit tests the design requires for the JSON serde plans:

- JsonEncodePlanProviderTest: structure/list/map/timestamp/blob/document
  descriptor variants, wire-name derivation, encode timestamp default
  (unixTimestamp), lazy child Shape retention, and per-shape caching in the
  JSON_ENCODE slot.
- JsonDecodePlanProviderTest: modeled member ordering, decode timestamp default
  (null), union flag, list/map value descriptors, and JSON_DECODE slot caching
  without occupying the encode slot.
- JsonPlanInvalidationTest: locationName mutation and members replacement
  rebuild both encode and decode plans; stable generation reuses the cached
  plan; mock-without-ShapeMap stays safe.

18 tests, 60 assertions, all green.
Add a decode test for a root-level timestamp shape (not a struct member),
covering the JsonShapeType::TIMESTAMP branch in JsonDecodePlanProvider. Brings
the JSON plan provider and shape-type classes to 100% line coverage.
Extend the invalidation tests to the remaining new foundation paths:

- offsetUnset on a shape trait bumps the generation and rebuilds the plan
  (the contract covers unset as well as set).
- ListShape and MapShape drop their resolved child (member / value) on
  mutation, so the rebuilt plan reflects the new child definition.

Covers AbstractModel::offsetUnset and the clearResolvedModelCache overrides on
StructureShape, ListShape, and MapShape. The base AbstractModel hook stays
uncovered by design (abstract; every concrete shape overrides it).
Replace XmlBody's per-request model reads with a compiled XmlEncodePlan cached
per shape via ShapePlanCache::XML_ENCODE (slot 4). XmlShapeType maps model
strings to integer tags; XmlEncodePlanProvider precomputes per shape the
namespace attribute, structure attribute-vs-element partition and element-name
resolution (honoring locationNameAtStructureLevel), flattened decisions, and
list/map element naming. XML timestamp default stays iso8601 (not JSON's
unixTimestamp). build() runs the plan path; buildLegacy() retains the old path
for the benchmark.

XMLWriter output is byte-identical: all tests/Api/Serializer pass (748 tests)
including RestXmlSerializerTest escaping and rest-xml ComplianceTest. 9 XML plan
unit tests; provider and shape-type at 100% line coverage.

Local x86 warm p50 improves 37-45% (SmallNoList -37.4%, NestedLarge -45.1%,
MapHeavy -42.8%), output identical. Add benchmark/xml-serde.php. Report in
php-fork-dev serde-plans results.
Replace XmlParser's per-response model reads with a compiled XmlDecodePlan
cached per shape via ShapePlanCache::XML_DECODE (slot 5). XmlDecodePlanProvider
precomputes per shape: element read-names (honoring the getOriginalDefinition
structure-level locationName special case), attribute key + namespace for
attribute members, flattened decisions, list/map element naming, union status,
timestamp decode default (null, not encode's iso8601), and a scalar coercion
tag. parse() runs the plan path; parseLegacy() retains the old dispatch for the
benchmark. Also benefits AWS Query response parsing (shared XmlParser).

A precomputed scalar coercion tag (COERCE_STRING/INT/FLOAT) removes a per-element
$shape['type'] read from the leaf path; without it the map-heavy case regressed
+6.8%, with it decode gains -10% to -23% warm p50 on x86.

Result arrays byte-identical: all 677 parser + 63 error-parser + 118 S3 parser
tests pass. 18 XML plan unit tests; XmlDecodePlanProvider at 100% coverage. Add
--direction=decode to benchmark/xml-serde.php.

Deviations from XML-Serde-Plans.md (extra descriptor fields, COERCE tag, doc not
yet updated) are documented in the XML report. Keep as a separate CR after XML
encode per the design.
Rename AbstractModel::getCachedPlan/setCachedPlan to getSerdePlan/cacheSerdePlan
to match the names specified in docs/serde/architecture.md. Updates all four
plan providers (JSON/XML encode/decode) and the serde tests. No behavior change;
all 1425 serializer + parser tests pass.
Restore two comments (XML node name extraction, locationName-from-definition
check) inadvertently dropped during the XmlParser plan-path rewrite. Comments
only; no behavior change. The legacy parse methods now match upstream.
Precompute the root element name (three-level precedence: ShapeMap original
locationName, resolved locationName, shape name) into XmlEncodePlan::rootName in
the provider, and read it in XmlBody::build(). This removes the per-request
determineRootElementName() metadata inspection from the plan path, satisfying
the XML design doc requirement that the runtime serializer not inspect shape
metadata to open the document root. determineRootElementName() is retained for
the legacy benchmark path.

Output byte-identical: all 748 serializer tests pass.
Add benchmark/ to .gitattributes export-ignore, matching how tests/, docs/, and
features/ are already handled. The serde benchmark harnesses live in the repo
for developers (and to sync to the x86 instance per the runbook) but should not
ship in the Composer-distributed package to customers.
@rohangavankar
rohangavankar requested a review from a team as a code owner September 29, 2026 19:52
Run phpcbf with phpcs.xml.dist on the files this change modifies.
Fixes control-structure spacing, missing instantiation parentheses,
foreach keyword spacing and one indentation error. Formatting only,
no behavior change.
XmlParser and XmlBody resolve their per-type methods at runtime by concatenating the shape type, so the snake_case names are part of the dispatch contract and cannot be renamed without changing that lookup. The names predate this change. PHPCS reports them only because the check scans whole touched files. Annotate each method with a targeted phpcs:ignore.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant