perf(serde): Compile per-shape JSON and XML serde plans with caching - #3363
Open
rohangavankar wants to merge 26 commits into
Open
rohangavankar wants to merge 26 commits into
rohangavankar wants to merge 26 commits into
Conversation
Add the shared cache foundation for serde plans (PR 0). Serializers and parsers can compile protocol instructions once per shape or operation and reuse them, replacing repeated model-array interpretation on every call. - ShapePlanCache: central slot registry (protocol + direction). - AbstractModel::getCachedPlan/setCachedPlan: per-slot plan storage. - ShapeMap: monotonic graph generation for cross-object invalidation. - clearResolvedModelCache hook on StructureShape, ListShape, MapShape, Operation, and Service; invoked on offsetSet/offsetUnset before the generation bump so plans derived from mutated shapes rebuild. - Service::setDefinition clears the operation cache and getOperation now returns stable instances so operation-level plans can warm. No protocol providers and no public API changes. Plan accessors are internal.
Add the json-serde-plans spec (requirements, design, tasks) grounded in the committed plan-cache foundation and the docs/serde design docs. Add a focused benchmark/json-serde.php that compares the legacy JsonBody::format() path against the compiled-plan path in one process, reporting first-use, p50, and p90 on small, nested-large, and map-heavy payloads. Local benchmark numbers are dev-grade. Merge evidence requires the x86 m7i.xlarge runbook flow.
Replace JsonBody's per-request model reads with a compiled JsonEncodePlan cached per shape via ShapePlanCache::JSON_ENCODE. JsonShapeType maps model strings to integer tags, and JsonEncodePlanProvider compiles one shape level lazily, fetching child plans when traversal reaches composite members. build() runs the plan path; buildLegacy() retains the old format() path so the benchmark can compare both in one process. Wire output is byte-identical: all tests/Api/Serializer + ComplianceTest pass (1252 tests). Local dev-grade encode numbers (200k iters, Xdebug off): warm p50 improves 27-47% (NestedLarge -26.8%, MapHeavy -36.8%, SmallNoList -46.7%). Small scalar-only payloads pay a ~2us first-use plan-compile cost. The PutItemRequest_Nested_L regression from the design doc did not reproduce here. Numbers in benchmark/results/json-encode-2026-09-21.md. Merge evidence still requires the x86 m7i.xlarge runbook flow.
Replace JsonParser's per-response model reads with a compiled JsonDecodePlan cached per shape via ShapePlanCache::JSON_DECODE. Structure members are stored as an ordered list to preserve V3 result order; the plan carries a union flag for the Unknown fallback. JsonDecodePlanProvider compiles one shape level lazily, fetching child plans when traversal reaches composite members. parse() runs the plan path; parseLegacy() retains the old path for the benchmark. Decode timestamp default stays null (DateTimeResult), null map values are skipped, and blob base64_decode is preserved. Output is identical: all tests/Api/Parser pass (677 tests). Extend benchmark/json-serde.php with --direction=decode. Local dev-grade decode numbers (200k iters, Xdebug off): warm p50 improves 21-51% (MapHeavy -50.7%, NestedLarge -31.5%, SmallNoList -21.0%), no small-payload first-use regression. Numbers in benchmark/results/json-decode-2026-09-21.md. Merge evidence still requires the x86 m7i.xlarge runbook flow.
Add --items option to the focused harness to match the runbook JSON command (--items=N sizes collections; 0 exercises the small scalar-only path). Capture x86 before/after encode numbers on benchmark-v4 (m7i.xlarge, PHP 8.1.34, JIT tracing, OPcache on). Warm p50 improves 27-52% across SmallNoList, NestedLarge, and MapHeavy at both items=50 and items=0. Small/empty payloads pay a few-microsecond first-use plan-compile cost that repays within a few warm calls. Wire output identical. The PutItemRequest_Nested_L regression did not reproduce on x86 (nested-large -37.6% p50). Full data in benchmark/results/json-encode-x86-2026-09-21.md.
Capture x86 before/after decode numbers on benchmark-v4 (m7i.xlarge, PHP 8.1.34, JIT tracing, OPcache on), comparing JsonParser::parseLegacy against the plan path. Warm p50 improves 13-44% across SmallNoList, NestedLarge, and MapHeavy at both items=50 and items=0. Empty-collection first-use pays a few-microsecond plan-compile cost. A one-off MapHeavy first-use spike was confirmed a measurement artifact across three isolated re-runs. Decoded output identical. Full data in benchmark/results/json-decode-x86-2026-09-21.md.
Add --mode=memory to the focused harness (retained plan heap, both directions on one graph) and capture the remaining evidence the design docs require. Retained memory: bounded by the model, not payload size (identical at items=50 and items=2000). NestedLarge ~11.7 KB, MapHeavy ~3.9 KB per graph. Large payload (items=2000, ~3 ms): encode warm p50 -37.6%, decode -40.7%, output identical. Full compliance corpus (scripts/benchmarks/serde_benchmark.php, 70 cases, all 5 protocols, foundation baseline vs serde-cache candidate): JSON RPC weighted p50 -35.9%; REST-JSON flat (HTTP-binding bound); REST-XML/Query/CBOR controls within +-0.3% (no regression). Raw result JSON preserved under x86-compliance-corpus/. Docs: json-memory-and-large-x86-2026-09-22.md, compliance-corpus-x86-2026-09-22.md.
Remove the JSON serde benchmark result docs and raw result JSON from the SDK tree and gitignore benchmark/results/. These are internal evidence and live in php-fork-dev, not the public aws-sdk-php package. Only benchmark/json-serde.php (the harness code, which needs the SDK autoloader) stays in the SDK.
Add the focused unit tests the design requires for the JSON serde plans: - JsonEncodePlanProviderTest: structure/list/map/timestamp/blob/document descriptor variants, wire-name derivation, encode timestamp default (unixTimestamp), lazy child Shape retention, and per-shape caching in the JSON_ENCODE slot. - JsonDecodePlanProviderTest: modeled member ordering, decode timestamp default (null), union flag, list/map value descriptors, and JSON_DECODE slot caching without occupying the encode slot. - JsonPlanInvalidationTest: locationName mutation and members replacement rebuild both encode and decode plans; stable generation reuses the cached plan; mock-without-ShapeMap stays safe. 18 tests, 60 assertions, all green.
Add a decode test for a root-level timestamp shape (not a struct member), covering the JsonShapeType::TIMESTAMP branch in JsonDecodePlanProvider. Brings the JSON plan provider and shape-type classes to 100% line coverage.
Extend the invalidation tests to the remaining new foundation paths: - offsetUnset on a shape trait bumps the generation and rebuilds the plan (the contract covers unset as well as set). - ListShape and MapShape drop their resolved child (member / value) on mutation, so the rebuilt plan reflects the new child definition. Covers AbstractModel::offsetUnset and the clearResolvedModelCache overrides on StructureShape, ListShape, and MapShape. The base AbstractModel hook stays uncovered by design (abstract; every concrete shape overrides it).
Replace XmlBody's per-request model reads with a compiled XmlEncodePlan cached per shape via ShapePlanCache::XML_ENCODE (slot 4). XmlShapeType maps model strings to integer tags; XmlEncodePlanProvider precomputes per shape the namespace attribute, structure attribute-vs-element partition and element-name resolution (honoring locationNameAtStructureLevel), flattened decisions, and list/map element naming. XML timestamp default stays iso8601 (not JSON's unixTimestamp). build() runs the plan path; buildLegacy() retains the old path for the benchmark. XMLWriter output is byte-identical: all tests/Api/Serializer pass (748 tests) including RestXmlSerializerTest escaping and rest-xml ComplianceTest. 9 XML plan unit tests; provider and shape-type at 100% line coverage. Local x86 warm p50 improves 37-45% (SmallNoList -37.4%, NestedLarge -45.1%, MapHeavy -42.8%), output identical. Add benchmark/xml-serde.php. Report in php-fork-dev serde-plans results.
Replace XmlParser's per-response model reads with a compiled XmlDecodePlan cached per shape via ShapePlanCache::XML_DECODE (slot 5). XmlDecodePlanProvider precomputes per shape: element read-names (honoring the getOriginalDefinition structure-level locationName special case), attribute key + namespace for attribute members, flattened decisions, list/map element naming, union status, timestamp decode default (null, not encode's iso8601), and a scalar coercion tag. parse() runs the plan path; parseLegacy() retains the old dispatch for the benchmark. Also benefits AWS Query response parsing (shared XmlParser). A precomputed scalar coercion tag (COERCE_STRING/INT/FLOAT) removes a per-element $shape['type'] read from the leaf path; without it the map-heavy case regressed +6.8%, with it decode gains -10% to -23% warm p50 on x86. Result arrays byte-identical: all 677 parser + 63 error-parser + 118 S3 parser tests pass. 18 XML plan unit tests; XmlDecodePlanProvider at 100% coverage. Add --direction=decode to benchmark/xml-serde.php. Deviations from XML-Serde-Plans.md (extra descriptor fields, COERCE tag, doc not yet updated) are documented in the XML report. Keep as a separate CR after XML encode per the design.
Rename AbstractModel::getCachedPlan/setCachedPlan to getSerdePlan/cacheSerdePlan to match the names specified in docs/serde/architecture.md. Updates all four plan providers (JSON/XML encode/decode) and the serde tests. No behavior change; all 1425 serializer + parser tests pass.
Restore two comments (XML node name extraction, locationName-from-definition check) inadvertently dropped during the XmlParser plan-path rewrite. Comments only; no behavior change. The legacy parse methods now match upstream.
Precompute the root element name (three-level precedence: ShapeMap original locationName, resolved locationName, shape name) into XmlEncodePlan::rootName in the provider, and read it in XmlBody::build(). This removes the per-request determineRootElementName() metadata inspection from the plan path, satisfying the XML design doc requirement that the runtime serializer not inspect shape metadata to open the document root. determineRootElementName() is retained for the legacy benchmark path. Output byte-identical: all 748 serializer tests pass.
Add benchmark/ to .gitattributes export-ignore, matching how tests/, docs/, and features/ are already handled. The serde benchmark harnesses live in the repo for developers (and to sync to the x86 instance per the runbook) but should not ship in the Composer-distributed package to customers.
Run phpcbf with phpcs.xml.dist on the files this change modifies. Fixes control-structure spacing, missing instantiation parentheses, foreach keyword spacing and one indentation error. Formatting only, no behavior change.
XmlParser and XmlBody resolve their per-type methods at runtime by concatenating the shape type, so the snake_case names are part of the dispatch contract and cannot be renamed without changing that lookup. The names predate this change. PHPCS reports them only because the check scans whole touched files. Annotate each method with a targeted phpcs:ignore.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Compile per-shape JSON and XML serialize/deserialize instructions once, cache them on
the shape, and reuse them on later calls. No public API change; wire output stays
byte-identical. Serialization runs in the
buildstep; parsing runs as the handlerresult resolves.
This is the JSON + XML half of the serde-plan work (shared
ShapePlanCachefoundationseparate workstream and are not in this PR.
What's included
ShapePlanCacheslot registry; plan cache onAbstractModel/ShapeMapkeyed off a monotonic generation, with invalidation on shape mutation.
JsonEncodePlanProvider/JsonDecodePlanProvider, wired intoJsonBodyandJsonParser.XmlEncodePlanProvider/XmlDecodePlanProvider, wired intoXmlBodyandXmlParser. Also benefits AWS Query response parsing (sharedXmlParser).Correctness
compliance tests. Output byte-identical (encode) / result-array-identical (decode) in
all cases.
aws/aws-sdk-phpmaster with no conflicts.Performance (x86 m7i.xlarge, PHP 8.1.34, OPcache on, JIT tracing)
Isolated body-codec, warm p50, median of 2 runs (SmallNoList / NestedLarge / MapHeavy):
All output byte-identical (encode) / result-array-identical (decode). JSON encode
SmallNoList shows a first-use cost (+63%, plan compiled on first call) that the warm
path recovers many times over.
Whole-operation REST-XML gains are diluted until HTTP-binding plans land (separate
workstream), as anticipated in the design.
Note on the XML decode
COERCEscalar tagThe XML decode plan precomputes a scalar coercion kind (
COERCE_STRING/INT/FLOAT) ratherthan re-reading
$shape['type']per leaf element. This was added after benchmarkingshowed the per-element model read regressed the map-heavy case. Re-verified via an
ablation on x86 (two runs each, same baseline):
The tag is the same "precompute metadata into the plan" pattern used throughout; it
turns a regression into a solid gain on leaf-heavy payloads.
Design-doc conformance
Follows the JSON and XML serde-plan design. The XML decode descriptor carries a few
fields beyond the original proposal (
M_ATTRKEY,M_TSFORMAT,M_COERCE) to removeper-element model reads from the hot path, consistent with the design's acceptance
criterion ("removes repeated metadata interpretation from the recursive hot path"). The
COERCE ablation table above is the evidence those fields earn their keep.