Skip to content

fix(python): Avoid exceptions when decoding non-UTF-8 data from string attrs - #5485

Merged
lgritz merged 3 commits into
AcademySoftwareFoundation:mainfrom
nrusch:py_string_safe_decode
Sep 21, 2026
Merged

lgritz merged 3 commits into
AcademySoftwareFoundation:mainfrom
nrusch:py_string_safe_decode

Conversation

@nrusch

@nrusch nrusch commented Sep 19, 2026

Copy link
Copy Markdown
Contributor

Add a new py_str_escaped helper function to handle conversion of string-like C++ types to py:str without raising a UnicodeDecodeError for non-UTF-8 input data.

This function is a fairly simple wrapper that calls the PyUnicode_DecodeUTF8 with the 'surrogateescape' error handling scheme (described at https://docs.python.org/3/library/codecs.html#error-handlers) and transfers ownership to the binding-specific object type.

Fixes #3856

Tests

Only manual tests so far, but I plan to add some test cases to the harness.

Checklist:

  • I have read the guidelines on contributions and code review procedures.
  • I have read the Policy on AI Coding Assistants
    and if I used AI coding assistants, I have an Assisted-by: TOOL / MODEL
    line in the pull request description above.
  • I have updated the documentation if my PR adds features or changes
    behavior.
  • I am sure that this PR's changes are tested in the testsuite.
  • I have run and passed the testsuite in CI before submitting the
    PR, by pushing the changes to my fork and seeing that the automated CI
    passed there. (Exceptions: If most tests pass and you can't figure out why
    the remaining ones fail, it's ok to submit the PR and ask for help. Or if
    any failures seem entirely unrelated to your change; sometimes things break
    on the GitHub runners.)
  • My code follows the prevailing code style of this project and I
    fixed any problems reported by the clang-format CI test.
  • If I added or modified a public C++ API call, I have also amended the
    corresponding Python bindings. If altering ImageBufAlgo functions, I also
    exposed the new functionality as oiiotool options.

Comment thread src/python/py_oiio.h
Comment on lines +481 to +491
#if defined(OIIO_PY_BACKEND_NANOBIND)
if (!py_str) {
py::raise_python_error();
}
return py::steal<py::str>(py_str);
#else
if (!py_str) {
throw py::error_already_set();
}
return py::reinterpret_steal<py::str>(py_str);
#endif

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I went with a single function definition containing implementation details for both backends, but it would be easy enough to just have two definitions in the oiio_py namespace with the Python API call duplicated if you prefer. Or we could split the difference and alias the exception type and handle stealing functions in py_backend.h in order to collapse this into one code path.

.OIIO_PY_PROP_RO("name",
[](const ParamValue& self) {
return oiio_py::str(self.name().string());
return py_str_escaped(self.name().string());

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Flagging that I've also wrapped the ParamValue.name accessor, because there's nothing that normalizes/sanitizes attribute names in the C++ API.

@nrusch nrusch changed the title python: Avoid exceptions when decoding non-UTF-8 data from string attrs fix(python): Avoid exceptions when decoding non-UTF-8 data from string attrs Sep 19, 2026
@lgritz

lgritz commented Sep 19, 2026

Copy link
Copy Markdown
Collaborator

Those wheel failures are unrelated. I just merged a fix that should take care of them. If you rebase on the current main, they should go away.

Signed-off-by: Nathan Rusch <nrusch@users.noreply.github.com>
@nrusch
nrusch force-pushed the py_string_safe_decode branch from 4c4f80b to 7137488 Compare September 19, 2026 23:49
@lgritz lgritz added the python Python APIs label Sep 21, 2026
Comment thread src/python/py_oiio.h
Signed-off-by: Larry Gritz <lg@larrygritz.com>

@lgritz lgritz left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

I added a couple lines to a comment, but otherwise, all looks great and I am merging.

@nrusch

nrusch commented Sep 21, 2026

Copy link
Copy Markdown
Contributor Author

Oh OK cool. I can come back and address the spec serialization later, but this should cover most of the potential user-facing issues.

@nrusch
nrusch marked this pull request as ready for review September 21, 2026 19:00
@lgritz

lgritz commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator

Oh, sorry, did I get ahead of you? If there are more changes you want to make before merging this PR, I can certainly wait.

@nrusch

nrusch commented Sep 21, 2026

Copy link
Copy Markdown
Contributor Author

Well, the outstanding tripping hazards here are just around ImageSpec serialization, in particular to XML: It's easy enough to route the ImageSpec.serialize and .to_xml wrappers through this same string decoding pathway, but even though you'll get back a valid Python string, the generated XML data will not be parseable.

I'd appreciate your thoughts on how to proceed here, since you have the context on how you'd like to dovetail this into the impending 3.2 cut (or whether that matters in practice for the Python bindings):

  1. Prioritize expediency by wrapping the serialization methods, accepting that the XML will be "broken" and I will try to fix it in a follow-up.
  2. Take some more time to find a solution for the XML serialization. So far, it looks like the best "simple" option for this might be converting invalid bytes to hex strings, but at least we would end up with valid XML (though the major caveat there is that string data would be different than the surrogateescape string decoding applied elsewhere). I'm also not quite sure how complicated this would be to implement yet.
  3. Leave the XML serialization broken (I can't really think of a good reason to do this).

@lgritz

lgritz commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator

So the problem is that is we turn the xml serialization into valid UTF-8 for Python's sake, that XML won't correctly encode the actual string of char 255's? So there is a discrepancy between what the string contains in the file, what the python str can contain, and how the xml needs to be expressed to properly round-trip, right?

Can't we just express it as hex/oct strings, like "\xff\xff"? I think that is your suggestion 2. I'd be fine with that. If we re-consume the resulting XML, will we end up where we started, with a string full of 0xff's?

Totally your choice whether to do that as part of this PR, or whether to commit what we have and do that in isolation next.

Signed-off-by: Nathan Rusch <nrusch@users.noreply.github.com>
@nrusch

nrusch commented Sep 21, 2026

Copy link
Copy Markdown
Contributor Author

OK. I'll push up the change to wrap those methods with this new helper as-is, and then look at a better solution for the XML data as a follow-up.

@nrusch

nrusch commented Sep 21, 2026

Copy link
Copy Markdown
Contributor Author

Done

@lgritz lgritz left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Follow-up to address the xml issue is expected separately.

@lgritz
lgritz merged commit 17fbee2 into AcademySoftwareFoundation:main Sep 21, 2026
64 checks passed
lgritz pushed a commit to lgritz/OpenImageIO that referenced this pull request Sep 21, 2026
…g attrs (AcademySoftwareFoundation#5485)

Add a new `py_str_escaped` helper function to handle conversion of
string-like C++ types to `py:str` without raising a `UnicodeDecodeError`
for non-UTF-8 input data.

This function is a fairly simple wrapper that calls the
`PyUnicode_DecodeUTF8` with the `'surrogateescape'` error handling
scheme (described at
https://docs.python.org/3/library/codecs.html#error-handlers) and
transfers ownership to the binding-specific object type.

Fixes AcademySoftwareFoundation#3856 

Only manual tests so far, but I plan to add some test cases to the
harness.

---------

Signed-off-by: Nathan Rusch <nrusch@users.noreply.github.com>
@lgritz lgritz added the devdays26 Dev Days 2026 label Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

devdays26 Dev Days 2026 python Python APIs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] python3.10 spec.get_string_attribute() throws a runtime exception with non-utf-8 string headers

2 participants