Preserve Pseudo C string contents and encoding - #8579
Merged
Merged
Conversation
Keep complete escape and format atoms, preserve whitespace inside literals, and avoid replaying text at the annotation parsing limit. Fixes Vector35/binaryninja#2063 and Vector35/binaryninja#2074.
Use the detected UTF-16 or UTF-32 literal prefix when rendering constant string data. Fixes Vector35/binaryninja#2064.
Render memcpy constant data as complete byte escapes. Embedded NULs and display annotation limits must not truncate a counted byte copy. Fixes Vector35/binaryninja#2065.
plafosse
force-pushed
the
test_fix_pseudoc_string_formatting
branch
from
October 1, 2026 20:03
3de0667 to
4925d41
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Wrapped string literals could lose format specifiers, split escape sequences, or drop literal whitespace. Outlined copies could also lose their Unicode encoding or omit bytes after an embedded NUL, changing the copied data when the Pseudo C was recompiled.
Keep complete escape and format sequences when wrapping string tokens, preserve literal whitespace, and retain the detected Unicode prefix. Render memcpy constant data as complete byte escapes so embedded NULs and annotation limits cannot shorten the copy.
Related issues: Vector35/binaryninja#2063, Vector35/binaryninja#2064, Vector35/binaryninja#2065, and Vector35/binaryninja#2074.
Companion core PR: Vector35/binaryninja#2078 contains the regression tests, minimized AArch64 binaries with C sources, and four corpus oracle updates. Merge this API PR before the core PR.
Local validation:
-O0and-O2; the two counted-copy cases also pass with AddressSanitizer.Existing failures in broader suites: CTest registers one test twice that fails in isolation but passes in the complete core suite; 24 Rust unit/integration tests also fail with the unmodified dev plugins; three unchanged Rust documentation examples fail to compile.
CI status: the branch build is currently unstable. The reported extension-manager network test fails because the live manifest now returns an official repository URL ending in
/extensions/, while the test expectsextensionswithout the slash. This follows the September 23 extension-server router change, and the same assertion failure reproduces with the original dev plugins. The local pytest total above excludes network tests under the default pytest configuration.