Skip to content

Lexer: no-break space and BOM are whitespace; error text stays UTF-8 - #13

Merged
revarbat merged 1 commit into
mainfrom
fix/nbsp-whitespace
Oct 5, 2026
Merged

revarbat merged 1 commit into
mainfrom
fix/nbsp-whitespace

Conversation

@revarbat

@revarbat revarbat commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Refs BelfrySCAD/BelfrySCAD#679

  • Whitespace, as in OpenSCAD's lexer.l: U+00A0 NO-BREAK SPACE (C2 A0), U+FEFF (EF BB BF, anywhere in the file) and a bare Latin-1 A0. OpenSCAD 2026.02 renders cube(1);<NBSP> without a warning; we failed to parse it.
  • unexpected character messages are valid UTF-8. The catch-all quoted one byte, so é or a no-break space in code produced a message nanobind could not convert (nanobind::str(): conversion error, shown in BelfrySCAD as 'utf-8' codec can't decode byte 0xc2). It now quotes the whole character, and shows a byte that starts no character as \xNN.

Tests: LexicalWhitespace.* (both fail without the lexer change); full suite 661/661.

🤖 Generated with Claude Code

As in OpenSCAD's lexer.l, U+00A0 (C2 A0), U+FEFF (EF BB BF) and a bare
Latin-1 A0 are whitespace. Qt's editor turned no-break spaces into
spaces, so a file holding one parsed in BelfrySCAD's editor and failed
everywhere else; once the editor kept them exactly (BelfrySCAD#681) it
failed there too (BelfrySCAD#679).

The catch-all quoted the single offending byte, so a non-ASCII
character produced an invalid-UTF-8 message that nanobind could not
convert: the user saw a codec error instead of the parse error. It now
quotes a whole UTF-8 character, and names a byte that starts none as
\xNN.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant