You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Align ANTLR lexer and parser with SPARQL 1.1 §19.2 grammar for Unicode escapes (\uXXXX and \UXXXXXXXX) across IRIs and string literals, accepting valid hex encodings while strictly rejecting malformed sequences, isolated surrogates, and code points exceeding 0x10FFFF.
Target Branch
feature/corese-next
Specification & Semantics
According to W3C SPARQL 1.1 Query (§19.2):
\u must be followed by exactly 4 hexadecimal characters representing a valid Unicode scalar value.
\U must be followed by exactly 8 hexadecimal characters representing a valid Unicode scalar value.
Code points greater than U+10FFFF must be rejected with a QuerySyntaxException.
Isolated surrogate code points in the range [U+D800, U+DFFF] must be rejected with a QuerySyntaxException.
Hexadecimal escapes inside IRI references (e.g. <\u0078>) and literals must be decoded properly to their Unicode code points during lexical analysis / parsing.
In Sparql.g4 (or ANTLR grammar / lexer tokens for IRI_REF, STRING_LITERAL*), check how escape sequences are recognized and whether \u / \U escapes in IRI brackets <...> are currently rejected or untreated.
In lexical decoding utilities or AST builders, verify validation of scalar ranges (> 0x10FFFF) and surrogate rejection (0xD800 - 0xDFFF).
Acceptance Criteria
[ ] Unit tests covering valid and malformed \u / \U escapes in strings, multiline strings, and IRIs.
[ ] Valid Unicode escapes in IRIs and literals compile without syntax errors.
[ ] Out-of-range values (> 0x10FFFF), isolated surrogates, and malformed hex lengths throw QuerySyntaxException.
[ ] Targeted W3C syntax-esc-04.rq and syntax-esc-05.rq tests pass.
[ ] 0 SonarLint issues and cognitive complexity < 15.
Summary
Align ANTLR lexer and parser with SPARQL 1.1 §19.2 grammar for Unicode escapes (
\uXXXXand\UXXXXXXXX) across IRIs and string literals, accepting valid hex encodings while strictly rejecting malformed sequences, isolated surrogates, and code points exceeding0x10FFFF.Target Branch
feature/corese-nextSpecification & Semantics
According to W3C SPARQL 1.1 Query (§19.2):
\umust be followed by exactly 4 hexadecimal characters representing a valid Unicode scalar value.\Umust be followed by exactly 8 hexadecimal characters representing a valid Unicode scalar value.U+10FFFFmust be rejected with aQuerySyntaxException.[U+D800, U+DFFF]must be rejected with aQuerySyntaxException.<\u0078>) and literals must be decoded properly to their Unicode code points during lexical analysis / parsing.GRAPH,OPTIONAL, andUNIONblocks are explicitly tracked in [Query] Enforce SPARQL 1.1 blank node scope validation across graph patterns #571 and are excluded from this task.Targeted W3C SPARQL 1.0 / 1.1 Test Cases
syntax-esc-04.rq(Positive syntax test:<\u0078> :p "xx\u0078")syntax-esc-05.rq(Positive syntax test:<\u0078> :p "xx\u0078")Technical Context & Points to Investigate
Sparql.g4(or ANTLR grammar / lexer tokens forIRI_REF,STRING_LITERAL*), check how escape sequences are recognized and whether\u/\Uescapes in IRI brackets<...>are currently rejected or untreated.> 0x10FFFF) and surrogate rejection (0xD800 - 0xDFFF).Acceptance Criteria
[ ] Unit tests covering valid and malformed
\u/\Uescapes in strings, multiline strings, and IRIs.[ ] Valid Unicode escapes in IRIs and literals compile without syntax errors.
[ ] Out-of-range values (>
0x10FFFF), isolated surrogates, and malformed hex lengths throwQuerySyntaxException.[ ] Targeted W3C
syntax-esc-04.rqandsyntax-esc-05.rqtests pass.[ ] 0 SonarLint issues and cognitive complexity < 15.