[ 
https://issues.apache.org/jira/browse/THRIFT-6122?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Dmytro Shteflyuk resolved THRIFT-6122.
--------------------------------------
    Fix Version/s: 0.25.0
       Resolution: Fixed

> Ruby JSONProtocol mishandles Unicode strings and surrogate pairs
> ----------------------------------------------------------------
>
>                 Key: THRIFT-6122
>                 URL: https://issues.apache.org/jira/browse/THRIFT-6122
>             Project: Thrift
>          Issue Type: Bug
>          Components: Ruby - Library
>            Reporter: Dmytro Shteflyuk
>            Assignee: Dmytro Shteflyuk
>            Priority: Major
>             Fix For: 0.25.0
>
>          Time Spent: 20m
>  Remaining Estimate: 0h
>
> h3. Problem
> The Ruby JSON protocol does not consistently apply JSON's UTF-8 string rules.
> When reading a supplementary Unicode character represented as a UTF-16 
> surrogate pair, such as {{\uD83D\uDE00}}, the protocol converts the high 
> surrogate independently. Ruby rejects that standalone surrogate with 
> {{RangeError}} before the low surrogate can be combined.
> When writing, the protocol emits bytes from the supplied Ruby string without 
> first normalizing declared character encodings to UTF-8. A valid string in an 
> encoding such as ISO-8859-1 can therefore produce a JSON payload containing 
> non-UTF-8 bytes.
> h3. Client impact
> Ruby clients and servers cannot read valid JSON payloads that use 
> surrogate-pair escapes for characters outside the Basic Multilingual Plane. 
> They can also emit JSON string values that other JSON implementations reject 
> or misinterpret because the wire payload is not valid UTF-8.
> h3. Reproduction
> With a {{Thrift::JsonProtocol}} backed by {{Thrift::MemoryBufferTransport}}:
> {code:ruby}
> reader = Thrift::JsonProtocol.new(
>   Thrift::MemoryBufferTransport.new('"\\uD83D\\uDE00"')
> )
> reader.read_string
> # RangeError: invalid codepoint 0xD83D in UTF-8
> value = +"caf\xE9"
> value.force_encoding(Encoding::ISO_8859_1)
> transport = Thrift::MemoryBufferTransport.new
> writer = Thrift::JsonProtocol.new(transport)
> writer.write_string(value)
> transport.read(transport.available).bytes
> # [34, 99, 97, 102, 233, 34]
> {code}
> The second payload contains byte {{0xE9}} directly rather than its UTF-8 
> representation {{0xC3 0xA9}}.
> h3. Expected behavior
> The reader should combine a valid high/low UTF-16 surrogate pair into one 
> Unicode scalar value and reject malformed escapes or unpaired surrogates with 
> a typed {{ProtocolException}}.
> The writer should preserve valid UTF-8 strings, transcode strings with a 
> declared source encoding to UTF-8, reject invalid byte sequences before 
> emitting partial JSON, and avoid mutating the caller's string.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to