Aias00 opened a new issue, #6515: URL: https://github.com/apache/shenyu/issues/6515
### Search before asking - [x] I had searched in the [issues](https://github.com/apache/shenyu/issues) and found no similar issues. ### Apache ShenYu Component shenyu-plugin ### What happened `AiTokenLimiterPlugin` records token usage by reading response chunks and parsing `completion_tokens`. For gzip responses it creates a raw `Inflater` and manually skips the gzip header only in the first buffer: ```java final Inflater inflater = isGzip ? new Inflater(true) : null; final AtomicBoolean headerSkipped = new AtomicBoolean(!isGzip); ... if (headerSkipped.compareAndSet(false, true)) { offset = skipGzipHeader(inBytes); } inflater.setInput(inBytes, offset, inBytes.length - offset); ``` `skipGzipHeader(...)` assumes the entire gzip header is available in that first buffer: ```java private int skipGzipHeader(final byte[] b) { int pos = 10; int flg = b[3] & 0xFF; ... return pos; } ``` A gzip stream can be split across arbitrary `DataBuffer`s. If the first buffer is shorter than the minimum gzip header, or if optional gzip header fields are split across buffers, `skipGzipHeader(...)` throws or returns an offset based on incomplete data. The exception is caught by the outer `doOnNext` handler: ```java } catch (Exception e) { LOG.error("read dataBuffer error", e); } ``` but `headerSkipped` has already been set to `true`, so later buffers are fed to the raw inflater as if the header had been consumed correctly. The decompressed text is then missing or corrupt, `completion_tokens` is not found, and `doFinally(...)` records `0` tokens from the empty/incomplete writer output. This means gzip response chunking can cause AI token usage to be undercounted, allowing later requests that should have been limited. ### Expected behavior The token limiter should handle gzip streams across buffer boundaries. It should either use a streaming gzip decoder, aggregate the bytes needed for decompression, or otherwise wait until the full gzip header is available before setting `headerSkipped` and feeding compressed bytes to the inflater. ### How to reproduce 1. Enable `ai-token-limiter` for an AI streaming endpoint. 2. Have the upstream return `Content-Encoding: gzip` and include usage data such as `"completion_tokens": 100`. 3. Make the gzip response arrive with the gzip header split across multiple `DataBuffer`s. 4. `skipGzipHeader(...)` fails or consumes an incomplete first chunk, later chunks cannot be decompressed correctly, and the plugin records `0` or incomplete token usage. ### Debug logs _No response_ ### Environment Current `master` branch. ### Are you willing to submit a PR? - [ ] Yes I am willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
