Author: ggregory
Date: Tue Sep 22 17:22:38 2026
New Revision: 1938440

Log:
Add thread model docuemnts (version 0.1.0)

Added:
   commons/cms-site/trunk/content/markdown/threat-model-0.1.0.md
   commons/cms-site/trunk/content/markdown/threat-model-short-0.1.0.md
Deleted:
   commons/cms-site/trunk/content/markdown/threat_model.md
Modified:
   commons/cms-site/trunk/content/site.xml
   commons/cms-site/trunk/content/xdoc/security.xml

Added: commons/cms-site/trunk/content/markdown/threat-model-0.1.0.md
==============================================================================
--- /dev/null   00:00:00 1970   (empty, because file is newly added)
+++ commons/cms-site/trunk/content/markdown/threat-model-0.1.0.md       Tue Sep 
22 17:22:38 2026        (r1938440)
@@ -0,0 +1,843 @@
+<!---
+ Licensed to the Apache Software Foundation (ASF) under one or more
+ contributor license agreements.  See the NOTICE file distributed with
+ this work for additional information regarding copyright ownership.
+ The ASF licenses this file to You under the Apache License, Version 2.0
+ (the "License"); you may not use this file except in compliance with
+ the License.  You may obtain a copy of the License at
+
+      https://www.apache.org/licenses/LICENSE-2.0
+
+ Unless required by applicable law or agreed to in writing, software
+ distributed under the License is distributed on an "AS IS" BASIS,
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ See the License for the specific language governing permissions and
+ limitations under the License.
+-->
+# Apache Commons Threat Model version 0.1.0
+
+## Introduction
+
+This page amends the [Apache Commons Security 
page](https://commons.apache.org/security.html).
+A condensed version is available in the
+[threat model summary](threat-model-short-0.1.0.html).
+
+Security researchers should use this document 
([Markdown](https://svn.apache.org/repos/asf/commons/cms-site/trunk/content/markdown/hreat-model-0.1.0.md)).
+Developers should use the [threat model 
summary](threat-model-short-0.1.0.html) 
([Markdown](https://svn.apache.org/repos/asf/commons/cms-site/trunk/content/markdown/hreat-model-short-0.1.0.md)).
+
+For information about reporting or asking questions about security,
+please see the [Apache Commons Security 
page](https://commons.apache.org/security.html).
+
+Please note that binary patches are never provided.
+If you need to apply a source code patch,
+use the building instructions for the component version that you are using.
+
+If you need help building a component,
+or other help following the instructions to mitigate the known vulnerabilities 
listed here,
+please send your questions to the public [user mailing list](mail-lists.html).
+
+## Scope and intended use
+
+Apache Commons is a family of small, focused, reusable Java libraries.
+Each component is added to an application's classpath
+and invoked in-process through its public Java API.
+Most components ship no daemons, no network services, and no long-running 
processes;
+a component runs inside the host application's JVM, on data the host handed it.
+
+Here "public Java API" means the API elements whose accessibility (JLS 6.6.1)
+makes them reachable by an unrelated caller:
+roughly, `public` members of public types
+and `protected` members exposed to subclasses.
+Elements that are documented as internal or private
+but remain reachable for purely technical reasons
+are not part of the public API.
+
+The notable exception is Commons Daemon,
+which ships native launchers (Unix `jsvc`, Windows `procrun`)
+that run a JVM as a system service managing host-supplied user code.
+Its profile diverges from the pure-library baseline;
+this model calls the divergence out where it matters below.
+
+### Two kinds of component
+
+Commons components fall into two kinds,
+and this model keeps them distinct:
+
+- **Decorators of JDK functionality.**
+  The JDK-parity principle applies directly:
+  their behavior follows that of the JDK features they expose.
+- **Low-level building blocks of additional functionality**
+  (parsing, scripting, configuration, server bootstrap, and so on).
+  Unless a component specifies otherwise,
+  these are not meant to be exposed directly to untrusted data.
+  In general they hand raw data back to the caller without any sanitization.
+  For example, an application that uses Compress
+  to extract an archive to the file system
+  is responsible for preventing symbolic-link and path-traversal slips;
+  the raw data is provided as-is.
+
+### User types
+
+This model distinguishes two types of user:
+
+- **Trusted users.**
+  Application developers,
+  who put a component on the classpath,
+  write the calling code, and wire its input and output,
+  and operators or administrators,
+  who supply runtime configuration where a component reads it
+  (configuration files, Daemon service definitions,
+  connection-pool settings, upload limits).
+  They have unrestricted access to the features of the component
+  and of the environment it runs in;
+  an attacker who controls either role
+  has won at a layer below Commons.
+- **Untrusted users.**
+  All other users,
+  in particular whoever authors the content
+  a host application may pass to a component:
+  archives, images, CSV files, multipart request bodies, templates.
+  It is generally unsafe to pass their data to Commons APIs.
+  In some cases the host can validate or sanitize the data
+  before passing it to Commons;
+  if this is not possible, unless otherwise specified,
+  it is not possible to securely process this data with Commons libraries.
+  The exceptions are spelled out per component:
+  the narrow default-configuration promise (see "Assumptions about inputs"),
+  the "secure parsing" guarantee where a component claims it,
+  and FileUpload,
+  whose input is untrusted network data by construction.
+
+The "Recurring untrusted-input surface" column of the family table below
+identifies, per component group,
+the data that untrusted users typically author.
+
+### Caller trust and the JDK-parity principle
+
+A component is intended to be called in-process,
+through its public Java API,
+by a host JVM application,
+on input a trusted user supplied, validated, or sanitized:
+the Commons libraries are low-level libraries
+designed to work with input that is either trusted
+or validated/sanitized by the application using the library.
+At its API a component sees a single actor, the in-process caller;
+it cannot tell which user type originated a given byte,
+so keeping untrusted content away from an API
+that is not documented to accept it
+is the trusted users' responsibility.
+If a host wraps a component in a service,
+the service (not Commons) owns the network trust boundary.
+
+The JDK-parity principle is the family's central test
+for whether a report describes a vulnerability.
+Calls to the Commons API come with the same caveat as calls to the JDK:
+they will usually do what the caller asks.
+A component that reads the file, opens the connection,
+or instantiates the class its arguments name
+is behaving like its JDK counterpart,
+not exhibiting a vulnerability.
+Correspondingly,
+do not report a behavior as a Commons component's vulnerability
+if the same behavior would be considered legitimate for the JDK.
+
+### Component family table
+
+Components are grouped by the recurring untrusted-input surface they expose,
+because that surface, not the component name,
+determines the threat profile.
+Every group is in model.
+
+| Family group                                         | Representative 
components                                                                      
        | Touches outside the process?                                          
                                           | Recurring untrusted-input surface  
                                                              |
+|------------------------------------------------------|--------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------|--------------------------------------------------------------------------------------------------|
+| Pure computation / data structures                   | Lang, Collections, 
Math, Numbers, Primitives, Statistics, Text (core)                              
    | No                                                                        
                                       | None beyond CPU and memory of the 
input given to it                                              |
+| Parsers / format readers                             | Compress, CSV, 
Configuration, Digester, DBCP/Pool (configuration), Imaging, Codec, CLI, BCEL, 
RDF, XML | Reads byte streams, files, or text the host passes in; performs 
external I/O where the parsed format requires it | Malformed or adversarial 
archives, configuration files, images, CSV, class files, command lines  |
+| Class-instantiating configuration / scripting blocks | BeanUtils, 
Configuration, Digester                                                         
            | No (instantiates classes named in a text configuration)           
                                               | Class names in a text 
configuration, as used for server bootstrapping; the input must be trusted |
+| Interpolation evaluators                             | Text 
(`StringSubstitutor`), Configuration (interpolation)                            
                  | Can reach out via lookups                                   
                                                     | Interpolation strings 
that reach a substitutor with active outbound lookups                      |
+| Script / command execution engines                   | JEXL, JXPath, Jelly 
(evaluate scripts and expressions), Exec (runs OS commands)                     
   | Yes: evaluate caller-supplied scripts or spawn child processes             
                                      | The script, expression, or command 
itself; executing it is the feature                           |
+| Upload / streaming input handlers                    | FileUpload            
                                                                                
 | Reads multipart streams the host feeds it; may spool to disk                 
                                    | Multipart bodies, part counts and sizes   
                                                       |
+| OS / network convenience wrappers                    | Net, Email, IO, VFS   
                                                                                
 | Yes: sockets, SMTP, file system                                              
                                    | Hostnames, paths, and addresses the host 
supplies                                                |
+| Crypto wrappers                                      | Crypto                
                                                                                
 | Yes: JNI to native cipher libraries                                          
                                    | Keys, IVs, and ciphertext the host 
supplies                                                      |
+| Native service launcher                              | Daemon (`jsvc`, 
`procrun`)                                                                      
       | Yes: runs a JVM as an OS service and manages a child process lifecycle 
                                          | Native launcher input and service 
configuration the operator supplies                            |
+
+Notes on the groupings:
+
+- Commons Text absorbed Commons Lang's deprecated 
`org.apache.commons.lang3.text` package
+  (for example, Text `StringSubstitutor` supersedes Lang `StrSubstitutor`).
+  Those Lang classes are deprecated but still maintained and remain in scope:
+  whatever this model says about a Text class
+  applies equally to its deprecated Lang counterpart.
+
+### Script and command execution engines: no trust boundary with the caller
+
+JEXL, JXPath, Jelly, and Exec
+exist to evaluate caller-supplied expressions or scripts,
+or to run caller-supplied OS commands.
+For these components the usual code/data distinction collapses by design:
+the script, expression, or command *is* the code being run,
+so "the input got executed" describes the feature, not an injection bug.
+These engines have no trust boundary with their caller
+and execute their input as-is;
+the trust boundary stays entirely in the host, permanently.
+Unlike the parsers,
+no future "secure parsing" guarantee can apply to them:
+the requirement that their input be trusted
+is intrinsic to what they do, not a provisional default.
+A report of the shape "JEXL evaluated my expression" or "Exec ran my command"
+is by design and will be closed as a disclaimed property.
+
+One exception is claimed:
+**JEXL's permission mechanism (`JexlPermissions`)
+is a documented security boundary since JEXL 3.7.0.**
+A demonstrated escape from a properly configured JEXL sandbox
+is a valid finding for JEXL 3.7.0 and later.
+The other three engines claim no sandbox of any kind.
+
+### Commons Daemon
+
+Commons Daemon is itself a process launcher that runs a JVM as an OS service,
+so its native launcher and service-management surface
+are in model as a distinct profile.
+Service configuration is trusted operator input.
+However, the privilege transitions performed by `jsvc` and `procrun`
+(dropping privileges, managing the service lifecycle)
+are a claimed boundary:
+a bug that leaves the child process running with unintended privileges
+is a valid finding.
+
+## Out of scope
+
+- **Validation of untrusted input.**
+  It is unsafe to provide possibly malicious input to Commons libraries
+  unless otherwise specified.
+  Input validation at the application's trust boundary
+  is the host's job, not the component's.
+- **Deserialization code of `Serializable` classes.**
+  When applications use Java serialization to deserialize untrusted input,
+  that is a vulnerability in the application.
+  Although several Commons classes provide custom `readObject` methods
+  and can serve as deserialization gadgets,
+  that is not a vulnerability in Commons.
+  This is the single most important out-of-scope statement in the family.
+- **Behaviors that would be legitimate for the JDK.**
+  Anything a Commons API does
+  that the equivalent JDK call would also do
+  is out of model by the JDK-parity principle.
+- **Network and service-level threats.**
+  Commons ships no service;
+  TLS, authentication, authorization, rate limiting, and network exposure
+  belong to the host application that wraps a component.
+- **Documentation, example, test, and benchmark code shipped in the 
repositories.**
+  Javadoc, `src/test`, examples, and JMH benchmarks
+  are not part of the supported API surface.
+
+## Trust boundaries and data flow
+
+A trust boundary is the point
+where data from a less-trusted party enters code that must not believe it:
+the line at which responsibility for validating data changes hands.
+Commons places that point in the host application,
+not at the Commons API,
+except where a component explicitly takes on a "secure parsing" guarantee (see 
below).
+
+- **The public Java API is an interface, not a trust boundary.**
+  Every byte a Commons method receives
+  arrived because the host application chose to pass it;
+  inside the call, data is treated as already vetted by the caller,
+  and responsibility for the arguments stays with the caller.
+  The host and the library are one trust zone;
+  the host's own boundary
+  (its network endpoint, upload form, or downloaded file)
+  is where untrusted data is supposed to stop.
+- **No trust transition happens inside a Commons component.**
+  A component does not re-authenticate, re-authorize, or sanitize
+  on the host's behalf.
+- **Deserialization is no different.**
+  The custom `readObject` methods of Commons `Serializable` classes
+  are entry points of the same kind:
+  they assume the stream they are handed comes from a trusted source.
+  Deserializing does not move the boundary;
+  a host that feeds an untrusted stream to `ObjectInputStream`
+  has carried untrusted data across its own boundary unvalidated
+  (see "Out of scope").
+- **A "secure parsing" guarantee moves the boundary into the component.**
+  Where a component claims it,
+  the component volunteers to accept data straight from the untrusted side
+  and contain the harms that guarantee names;
+  everywhere else the boundary stays in the host.
+
+## Assumptions about the environment
+
+- **Runtime:**
+  a conformant JVM at the component's documented minimum Java baseline.
+  Behavior on unsupported or end-of-life JVMs is undefined.
+- **Concurrency:**
+  thread safety is per-class, not a family guarantee.
+  Treat a Commons class as not thread-safe
+  unless its Javadoc explicitly states that it is;
+  the host must not share a non-thread-safe instance across threads.
+- **Memory:**
+  the host provides adequate heap;
+  a component makes no promise about behavior under host-imposed memory 
pressure.
+- **Side effects:**
+  - The pure-computation group performs no I/O
+    beyond consuming the values the caller handed it,
+    opens no sockets, spawns no processes, and installs no signal handlers.
+  - Parser groups do perform external I/O
+    when the parsed format requires it.
+    For example, Configuration can include additional properties files via an 
`include` key
+    (bounded to local resources in current releases).
+    Such format-driven I/O is a feature, not a vulnerability.
+  - Environment variables are read by several components,
+    including Configuration, Exec, JCS, Lang, Net, and Text;
+    "reads no environment variables" is not a family-wide property.
+  - The OS/network convenience group (Net, Email, IO, VFS)
+    exists precisely to do I/O with host-supplied targets;
+    the side effects are the feature,
+    and their safety is the host's responsibility.
+  - Temporary files are created by FileUpload, Compress, Exec, and IO.
+  - Native code is called by Crypto (JNI to native cipher libraries)
+    and by Daemon (native launchers).
+    Exec spawns child processes but is pure Java itself.
+
+## Configuration variants that matter to security
+
+Commons components are plain Maven-built JARs with few compile-time knobs;
+the security-relevant variation is runtime configuration,
+concentrated in a handful of components.
+
+| Knob or mode                                                                 
               | Component                                 | Default            
                                 | Effect on the model                          
                                                                                
                                              |
+|---------------------------------------------------------------------------------------------|-------------------------------------------|-----------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
+| Interpolation lookups that reach out: `script`, `url`, `dns`                 
               | Text (`StringSubstitutor`), Configuration | Disabled by 
default                                 | Re-enabling such a lookup re-opens a 
remote-code-execution / SSRF surface if untrusted strings reach the 
substitutor; findings against that configuration are out of model |
+| Ambient-state lookups: `env`, `sys`, `file`, `xml`, `properties`, ...        
               | Text (`StringSubstitutor`), Configuration | Enabled by default 
                                 | Untrusted templates can embed ambient 
process state in output; see "Assumptions about outputs"                        
                                                     |
+| Upload limits (`sizeMax`, `fileSizeMax`, `fileCountMax`)                     
               | FileUpload                                | Unset (unbounded) 
unless the host sets them         | Internal bounds are a configuration feature 
the host owns, more than a security property                                    
                                               |
+| Per-format decompression limits                                              
               | Compress                                  | No global output 
cap; some formats expose limits    | Same framing: internal bounds are a 
feature; safety against malicious archives is the separate "secure parsing" 
guarantee                                                  |
+
+FileUpload and Compress are building blocks of pipelines
+that process both trusted and untrusted data.
+The mechanism that limits output is normally left to the host,
+so it can pick the limiting logic that best fits the pipeline.
+Limiting internal sizes is exposed as a configuration option,
+but as a feature more than a security property:
+controlling the size of the input,
+and in most cases the size of the output,
+is up to the host.
+
+## Assumptions about inputs
+
+**The general rule:
+all input is assumed trusted (not attacker-controllable).**
+The trusted users decide whether to pass attacker-controllable input to a 
component;
+the family does not take on an adversary the host itself chose to admit.
+Commons accepts input only through public API parameters.
+
+**The exception, a narrow default-configuration promise:**
+for **Compress, Configuration, Text, and FileUpload**,
+passing any data to them in their default configuration will not:
+
+- cause excessive CPU consumption,
+- cause excessive memory allocation,
+- execute remote code.
+
+This promise covers availability and integrity only.
+It does not cover confidentiality (see "Assumptions about outputs"),
+and these components still do not sanitize their output:
+malicious input can produce malicious output;
+the garbage-in / garbage-out principle applies.
+
+Representative entry points by family group:
+
+| Component / entry point | Parameter | Attacker-controllable? | Caller must 
enforce |
+| --- | --- | --- | --- |
+| `StringUtils`, Math (pure computation) | value arguments | Effectively no: 
data-in / data-out on trusted input | Nothing beyond the documented contract |
+| `CompressorInputStream` / `ArchiveInputStream` (Compress) | compressed or 
archive bytes | Yes, if the host feeds untrusted archives | On versions without 
the "secure parsing" guarantee: cap decompressed size, entry count, and wall 
clock. Always: validate entry paths before writing extracted files |
+| `StringSubstitutor.replace(...)` (Text) | template string | Yes, if the host 
passes untrusted templates | Do not pass untrusted strings to a substitutor 
with active outbound lookups; mind ambient-state lookups in output |
+| Configuration interpolation | configuration values | Yes, if the 
configuration is untrusted | Treat configuration as trusted, or disable 
interpolation lookups |
+| FileUpload parse | multipart stream, part count, sizes | Yes; this is 
network input by construction | Set `sizeMax`, `fileSizeMax`, `fileCountMax`; 
sanitize filenames |
+| Net, Email | hostname, address, path | Yes, if host-supplied from an 
untrusted source | Validate or allow-list targets; the wrapper does not |
+| Crypto | key, IV, ciphertext | Host-supplied | Key management, IV 
uniqueness, integrity (AEAD or MAC) |
+
+No family-wide input bound exists;
+individual components may document format-specific limits.
+Parsers are the acute case:
+a small input can drive large CPU or memory use
+(decompression bombs, deeply nested structures, catastrophic regular 
expressions).
+
+## Assumptions about outputs
+
+Commons components are as neutral on the way out as on the way in.
+Four rules govern outputs:
+
+1. **Output derived from untrusted input is itself untrusted.**
+   Components do not sanitize output.
+   Malicious input can produce malicious output;
+   the garbage-in / garbage-out principle applies.
+   Where a component offers both a raw accessor and a safe consumer,
+   choosing is the host's job:
+   for example, Compress exposes `ArchiveEntry.resolveIn`
+   so the host can consume entry names without path traversal,
+   while `getName` returns the raw, possibly hostile string.
+   Path validation before writing extracted files is,
+   and will remain, the host's responsibility.
+2. **Outputs may embed ambient process state.**
+   Wherever an enabled feature reads ambient state,
+   that state can appear in output.
+   The acute case is interpolation:
+   the default lookup sets of Text and Configuration
+   still resolve prefixes such as `env`, `sys`, `file`, `xml`, and 
`properties`,
+   so an untrusted template like `${env:SECRET}`
+   embeds process state in the result even in default configuration.
+   The component promises nothing about where output flows.
+   A host that passes untrusted templates
+   and reflects the result to a requester
+   has built the information disclosure itself;
+   confidentiality of output is an output-flow responsibility of the host.
+3. **Writers and encoders aim to preserve output structure.**
+   Where a component generates a structured format,
+   it aims to prevent untrusted content
+   from breaking the structure of the generated document:
+   CSV's printer quotes cell content
+   so that a cell value does not inject a record or field;
+   Codec's encoders emit well-formed output for their format;
+   Text's `StringEscapeUtils` escape functions
+   produce output that is inert in the target syntax they document.
+   This is an aim, not a claimed security property:
+   a quoting or escaping bug
+   that lets content break the generated structure
+   is an ordinary bug, not a security vulnerability,
+   unless a component specifies otherwise.
+4. **Passive sinks are in model; active sinks are not.**
+   A passive sink consumes output in its original format.
+   An active sink reinterprets output in another language:
+   a spreadsheet evaluating a CSV cell that begins with `=` as a formula,
+   a terminal executing ANSI escape sequences embedded in a string.
+   Commons aims to preserve structure,
+   not safety in every language in which a downstream consumer might 
reinterpret the output.
+   Defending active sinks is the host's responsibility.
+
+Side-effect outputs
+(spooled temporary files, extracted files, archives the component writes)
+follow the same rules:
+the component creates what it was asked to create,
+and the host owns where those artifacts land and who can read them.
+
+## Adversary model
+
+- **In scope:**
+  whoever supplies the content a host passes into a component
+  through a channel the component's own documentation says is safe for 
untrusted content;
+  primarily, malicious input aimed at a component
+  that claims the "secure parsing" guarantee,
+  and malformed input aimed at APIs documented to tolerate it.
+- **The default position is that malicious input is out of scope:**
+  it is unsafe to provide possibly malicious input to Commons libraries
+  unless otherwise specified.
+  The in-scope adversary is narrow.
+- **Explicitly out of scope:**
+  - An attacker who already controls the calling code;
+    they have won at a layer below Commons.
+  - An attacker whose capability is "the host deserialized my bytes";
+    that is the host's vulnerability, not Commons's.
+  - An attacker exploiting a behavior the JDK would also exhibit.
+  - Side-channel, timing, and co-tenant adversaries.
+- **Attacker goals considered:**
+  availability (CPU or memory exhaustion via crafted parser input)
+  and ambient I/O (making a parser touch resources the caller did not 
designate)
+  against components claiming the "secure parsing" guarantee;
+  escape from the JEXL sandbox (JEXL 3.7.0+);
+  privilege retention across Daemon's privilege transitions.
+
+## Security properties the project provides
+
+These are **not** security properties:
+
+- **Functional correctness** is **not** a security property.
+  Wrong output on valid input is an ordinary bug.
+- **Rejecting malformed input** is **not** a security property.
+  Malformed input should not drive excessive CPU or memory use,
+  but beyond that the garbage-in / garbage-out principle applies.
+
+The claimed properties are:
+
+1. **Memory safety.**
+   Pure-Java components inherit memory safety from the JVM.
+   The two native components,
+   Crypto (JNI to native cipher libraries)
+   and Daemon (native launchers),
+   claim memory safety for **Commons's own JNI and native glue code**:
+   a reproducible crash or memory corruption in that code,
+   reachable through the component's API,
+   is security-critical and a valid report.
+   Crashes originating inside the underlying native library itself
+   (for example, OpenSSL)
+   are findings against that upstream project, not against Commons.
+2. **Expected resource use for the requested operation.**
+   A component uses the memory and CPU
+   one would expect for the specific usage,
+   with no surprises.
+   Often that is linear in the input,
+   but the guarantee is per-operation:
+   Statistics, for example,
+   is expected to use constant memory for degree moments,
+   linear memory for rank statistics,
+   and CPU in line with a decent sorting algorithm.
+   Violations on in-contract input are correctness bugs;
+   for a component that claims the "secure parsing" guarantee below,
+   the same blowup on malicious input is a valid security finding.
+   This property exists chiefly so that the numeric components,
+   Math and Statistics,
+   have a property to be evaluated against;
+   without it those libraries would have
+   no trust boundary or security property to check at all.
+   The math-related components
+   (Math and Statistics, and also RNG, Numbers, and Geometry)
+   should additionally follow Oracle's Secure Coding Guidelines for Java SE,
+   Guideline 3-9 / INJECT-9
+   ("Prevent injection of exceptional floating point values"):
+   check for infinities and NaN
+   and raise an exception where the value does not make sense.
+   This property is under discussion on the security list
+   and may be revised or withdrawn.
+3. **"Secure parsing" (per-component property).**
+   This is **not** a family-wide property.
+   Where a component takes it on,
+   passing arbitrary (including malicious) input to it will not:
+   - cause excessive CPU consumption,
+   - cause excessive memory allocation,
+   - execute code,
+   - cause reads from, or connections to,
+     anything the caller did not explicitly designate
+     (the no-ambient-I/O clause,
+     which covers XML external entities,
+     include-file chasing,
+     and server-side request forgery driven through a parser).
+
+   Catastrophic regular-expression backtracking triggered by malicious input
+   counts against the excessive-CPU clause for components claiming this 
guarantee.
+   The guarantee is introduced component by component, in waves;
+   each component's threat model documents it
+   with a "since" value that states the initial version that guarantees it.
+
+   **Experimental scan configuration (2026-08-12):**
+   for the current LLM security scan only,
+   the parser components
+   (BCEL, CLI, Codec, Compress, Configuration, CSV,
+   DBCP/Pool configuration parsing, Digester, FileUpload, Imaging, RDF, XML)
+   and the interpolation/expression engines
+   (Text's `StringSubstitutor`, Configuration interpolation,
+   JEXL, JXPath, Jelly)
+   are treated as if they claimed this guarantee,
+   to maximize findings.
+   For the expression engines the code-execution clause
+   does not apply to by-design evaluation of the supplied expression;
+   the resource clauses, the no-ambient-I/O clause
+   (the documented default lookup set counts as caller-designated),
+   and the JEXL 3.7.0+ sandbox do the work there.
+   This is a scan-harness setting, not a commitment;
+   it may be narrowed or withdrawn per component
+   after the results are reviewed,
+   and findings that violate only this experimental claim
+   are hardening input, not vulnerabilities against the baseline.
+4. **The JEXL sandbox, since JEXL 3.7.0.**
+   `JexlPermissions` is a documented security boundary;
+   an escape from a properly configured sandbox
+   is a valid finding for JEXL 3.7.0 and later.
+5. **Daemon privilege transitions.**
+   `jsvc` and `procrun` performing a documented privilege drop
+   must not leave the managed child with unintended privileges.
+
+There is deliberately no family-level guarantee of:
+constant-time behavior,
+collision resistance of non-cryptographic hashes and checksums,
+cryptographic randomness from non-cryptographic RNGs,
+or safe handling of untrusted serialized, expression, or archive input
+beyond the per-component "secure parsing" waves.
+
+## Security properties the project does not provide
+
+This is the load-bearing section for the family,
+and one rule quantifies over all of it:
+**no safety against untrusted or malicious input, by default.**
+It is unsafe to provide possibly malicious input to Commons libraries
+unless otherwise specified;
+the host must validate or sanitize at its boundary.
+The "secure parsing" guarantee above
+is the per-component, versioned exception, not the rule.
+The subsections that follow name the recurring weakness classes
+by their [CWE](https://cwe.mitre.org/) identifiers;
+each is an instance of the general rule,
+recorded here with the refinements that apply to it.
+
+### [CWE-502](https://cwe.mitre.org/data/definitions/502.html): 
Deserialization of Untrusted Data
+
+**Commons provides no deserialization safety.**
+Java deserialization was never designed with security in mind.
+It is de facto a parser
+whose behavior is driven by whatever classes happen to be on the classpath.
+Commons itself does not deserialize anything.
+Its `Serializable` classes carry deserialization code
+(custom `readObject` methods included)
+only so that hosts that already chose Java serialization can use them;
+that code is not hardened against untrusted streams
+and is out of scope (see "Out of scope").
+
+Known deserialization gadgets **may** be documented as warnings,
+but removing them is not a goal
+and would break backward compatibility;
+fully hardening Java serialization is a lost cause.
+Safer mechanisms have existed for a long time
+(Commons Digester introduced over twenty years ago).
+If you must deserialize semi-trusted data,
+Commons IO's `ValidatingObjectInputStream`
+provides an allow/deny-list wrapper.
+The safe default remains:
+do not deserialize untrusted data at all.
+
+### [CWE-22](https://cwe.mitre.org/data/definitions/22.html): Path Traversal 
("zip-slip")
+
+**Commons provides no path-traversal protection.**
+Validating extraction targets is,
+and will remain, the host's job,
+including for components that claim "secure parsing".
+Compress supports this with a safe API (`ArchiveEntry.resolveIn`)
+but does not enforce it;
+the raw entry name remains available via `getName`.
+
+### [CWE-917](https://cwe.mitre.org/data/definitions/917.html): Expression 
Language Injection (interpolation and expression evaluation)
+
+**Commons provides no expression or interpolation sandbox
+for untrusted templates**,
+with the sole exception of the JEXL permission mechanism (3.7.0+).
+Passing untrusted strings into Text's `StringSubstitutor`,
+Configuration's interpolation, JXPath, or Jelly
+can reach outbound lookups or code evaluation
+if the host has enabled them,
+and can read ambient process state through the default lookups
+(see "Assumptions about outputs").
+
+### [CWE-611](https://cwe.mitre.org/data/definitions/611.html): XML External 
Entity Reference (XXE)
+
+**Commons provides no XXE hardening guarantee at the baseline.**
+Components that claim "secure parsing"
+cover XXE through the no-ambient-I/O clause.
+For all other XML-consuming components,
+an inventory of which components configure their underlying parser defensively
+is in progress;
+until a component's documentation states otherwise,
+assume that safe external-entity and DTD handling
+is your parser configuration's job.
+
+### [CWE-1333](https://cwe.mitre.org/data/definitions/1333.html): Inefficient 
Regular Expression Complexity (ReDoS)
+
+**Commons provides no ReDoS protection at the baseline.**
+Catastrophic regular-expression backtracking on untrusted input
+is a disclaimed denial-of-service surface,
+except against components claiming "secure parsing",
+where it counts as excessive CPU.
+
+### [CWE-918](https://cwe.mitre.org/data/definitions/918.html): Server-Side 
Request Forgery (SSRF)
+
+**The OS and network wrappers provide no transport security,
+authentication, or target validation.**
+Net, Email, IO, and VFS
+act on the hosts, addresses, and paths they are given;
+they do not allow-list targets and provide no SSRF defense.
+This is by design and permanent;
+target validation belongs to the host.
+
+### False friends (features mistaken for security primitives)
+
+- **Codec digests, CRCs, and checksums
+  are integrity-error detection, not MACs.**
+  A CRC32 or a plain MD5/SHA digest over data
+  is not tamper evidence against an adversary
+  ([CWE-345](https://cwe.mitre.org/data/definitions/345.html)).
+- **Lang's `RandomStringUtils` and `RandomUtils` are not CSPRNGs.**
+  Do not use them to generate secrets or tokens
+  ([CWE-338](https://cwe.mitre.org/data/definitions/338.html)).
+- **`StringSubstitutor` and interpolation
+  are templating conveniences, not a sandbox.**
+  They look like harmless string substitution but evaluate lookups;
+  this was the root of CVE-2022-42889 ("Text4Shell");
+  see CWE-917 above.
+- **Crypto is a thin wrapper over native ciphers,
+  not a misuse-resistant cryptography framework.**
+  It will happily let you reuse an IV
+  ([CWE-323](https://cwe.mitre.org/data/definitions/323.html),
+  [CWE-329](https://cwe.mitre.org/data/definitions/329.html)),
+  use ECB mode
+  ([CWE-327](https://cwe.mitre.org/data/definitions/327.html)),
+  or omit a MAC
+  ([CWE-353](https://cwe.mitre.org/data/definitions/353.html)).
+
+## Downstream responsibilities
+
+For the assumptions above to hold,
+the host application must:
+
+1. **Validate or sanitize untrusted input
+   before passing it to any Commons component**,
+   except where a component version claims the "secure parsing" guarantee.
+2. **Never deserialize untrusted data**,
+   regardless of which Commons components are on the classpath.
+3. **Cap decompression and parse resource use**
+   (decompressed size, entry count, nesting depth, wall clock)
+   when feeding untrusted streams to a component
+   that does not claim "secure parsing".
+4. **Sanitize archive-entry paths against traversal**
+   before writing extracted files;
+   prefer `ArchiveEntry.resolveIn` over raw `getName`.
+5. **Keep untrusted strings out of interpolation and expression evaluators**,
+   leave outbound lookups disabled unless the input is trusted,
+   and treat interpolated output as capable of embedding ambient process state.
+6. **Control where output flows.**
+   Output derived from untrusted input is untrusted;
+   do not reflect it to less-trusted parties
+   without encoding for the destination,
+   and account for active sinks (spreadsheets, terminals)
+   that reinterpret content.
+7. **Set FileUpload limits**
+   (`sizeMax`, `fileSizeMax`, `fileCountMax`);
+   these are configuration features the host owns.
+8. **Configure XML parsers**
+   to disable external entities and DTDs
+   where a component consumes host-supplied XML,
+   unless the component documents that it does so itself.
+9. **Allow-list network, SMTP, and file-system targets**
+   given to Net, Email, and VFS,
+   and never pass Exec (or JEXL, JXPath, Jelly)
+   input you would not run yourself.
+10. **Own key management, IV uniqueness, and integrity (AEAD or MAC)**
+    when using Crypto.
+11. **Direct "does this published CVE affect me?" questions
+    to the user mailing list**, not the security list.
+
+## Known misuse patterns
+
+- **Deserializing untrusted data
+  with Collections or BeanUtils on the classpath.**
+  Looks like ordinary object I/O;
+  is a remote-code-execution gadget surface.
+  Instead: never deserialize untrusted bytes;
+  use an allow-list filter or a safe format.
+- **Extracting untrusted archives and writing entries verbatim.**
+  Enables zip-slip path traversal
+  and, on versions without "secure parsing", decompression bombs.
+  Instead: resolve entry paths safely (`ArchiveEntry.resolveIn`),
+  cap output size and entry count.
+- **Interpolating untrusted strings
+  via `StringSubstitutor` or Configuration with active lookups.**
+  Looks like string templating;
+  can be remote code execution or SSRF with outbound lookups enabled,
+  and information disclosure through ambient-state lookups
+  even in default configuration.
+  Instead: never interpolate untrusted input;
+  restrict the lookup set.
+- **Parsing untrusted XML without hardening the parser.**
+  XXE, SSRF, denial of service.
+  Instead: disable external entities and DTDs.
+- **Running FileUpload with no size or count limits on a public endpoint.**
+  Memory exhaustion, disk fill.
+  Instead: set the limits.
+- **Using a non-cryptographic digest or RNG as a security primitive.**
+  See the false friends above.
+- **Serving generated CSV to spreadsheet users
+  without considering formula injection.**
+  The CSV is structurally valid;
+  the spreadsheet is an active sink.
+  Instead: encode or prefix risky cell values for the destination.
+
+## Known non-findings (recurring false positives)
+
+This list can be fed to scanners and automated triage
+as a suppression list.
+
+- **"Collections `InvokerTransformer` enables RCE via deserialization."**
+  Not a Commons finding:
+  the host deserializing untrusted data is the vulnerability;
+  Commons classes being usable as gadgets is out of model.
+- **"Class X is `Serializable` and could be a gadget"** or
+  **"the custom `readObject` method of class X misbehaves on a crafted 
stream."**
+  Same disposition;
+  containing `Serializable` classes is an implementation detail,
+  not a component's purpose,
+  and the deserialization code of those classes is out of scope:
+  a crafted stream can only reach it
+  because the host deserialized untrusted data.
+- **"API Y does what its arguments tell it to
+  (reflection, class loading, file access), therefore dangerous."**
+  If the equivalent JDK call would be considered legitimate,
+  it is not a Commons vulnerability,
+  by the JDK-parity principle.
+- **"Component throws or runs out of memory on fuzzed malicious input."**
+  Out of model
+  unless the component version claims the "secure parsing" guarantee.
+- **"A functional-correctness bug reported as a vulnerability."**
+  Wrong output on valid input is an ordinary bug,
+  not a security finding.
+- **"`StringSubstitutor` can execute lookups."**
+  With outbound lookups disabled (the default)
+  and trusted templates,
+  this is by design.
+- **"Generated CSV opens as formulas in a spreadsheet" /
+  "output contains terminal escape sequences."**
+  Active-sink reinterpretation;
+  writers and encoders aim to preserve structure only,
+  not safety in every language in which a downstream consumer may evaluate the 
output.
+- **"Non-constant-time comparison, non-cryptographic hash, or non-CSPRNG."**
+  By design;
+  these are not security primitives.
+
+## Conditions that would change this model
+
+- A component starts accepting a new input format
+  or gains a network surface
+  (for example, a server mode):
+  re-scope the family table and boundaries.
+- A component changes a security-relevant default
+  (for example, FileUpload ships a default size cap,
+  or a lookup is re-enabled by default):
+  re-rule the relevant cases.
+- A component takes on the "secure parsing" guarantee
+  (the next wave lands):
+  promote the relevant properties from disclaimed to claimed,
+  with a "since" tag.
+- A component publishes or updates a per-component security page
+  that diverges from this baseline:
+  that page wins for that component
+  and is reflected back here.
+  The per-component pass proceeds wave-parsers first:
+  Compress, then Configuration, FileUpload, and Text.
+- The XML-consumer inventory completes:
+  the XXE baseline statement is finalized per component.
+- Any report that cannot be routed to a single disposition below
+  is a model gap
+  and triggers a revision of this document rather than an ad-hoc call.
+
+## Triage dispositions
+
+Every incoming report receives exactly one of these dispositions,
+citing the licensing section of this model.
+
+| Disposition | Meaning | Licensed by |
+| --- | --- | --- |
+| `VALID` | Violates a claimed property (including "secure parsing" on a 
claiming version, the JEXL 3.7.0+ sandbox, or Daemon privilege transitions), 
reachable from attacker-controllable input by an in-scope adversary, and would 
not be legitimate for the JDK. | Security properties provided; Assumptions 
about inputs/outputs; Adversary model |
+| `VALID-HARDENING` | No claimed property broken, but the API makes a known 
misuse easy enough that the component elects to harden (for example, disabling 
a risky lookup by default). Private report; maintainer discretion; usually no 
CVE. | Known misuse patterns |
+| `OUT-OF-MODEL: trusted-input` | Requires attacker control of input the model 
treats as host-validated or trusted. | Assumptions about inputs; Out of scope |
+| `OUT-OF-MODEL: adversary-not-in-scope` | Requires a capability the adversary 
model excludes, including "the host deserialized untrusted bytes." | Adversary 
model |
+| `OUT-OF-MODEL: jdk-parity` | The behavior would be considered legitimate for 
the JDK. | Scope; Known non-findings |
+| `OUT-OF-MODEL: unsupported-component` | Lands in test, example, or benchmark 
code. | Out of scope |
+| `OUT-OF-MODEL: non-default-build` | Only manifests under a discouraged, 
non-default configuration knob (for example, a re-enabled outbound lookup). | 
Configuration variants |
+| `OUT-OF-MODEL: no-secure-parsing-yet` | Malicious-input finding against a 
parser version that does not claim the "secure parsing" guarantee. | Security 
properties provided/not provided |
+| `OUT-OF-MODEL: upstream` | The defect lies in an underlying native or 
third-party library (for example, OpenSSL), not in Commons's own code. | 
Security properties provided (memory safety) |
+| `BY-DESIGN: property-disclaimed` | Concerns a property this model explicitly 
does not provide (deserialization safety, interpolation sandbox, active-sink 
safety, false friends, functional correctness as security). | Security 
properties not provided; Assumptions about outputs |
+| `KNOWN-NON-FINDING` | Matches a recurring false positive above. | Known 
non-findings |
+| `MODEL-GAP` | Cannot be cleanly routed above. | Triggers a model revision |
+
+The family-specific `OUT-OF-MODEL: jdk-parity` disposition
+encodes the Commons rubric's central test
+and is the correct call for a large share of automated findings.
+The `OUT-OF-MODEL: no-secure-parsing-yet` disposition
+tracks the wave rollout:
+as each component takes the guarantee on,
+its findings migrate from this disposition to `VALID`.
\ No newline at end of file

Added: commons/cms-site/trunk/content/markdown/threat-model-short-0.1.0.md
==============================================================================
--- /dev/null   00:00:00 1970   (empty, because file is newly added)
+++ commons/cms-site/trunk/content/markdown/threat-model-short-0.1.0.md Tue Sep 
22 17:22:38 2026        (r1938440)
@@ -0,0 +1,154 @@
+<!---
+ Licensed to the Apache Software Foundation (ASF) under one or more
+ contributor license agreements.  See the NOTICE file distributed with
+ this work for additional information regarding copyright ownership.
+ The ASF licenses this file to You under the Apache License, Version 2.0
+ (the "License"); you may not use this file except in compliance with
+ the License.  You may obtain a copy of the License at
+
+      https://www.apache.org/licenses/LICENSE-2.0
+
+ Unless required by applicable law or agreed to in writing, software
+ distributed under the License is distributed on an "AS IS" BASIS,
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ See the License for the specific language governing permissions and
+ limitations under the License.
+-->
+# Apache Commons Threat Model (summary) version 0.1.0
+
+## Introduction
+
+This page is the short version of the
+[Apache Commons Threat Model](threat-model-0.1.0.html),
+which is authoritative where the two differ.
+Both amend the
+[Apache Commons Security page](https://commons.apache.org/security.html).
+
+Security researchers should use the [Apache Commons Threat 
Model](threat-model-0.1.0.html) 
([Markdown](https://svn.apache.org/repos/asf/commons/cms-site/trunk/content/markdown/hreat-model-0.1.0.md)).
+Developers should use this page 
([Markdown](https://svn.apache.org/repos/asf/commons/cms-site/trunk/content/markdown/hreat-model-short-0.1.0.md)).
+
+## The model in brief
+
+Apache Commons components are small, reusable Java libraries
+that run inside the host application's JVM
+and are invoked in-process through their public Java API.
+The caller is trusted:
+the API is an interface, not a trust boundary,
+and validating untrusted data is the host application's job,
+at the host's own perimeter.
+Calls to the Commons API come with the same caveat as calls to the JDK:
+they will usually do what the caller asks,
+and a behavior that would be legitimate for the JDK
+is not a Commons vulnerability.
+Components do not sanitize output:
+malicious input can produce malicious output,
+and output derived from untrusted input is itself untrusted.
+A component can explicitly claim a "secure parsing" guarantee
+from a documented version onward;
+where none is claimed (currently: no component claims one,
+although for the current security scan
+all parser and interpolation/expression components
+are experimentally treated as claiming it, see below),
+harm from malicious input is not a Commons vulnerability.
+
+## User types
+
+- **Trusted users:** application developers and operators/administrators.
+  They have unrestricted access to the component and its environment;
+  an attacker who controls either role has won at a layer below Commons.
+- **Untrusted users:** everyone else,
+  in particular whoever authors the content a host may pass to a component
+  (archives, images, CSV files, multipart request bodies, templates).
+  It is generally unsafe to pass their data to Commons APIs.
+  In some cases the host can validate or sanitize the data
+  before passing it to Commons.
+  If this is not possible, unless otherwise specified,
+  it is not possible to securely process this data with Commons libraries.
+
+## What Commons guarantees
+
+- **Memory safety** of its own code:
+  managed Java under the JVM,
+  plus Commons's own JNI and native glue in Crypto and Daemon
+  (defects inside underlying native libraries such as OpenSSL go upstream).
+- **Expected resource use** for the requested operation on in-contract input
+  (a correctness expectation, chiefly for the numeric components;
+  excessive use on malicious input is only a security finding
+  under a "secure parsing" guarantee;
+  this property is under discussion and may be revised).
+- **"Secure parsing"** where a component's threat model claims it,
+  with a "since" version:
+  malicious input cannot cause excessive CPU consumption,
+  excessive memory allocation,
+  code execution,
+  or reads from and connections to anything the caller did not designate
+  (for the current security scan,
+  experimentally treated as claimed by all parser
+  and interpolation/expression components, to maximize findings;
+  a scan-harness setting, not a final commitment).
+- **The JEXL sandbox** (`JexlPermissions`), since JEXL 3.7.0.
+- **Daemon privilege transitions:**
+  `jsvc` and `procrun` must not leave the managed child
+  with unintended privileges.
+
+## What Commons does not guarantee
+
+One rule covers it all:
+**no safety against untrusted or malicious input, by default.**
+Recurring instances:
+
+- **[CWE-502](https://cwe.mitre.org/data/definitions/502.html)
+  (deserialization):**
+  Commons never deserializes anything;
+  its `readObject` code assumes trusted streams,
+  and gadget classes are not Commons vulnerabilities.
+  Never deserialize untrusted data.
+- **[CWE-22](https://cwe.mitre.org/data/definitions/22.html)
+  (path traversal, "zip-slip"):**
+  validating extraction targets is permanently the host's job;
+  Compress offers `ArchiveEntry.resolveIn` but does not enforce it.
+- **[CWE-917](https://cwe.mitre.org/data/definitions/917.html)
+  (expression language injection):**
+  no sandbox for untrusted templates
+  (the JEXL permission mechanism, 3.7.0+, excepted);
+  default lookups can embed environment variables and files in output.
+- **[CWE-611](https://cwe.mitre.org/data/definitions/611.html) (XXE):**
+  XML parser hardening is the host's configuration,
+  unless a component documents otherwise.
+- **[CWE-1333](https://cwe.mitre.org/data/definitions/1333.html) (ReDoS):**
+  disclaimed at the baseline;
+  counts as excessive CPU under "secure parsing".
+- **[CWE-918](https://cwe.mitre.org/data/definitions/918.html) (SSRF):**
+  Net, Email, IO, and VFS contact whatever targets they are given;
+  allow-listing is the host's job.
+- **False friends:**
+  Codec digests and CRCs are not MACs;
+  Lang's random utilities are not CSPRNGs;
+  `StringSubstitutor` is not a sandbox;
+  Crypto is a thin cipher wrapper that will not stop misuse.
+- **Writer and encoder output structure:**
+  components aim to keep untrusted content
+  from breaking the structure of generated output
+  (CSV printing, Codec encoding, Text escaping),
+  but a failure there is an ordinary bug,
+  not a security vulnerability,
+  unless a component specifies otherwise.
+- **Active sinks:**
+  structurally valid output can still be reinterpreted downstream
+  (CSV formulas in spreadsheets, ANSI escapes in terminals);
+  defending those is the host's job.
+
+## Reporting
+
+A report is valid if it violates a guarantee above,
+is reachable from input the model treats as attacker-controllable,
+and describes behavior that would not be legitimate for the JDK.
+Report undisclosed vulnerabilities privately,
+as described on the
+[Apache Commons Security page](https://commons.apache.org/security.html);
+questions about published vulnerabilities
+go to the public [user mailing list](mail-lists.html).
+Everything else,
+including the component family table, triage dispositions,
+misuse patterns, and known non-findings,
+is in the [full threat model](threat_model.html).

Modified: commons/cms-site/trunk/content/site.xml
==============================================================================
--- commons/cms-site/trunk/content/site.xml     Tue Sep 22 16:19:36 2026        
(r1938439)
+++ commons/cms-site/trunk/content/site.xml     Tue Sep 22 17:22:38 2026        
(r1938440)
@@ -160,7 +160,10 @@
 
     <!-- Should agree with commons-parent -->
     <menu name="General Information">
-      <item name="Security"               href="/security.html"/>
+      <item name="Security"                       href="/security.html">
+        <item name="Thread Model for Developers"  
href="/threat-model-short-0.1.0.html" />
+        <item name="Thread Model for Researchers" 
href="/threat-model-0.1.0.html" />
+      </item>
       <item name="Volunteering"           href="/volunteering.html" />
       <item name="Contributing Patches"   href="/patches.html" />
       <item name="Building Components"    href="/building.html" />

Modified: commons/cms-site/trunk/content/xdoc/security.xml
==============================================================================
--- commons/cms-site/trunk/content/xdoc/security.xml    Tue Sep 22 16:19:36 
2026        (r1938439)
+++ commons/cms-site/trunk/content/xdoc/security.xml    Tue Sep 22 17:22:38 
2026        (r1938440)
@@ -87,6 +87,12 @@
           in the Java SE Core Libraries documentation.
         </p>  
       </section>
+      <section name="Threat Model">
+        <ul>
+          <li><a href="threat-model-short-0.1.0.html">Thread Model 
(short)</a></li>
+          <li><a href="threat-model-0.1.0.html">Thread Model</a></li>
+        </ul>
+      </section>
 
       <section name="Known Security Vulnerabilities">
         <p>Known security vulnerabilities fixed in released versions of

Reply via email to