This is an automated email from the ASF dual-hosted git repository.

eldenmoon pushed a commit to branch master
in repository https://gitbox.apache.org/repos/asf/doris.git


The following commit(s) were added to refs/heads/master by this push:
     new 64d2911f64d [fix](inverted index) Add a norms index property and a BE 
config to skip norms on variant paths (#68039)
64d2911f64d is described below

commit 64d2911f64d0620adb7c79881a5ec3a716ed19aa
Author: lihangyu <[email protected]>
AuthorDate: Thu Sep 17 16:37:23 2026 +0800

    [fix](inverted index) Add a norms index property and a BE config to skip 
norms on variant paths (#68039)
    
    ### What problem does this PR solve?
    
    Problem Summary:
    
    An inverted index on a VARIANT column is copied to every extracted
    subcolumn, so a segment holds one
    copy of it per path. Since #53980 every analyzed index calls
    `setOmitNorms(false)`, and CLucene writes
    norms densely: 12 header bytes plus **one byte per segment row for every
    indexed field**, including
    the rows that hold no value for that field. For a VARIANT column that
    cost is multiplied by the
    number of paths.
    
    On a real table (~1400 paths, 7 non-BKD indexes per path) a single
    base-compacted segment held 9891
    `.nrm` files of 4,055,884 bytes each: **37.36 GiB of norms out of a
    38.47 GiB `.idx`, 97.1%**.
    Segments written before the upgrade had no `.nrm` at all, and the
    non-norms bytes had not grown.
    
    Reproduced end to end on 1M rows / 1000 paths / 11 `field_pattern`
    indexes, by loading with 3.1 and
    then upgrading the same data in place to 4.1:
    
    | | index size after full compaction | `.nrm` |
    | --- | --- | --- |
    | 3.1 | 99.6 MiB | none |
    | 4.1, same data | 934.1 MiB | 875 files x 1,000,012 B = 834.5 MiB |
    
    **9.4x**, and the norms are the whole difference. A/B on the same data:
    norms on 457.3 MiB vs norms
    off 99.6 MiB (**4.6x**), with identical `MATCH` row counts and sums.
    
    #### What this PR changes
    
    This PR gives operators a way to reclaim that space without changing
    what anyone gets by default,
    and applies it the same way to the CLucene formats (V2/V3) and to SNII.
    
    1. **A `norms` index property.** An analyzed index writes norms unless
    it sets `"norms" = "false"`,
    on a VARIANT path exactly as on an ordinary column. A subcolumn copy
    inherits the properties of
    the index it came from, so the property on a whole-column VARIANT index
    also settles every one of
       its subcolumn copies.
    2. **A mutable BE config `inverted_index_skip_norms_for_variant`, off by
    default.** Turning it on
    drops norms for every index on a VARIANT path (an index declared with a
    `field_pattern`, or the
    copy an extracted subcolumn inherits, which carries the path as its
    index suffix), whatever that
    index's property says, so a cluster hit by the size regression can
    recover the space without
       rewriting its index definitions.
    3. **One policy function.** `should_write_index_norms(const
    TabletIndex&)` holds the rule, and the
    three places that decide whether norms get written all call it: the
    CLucene writer, the SNII
    writer, and SNII direct compaction. SNII direct compaction rebuilds
    norms from the merged postings
    rather than copying them; the other SNII compaction path rebuilds the
    index through the SNII
       writer.
    4. **BM25 scoring needs norms from every segment it reads, in both
    formats.** `MATCH` filtering is
    unaffected; a query that computes `score()` on an analyzed index fails
    with
    `INVERTED_INDEX_NOT_SUPPORTED` when any segment of the field has no
    norms.
    - SNII already worked this way (`resolve_snii_scoring_segment`); its
    message now names the
    property and the config instead of suggesting that compaction will bring
    the norms back.
    - CLucene did not. It keeps a field's token count in the `.nrm` header,
    and a row without norms
    reads the fake norm `encodeNorm(0)`, i.e. length 0. With no segment
    carrying norms, `avgdl` was
    0 and `score()` came out NaN. With only some segments carrying norms, as
    while the config is
    being turned on, `avgdl` stayed positive and the rows without norms
    ranked as the shortest
    possible documents, above otherwise identical rows.
    `CollectionStatistics` now refuses such a
    segment before any score is computed. An index that is not analyzed
    never writes norms and is
    collected as before, so `SEARCH` scoring over keyword fields is
    unchanged.
    
    #### Compatibility
    
    - Defaults are unchanged: norms are still written everywhere unless
    someone opts out, and an index
      that has norms scores exactly as before.
    - No storage format change.
    - Turning the config on, or setting `"norms" = "false"`, is for indexes
    that are never ranked with
    `score()`. For such an index, `score()` fails as soon as one segment
    without norms exists.
    - A table upgraded from a version that did not write norms (before 4.0)
    now fails `score()` on its
    analyzed indexes until compaction has rewritten the old segments, where
    it used to return NaN or
      skewed scores.
    - Change the property or the config only after every BE runs this
    version.
    
    ### Release note
    
    Added the `norms` inverted index property and the BE config
    `inverted_index_skip_norms_for_variant`
    (off by default, mutable), which drops BM25 norms for inverted indexes
    on VARIANT paths and can cut
    the index size of a VARIANT column with many indexed paths by several
    times. Both apply to the V2,
    V3 and SNII index storage formats. A query that computes `score()` on an
    analyzed index now fails
    when a segment it reads has no norms, instead of returning NaN or
    inconsistent scores.
    
    ### Check List (For Author)
    
    - Test: Unit Test / Regression test
    - `be/test/storage/segment/inverted_index_writer_test.cpp`:
    `NormsFollowIndexNormsProperty`
    - `be/test/storage/index/snii_writer_test.cpp`:
    `SniiWriterNorms.WritesNormsFollowSharedNormsPolicy`
    -
    
`be/test/storage/index/snii/compaction/snii_compaction_eligibility_test.cpp`:
          `DestinationWritesNormsFollowSharedNormsPolicy`
    -
    `be/test/storage/index/inverted/similarity/collection_statistics_test.cpp`:
    `LegacyV3RejectsSegmentWrittenWithoutNorms` (alone and next to a segment
    with norms),
    `LegacyV3KeywordIndexWithoutNormsIsStillCollected`, and the existing
    SNII rejection tests
    -
    
`fe/fe-core/src/test/java/org/apache/doris/analysis/InvertedIndexPropertiesTest.java`
    - Regression: `inverted_index_p0/test_variant_subcolumn_index_norms`
    (checks the `.nrm` files)
    and `inverted_index_p0/storage_format/test_storage_format_snii_norms`
    (reads norms off the
    scores): `MATCH` still filters, `score()` fails without norms, including
    on a mixed-generation
    table. Both are nonConcurrent, since they flip the BE config. Also
    rerun:
    `test_bm25_score_variant`, `test_omit_norms`,
    `test_search_score_topn_predicates`,
    `test_search_score_cache`, `test_search_score_topn_delete_predicate`,
          `test_storage_format_snii`
    - Behavior changed: No for defaults. `score()` on an analyzed index with
    a segment that has no norms
      now fails instead of returning NaN or skewed scores.
    - Does this need documentation: Yes
    (https://github.com/apache/doris-website/pull/4146)
---
 be/src/common/config.cpp                           |   5 +
 be/src/common/config.h                             |   5 +
 .../index/inverted/inverted_index_parser.cpp       |  19 ++
 .../storage/index/inverted/inverted_index_parser.h |  12 +
 .../index/inverted/inverted_index_writer.cpp       |   2 +-
 .../inverted/similarity/collection_statistics.cpp  |  22 +-
 .../storage/index/snii/compaction/eligibility.cpp  |   8 +-
 be/src/storage/index/snii/snii_index_writer.cpp    |  10 +-
 be/src/storage/index/snii/snii_index_writer.h      |   3 +-
 .../similarity/collection_statistics_test.cpp      |  64 ++++-
 .../snii_compaction_eligibility_test.cpp           |  36 +++
 be/test/storage/index/snii_writer_test.cpp         |  44 ++++
 .../storage/segment/inverted_index_writer_test.cpp | 104 +++++++-
 .../doris/analysis/InvertedIndexProperties.java    |   6 +
 .../apache/doris/analysis/InvertedIndexUtil.java   |  10 +
 .../analysis/InvertedIndexPropertiesTest.java      |  25 ++
 .../test_storage_format_snii_norms.out             |  25 ++
 .../test_variant_subcolumn_index_norms.out         |  17 ++
 .../test_storage_format_snii_norms.groovy          | 230 +++++++++++++++++
 .../test_variant_subcolumn_index_norms.groovy      | 285 +++++++++++++++++++++
 20 files changed, 917 insertions(+), 15 deletions(-)

diff --git a/be/src/common/config.cpp b/be/src/common/config.cpp
index 96eb4362142..a0b479a4eea 100644
--- a/be/src/common/config.cpp
+++ b/be/src/common/config.cpp
@@ -1408,6 +1408,11 @@ DEFINE_mBool(debug_inverted_index_compaction, "false");
 DEFINE_mBool(inverted_index_ram_dir_enable, "true");
 // wheather index by RAM directory when base compaction
 DEFINE_mBool(inverted_index_ram_dir_enable_when_base_compaction, "true");
+// Norms cost one byte per segment row, including rows that hold no value for 
the field. A segment
+// holds one index per variant path, so writing norms for them costs rows * 
paths bytes. Turn this on
+// to leave norms out of every index on a variant path, whatever its "norms" 
property says; BM25
+// scoring (score()) on those indexes then fails.
+DEFINE_mBool(inverted_index_skip_norms_for_variant, "false");
 // use num_broadcast_buffer blocks as buffer to do broadcast
 DEFINE_Int32(num_broadcast_buffer, "32");
 
diff --git a/be/src/common/config.h b/be/src/common/config.h
index f5399093dd5..6bda64a1455 100644
--- a/be/src/common/config.h
+++ b/be/src/common/config.h
@@ -1462,6 +1462,11 @@ DECLARE_mBool(debug_inverted_index_compaction);
 DECLARE_mBool(inverted_index_ram_dir_enable);
 // wheather index by RAM directory when base compaction
 DECLARE_mBool(inverted_index_ram_dir_enable_when_base_compaction);
+// Norms cost one byte per segment row, including rows that hold no value for 
the field. A segment
+// holds one index per variant path, so writing norms for them costs rows * 
paths bytes. Turn this on
+// to leave norms out of every index on a variant path, whatever its "norms" 
property says; BM25
+// scoring (score()) on those indexes then fails.
+DECLARE_mBool(inverted_index_skip_norms_for_variant);
 // use num_broadcast_buffer blocks as buffer to do broadcast
 DECLARE_Int32(num_broadcast_buffer);
 
diff --git a/be/src/storage/index/inverted/inverted_index_parser.cpp 
b/be/src/storage/index/inverted/inverted_index_parser.cpp
index 3cfc3d4b0e9..f92ae7eefc4 100644
--- a/be/src/storage/index/inverted/inverted_index_parser.cpp
+++ b/be/src/storage/index/inverted/inverted_index_parser.cpp
@@ -17,6 +17,8 @@
 
 #include "storage/index/inverted/inverted_index_parser.h"
 
+#include "common/config.h"
+#include "storage/tablet/tablet_schema.h"
 #include "util/string_util.h"
 
 namespace doris {
@@ -117,6 +119,23 @@ std::string 
get_parser_phrase_support_string_from_properties(
     return INVERTED_INDEX_PARSER_PHRASE_SUPPORT_NO;
 }
 
+bool should_write_index_norms(const TabletIndex& index_meta) {
+    // A variant path index (a field_pattern index, or the copy inherited by 
one extracted
+    // subcolumn, which carries the path as its index suffix) is one of 
possibly thousands in a
+    // segment, so its norms can dwarf the data. The config drops them 
whatever the property says,
+    // so that a cluster can reclaim that space without rewriting its index 
definitions.
+    const bool variant_path_index =
+            !index_meta.get_index_suffix().empty() || 
!index_meta.field_pattern().empty();
+    if (variant_path_index && config::inverted_index_skip_norms_for_variant) {
+        return false;
+    }
+    const auto& properties = index_meta.properties();
+    if (auto it = properties.find(INVERTED_INDEX_NORMS_KEY); it != 
properties.end()) {
+        return it->second == INVERTED_INDEX_PARSER_TRUE;
+    }
+    return true;
+}
+
 CharFilterMap get_parser_char_filter_map_from_properties(
         const std::map<std::string, std::string>& properties) {
     if (!properties.contains(INVERTED_INDEX_PARSER_CHAR_FILTER_TYPE)) {
diff --git a/be/src/storage/index/inverted/inverted_index_parser.h 
b/be/src/storage/index/inverted/inverted_index_parser.h
index 6fd5e34d466..329a22570e1 100644
--- a/be/src/storage/index/inverted/inverted_index_parser.h
+++ b/be/src/storage/index/inverted/inverted_index_parser.h
@@ -34,6 +34,8 @@ class Analyzer;
 
 namespace doris {
 
+class TabletIndex;
+
 enum class InvertedIndexParserType {
     PARSER_UNKNOWN = 0,
     PARSER_NONE = 1,
@@ -89,6 +91,9 @@ const std::string INVERTED_INDEX_PARSER_PHRASE_SUPPORT_KEY = 
"support_phrase";
 const std::string INVERTED_INDEX_PARSER_PHRASE_SUPPORT_YES = "true";
 const std::string INVERTED_INDEX_PARSER_PHRASE_SUPPORT_NO = "false";
 
+// Whether an analyzed index stores BM25 norms, which take one byte per row of 
the segment.
+const std::string INVERTED_INDEX_NORMS_KEY = "norms";
+
 const std::string INVERTED_INDEX_PARSER_CHAR_FILTER_TYPE = "char_filter_type";
 const std::string INVERTED_INDEX_PARSER_CHAR_FILTER_PATTERN = 
"char_filter_pattern";
 const std::string INVERTED_INDEX_PARSER_CHAR_FILTER_REPLACEMENT = 
"char_filter_replacement";
@@ -152,6 +157,13 @@ std::string get_parser_mode_string_from_properties(
 std::string get_parser_phrase_support_string_from_properties(
         const std::map<std::string, std::string>& properties);
 
+// Whether an analyzed index writes BM25 norms: the one policy shared by every 
index storage format
+// and by index compaction. Norms cost one byte per row of the segment, 
including rows that have no
+// value for the field. An index writes them unless its "norms" property is 
"false", or unless it
+// is on a variant path while inverted_index_skip_norms_for_variant is on, 
which wins over the
+// property.
+bool should_write_index_norms(const TabletIndex& index_meta);
+
 CharFilterMap get_parser_char_filter_map_from_properties(
         const std::map<std::string, std::string>& properties);
 
diff --git a/be/src/storage/index/inverted/inverted_index_writer.cpp 
b/be/src/storage/index/inverted/inverted_index_writer.cpp
index 519568c28eb..ed03225fae6 100644
--- a/be/src/storage/index/inverted/inverted_index_writer.cpp
+++ b/be/src/storage/index/inverted/inverted_index_writer.cpp
@@ -162,7 +162,7 @@ Status 
InvertedIndexColumnWriter<field_type>::create_field(lucene::document::Fie
     (*field)->setOmitTermFreqAndPositions(
             
!(get_parser_phrase_support_string_from_properties(_index_meta->properties()) ==
               INVERTED_INDEX_PARSER_PHRASE_SUPPORT_YES));
-    if (_should_analyzer) {
+    if (_should_analyzer && should_write_index_norms(*_index_meta)) {
         (*field)->setOmitNorms(false);
     }
 
diff --git a/be/src/storage/index/inverted/similarity/collection_statistics.cpp 
b/be/src/storage/index/inverted/similarity/collection_statistics.cpp
index 69bd4155389..06f88bacbb1 100644
--- a/be/src/storage/index/inverted/similarity/collection_statistics.cpp
+++ b/be/src/storage/index/inverted/similarity/collection_statistics.cpp
@@ -47,7 +47,9 @@ Result<SniiScoringSegmentStats> 
resolve_snii_scoring_segment(uint64_t index_doc_
     if (!has_positions || !has_norms) {
         return 
ResultError(Status::Error<ErrorCode::INVERTED_INDEX_NOT_SUPPORTED, false>(
                 "SNII scoring requires positions and norms; this segment was 
written without "
-                "norms -- rebuild the index or wait for compaction to rewrite 
it"));
+                "positions or without norms. Norms are left out when the index 
sets \"norms\" = "
+                "\"false\" or, for a variant path, when 
inverted_index_skip_norms_for_variant is "
+                "on"));
     }
     return SniiScoringSegmentStats {.doc_count = index_doc_count,
                                     .token_count = sum_total_term_freq};
@@ -313,8 +315,22 @@ Status CollectionStatistics::process_segment(const 
RowsetSharedPtr& rowset,
         index_reader = index_searcher->getReader();
 #endif
         total_segment_docs = std::max(total_segment_docs, 
index_reader->maxDoc());
-        _total_num_tokens[ws_field_name] +=
-                
index_reader->sumTotalTermFreq(ws_field_name.c_str()).value_or(0);
+        // BM25 on an analyzed index needs the record length of every row, and 
CLucene keeps
+        // them, together with the field's token count, in the norms. A 
segment written without
+        // norms would feed a zero avgdl, or rank its rows as zero-length 
documents next to the
+        // segments that have norms, so refuse to score the collection, as 
SNII does. An index
+        // that is not analyzed never writes norms and is left as it is.
+        const auto token_count = 
index_reader->sumTotalTermFreq(ws_field_name.c_str());
+        if (!token_count.has_value() &&
+            segment_v2::inverted_index::InvertedIndexAnalyzer::should_analyzer(
+                    collect_info.index_meta->properties())) {
+            return Status::Error<ErrorCode::INVERTED_INDEX_NOT_SUPPORTED>(
+                    "BM25 scoring requires norms, but segment {} was written 
without norms for "
+                    "field {}. Norms are left out when the index sets 
\"norms\" = \"false\" or, "
+                    "for a variant path, when 
inverted_index_skip_norms_for_variant is on",
+                    seg_path, StringHelper::to_string(ws_field_name));
+        }
+        _total_num_tokens[ws_field_name] += token_count.value_or(0);
 
         for (const auto& logical_term_bytes : collect_info.unique_terms) {
             const auto logical_term =
diff --git a/be/src/storage/index/snii/compaction/eligibility.cpp 
b/be/src/storage/index/snii/compaction/eligibility.cpp
index 7ce91bd3ab4..9f958d9ebd5 100644
--- a/be/src/storage/index/snii/compaction/eligibility.cpp
+++ b/be/src/storage/index/snii/compaction/eligibility.cpp
@@ -27,6 +27,7 @@
 #include "common/config.h"
 #include "common/exception.h"
 #include "storage/index/inverted/analyzer/analyzer.h"
+#include "storage/index/inverted/inverted_index_parser.h"
 #include "storage/index/snii/format/format_constants.h"
 #include "storage/index/snii/format/phrase_bigram.h"
 #include "storage/index/snii/reader/logical_index_reader.h"
@@ -244,8 +245,11 @@ Status validate_snii_compaction_eligibility(
                                                              source_ordinal));
     }
     RETURN_IF_ERROR(validate_destination_policy(destination_index, 
analyzer_provider_factory));
-    out->destination_writes_norms =
-            
inverted_index::InvertedIndexAnalyzer::should_analyzer(destination_index.properties());
+    // Merged norms are rebuilt from the postings, so the destination follows 
the same norms
+    // policy as a fresh write whether or not the sources carry norms.
+    out->destination_writes_norms = 
inverted_index::InvertedIndexAnalyzer::should_analyzer(
+                                            destination_index.properties()) &&
+                                    
should_write_index_norms(destination_index);
     return Status::OK();
 }
 
diff --git a/be/src/storage/index/snii/snii_index_writer.cpp 
b/be/src/storage/index/snii/snii_index_writer.cpp
index 94ec6610c0b..5d6d3094266 100644
--- a/be/src/storage/index/snii/snii_index_writer.cpp
+++ b/be/src/storage/index/snii/snii_index_writer.cpp
@@ -29,6 +29,7 @@
 #include "common/logging.h"
 #include "storage/index/index_file_writer.h"
 #include "storage/index/inverted/analyzer/analyzer.h"
+#include "storage/index/inverted/inverted_index_parser.h"
 #include "storage/index/inverted/query/query_info.h"
 #include "storage/index/snii/query/bm25_scorer.h"
 #include "storage/index/snii/writer/global_memory_limiter.h"
@@ -115,10 +116,11 @@ Status SniiIndexColumnWriter::init() {
         return Status::Error<ErrorCode::INVERTED_INDEX_ANALYZER_ERROR>(
                 "SNII create analyzer failed: {}", e.what());
     }
-    // A2: Analyzed indexes with positions always write norms (tokens per 
document, clamped to
-    // 1..255), matching CLucene's scoring capabilities. Keyword or 
positionless indexes omit
-    // them. Norms are an optional core-metadata region ignored by older 
readers.
-    _writes_norms = _should_analyzer && _has_positions;
+    // A2: Analyzed indexes with positions write norms (tokens per document, 
clamped to 1..255),
+    // matching CLucene's scoring capabilities, unless the shared norms policy 
turns them off.
+    // Keyword or positionless indexes omit them. Norms are an optional 
core-metadata region
+    // ignored by older readers.
+    _writes_norms = _should_analyzer && _has_positions && 
should_write_index_norms(*_index_meta);
     return Status::OK();
 }
 
diff --git a/be/src/storage/index/snii/snii_index_writer.h 
b/be/src/storage/index/snii/snii_index_writer.h
index b17624dbfbf..ed54a4657bb 100644
--- a/be/src/storage/index/snii/snii_index_writer.h
+++ b/be/src/storage/index/snii/snii_index_writer.h
@@ -93,7 +93,8 @@ private:
     bool _should_analyzer = false;
     bool _has_positions = false;
     const bool _is_char;
-    // A2: Analyzed indexes with positions always write BM25 norms, matching 
CLucene.
+    // A2: Analyzed indexes with positions write BM25 norms, matching CLucene, 
unless the shared
+    // norms policy (should_write_index_norms) turns them off.
     bool _writes_norms = false;
     // Latch: set_direct_load() ran. The first call wins; a repeat or late call
     // is ignored (and logged) so one index keeps one stable compression-tier
diff --git 
a/be/test/storage/index/inverted/similarity/collection_statistics_test.cpp 
b/be/test/storage/index/inverted/similarity/collection_statistics_test.cpp
index 21b62a02f18..f9c7693b67b 100644
--- a/be/test/storage/index/inverted/similarity/collection_statistics_test.cpp
+++ b/be/test/storage/index/inverted/similarity/collection_statistics_test.cpp
@@ -375,7 +375,9 @@ protected:
         return splits;
     }
 
-    TabletSchemaSPtr create_legacy_v3_schema() {
+    TabletSchemaSPtr create_legacy_v3_schema(std::map<std::string, 
std::string> properties = {
+                                                     {"parser", "standard"},
+                                                     {"support_phrase", 
"true"}}) {
         TabletSchemaPB schema_pb;
         schema_pb.set_keys_type(DUP_KEYS);
         
schema_pb.set_inverted_index_storage_format(InvertedIndexStorageFormatPB::V3);
@@ -392,8 +394,7 @@ protected:
         index._index_id = 1;
         index._index_type = IndexType::INVERTED;
         index._col_unique_ids.push_back(1);
-        index._properties["parser"] = "standard";
-        index._properties["support_phrase"] = "true";
+        index._properties = std::move(properties);
         tablet_schema->append_index(std::move(index));
         return tablet_schema;
     }
@@ -910,6 +911,63 @@ TEST_F(CollectionStatisticsTest, 
LegacyV3SkipsEmptySegmentAfterCollectingAvailab
     expect_collected_term(L"1", L"alpha", 1);
 }
 
+// BM25 needs norms from every segment of an analyzed index: a segment written 
without them makes
+// the whole collection refuse to score, on its own or next to segments that 
have norms.
+TEST_F(CollectionStatisticsTest, LegacyV3RejectsSegmentWrittenWithoutNorms) {
+    auto with_norms_schema = create_legacy_v3_schema();
+    auto without_norms_schema = create_legacy_v3_schema(
+            {{"parser", "standard"}, {"support_phrase", "true"}, {"norms", 
"false"}});
+    const std::string with_norms_path = test_dir_ + 
"/legacy_v3_with_norms_0.dat";
+    const std::string without_norms_path = test_dir_ + 
"/legacy_v3_without_norms_1.dat";
+    ASSERT_TRUE(write_legacy_v3_segment(with_norms_schema, 
with_norms_path).ok());
+    ASSERT_TRUE(write_legacy_v3_segment(without_norms_schema, 
without_norms_path).ok());
+
+    auto collect = [&](const std::vector<std::string>& segment_paths) {
+        auto rowset_meta = 
std::make_shared<collection_statistics::MockRowsetMeta>();
+        auto rowset = 
std::make_shared<collection_statistics::MockRowset>(without_norms_schema,
+                                                                          
rowset_meta);
+        rowset->set_num_segments(static_cast<int>(segment_paths.size()));
+        for (size_t i = 0; i < segment_paths.size(); ++i) {
+            rowset->set_segment_path(static_cast<int>(i), segment_paths[i]);
+        }
+        auto reader = 
std::make_shared<collection_statistics::MockRowsetReader>(rowset);
+        std::vector<RowSetSplits> splits {RowSetSplits(reader)};
+        return stats_->collect(runtime_state_.get(), splits, 
without_norms_schema,
+                               create_match_expr_contexts("alpha"), nullptr);
+    };
+
+    Status status = collect({without_norms_path});
+    EXPECT_EQ(status.code(), ErrorCode::INVERTED_INDEX_NOT_SUPPORTED) << 
status;
+    EXPECT_NE(status.to_string().find("written without norms"), 
std::string::npos) << status;
+    expect_no_collected_tokens(L"1");
+
+    status = collect({with_norms_path, without_norms_path});
+    EXPECT_EQ(status.code(), ErrorCode::INVERTED_INDEX_NOT_SUPPORTED) << 
status;
+    expect_no_collected_tokens(L"1");
+}
+
+// An index that is not analyzed never writes norms, and its scoring 
statistics are collected as
+// before.
+TEST_F(CollectionStatisticsTest, 
LegacyV3KeywordIndexWithoutNormsIsStillCollected) {
+    auto tablet_schema = create_legacy_v3_schema({});
+    const std::string segment_path = test_dir_ + "/legacy_v3_keyword_0.dat";
+    ASSERT_TRUE(write_legacy_v3_segment(tablet_schema, segment_path).ok());
+
+    auto rowset_meta = 
std::make_shared<collection_statistics::MockRowsetMeta>();
+    auto rowset = 
std::make_shared<collection_statistics::MockRowset>(tablet_schema, rowset_meta);
+    rowset->set_num_segments(1);
+    rowset->set_segment_path(0, segment_path);
+    auto reader = 
std::make_shared<collection_statistics::MockRowsetReader>(rowset);
+    std::vector<RowSetSplits> splits {RowSetSplits(reader)};
+
+    const Status status = stats_->collect(runtime_state_.get(), splits, 
tablet_schema,
+                                          create_search_contexts("TERM", 
"alpha beta"), nullptr);
+
+    ASSERT_TRUE(status.ok()) << status;
+    expect_collected_stats(L"1", 1, 0);
+    expect_collected_term(L"1", L"alpha beta", 1);
+}
+
 TEST_F(CollectionStatisticsTest, SniiScoringUsesPhysicalStatistics) {
     auto tablet_schema = create_snii_schema();
     auto expr_contexts = create_match_expr_contexts("alpha");
diff --git 
a/be/test/storage/index/snii/compaction/snii_compaction_eligibility_test.cpp 
b/be/test/storage/index/snii/compaction/snii_compaction_eligibility_test.cpp
index bbf5c0493d3..ab912a45557 100644
--- a/be/test/storage/index/snii/compaction/snii_compaction_eligibility_test.cpp
+++ b/be/test/storage/index/snii/compaction/snii_compaction_eligibility_test.cpp
@@ -248,6 +248,42 @@ TEST(SniiCompactionEligibilityTest, 
DestinationWritesNormsExactlyWhenAnalyzed) {
     EXPECT_FALSE(keyword.destination_writes_norms);
 }
 
+// The destination follows the norms policy shared with fresh writes: the 
"norms" property, and on
+// a variant path inverted_index_skip_norms_for_variant, which wins over the 
property.
+TEST(SniiCompactionEligibilityTest, 
DestinationWritesNormsFollowSharedNormsPolicy) {
+    const bool original_skip_norms_for_variant =
+            doris::config::inverted_index_skip_norms_for_variant;
+    auto legacy = open_index({});
+    auto writes_norms = [&legacy](const std::map<std::string, std::string>& 
properties,
+                                  const std::string& index_suffix) {
+        auto source_meta = make_index(properties, doris::IndexType::INVERTED, 
7, index_suffix);
+        auto destination = make_index(properties, doris::IndexType::INVERTED, 
7, index_suffix);
+        std::vector sources {source(*legacy, *source_meta)};
+        compaction::SniiCompactionEligibility eligibility;
+        const Status status = compaction::validate_snii_compaction_eligibility(
+                sources, *destination, &eligibility);
+        EXPECT_TRUE(status.ok()) << status.to_string();
+        return eligibility.destination_writes_norms;
+    };
+    auto norms_off = plain_properties();
+    norms_off["norms"] = "false";
+    auto norms_on = plain_properties();
+    norms_on["norms"] = "true";
+
+    doris::config::inverted_index_skip_norms_for_variant = false;
+    EXPECT_TRUE(writes_norms(plain_properties(), ""));
+    EXPECT_FALSE(writes_norms(norms_off, ""));
+    EXPECT_TRUE(writes_norms(plain_properties(), "v.s_host"));
+    EXPECT_FALSE(writes_norms(norms_off, "v.s_host"));
+
+    doris::config::inverted_index_skip_norms_for_variant = true;
+    EXPECT_TRUE(writes_norms(plain_properties(), ""));
+    EXPECT_FALSE(writes_norms(plain_properties(), "v.s_host"));
+    EXPECT_FALSE(writes_norms(norms_on, "v.s_host"));
+
+    doris::config::inverted_index_skip_norms_for_variant = 
original_skip_norms_for_variant;
+}
+
 TEST(SniiCompactionEligibilityTest, 
RejectsLegacyBigramMarkerBeforeMergeExecution) {
     snii_test::MemoryFile file;
     writer::SniiIndexInput input;
diff --git a/be/test/storage/index/snii_writer_test.cpp 
b/be/test/storage/index/snii_writer_test.cpp
index 2451c8aa578..bae359c434d 100644
--- a/be/test/storage/index/snii_writer_test.cpp
+++ b/be/test/storage/index/snii_writer_test.cpp
@@ -333,6 +333,50 @@ TEST(SniiWriterFailureLatch, 
AnalyzerFailureDiscardsStateAndBlocksFinish) {
     EXPECT_EQ(writer.memory_reporter_for_test(), nullptr);
 }
 
+// The writer follows the norms policy it shares with the CLucene writer and 
SNII compaction: the
+// "norms" property, and on a variant path 
inverted_index_skip_norms_for_variant, which wins over
+// the property.
+TEST(SniiWriterNorms, WritesNormsFollowSharedNormsPolicy) {
+    const bool original_skip_norms_for_variant =
+            doris::config::inverted_index_skip_norms_for_variant;
+    auto writes_norms = [](const std::map<std::string, std::string>& 
extra_properties,
+                           const std::string& index_suffix) {
+        doris::TabletIndexPB index_pb;
+        index_pb.set_index_type(doris::IndexType::INVERTED);
+        index_pb.set_index_id(92);
+        index_pb.set_index_name("norms_policy");
+        index_pb.add_col_unique_id(0);
+        index_pb.set_index_suffix_name(index_suffix);
+        index_pb.mutable_properties()->insert({"parser", "english"});
+        index_pb.mutable_properties()->insert({"support_phrase", "true"});
+        for (const auto& [key, value] : extra_properties) {
+            index_pb.mutable_properties()->insert({key, value});
+        }
+        doris::TabletIndex index_meta;
+        index_meta.init_from_pb(index_pb);
+        doris::segment_v2::SniiIndexColumnWriter writer(nullptr, &index_meta,
+                                                        
doris::FieldType::OLAP_FIELD_TYPE_VARCHAR);
+        const doris::Status status = writer.init();
+        EXPECT_TRUE(status.ok()) << status.to_string();
+        return writer.writes_norms_for_test();
+    };
+
+    doris::config::inverted_index_skip_norms_for_variant = false;
+    EXPECT_TRUE(writes_norms({}, ""));
+    EXPECT_FALSE(writes_norms({{"norms", "false"}}, ""));
+    EXPECT_TRUE(writes_norms({}, "v.s_host"));
+    EXPECT_FALSE(writes_norms({{"norms", "false"}}, "v.s_host"));
+    EXPECT_TRUE(writes_norms({{"field_pattern", "s_*"}}, ""));
+
+    doris::config::inverted_index_skip_norms_for_variant = true;
+    EXPECT_TRUE(writes_norms({}, ""));
+    EXPECT_FALSE(writes_norms({}, "v.s_host"));
+    EXPECT_FALSE(writes_norms({{"norms", "true"}}, "v.s_host"));
+    EXPECT_FALSE(writes_norms({{"field_pattern", "s_*"}}, ""));
+
+    doris::config::inverted_index_skip_norms_for_variant = 
original_skip_norms_for_variant;
+}
+
 TEST(SniiDocIdSinkGrowth, AppendRangeGrowsGeometrically) {
     std::vector<uint32_t> docids;
     doris::snii::query::VectorDocIdSink sink(docids);
diff --git a/be/test/storage/segment/inverted_index_writer_test.cpp 
b/be/test/storage/segment/inverted_index_writer_test.cpp
index 1994e226182..abcb235eeb2 100644
--- a/be/test/storage/segment/inverted_index_writer_test.cpp
+++ b/be/test/storage/segment/inverted_index_writer_test.cpp
@@ -501,7 +501,9 @@ public:
     }
 
     // Helper method to create an inverted index with tokenization enabled
-    void create_tokenized_index(std::string_view rowset_id, int seg_id, bool 
enable_analyzer) {
+    void create_tokenized_index(std::string_view rowset_id, int seg_id, bool 
enable_analyzer,
+                                const std::string& index_suffix = "",
+                                const std::map<std::string, std::string>& 
extra_properties = {}) {
         auto tablet_schema = create_schema();
 
         // Create index meta with tokenization setting
@@ -522,9 +524,15 @@ public:
             // This will make should_analyzer() return true
             (*properties)["parser"] = "standard";
         }
+        for (const auto& [key, value] : extra_properties) {
+            (*properties)[key] = value;
+        }
 
         TabletIndex idx_meta;
         idx_meta.init_from_pb(*index_meta_pb.get());
+        if (!index_suffix.empty()) {
+            idx_meta.set_escaped_escaped_index_suffix_path(index_suffix);
+        }
 
         std::string index_path_prefix 
{InvertedIndexDescriptor::get_index_file_path_prefix(
                 local_segment_path(kTestDir, rowset_id, seg_id))};
@@ -1825,4 +1833,98 @@ TEST_F(InvertedIndexWriterTest, 
NormsFileCreationWithTokenization) {
             << "inverted_index_writer.cpp where .nrm file creation depends on 
_should_analyzer.";
 }
 
+// Norms take one byte per segment row for every indexed path, so an index on 
a variant path (a
+// field_pattern index, or the copy inherited by one extracted subcolumn, 
which carries the path as
+// its index suffix) writes none by default. The "norms" property overrides 
that per index.
+TEST_F(InvertedIndexWriterTest, NormsFollowIndexNormsProperty) {
+    auto make_index_meta = [](const std::string& index_suffix,
+                              const std::map<std::string, std::string>& 
extra_properties) {
+        TabletIndexPB index_meta_pb;
+        index_meta_pb.set_index_type(IndexType::INVERTED);
+        index_meta_pb.set_index_id(1);
+        index_meta_pb.set_index_name("test");
+        index_meta_pb.add_col_unique_id(1); // c2 column id
+        (*index_meta_pb.mutable_properties())["parser"] = "standard";
+        for (const auto& [key, value] : extra_properties) {
+            (*index_meta_pb.mutable_properties())[key] = value;
+        }
+        TabletIndex index_meta;
+        index_meta.init_from_pb(index_meta_pb);
+        if (!index_suffix.empty()) {
+            index_meta.set_escaped_escaped_index_suffix_path(index_suffix);
+        }
+        return index_meta;
+    };
+    auto path_prefix = [this](const std::string& rowset_id, int seg_id) {
+        return std::string 
{InvertedIndexDescriptor::get_index_file_path_prefix(
+                local_segment_path(kTestDir, rowset_id, seg_id))};
+    };
+
+    bool original_skip_norms_for_variant = 
config::inverted_index_skip_norms_for_variant;
+
+    // an analyzed index writes norms wherever it sits, and only "norms" = 
"false" drops them
+    config::inverted_index_skip_norms_for_variant = false;
+
+    create_tokenized_index("plain_column_default", 0, true, "");
+    TabletIndex plain_default = make_index_meta("", {});
+    EXPECT_TRUE(check_norms_file_exists(path_prefix("plain_column_default", 
0), &plain_default))
+            << "an analyzed index must write .nrm by default";
+
+    create_tokenized_index("plain_column_norms_off", 1, true, "", {{"norms", 
"false"}});
+    TabletIndex plain_norms_off = make_index_meta("", {{"norms", "false"}});
+    EXPECT_FALSE(
+            check_norms_file_exists(path_prefix("plain_column_norms_off", 1), 
&plain_norms_off))
+            << "norms = false must drop .nrm for an ordinary column index";
+
+    create_tokenized_index("variant_subcolumn_default", 2, true, "v.s_host");
+    TabletIndex subcolumn_default = make_index_meta("v.s_host", {});
+    
EXPECT_TRUE(check_norms_file_exists(path_prefix("variant_subcolumn_default", 2),
+                                        &subcolumn_default))
+            << "a variant subcolumn index must write .nrm by default too";
+
+    create_tokenized_index("variant_subcolumn_norms_off", 3, true, "v.s_host",
+                           {{"norms", "false"}});
+    TabletIndex subcolumn_norms_off = make_index_meta("v.s_host", {{"norms", 
"false"}});
+    
EXPECT_FALSE(check_norms_file_exists(path_prefix("variant_subcolumn_norms_off", 
3),
+                                         &subcolumn_norms_off))
+            << "norms = false must drop .nrm for a variant subcolumn index";
+
+    create_tokenized_index("field_pattern_default", 4, true, "", 
{{"field_pattern", "s_*"}});
+    TabletIndex field_pattern_default = make_index_meta("", {{"field_pattern", 
"s_*"}});
+    EXPECT_TRUE(check_norms_file_exists(path_prefix("field_pattern_default", 
4),
+                                        &field_pattern_default))
+            << "a field_pattern index must write .nrm by default too";
+
+    // the config drops norms for a variant path index whatever its property 
says, and leaves every
+    // other index alone
+    config::inverted_index_skip_norms_for_variant = true;
+
+    create_tokenized_index("variant_subcolumn_skipped", 5, true, "v.s_host");
+    TabletIndex subcolumn_skipped = make_index_meta("v.s_host", {});
+    
EXPECT_FALSE(check_norms_file_exists(path_prefix("variant_subcolumn_skipped", 
5),
+                                         &subcolumn_skipped))
+            << "the config must drop .nrm for a variant subcolumn index";
+
+    create_tokenized_index("variant_subcolumn_norms_on_skipped", 6, true, 
"v.s_host",
+                           {{"norms", "true"}});
+    TabletIndex subcolumn_norms_on_skipped = make_index_meta("v.s_host", 
{{"norms", "true"}});
+    
EXPECT_FALSE(check_norms_file_exists(path_prefix("variant_subcolumn_norms_on_skipped",
 6),
+                                         &subcolumn_norms_on_skipped))
+            << "the config must win over norms = true on a variant subcolumn 
index";
+
+    create_tokenized_index("field_pattern_skipped", 7, true, "", 
{{"field_pattern", "s_*"}});
+    TabletIndex field_pattern_skipped = make_index_meta("", {{"field_pattern", 
"s_*"}});
+    EXPECT_FALSE(check_norms_file_exists(path_prefix("field_pattern_skipped", 
7),
+                                         &field_pattern_skipped))
+            << "the config must drop .nrm for a field_pattern index";
+
+    create_tokenized_index("plain_column_not_skipped", 8, true, "");
+    TabletIndex plain_not_skipped = make_index_meta("", {});
+    EXPECT_TRUE(
+            check_norms_file_exists(path_prefix("plain_column_not_skipped", 
8), &plain_not_skipped))
+            << "the config must leave an ordinary column index alone";
+
+    config::inverted_index_skip_norms_for_variant = 
original_skip_norms_for_variant;
+}
+
 } // namespace doris::segment_v2
diff --git 
a/fe/fe-catalog/src/main/java/org/apache/doris/analysis/InvertedIndexProperties.java
 
b/fe/fe-catalog/src/main/java/org/apache/doris/analysis/InvertedIndexProperties.java
index b9b5756f17d..451e65d9c52 100644
--- 
a/fe/fe-catalog/src/main/java/org/apache/doris/analysis/InvertedIndexProperties.java
+++ 
b/fe/fe-catalog/src/main/java/org/apache/doris/analysis/InvertedIndexProperties.java
@@ -53,6 +53,12 @@ public class InvertedIndexProperties {
 
     public static String INVERTED_INDEX_SUPPORT_PHRASE_KEY = "support_phrase";
 
+    // Whether an analyzed index stores BM25 norms, which default to being 
stored. Norms cost one
+    // byte per row of the segment for every indexed path, so the BE config
+    // inverted_index_skip_norms_for_variant (off by default) can leave them 
out for every index on
+    // a variant path, which it does whatever this property says.
+    public static String INVERTED_INDEX_NORMS_KEY = "norms";
+
     public static String INVERTED_INDEX_PARSER_IGNORE_ABOVE_KEY = 
"ignore_above";
 
     public static String INVERTED_INDEX_PARSER_LOWERCASE_KEY = "lower_case";
diff --git 
a/fe/fe-core/src/main/java/org/apache/doris/analysis/InvertedIndexUtil.java 
b/fe/fe-core/src/main/java/org/apache/doris/analysis/InvertedIndexUtil.java
index 9f58bde9248..57da08fe994 100644
--- a/fe/fe-core/src/main/java/org/apache/doris/analysis/InvertedIndexUtil.java
+++ b/fe/fe-core/src/main/java/org/apache/doris/analysis/InvertedIndexUtil.java
@@ -70,6 +70,9 @@ public class InvertedIndexUtil {
     public static String INVERTED_INDEX_SUPPORT_PHRASE_KEY =
             InvertedIndexProperties.INVERTED_INDEX_SUPPORT_PHRASE_KEY;
 
+    public static String INVERTED_INDEX_NORMS_KEY =
+            InvertedIndexProperties.INVERTED_INDEX_NORMS_KEY;
+
     public static String INVERTED_INDEX_PARSER_IGNORE_ABOVE_KEY =
             InvertedIndexProperties.INVERTED_INDEX_PARSER_IGNORE_ABOVE_KEY;
 
@@ -202,6 +205,7 @@ public class InvertedIndexUtil {
                 INVERTED_INDEX_PARSER_KEY_ALIAS,
                 INVERTED_INDEX_PARSER_MODE_KEY,
                 INVERTED_INDEX_SUPPORT_PHRASE_KEY,
+                INVERTED_INDEX_NORMS_KEY,
                 INVERTED_INDEX_PARSER_CHAR_FILTER_TYPE,
                 INVERTED_INDEX_PARSER_CHAR_FILTER_PATTERN,
                 INVERTED_INDEX_PARSER_CHAR_FILTER_REPLACEMENT,
@@ -288,6 +292,12 @@ public class InvertedIndexUtil {
                     + ", support_phrase must be true or false");
         }
 
+        String norms = properties.get(INVERTED_INDEX_NORMS_KEY);
+        if (norms != null && !norms.matches("true|false")) {
+            throw new AnalysisException("Invalid inverted index 'norms' value: 
" + norms
+                    + ", norms must be true or false");
+        }
+
         checkCharFilterProperties(properties);
 
         if (ignoreAbove != null) {
diff --git 
a/fe/fe-core/src/test/java/org/apache/doris/analysis/InvertedIndexPropertiesTest.java
 
b/fe/fe-core/src/test/java/org/apache/doris/analysis/InvertedIndexPropertiesTest.java
index a482685cb3b..e439360a026 100644
--- 
a/fe/fe-core/src/test/java/org/apache/doris/analysis/InvertedIndexPropertiesTest.java
+++ 
b/fe/fe-core/src/test/java/org/apache/doris/analysis/InvertedIndexPropertiesTest.java
@@ -425,6 +425,31 @@ public class InvertedIndexPropertiesTest {
         }
     }
 
+    // --- norms ---
+
+    @Test
+    public void testNormsPropertyAccepted() throws AnalysisException {
+        for (String value : new String[] {"true", "false"}) {
+            Map<String, String> props = new HashMap<>();
+            props.put("parser", "english");
+            props.put("norms", value);
+            InvertedIndexUtil.checkInvertedIndexParser("col1", 
PrimitiveType.STRING, props,
+                    TInvertedIndexFileStorageFormat.V2);
+        }
+    }
+
+    @Test
+    public void testNormsPropertyRejectsOtherValues() {
+        Map<String, String> props = new HashMap<>();
+        props.put("parser", "english");
+        props.put("norms", "yes");
+        AnalysisException exception = 
Assertions.assertThrows(AnalysisException.class,
+                () -> InvertedIndexUtil.checkInvertedIndexParser("col1", 
PrimitiveType.STRING, props,
+                        TInvertedIndexFileStorageFormat.V2));
+        Assertions.assertTrue(exception.getMessage().contains("norms must be 
true or false"),
+                exception.getMessage());
+    }
+
     // The SNII gate in checkInvertedIndexParser sees the parent VARIANT type 
on a whole-column
     // index and the sub-column type on a field_pattern index. Only the latter 
can be judged,
     // so VARIANT itself must pass.
diff --git 
a/regression-test/data/inverted_index_p0/storage_format/test_storage_format_snii_norms.out
 
b/regression-test/data/inverted_index_p0/storage_format/test_storage_format_snii_norms.out
new file mode 100644
index 00000000000..5f52500503f
--- /dev/null
+++ 
b/regression-test/data/inverted_index_p0/storage_format/test_storage_format_snii_norms.out
@@ -0,0 +1,25 @@
+-- This file is automatically generated. You should know what you did if you 
want to edit this
+-- !snii_plain_column_score --
+1      0.3541
+2      0.562
+
+-- !snii_field_pattern_score --
+1      0.311
+2      0.5235
+
+-- !snii_field_pattern_no_norms_match --
+1
+2
+
+-- !snii_whole_column_no_norms_match --
+1
+2
+
+-- !snii_config_plain_column_score --
+1      0.3541
+2      0.562
+
+-- !snii_mixed_match --
+1
+2
+
diff --git 
a/regression-test/data/inverted_index_p0/test_variant_subcolumn_index_norms.out 
b/regression-test/data/inverted_index_p0/test_variant_subcolumn_index_norms.out
new file mode 100644
index 00000000000..9b2cf745293
--- /dev/null
+++ 
b/regression-test/data/inverted_index_p0/test_variant_subcolumn_index_norms.out
@@ -0,0 +1,17 @@
+-- This file is automatically generated. You should know what you did if you 
want to edit this
+-- !variant_subcolumn_score --
+2      0.6931
+3      0.61
+
+-- !variant_subcolumn_match_no_norms --
+2
+3
+
+-- !plain_column_score --
+1      0.5754
+3      0.8714
+
+-- !mixed_match --
+1
+2
+
diff --git 
a/regression-test/suites/inverted_index_p0/storage_format/test_storage_format_snii_norms.groovy
 
b/regression-test/suites/inverted_index_p0/storage_format/test_storage_format_snii_norms.groovy
new file mode 100644
index 00000000000..e8888a8d41b
--- /dev/null
+++ 
b/regression-test/suites/inverted_index_p0/storage_format/test_storage_format_snii_norms.groovy
@@ -0,0 +1,230 @@
+// Licensed to the Apache Software Foundation (ASF) under one
+// or more contributor license agreements.  See the NOTICE file
+// distributed with this work for additional information
+// regarding copyright ownership.  The ASF licenses this file
+// to you under the Apache License, Version 2.0 (the
+// "License"); you may not use this file except in compliance
+// with the License.  You may obtain a copy of the License at
+//
+//   http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing,
+// software distributed under the License is distributed on an
+// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+// KIND, either express or implied.  See the License for the
+// specific language governing permissions and limitations
+// under the License.
+
+// SNII follows the same norms policy as the CLucene formats: an analyzed 
index writes BM25 norms
+// unless "norms" = "false", and inverted_index_skip_norms_for_variant drops 
them for every index on
+// a variant path whatever the property says. BM25 scoring needs norms, so 
score() on an index
+// without them fails, while MATCH filtering keeps working. Norms are not 
visible from outside a
+// SNII file, so this suite reads them off the queries: with norms, rows 1 and 
2 match "alpha" once
+// each and the three-token row 1 scores lower than the one-token row 2; 
without norms, score()
+// is refused.
+// It flips a BE config, so it must not share the cluster with other suites.
+suite("test_storage_format_snii_norms", "p0,nonConcurrent") {
+    sql """ set enable_match_without_inverted_index = false """
+    sql """ set default_variant_enable_typed_paths_to_sparse = false """
+    sql """ set default_variant_enable_doc_mode = false """
+
+    sql "DROP TABLE IF EXISTS test_storage_format_snii_norms"
+    sql """
+        CREATE TABLE test_storage_format_snii_norms (
+            id INT,
+            content TEXT,
+            v variant<
+                's_*' : text,
+                't_*' : text,
+                PROPERTIES("variant_max_subcolumns_count"="0")
+            >,
+            vn variant<
+                'b_*' : text,
+                PROPERTIES("variant_max_subcolumns_count"="0")
+            >,
+            INDEX idx_content (content) USING INVERTED PROPERTIES(
+                "parser"="english",
+                "support_phrase"="true"
+            ),
+            INDEX idx_v_s (v) USING INVERTED PROPERTIES(
+                "parser"="english",
+                "support_phrase"="true",
+                "field_pattern"="s_*"
+            ),
+            INDEX idx_v_t (v) USING INVERTED PROPERTIES(
+                "parser"="english",
+                "support_phrase"="true",
+                "field_pattern"="t_*",
+                "norms"="false"
+            ),
+            INDEX idx_vn (vn) USING INVERTED PROPERTIES(
+                "parser"="english",
+                "support_phrase"="true",
+                "norms"="false"
+            )
+        ) ENGINE=OLAP DUPLICATE KEY(id)
+        DISTRIBUTED BY HASH(id) BUCKETS 1
+        PROPERTIES (
+            "replication_allocation" = "tag.location.default: 1",
+            "disable_auto_compaction" = "true",
+            "inverted_index_storage_format" = "SNII"
+        )
+    """
+    sql """ insert into test_storage_format_snii_norms values
+            (1, 'alpha database server',
+                parse_to_variant('{"s_note":"alpha database server", 
"t_note":"alpha database server"}'),
+                parse_to_variant('{"b_note":"alpha database server"}')),
+            (2, 'alpha',
+                parse_to_variant('{"s_note":"alpha", "t_note":"alpha"}'),
+                parse_to_variant('{"b_note":"alpha"}')),
+            (3, 'gamma', parse_to_variant('{"other":"gamma"}'), 
parse_to_variant('{"other":"gamma"}'))
+    """
+    sql " sync "
+
+    // an ordinary column and a field_pattern index keep norms by default: row 
1 scores lower
+    order_qt_snii_plain_column_score """
+        select id, round(score(), 4) from test_storage_format_snii_norms
+        where content match_any "alpha" order by score() desc limit 10
+    """
+    order_qt_snii_field_pattern_score """
+        select id, round(score(), 4) from test_storage_format_snii_norms
+        where cast(v["s_note"] as string) match_any "alpha" order by score() 
desc limit 10
+    """
+
+    // "norms" = "false" drops them, on a field_pattern index and on the 
copies a whole-column
+    // index hands to its subcolumns: MATCH still filters, score() is refused
+    order_qt_snii_field_pattern_no_norms_match """
+        select id from test_storage_format_snii_norms
+        where cast(v["t_note"] as string) match_any "alpha"
+    """
+    test {
+        sql """
+            select id, score() from test_storage_format_snii_norms
+            where cast(v["t_note"] as string) match_any "alpha" order by 
score() desc limit 10
+        """
+        exception "written without"
+    }
+    order_qt_snii_whole_column_no_norms_match """
+        select id from test_storage_format_snii_norms
+        where cast(vn["b_note"] as string) match_any "alpha"
+    """
+    test {
+        sql """
+            select id, score() from test_storage_format_snii_norms
+            where cast(vn["b_note"] as string) match_any "alpha" order by 
score() desc limit 10
+        """
+        exception "written without"
+    }
+
+    // with the config on, every index on a variant path leaves norms out, 
even one that asks for
+    // them, while an ordinary column index keeps them
+    setBeConfigTemporary([inverted_index_skip_norms_for_variant: true]) {
+        sql "DROP TABLE IF EXISTS test_storage_format_snii_norms_config"
+        sql """
+            CREATE TABLE test_storage_format_snii_norms_config (
+                id INT,
+                content TEXT,
+                v variant<
+                    's_*' : text,
+                    PROPERTIES("variant_max_subcolumns_count"="0")
+                >,
+                vf variant<
+                    'c_*' : text,
+                    PROPERTIES("variant_max_subcolumns_count"="0")
+                >,
+                INDEX idx_content (content) USING INVERTED PROPERTIES(
+                    "parser"="english",
+                    "support_phrase"="true"
+                ),
+                INDEX idx_v_s (v) USING INVERTED PROPERTIES(
+                    "parser"="english",
+                    "support_phrase"="true",
+                    "field_pattern"="s_*"
+                ),
+                INDEX idx_vf (vf) USING INVERTED PROPERTIES(
+                    "parser"="english",
+                    "support_phrase"="true",
+                    "norms"="true"
+                )
+            ) ENGINE=OLAP DUPLICATE KEY(id)
+            DISTRIBUTED BY HASH(id) BUCKETS 1
+            PROPERTIES (
+                "replication_allocation" = "tag.location.default: 1",
+                "disable_auto_compaction" = "true",
+                "inverted_index_storage_format" = "SNII"
+            )
+        """
+        sql """ insert into test_storage_format_snii_norms_config values
+                (1, 'alpha database server',
+                    parse_to_variant('{"s_note":"alpha database server"}'),
+                    parse_to_variant('{"c_note":"alpha database server"}')),
+                (2, 'alpha', parse_to_variant('{"s_note":"alpha"}'), 
parse_to_variant('{"c_note":"alpha"}')),
+                (3, 'gamma', parse_to_variant('{"other":"gamma"}'), 
parse_to_variant('{"other":"gamma"}'))
+        """
+        sql " sync "
+    }
+
+    order_qt_snii_config_plain_column_score """
+        select id, round(score(), 4) from test_storage_format_snii_norms_config
+        where content match_any "alpha" order by score() desc limit 10
+    """
+    test {
+        sql """
+            select id, score() from test_storage_format_snii_norms_config
+            where cast(v["s_note"] as string) match_any "alpha" order by 
score() desc limit 10
+        """
+        exception "written without"
+    }
+    test {
+        sql """
+            select id, score() from test_storage_format_snii_norms_config
+            where cast(vf["c_note"] as string) match_any "alpha" order by 
score() desc limit 10
+        """
+        exception "written without"
+    }
+
+    // segments with and without norms side by side, as while the config is 
being turned on:
+    // MATCH still filters, and score() is refused rather than ranking the two 
kinds differently
+    sql "DROP TABLE IF EXISTS test_storage_format_snii_norms_mixed"
+    sql """
+        CREATE TABLE test_storage_format_snii_norms_mixed (
+            id INT,
+            v variant<
+                's_*' : text,
+                PROPERTIES("variant_max_subcolumns_count"="0")
+            >,
+            INDEX idx_v_s (v) USING INVERTED PROPERTIES(
+                "parser"="english",
+                "support_phrase"="true",
+                "field_pattern"="s_*"
+            )
+        ) ENGINE=OLAP DUPLICATE KEY(id)
+        DISTRIBUTED BY HASH(id) BUCKETS 1
+        PROPERTIES (
+            "replication_allocation" = "tag.location.default: 1",
+            "disable_auto_compaction" = "true",
+            "inverted_index_storage_format" = "SNII"
+        )
+    """
+    sql """ insert into test_storage_format_snii_norms_mixed values
+            (1, parse_to_variant('{"s_note":"alpha database server"}'))
+    """
+    setBeConfigTemporary([inverted_index_skip_norms_for_variant: true]) {
+        sql """ insert into test_storage_format_snii_norms_mixed values
+                (2, parse_to_variant('{"s_note":"alpha"}'))
+        """
+    }
+    sql " sync "
+
+    order_qt_snii_mixed_match """
+        select id from test_storage_format_snii_norms_mixed
+        where cast(v["s_note"] as string) match_any "alpha"
+    """
+    test {
+        sql """
+            select id, score() from test_storage_format_snii_norms_mixed
+            where cast(v["s_note"] as string) match_any "alpha" order by 
score() desc limit 10
+        """
+        exception "written without"
+    }
+}
diff --git 
a/regression-test/suites/inverted_index_p0/test_variant_subcolumn_index_norms.groovy
 
b/regression-test/suites/inverted_index_p0/test_variant_subcolumn_index_norms.groovy
new file mode 100644
index 00000000000..5b8d65debc2
--- /dev/null
+++ 
b/regression-test/suites/inverted_index_p0/test_variant_subcolumn_index_norms.groovy
@@ -0,0 +1,285 @@
+// Licensed to the Apache Software Foundation (ASF) under one
+// or more contributor license agreements.  See the NOTICE file
+// distributed with this work for additional information
+// regarding copyright ownership.  The ASF licenses this file
+// to you under the Apache License, Version 2.0 (the
+// "License"); you may not use this file except in compliance
+// with the License.  You may obtain a copy of the License at
+//
+//   http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing,
+// software distributed under the License is distributed on an
+// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+// KIND, either express or implied.  See the License for the
+// specific language governing permissions and limitations
+// under the License.
+
+// An analyzed index writes dense BM25 norms (.nrm, one byte per segment row), 
on a variant path as
+// on any other column, and "norms" = "false" drops them. Norms on a variant 
path cost rows * paths
+// bytes, so inverted_index_skip_norms_for_variant leaves them out there 
whatever the property says.
+// This covers both an index declared with a field_pattern and a whole-column
+// index on a VARIANT column, whose per-subcolumn copies inherit the 
properties of the index they
+// come from. BM25 scoring needs norms: score() on an index without them 
fails, also while only
+// some segments lack them, and MATCH filtering keeps working.
+// It flips a BE config, so it must not share the cluster with other suites.
+suite("test_variant_subcolumn_index_norms", "p0,nonConcurrent") {
+    if (isCloudMode()) {
+        return
+    }
+
+    sql """ set enable_segment_limit_pushdown = true """
+    sql """ set enable_match_without_inverted_index = false """
+    sql """ set default_variant_enable_typed_paths_to_sparse = false """
+    sql """ set default_variant_enable_doc_mode = false """
+
+    sql "DROP TABLE IF EXISTS test_variant_subcolumn_index_norms"
+    sql """
+        CREATE TABLE test_variant_subcolumn_index_norms (
+            id INT,
+            content TEXT,
+            v variant<
+                's_*' : text,
+                't_*' : text,
+                PROPERTIES("variant_max_subcolumns_count"="0")
+            >,
+            vd variant<
+                'a_*' : text,
+                PROPERTIES("variant_max_subcolumns_count"="0")
+            >,
+            vn variant<
+                'b_*' : text,
+                PROPERTIES("variant_max_subcolumns_count"="0")
+            >,
+            INDEX idx_content (content) USING INVERTED PROPERTIES(
+                "parser"="english",
+                "support_phrase"="true"
+            ),
+            INDEX idx_v_s (v) USING INVERTED PROPERTIES(
+                "parser"="english",
+                "support_phrase"="true",
+                "field_pattern"="s_*"
+            ),
+            INDEX idx_v_t (v) USING INVERTED PROPERTIES(
+                "parser"="english",
+                "support_phrase"="true",
+                "field_pattern"="t_*",
+                "norms"="false"
+            ),
+            INDEX idx_vd (vd) USING INVERTED PROPERTIES(
+                "parser"="english",
+                "support_phrase"="true"
+            ),
+            INDEX idx_vn (vn) USING INVERTED PROPERTIES(
+                "parser"="english",
+                "support_phrase"="true",
+                "norms"="false"
+            )
+        ) ENGINE=OLAP DUPLICATE KEY(id)
+        DISTRIBUTED BY HASH(id) BUCKETS 1
+        PROPERTIES (
+            "replication_allocation" = "tag.location.default: 1",
+            "disable_auto_compaction" = "true",
+            "inverted_index_storage_format" = "V2"
+        )
+    """
+    sql """ insert into test_variant_subcolumn_index_norms values
+            (1, 'alpha database server', parse_to_variant('{"s_host":"alpha 
database server"}'),
+                parse_to_variant('{"a_host":"alpha database server"}'),
+                parse_to_variant('{"b_host":"alpha database server"}')),
+            (2, 'beta server cluster', parse_to_variant('{"s_host":"beta 
server cluster", "s_note":"alpha", "t_note":"alpha"}'),
+                parse_to_variant('{"a_host":"beta server cluster"}'),
+                parse_to_variant('{"b_host":"beta server cluster"}')),
+            (3, 'alpha', parse_to_variant('{"s_note":"alpha alpha beta", 
"t_note":"alpha beta"}'),
+                parse_to_variant('{"a_host":"alpha"}'),
+                parse_to_variant('{"b_host":"alpha"}')),
+            (4, 'gamma', parse_to_variant('{"other":"alpha"}'),
+                parse_to_variant('{"other":"alpha"}'),
+                parse_to_variant('{"other":"alpha"}'))
+    """
+    sql " sync "
+
+    // scores are printed so that a NaN (which compares greater than 0) cannot 
slip through
+    order_qt_variant_subcolumn_score """
+        select id, round(score(), 4)
+        from test_variant_subcolumn_index_norms
+        where cast(v["s_note"] as string) match_phrase "alpha"
+        order by score() desc
+        limit 10
+    """
+    // without norms MATCH still filters, and score() is refused
+    order_qt_variant_subcolumn_match_no_norms """
+        select id
+        from test_variant_subcolumn_index_norms
+        where cast(v["t_note"] as string) match_phrase "alpha"
+    """
+    test {
+        sql """
+            select id, score()
+            from test_variant_subcolumn_index_norms
+            where cast(v["t_note"] as string) match_phrase "alpha"
+            order by score() desc
+            limit 10
+        """
+        exception "written without norms"
+    }
+    order_qt_plain_column_score """
+        select id, round(score(), 4)
+        from test_variant_subcolumn_index_norms
+        where content match_phrase "alpha"
+        order by score() desc
+        limit 10
+    """
+
+    def backendIdToIp = [:]
+    def backendIdToHttpPort = [:]
+    getBackendIpHttpPort(backendIdToIp, backendIdToHttpPort)
+    def normsBySuffixOf = { tableName ->
+        def tablet = sql_return_maparray("show tablets from ${tableName}")[0]
+        def (code, out, err) = http_client("GET", String.format(
+                "http://%s:%s/api/show_nested_index_file?tablet_id=%s";,
+                backendIdToIp.get(tablet.BackendId), 
backendIdToHttpPort.get(tablet.BackendId),
+                tablet.TabletId))
+        logger.info("show_nested_index_file of ${tableName} code=${code}, 
out=${out}, err=${err}")
+        assertEquals(0, code)
+        def norms = [:]
+        for (def rowset in parseJson(out.trim()).rowsets) {
+            for (def segment in rowset.segments) {
+                for (def index in segment.indices) {
+                    norms[index.index_suffix] = index.files.any { file -> 
file.name.endsWith(".nrm") }
+                }
+            }
+        }
+        logger.info("norms by index suffix of ${tableName}: ${norms}")
+        return norms
+    }
+    // the suffix is the escaped variant path, e.g. v%2Es%5Fhost for v.s_host
+    def normsOf = { norms, path ->
+        norms.find { suffix, hasNorms ->
+            suffix.replace("%2E", ".").replace("%5F", "_").contains(path)
+        }?.value
+    }
+
+    def normsBySuffix = normsBySuffixOf("test_variant_subcolumn_index_norms")
+    // an analyzed index writes norms wherever it sits, and "norms" = "false" 
drops them
+    assertEquals(true, normsBySuffix[""])
+    assertEquals(true, normsOf(normsBySuffix, "s_host"))
+    assertEquals(true, normsOf(normsBySuffix, "s_note"))
+    assertEquals(false, normsOf(normsBySuffix, "t_note"))
+    // a whole-column index on a VARIANT column has no suffix of its own, but 
every subcolumn copy
+    // inherits its properties: idx_vd keeps norms, idx_vn drops them because 
it asks to
+    assertEquals(true, normsOf(normsBySuffix, "a_host"))
+    assertEquals(false, normsOf(normsBySuffix, "b_host"))
+
+    // the skip is a dynamic BE config: with it turned on, every index on a 
variant path leaves norms
+    // out, even one that asks for them, while an ordinary column index is 
untouched
+    setBeConfigTemporary([inverted_index_skip_norms_for_variant: true]) {
+        sql "DROP TABLE IF EXISTS test_variant_subcolumn_index_norms_config"
+        sql """
+            CREATE TABLE test_variant_subcolumn_index_norms_config (
+                id INT,
+                content TEXT,
+                v variant<
+                    's_*' : text,
+                    PROPERTIES("variant_max_subcolumns_count"="0")
+                >,
+                vf variant<
+                    'c_*' : text,
+                    PROPERTIES("variant_max_subcolumns_count"="0")
+                >,
+                INDEX idx_content (content) USING INVERTED PROPERTIES(
+                    "parser"="english",
+                    "support_phrase"="true"
+                ),
+                INDEX idx_v_s (v) USING INVERTED PROPERTIES(
+                    "parser"="english",
+                    "support_phrase"="true",
+                    "field_pattern"="s_*"
+                ),
+                INDEX idx_vf (vf) USING INVERTED PROPERTIES(
+                    "parser"="english",
+                    "support_phrase"="true",
+                    "norms"="true"
+                )
+            ) ENGINE=OLAP DUPLICATE KEY(id)
+            DISTRIBUTED BY HASH(id) BUCKETS 1
+            PROPERTIES (
+                "replication_allocation" = "tag.location.default: 1",
+                "disable_auto_compaction" = "true",
+                "inverted_index_storage_format" = "V2"
+            )
+        """
+        sql """ insert into test_variant_subcolumn_index_norms_config values
+                (1, 'alpha database server', 
parse_to_variant('{"s_host":"alpha database server"}'),
+                    parse_to_variant('{"c_host":"alpha database server"}')),
+                (2, 'beta server cluster', parse_to_variant('{"s_host":"beta 
server cluster"}'),
+                    parse_to_variant('{"c_host":"beta server cluster"}'))
+        """
+        sql " sync "
+
+        def configNorms = 
normsBySuffixOf("test_variant_subcolumn_index_norms_config")
+        assertEquals(false, normsOf(configNorms, "s_host"))
+        assertEquals(false, normsOf(configNorms, "c_host"))
+        assertEquals(true, configNorms[""])
+    }
+    test {
+        sql """
+            select id, score()
+            from test_variant_subcolumn_index_norms_config
+            where cast(vf["c_host"] as string) match_phrase "alpha"
+            order by score() desc
+            limit 10
+        """
+        exception "written without norms"
+    }
+
+    // segments with and without norms side by side, as while the config is 
being turned on:
+    // MATCH still filters, and score() is refused rather than ranking the 
rows of the newer
+    // segment as zero-length documents
+    sql "DROP TABLE IF EXISTS test_variant_subcolumn_index_norms_mixed"
+    sql """
+        CREATE TABLE test_variant_subcolumn_index_norms_mixed (
+            id INT,
+            v variant<
+                's_*' : text,
+                PROPERTIES("variant_max_subcolumns_count"="0")
+            >,
+            INDEX idx_v_s (v) USING INVERTED PROPERTIES(
+                "parser"="english",
+                "support_phrase"="true",
+                "field_pattern"="s_*"
+            )
+        ) ENGINE=OLAP DUPLICATE KEY(id)
+        DISTRIBUTED BY HASH(id) BUCKETS 1
+        PROPERTIES (
+            "replication_allocation" = "tag.location.default: 1",
+            "disable_auto_compaction" = "true",
+            "inverted_index_storage_format" = "V2"
+        )
+    """
+    sql """ insert into test_variant_subcolumn_index_norms_mixed values
+            (1, parse_to_variant('{"s_note":"alpha database server"}'))
+    """
+    setBeConfigTemporary([inverted_index_skip_norms_for_variant: true]) {
+        sql """ insert into test_variant_subcolumn_index_norms_mixed values
+                (2, parse_to_variant('{"s_note":"alpha"}'))
+        """
+    }
+    sql " sync "
+
+    order_qt_mixed_match """
+        select id
+        from test_variant_subcolumn_index_norms_mixed
+        where cast(v["s_note"] as string) match_phrase "alpha"
+    """
+    test {
+        sql """
+            select id, score()
+            from test_variant_subcolumn_index_norms_mixed
+            where cast(v["s_note"] as string) match_phrase "alpha"
+            order by score() desc
+            limit 10
+        """
+        exception "written without norms"
+    }
+}


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to