laskoviymishka commented on code in PR #3066: URL: https://github.com/apache/iceberg-rust/pull/3066#discussion_r4169592290
########## crates/iceberg/src/spec/schema/assign_fresh_ids.rs: ########## @@ -0,0 +1,507 @@ +// Licensed to the Apache Software Foundation (ASF) under one +// or more contributor license agreements. See the NOTICE file +// distributed with this work for additional information +// regarding copyright ownership. The ASF licenses this file +// to you under the Apache License, Version 2.0 (the +// "License"); you may not use this file except in compliance +// with the License. You may obtain a copy of the License at +// +// http://www.apache.org/licenses/LICENSE-2.0 +// +// Unless required by applicable law or agreed to in writing, +// software distributed under the License is distributed on an +// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY +// KIND, either express or implied. See the License for the +// specific language governing permissions and limitations +// under the License. + +use super::utils::try_insert_field; +use super::*; + +/// Reassigns `schema`'s field ids, reusing ids from `base` for fields whose full name is unchanged +/// and drawing fresh ids from `start_from` upwards for everything else. +/// +/// `start_from` must be past every id the table has ever assigned, not merely past `base`'s ids: +/// pass `table_metadata.last_column_id() + 1`, which also reserves the ids of columns already +/// dropped from `base`. A reused id does not consume a fresh one, so a too-low `start_from` fails in +/// one of two ways. Colliding with an id still present in `base` yields duplicate ids and a generic +/// error from `build()`. Landing on the id of a column already dropped from `base` is worse: nothing +/// here or in `build()` notices, and a new column silently inherits a retired id. Seeding from a +/// schema's `highest_field_id() + 1` is the usual way to hit the latter. +/// +/// The returned `schema_id` is carried over unchanged and is not authoritative; it is arbitrated by +/// [`TableMetadataBuilder::add_schema`](crate::spec::TableMetadataBuilder::add_schema). +pub(crate) fn assign_fresh_ids(schema: Schema, base: &Schema, start_from: i32) -> Result<Schema> { + let Schema { + r#struct, + schema_id, + identifier_field_ids, + alias_to_id, + id_to_name, + .. + } = schema; + let mut assigner = AssignFreshIds::new(id_to_name, base, start_from); + let fields = assigner.assign_fields(r#struct.fields().to_vec())?; + let identifier_field_ids = assigner.apply_to_identifier_fields(identifier_field_ids)?; + let alias_to_id = assigner.apply_to_aliases(alias_to_id)?; + + Schema::builder() + .with_schema_id(schema_id) + .with_fields(fields) + .with_identifier_field_ids(identifier_field_ids) + .with_alias(alias_to_id) + .build() +} + +struct AssignFreshIds { + next_field_id: i32, + target_names: HashMap<i32, String>, + base_ids: HashMap<String, i32>, + old_to_new_id: HashMap<i32, i32>, +} + +impl AssignFreshIds { + fn new(target_names: HashMap<i32, String>, base: &Schema, start_from: i32) -> Self { + // Partial guard only: it trips the common `base.highest_field_id() + 1` mistake, but ids + // dropped above `base`'s highest still slip through, as that bound lives in `last_column_id`. + debug_assert!( Review Comment: Let's convert this to a returned `Err(DataInvalid)` rather than a `debug_assert`. Your own `seeded_too_low` test is the clincher: seed `base.highest_field_id() + 1 = 4` clears the assert (`4 > 3`) and still lands on retired id 4 — the dropped-id reuse from round 1 walks straight past the guard, and in release the assert compiles out and guards nothing (on top of panicking in library code, which we don't do elsewhere). I'd rather take `last_column_id` than `start_from` and derive the seed internally, so the contract lives in the signature and matches what the #3056 caller actually holds (`table_metadata.last_column_id()`): ```rust if last_column_id < base.highest_field_id() { return Err(Error::new( ErrorKind::DataInvalid, format!( "last_column_id ({last_column_id}) is below base.highest_field_id() ({})", base.highest_field_id() ), )); } let start_from = last_column_id.checked_add(1).ok_or_else(|| { /* overflow */ })?; ``` An in-place `if … return Err` that keeps `start_from` works too, but then the dropped-id case still rides on the caller seeding correctly — so I lean toward taking `last_column_id`. wdyt? (The comment just above is backwards, fwiw — it says the assert trips the `highest_field_id() + 1` mistake, but that's the exact seed your test shows passing.) ########## crates/iceberg/src/spec/schema/mod.rs: ########## @@ -25,6 +25,9 @@ mod utils; mod visitor; pub use self::visitor::*; pub(super) mod _serde; +// TODO: Use for table replacement in https://github.com/apache/iceberg-rust/issues/3056. +#[allow(dead_code)] Review Comment: Small thing while we're here — `#[expect(dead_code)]` instead of `#[allow]` (assuming MSRV ≥ 1.81) will warn and self-remove once the #3056 caller lands and uses the module; the `#[allow]` just lingers after it's no longer dead. ########## crates/iceberg/src/spec/schema/assign_fresh_ids.rs: ########## @@ -0,0 +1,507 @@ +// Licensed to the Apache Software Foundation (ASF) under one +// or more contributor license agreements. See the NOTICE file +// distributed with this work for additional information +// regarding copyright ownership. The ASF licenses this file +// to you under the Apache License, Version 2.0 (the +// "License"); you may not use this file except in compliance +// with the License. You may obtain a copy of the License at +// +// http://www.apache.org/licenses/LICENSE-2.0 +// +// Unless required by applicable law or agreed to in writing, +// software distributed under the License is distributed on an +// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY +// KIND, either express or implied. See the License for the +// specific language governing permissions and limitations +// under the License. + +use super::utils::try_insert_field; +use super::*; + +/// Reassigns `schema`'s field ids, reusing ids from `base` for fields whose full name is unchanged +/// and drawing fresh ids from `start_from` upwards for everything else. +/// +/// `start_from` must be past every id the table has ever assigned, not merely past `base`'s ids: +/// pass `table_metadata.last_column_id() + 1`, which also reserves the ids of columns already +/// dropped from `base`. A reused id does not consume a fresh one, so a too-low `start_from` fails in +/// one of two ways. Colliding with an id still present in `base` yields duplicate ids and a generic +/// error from `build()`. Landing on the id of a column already dropped from `base` is worse: nothing +/// here or in `build()` notices, and a new column silently inherits a retired id. Seeding from a +/// schema's `highest_field_id() + 1` is the usual way to hit the latter. +/// +/// The returned `schema_id` is carried over unchanged and is not authoritative; it is arbitrated by +/// [`TableMetadataBuilder::add_schema`](crate::spec::TableMetadataBuilder::add_schema). +pub(crate) fn assign_fresh_ids(schema: Schema, base: &Schema, start_from: i32) -> Result<Schema> { + let Schema { + r#struct, + schema_id, + identifier_field_ids, + alias_to_id, + id_to_name, + .. + } = schema; + let mut assigner = AssignFreshIds::new(id_to_name, base, start_from); + let fields = assigner.assign_fields(r#struct.fields().to_vec())?; + let identifier_field_ids = assigner.apply_to_identifier_fields(identifier_field_ids)?; + let alias_to_id = assigner.apply_to_aliases(alias_to_id)?; + + Schema::builder() + .with_schema_id(schema_id) + .with_fields(fields) + .with_identifier_field_ids(identifier_field_ids) + .with_alias(alias_to_id) + .build() +} + +struct AssignFreshIds { + next_field_id: i32, + target_names: HashMap<i32, String>, + base_ids: HashMap<String, i32>, + old_to_new_id: HashMap<i32, i32>, +} + +impl AssignFreshIds { + fn new(target_names: HashMap<i32, String>, base: &Schema, start_from: i32) -> Self { + // Partial guard only: it trips the common `base.highest_field_id() + 1` mistake, but ids + // dropped above `base`'s highest still slip through, as that bound lives in `last_column_id`. + debug_assert!( + start_from > base.highest_field_id(), + "start_from must exceed every id the table has assigned; pass last_column_id() + 1" + ); + + Self { + next_field_id: start_from, + target_names, + base_ids: base + .field_id_to_name_map() + .iter() + .map(|(id, name)| (name.clone(), *id)) + .collect(), + old_to_new_id: HashMap::new(), + } + } + + /// Returns `base`'s id when the field's full name is unchanged, otherwise consumes a fresh id by + /// advancing `next_field_id`. + fn resolve_or_assign_id(&mut self, old_id: i32) -> Result<i32> { + if let Some(id) = self + .target_names + .get(&old_id) + .and_then(|name| self.base_ids.get(name)) + { + return Ok(*id); + } + + let id = self.next_field_id; + self.next_field_id = self.next_field_id.checked_add(1).ok_or_else(|| { + Error::new( + ErrorKind::DataInvalid, + "Field ID overflowed, cannot add more fields", + ) + })?; + Ok(id) + } + + fn assign_fields(&mut self, fields: Vec<NestedFieldRef>) -> Result<Vec<NestedFieldRef>> { + let outer_fields = fields + .into_iter() + .map(|field| { + let new_id = self.resolve_or_assign_id(field.id)?; + try_insert_field(&mut self.old_to_new_id, field.id, new_id)?; + Ok(Arc::new(Arc::unwrap_or_clone(field).with_id(new_id))) + }) + .collect::<Result<Vec<_>>>()?; + + outer_fields + .into_iter() + .map(|field| { + if !field.field_type.is_nested() { + Ok(field) + } else { + let mut field = Arc::unwrap_or_clone(field); + *field.field_type = self.assign_type(*field.field_type)?; + Ok(Arc::new(field)) + } + }) + .collect() + } + + fn assign_type(&mut self, field_type: Type) -> Result<Type> { + match field_type { + Type::Primitive(primitive) => Ok(Type::Primitive(primitive)), + Type::Struct(r#struct) => Ok(Type::Struct(StructType::new( + self.assign_fields(r#struct.fields().to_vec())?, + ))), + Type::List(list) => { + let new_id = self.resolve_or_assign_id(list.element_field.id)?; + try_insert_field(&mut self.old_to_new_id, list.element_field.id, new_id)?; + let mut element_field = Arc::unwrap_or_clone(list.element_field); + element_field.id = new_id; + *element_field.field_type = self.assign_type(*element_field.field_type)?; + Ok(Type::List(ListType { + element_field: Arc::new(element_field), + })) + } + Type::Map(map) => { + // Key and value ids are resolved before recursing into either, matching Java and + // intentionally unlike `ReassignFieldIds`, which recurses the key first. + let new_key_id = self.resolve_or_assign_id(map.key_field.id)?; + let new_value_id = self.resolve_or_assign_id(map.value_field.id)?; + try_insert_field(&mut self.old_to_new_id, map.key_field.id, new_key_id)?; + try_insert_field(&mut self.old_to_new_id, map.value_field.id, new_value_id)?; + + let mut key_field = Arc::unwrap_or_clone(map.key_field); + key_field.id = new_key_id; + *key_field.field_type = self.assign_type(*key_field.field_type)?; + + let mut value_field = Arc::unwrap_or_clone(map.value_field); + value_field.id = new_value_id; + *value_field.field_type = self.assign_type(*value_field.field_type)?; + + Ok(Type::Map(MapType { + key_field: Arc::new(key_field), + value_field: Arc::new(value_field), + })) + } + Type::Variant(variant) => Ok(Type::Variant(variant)), + } + } + + fn apply_to_identifier_fields(&self, field_ids: HashSet<i32>) -> Result<HashSet<i32>> { + field_ids + .into_iter() + .map(|id| { + self.old_to_new_id.get(&id).copied().ok_or_else(|| { + Error::new( + ErrorKind::DataInvalid, + format!("identifier field id {id} not found"), + ) + }) + }) + .collect() + } + + fn apply_to_aliases(&self, aliases: BiHashMap<String, i32>) -> Result<BiHashMap<String, i32>> { + aliases + .into_iter() + .map(|(name, id)| { + self.old_to_new_id + .get(&id) + .copied() + .ok_or_else(|| { + Error::new( + ErrorKind::DataInvalid, + format!("Field with id {id} for alias {name} not found"), + ) + }) + .map(|new_id| (name, new_id)) + }) + .collect() + } +} + +#[cfg(test)] +mod tests { + use super::*; + use crate::spec::{Literal, VariantType}; + + fn empty_schema() -> Schema { + Schema::builder().build().unwrap() + } + + #[test] + fn test_assign_fresh_ids() { + let schema = Schema::builder() + .with_fields(vec![ + NestedField::required(0, "a", Type::Primitive(PrimitiveType::Int)).into(), + NestedField::required(1, "c", Type::Primitive(PrimitiveType::Int)) + .with_initial_default(Literal::int(23)) + .with_write_default(Literal::int(34)) + .into(), + NestedField::required(2, "B", Type::Primitive(PrimitiveType::Int)).into(), + ]) + .build() + .unwrap(); + let expected = Schema::builder() + .with_fields(vec![ + NestedField::required(11, "a", Type::Primitive(PrimitiveType::Int)).into(), + NestedField::required(12, "c", Type::Primitive(PrimitiveType::Int)) + .with_initial_default(Literal::int(23)) + .with_write_default(Literal::int(34)) + .into(), + NestedField::required(13, "B", Type::Primitive(PrimitiveType::Int)).into(), + ]) + .build() + .unwrap(); + + let assigned = assign_fresh_ids(schema, &empty_schema(), 11).unwrap(); + + assert_eq!(assigned.as_struct(), expected.as_struct()); + } + + #[test] + fn test_assign_fresh_ids_with_type() { + for test_type in [ + Type::Primitive(PrimitiveType::Boolean), + Type::Primitive(PrimitiveType::Int), + Type::Primitive(PrimitiveType::Long), + Type::Primitive(PrimitiveType::Float), + Type::Primitive(PrimitiveType::Double), + Type::Primitive(PrimitiveType::Decimal { + precision: 9, + scale: 2, + }), + Type::Primitive(PrimitiveType::Date), + Type::Primitive(PrimitiveType::Time), + Type::Primitive(PrimitiveType::Timestamp), + Type::Primitive(PrimitiveType::Timestamptz), + Type::Primitive(PrimitiveType::TimestampNs), + Type::Primitive(PrimitiveType::TimestamptzNs), + Type::Primitive(PrimitiveType::String), + Type::Primitive(PrimitiveType::Uuid), + Type::Primitive(PrimitiveType::Fixed(16)), + Type::Primitive(PrimitiveType::Binary), + Type::Variant(VariantType), + ] { + let schema = Schema::builder() + .with_fields(vec![ + NestedField::required(0, "id", Type::Primitive(PrimitiveType::Int)).into(), + NestedField::optional(1, "data", test_type.clone()).into(), + ]) + .build() + .unwrap(); + let expected = Schema::builder() + .with_fields(vec![ + NestedField::required(11, "id", Type::Primitive(PrimitiveType::Int)).into(), + NestedField::optional(12, "data", test_type.clone()).into(), + ]) + .build() + .unwrap(); + + let assigned = assign_fresh_ids(schema, &empty_schema(), 11).unwrap(); + + assert_eq!( + assigned.as_struct(), + expected.as_struct(), + "failed for type: {test_type:?}" + ); + } + } + + #[test] + fn test_assign_fresh_ids_reuses_full_names_and_assigns_new_ids_in_correct_order() { + let base = Schema::builder() + .with_fields(vec![ + NestedField::required( + 1, + "nested", + Type::Struct(StructType::new(vec![ + NestedField::optional(2, "a", Type::Primitive(PrimitiveType::Int)).into(), + ])), + ) + .into(), + NestedField::optional( + 3, + "items", + Type::List(ListType::new( + NestedField::list_element(4, Type::Primitive(PrimitiveType::String), false) + .into(), + )), + ) + .into(), + NestedField::optional( + 5, + "properties", + Type::Map(MapType::optional( + 6, + Type::Primitive(PrimitiveType::String), + 7, + Type::Primitive(PrimitiveType::Long), + )), + ) + .into(), + NestedField::required(8, "x", Type::Primitive(PrimitiveType::Long)).into(), + NestedField::optional(9, "dropped", Type::Primitive(PrimitiveType::Int)).into(), + ]) + .build() + .unwrap(); + let replacement = Schema::builder() + .with_schema_id(1) + .with_identifier_field_ids([18]) + .with_alias(BiHashMap::from_iter([("a_alias".to_string(), 12)])) + .with_fields(vec![ + NestedField::required( + 10, + "nested", + Type::Struct(StructType::new(vec![ + NestedField::optional(11, "b", Type::Primitive(PrimitiveType::Int)).into(), + NestedField::optional(12, "a", Type::Primitive(PrimitiveType::Int)).into(), + ])), + ) + .into(), + NestedField::optional( + 13, + "items", + Type::List(ListType::new( + NestedField::list_element( + 14, + Type::Primitive(PrimitiveType::String), + false, + ) + .into(), + )), + ) + .into(), + NestedField::optional( + 15, + "properties", + Type::Map(MapType::optional( + 16, + Type::Primitive(PrimitiveType::String), + 17, + Type::Primitive(PrimitiveType::Long), + )), + ) + .into(), + NestedField::required(18, "x", Type::Primitive(PrimitiveType::Long)).into(), + NestedField::optional(19, "z", Type::Primitive(PrimitiveType::Int)).into(), + ]) + .build() + .unwrap(); + + let assigned = assign_fresh_ids(replacement, &base, 10).unwrap(); + + assert_eq!(assigned.field_by_name("nested").unwrap().id, 1); + assert_eq!(assigned.field_by_name("nested.a").unwrap().id, 2); + assert_eq!(assigned.field_by_name("items").unwrap().id, 3); + assert_eq!(assigned.field_by_name("items.element").unwrap().id, 4); + assert_eq!(assigned.field_by_name("properties").unwrap().id, 5); + assert_eq!(assigned.field_by_name("properties.key").unwrap().id, 6); + assert_eq!(assigned.field_by_name("properties.value").unwrap().id, 7); + assert_eq!(assigned.field_by_name("x").unwrap().id, 8); + assert_eq!(assigned.field_by_name("z").unwrap().id, 10); + assert_eq!(assigned.field_by_name("nested.b").unwrap().id, 11); + assert_eq!( + assigned.identifier_field_ids().collect::<HashSet<_>>(), + HashSet::from([8]) + ); + assert_eq!(assigned.field_by_alias("a_alias").unwrap().id, 2); + assert_eq!(assigned.highest_field_id(), 11); + } + + /// The table once had ids 1..=5; only 1 and 3 survive in `base`, so 4 and 5 are retired and a + /// correct seed is 6, not `base.highest_field_id() + 1`. + fn schemas_with_dropped_column_ids() -> (Schema, Schema) { + let base = Schema::builder() + .with_fields(vec![ + NestedField::required(1, "a", Type::Primitive(PrimitiveType::Int)).into(), + NestedField::required(3, "b", Type::Primitive(PrimitiveType::Int)).into(), + ]) + .build() + .unwrap(); + let replacement = Schema::builder() + .with_fields(vec![ + NestedField::required(1, "a", Type::Primitive(PrimitiveType::Int)).into(), + NestedField::optional(2, "fresh", Type::Primitive(PrimitiveType::Int)).into(), + ]) + .build() + .unwrap(); + (base, replacement) + } + + #[test] + fn test_assign_fresh_ids_skips_dropped_column_ids_when_seeded_correctly() { + let (base, replacement) = schemas_with_dropped_column_ids(); + + let assigned = assign_fresh_ids(replacement, &base, 6).unwrap(); + + assert_eq!(assigned.field_by_name("a").unwrap().id, 1); + assert_eq!(assigned.field_by_name("fresh").unwrap().id, 6); + } + + #[test] + fn test_assign_fresh_ids_seeded_too_low_reuses_dropped_column_id() { + let (base, replacement) = schemas_with_dropped_column_ids(); + + // A seed of 4 clears the `debug_assert` in `new` (base's highest id is 3) and `build()` + // sees no duplicates, so the retired id 4 is reused silently: the caller's contract, and + // the limit of the assert. + let assigned = assign_fresh_ids(replacement, &base, base.highest_field_id() + 1).unwrap(); Review Comment: Once the guard returns an error, this test should assert that error instead of pinning the corruption. Right now it locks in `fresh.id == 4` (a retired id) as the expected result, so a future fix that returns `Err` for the too-low seed makes this test fail — the wrong direction. Flip it to `assign_fresh_ids(...).unwrap_err()` and check `err.kind() == ErrorKind::DataInvalid`, and the comment can describe the rejection rather than the silent reuse. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
