timsaucer commented on code in PR #1763:
URL:
https://github.com/apache/datafusion-python/pull/1763#discussion_r4122784203
##########
crates/core/src/dataframe.rs:
##########
@@ -1320,6 +1345,26 @@ impl PyDataFrame {
let df = self.df.as_ref().fill_null(&scalar_value.0, &cols)?;
Ok(Self::new(df))
}
+
+ /// Fill NaN values with a specified value for specific floating-point
columns
+ #[pyo3(signature = (value, columns=None))]
+ fn fill_nan(
+ &self,
+ value: Py<PyAny>,
+ columns: Option<Vec<PyBackedStr>>,
+ py: Python,
+ ) -> PyDataFusionResult<Self> {
+ let scalar_value: PyScalarValue = value.extract(py)?;
+
+ let cols = match columns {
+ Some(col_names) => col_names.iter().map(|c|
c.to_string()).collect(),
+ None => Vec::new(), // Empty vector means fill NaN for all columns
+ };
+
+ let cols = cols.iter().map(String::as_str).collect::<Vec<_>>();
+ let df = self.df.as_ref().fill_nan(&scalar_value.0, &cols)?;
Review Comment:
Confirmed in pure Rust on 55.1.0, for both `Name` and `a.b` column names,
and filed upstream as apache/datafusion#25829 with a suggested fix
(`Expr::Column(Column::from((qualifier, field)))`, which also keeps the
qualifier). I've left the wrapper as-is so `fill_nan` and `fill_null` keep
behaving the same way, and we'll pick up the upstream fix when it lands.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]