Skip to content

Native Iceberg scan fails with field not found when an equality delete is keyed on a nested field #6782

Description

@andygrove

Describe the bug

When a table has an equality delete keyed on a field nested in a struct, every query the native Iceberg scan runs on that table fails with Iceberg scan error: Unexpected => file scan task generate failed, source: Unexpected => field not found. It fails even when the query projects the key, where Spark returns the right answer.

iceberg-rust applies an equality delete by evolving the delete file's batches to the equality field ids (BasicDeleteFileLoader::evolve_schema, which builds a RecordBatchTransformer). The transformer looks up each id among the delete file's top-level columns only (build_field_id_to_arrow_schema_map), so a nested id such as s.k is never found. The check in CometScanRule (deleteFileTypesSupported) rejects struct and variant equality keys, but it lets a primitive nested inside a struct through. The iceberg-rust code is the same at Comet's current pin, at bb1e4a4 (the 1.1 pin), and on iceberg-rust's main.

Steps to reproduce

Spark SQL can't write equality deletes, so the delete file is written through Iceberg's API, the way Comet's CometEqualityDeletes test helper does, but with the nested field id passed explicitly.

spark.sql("""CREATE TABLE cat.db.t (id INT, s STRUCT<k: INT, a: INT>) USING iceberg
  TBLPROPERTIES ('format-version' = '2')""")
spark.sql("""INSERT INTO cat.db.t SELECT CAST(id AS INT),
  named_struct('k', CAST(id % 10 AS INT), 'a', CAST(id AS INT)) FROM range(100)""")

// In package org.apache.iceberg.data, because GenericFileWriterFactory.builderFor is package-private.
val deleteRowSchema = table.schema().select("s.k")
val s = GenericRecord.create(deleteRowSchema.findField("s").`type`().asStructType())
s.setField("k", Integer.valueOf(3))
val row = GenericRecord.create(deleteRowSchema)
row.setField("s", s)
val writer = GenericFileWriterFactory
  .builderFor(table)
  .equalityDeleteRowSchema(deleteRowSchema)
  .equalityFieldIds(Array(table.schema().findField("s.k").fieldId()))
  .build()
  .newEqualityDeleteWriter(
    EncryptedFiles.encryptedOutput(table.io().newOutputFile(path), EncryptionKeyMetadata.EMPTY),
    table.spec(),
    null)
try writer.write(java.util.Collections.singletonList[Record](row)) finally writer.close()
table.newRowDelta().addDeletes(writer.toDeleteFile()).commit()

On main at 7d29453 (Spark 4.1, Iceberg 1.11):

Query Spark Comet
SELECT count(*), sum(s.k), sum(s.a) FROM t [90,420,4470] field not found
SELECT count(*), sum(s.a) FROM t [0,null] field not found
SELECT count(*), sum(id) FROM t [0,null] field not found

The native scan fails the same way on Spark 3.4 with Iceberg 1.5.2.

Expected behavior

Comet falls back to Spark when an equality delete's field ids include a field nested below the top level, until iceberg-rust can apply such a delete.

A test can compare against Spark only for a query that projects the key. When the query doesn't project it, Spark is wrong too. Iceberg's DeleteFilter.fileProjection can add a missing equality column only at the top level. Iceberg 1.5.2 then throws Cannot find required field for ID, and Iceberg 1.11 deletes every row, which is where the two [0,null] results above come from.

Additional context

Found while reviewing #6725. It predates that PR, and fails the same way with that PR's nested schema pruning on or off.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions