Read an enum argument from its text through the public parser - #135
Merged
estebanzimanyi merged 1 commit intoOct 3, 2026
Merged
estebanzimanyi merged 1 commit into
estebanzimanyi merged 1 commit into
Conversation
An argument of a C enum travels in SQL as the text its name is, and both JVM SQL surfaces turn that text into the enum through the catalog function returning that enum from one string. The rule is one: SqlModel._enum_parsers in codegen_jvm.py and ENUM_PARSER in codegen_spark_udfs.py both admit a public parser alone (null_handle_type_from_string, interptype_from_string, raquet_pixtype_from_string). In the Spark arm an enum argument is a string whose conversion the four call builders now apply, as they pass every other scalar argument through its expression. Witness. Against MobilityDB d29d4ab293 and the catalog of MEOS-API 5ba0cd8492, the Spark arm refuses tjsonb_to_tbigint, public since MobilityDB 21a15203ba, and every other function taking a nullHandleType or an interpType, as an argument it cannot read, which the gaps ledger records; the Flink arm reads an interpType through interptype_from_string, which the catalog stated internal until MobilityDB 671ed23326. Why. A binding calls the public API alone, and a public parser tests its argument at entry, so a null or unknown name is reported to the engine rather than read. Flink already passed an enum as its text; Spark refused every function taking one, so the two surfaces differed on the constructors and conversions that take an interpolation or a null handling. Measured. From the same catalog and jar, against the generator that answers every overload of a Spark name whose argument classes differ, the Flink surface this branch generates is byte-identical, and Spark registers 3,724 names where that generator registers 3,646: 78 added and none removed, among them tintSeq, tfloatSeqSet and the other <type>Seq and <type>SeqSet constructors, setInterp, tsample, appendInstant, the tjsonb_to_* conversions and the JSON functions taking a null handling; tnpoint gains its sequence constructors from a base value and a time span. The gaps ledger loses 33 functions and lists 711, none new; the two parsers leave it as the conversions the surface calls.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An argument of a C enum travels in SQL as the text its name is, and both
JVM SQL surfaces turn that text into the enum through the catalog function
returning that enum from one string. The rule is one: SqlModel._enum_parsers
in codegen_jvm.py and ENUM_PARSER in codegen_spark_udfs.py both admit a
public parser alone (null_handle_type_from_string, interptype_from_string,
raquet_pixtype_from_string). In the Spark arm an enum argument is a string
whose conversion the four call builders now apply, as they pass every other
scalar argument through its expression.
Witness. Against MobilityDB d29d4ab293 and the catalog of MEOS-API
5ba0cd8492, the Spark arm refuses tjsonb_to_tbigint, public since
MobilityDB 21a15203ba, and every other function taking a nullHandleType or
an interpType, as an argument it cannot read, which the gaps ledger records;
the Flink arm reads an interpType through interptype_from_string, which the
catalog stated internal until MobilityDB 671ed23326.
Why. A binding calls the public API alone, and a public parser tests its
argument at entry, so a null or unknown name is reported to the engine
rather than read. Flink already passed an enum as its text; Spark refused
every function taking one, so the two surfaces differed on the constructors
and conversions that take an interpolation or a null handling.
Measured. From the same catalog and jar, against the generator that
answers every overload of a Spark name whose argument classes differ, the
Flink surface this branch generates is byte-identical, and Spark registers
3,724 names where that generator registers 3,646: 78 added and none removed, among them tintSeq,
tfloatSeqSet and the other Seq and SeqSet constructors,
setInterp, tsample, appendInstant, the tjsonb_to_* conversions and the
JSON functions taking a null handling; tnpoint gains its sequence
constructors from a base value and a time span. The gaps ledger loses 33
functions and lists 711, none new; the two parsers leave it as the
conversions the surface calls.