Apache Spark Master Fixes Executor Resize and ASCII String Lookups


Apache Spark is the JVM engine behind most large SQL, batch, and streaming jobs. The last seven days on Apache Spark master landed 103 commits across 489 files (29533 insertions, 4930 deletions). The operator visible slice is Kubernetes executor resize, hold and resume on the standalone master, and a few SQL paths that returned the wrong row.

Since Spark 4.2.0, spark.kubernetes.executor.resizeInterval has been documented as 0s, and that zero was the default in Config.scala. ExecutorResizePlugin ignored it. An unset key still ran the plugin every minute. Setting 0, as the Kubernetes guide said to do, threw IllegalArgumentException from scheduleAtFixedRate and SparkContext never started.

The interval fix sets the default to 1m and treats 0 as off. An absent key still resizes every minute. An explicit 0 now starts the context with the plugin disabled. Only the written default caught up.

The loop was also touching pods it should skip. Inactive pods left behind when spark.kubernetes.executor.deleteOnTermination is false (spark-exec-inactive=true) no longer get failing metrics calls. Placeholders stamped spark-exec-id=EXECID by the deployment and statefulset allocators no longer enter cappedExecutors. PVC reports in ExecutorPVCResizePlugin drop executors with no listed pod, so latestReports stops growing and a stale size is not applied to a reused volume. Maps keyed by PVC name stay, because a volume can outlive the executor.

Deployment support is the behavior change on top of that. In 4.2.0 the plugin refused the deployment allocator and only started for direct. The notes record a check on Kubernetes 1.37.0: pods grow in place up to spark.kubernetes.executor.resizeMaxMemory, the Deployment template stays, and one ReplicaSet remains. That path is for unreleased 4.4.0. A 4.2 cluster still refuses deployment.

Hold returns executors and keeps the driver plus shuffle files already written. The Master UI could already do it, but a script must send the per UI csrfToken or the response is 403. The documented calls also need the trailing slash. Without it the server redirects and nothing is held. The standalone docs now show both UI forms.

The REST actions skip the token:

curl -XPOST http://IP:PORT/v1/submissions/hold/app-20260930120000-0000

POST /v1/submissions/hold/<app-id> and POST /v1/submissions/resume/<app-id> take an application id, not a submission id. A client mode app has no submission id, so REST kill cannot see it. Hold and resume can. success only means the Master forwarded the request. The driver drains on its own. Read held and draining on /json/.

spark.ui.holdEnabled still gates both the Master and the application. A Master that is not ALIVE, or an app that has not reported hold support, rejects the call. No new config, and no ACLs, same as kill. Put JWSFilter in front, or set spark.master.rest.enabled to false, as the configuration reference already says. These actions are on master for 4.4.0. Released builds do not have them.

Position lookups on UTF8String sit under substring and getChar. Each call walked lead bytes with numBytesForFirstByte, because a UTF-8 code point is one to four bytes. For a fully ASCII string the indexes match, so the walk is overhead.

numChars() already reads those bytes. It now stores isFullAscii when the flag is still unknown. A lead byte at or above 0x80 is not ASCII, including invalid UTF-8 where the code point count equals the byte count. A lone 0x80 must not take the fast path. numChars == numBytes would mislabel that input.

if (isFullAscii == IsFullAscii.FULL_ASCII) {
  j = Math.max(start, 0);
  i = (until == Integer.MAX_VALUE) ? numBytes : Math.min(until, numBytes);
}

substring, getChar, charPosToByte, and bytePosToChar use that branch. The fast path only reads the flag, so a cold string stays on the old scan. getChar warms the flag through its bounds check. Results do not change, and there is no SQL conf. Multi byte text costs the same as before.

The forward ASOF scan in SortMergeAsOfJoinExec reused the nearest loop, which subtracts keys and keeps the smallest distance. The right side is sorted ascending, so the first row that satisfies the inequality is already the closest. The subtraction existed only to overflow.

SET spark.sql.join.asofJoin.enabled=true;
SELECT r.tag
FROM VALUES (-100) AS t(k)
ASOF JOIN VALUES (-50, 'near'), (2147483647, 'far') AS r(k, tag)
  MATCH_CONDITION (t.k <= r.k);

ANSI mode turns 2147483647 - (-100) into ARITHMETIC_OVERFLOW. With ANSI off, the int wraps and the join returns far. The row you want is near. The same pattern hits BIGINT, DECIMAL(38, 0), and both interval types. Forward now keeps the first match and stops once the condition goes false. Nearest still subtracts. With spark.sql.legacy.interval.enabled set, only nearest builds a calendar interval ordering, so the other directions no longer fail with DATATYPE_CANNOT_ORDER. Empty STRUCTs are now legal in MATCH_CONDITION, and analyzer tests moved to end to end tests.

dropDuplicates lost late row filtering when a later select removed the event time column. ColumnPruning stripped the attribute StreamingDeduplicateExec needs, so a late row with a new key was emitted unless the query still projected the timestamp. Deduplicate.references now keeps the watermark attributes. Batch dedup is unchanged.

Corrupt BYTE_STREAM_SPLIT pages now throw ParquetDecodingException instead of ArrayIndexOutOfBoundsException or a wrong stride. A required column with a mismatched value count fails at page open, even under LIMIT. The reader is not in a release yet.

from_protobuf gained convert.timestamp.duration.to.native, default true. Setting it to false keeps Timestamp and Duration as struct<seconds: bigint, nanos: int> (4.5 seconds becomes seconds 4 and nanos 500000000) instead of TimestampType and DayTimeIntervalType. to_protobuf is unchanged.

Reused Python workers counted idle time as initialization. The fix uses the later of task start and worker boot. Fresh workers stay the same. Counters now go out as JSON from worker.py. pythonProcessingTime still maps from pythonExecutionDurationMs.

Wide CTE planning got cheaper without a plan change. Attribute restore in PushdownPredicatesAndPruneColumnsForCTEDef uses AttributeMap instead of a linear search. The patch benchmark goes from 26.54 ms to 1.33 ms at 2000 columns, about 19.9 times faster, and is a few percent slower at one or two columns.

On unreleased master, dispatch no longer loads spark.udf.worker.dispatcherFactory into the Spark JVM. Bouncy Castle 1.85 and the npm fixes stay in test and dev dependencies.

If you build master toward 4.4.0, retest deployment resize, any hold script that posts to the UI without a token, and forward ASOF joins on large integer keys. On a released 4.2 line the resize crash is still there. Nothing in this window is a backport.