Great Expectations 1.23.1 was published on 18 September 2026. The release corrects Spark evaluation of ExpectColumnValuesToMatchRegexList when match_on="all", and it stops a validator on a shared datasource from reporting another validator’s batch. File context reload for spark_schema and UTF-8 config reads land in the same tag.
The full release notes and downloads are on the GitHub release page. GitHub lists this tag with prerelease set to false.
Spark regex lists check each pattern ¶
ExpectColumnValuesToMatchRegexList takes a regex_list and a match_on mode. With match_on="all", a value passes only when every pattern matches it. Pandas and SQL already applied each pattern on its own. Spark did not. Patterns anchored at different positions could fail a value that satisfied every pattern. A value that starts with A and ends in three digits should match both ^A and [0-9]{3}$.
12198 changes the Spark path so each regex is checked separately against each column value. The call site stays the same:
gxe.ExpectColumnValuesToMatchRegexList(
column="id",
regex_list=["^A", "[0-9]{3}$"],
match_on="all",
)
Rows that meet every pattern, and that Spark rejected only because the anchors sat at different positions, can now pass. The notes do not describe a flag that keeps the old Spark result.
Shared datasources keep batch identity ¶
Validators built on the same datasource were borrowing one another’s Batch. Two validation definitions running on threads, or two get_validator calls on one thread, could evaluate the wrong rows. The result could still carry the wrong batch_id, batch_spec, and batch_definition.
The notes describe a long standing bug. It stayed latent until 12148 in 1.23.0 made it reachable. 12211 stores batch identity on the validator even when the execution engine is shared. Each validator evaluates and reports the batch it loaded.
validator_a = context.get_validator(batch_request=request_a)
validator_b = context.get_validator(batch_request=request_b)
# validator_a still validates request_a's batch
result = validator_a.expect_table_row_count_to_equal(value=3)
If the configuration names a batch the engine does not hold, validation now raises. The previous path silently ran the most recently loaded batch. A success recorded on 1.23.0 for a reused datasource does not prove the reported batch was the requested batch.
Spark schema reload and config encoding ¶
A spark_schema stored in great_expectations.yml did not survive reopening a File Data Context. Reload failed inside PySpark. 12200 reads the persisted value with StructType.fromJson, so this sequence round trips the schema:
context = gx.get_context(mode="file")
asset = context.data_sources.get("spark_ds").get_asset("my_asset")
assert asset.spark_schema is not None
A value that is not an accepted schema form raises a validation error naming the field and the accepted types.
Config reads and writes are pinned to UTF-8 no matter what locale the process uses. 12182 covers config_variables.yml, including a variable whose value contains characters outside ASCII. 12204 covers the remaining reads and writes in the serializable data context: great_expectations.yml, and the .gitignore read while scaffolding a project. Scaffolding a .gitignore that contains characters outside ASCII now works on a host whose locale is not UTF-8. Project YAML that is not valid UTF-8 raises an error that names the file. A checkout saved in another encoding fails on first open until that file is converted.
Gallery tier and documentation ¶
12150 adds a gallery support tier. One measured case per registered expectation pairs a passing configuration with a failing configuration. Nine data sources declare the tier after that full gallery run: pandas in memory, pandas filesystem CSV, SQLite, MySQL, PostgreSQL, Trino, BigQuery, Databricks, and Redshift. Each member has a required CI lane. Membership and coverage guards sit on the tier, and a measurement mode evaluates a new candidate backend. The tier is a test claim about those backends.
12224 adds a deprecation timeline table to the changelog. A row names the deprecated item, the version that deprecated it, and the version that removes it. The 1.23.1 notes do not put those removals in this tag. 12167 documents Oracle as oracle+oracledb://...?service_name=..., with the tested database version and add_sql. That is connection documentation, not a new execution engine. 12177 points credential pages at gx/uncommitted/config_variables.yml and fixes a mistagged code fence, a snippet name, and spelling in the core docs, ADRs, gallery docs, and contrib READMEs.
The rest is internal. 12199 skips the Spark test-connection test when PySpark is installed. That test targets a missing PySpark install, so contributors who already have it no longer fail the job. 12223 requires exactly one current tag at the start of a pull request title. The contributor docs name only the four current tags. 12207 adds repository, documentation, and homepage URLs to the published package metadata. 12184, 12185, and 12178 extend mypy and remove those modules from the exclude list. Two guards that could never fail now test whether an optional dependency imported. No runtime validation path changes in that work.
Upgrade notes ¶
Three checks are worth running before a fleet moves to this tag.
- Retest any
ExpectColumnValuesToMatchRegexListsuite that usesmatch_on="all"on Spark. Failure counts can drop where anchors sit at different positions. - Retest jobs that build more than one validator from one datasource, including threaded validation definitions. A named batch the engine does not hold now raises.
- Confirm
great_expectations.yml,config_variables.yml, and any scaffolded.gitignoreare validUTF-8, and that a savedspark_schemais a formStructType.fromJsonaccepts.
Where to get it ¶
- Release page: Great Expectations 1.23.1
- Repository: great_expectations
- Tag:
1.23.1