petl v1.7.27 - Pandas Integer Precision and setheader Callables


petl v1.7.27 was published on 24 September 2026. fromdataframe now keeps large integers intact when a pandas DataFrame mixes them with floats, and setheader can apply a callable to every existing field name. Pick up the integer fix first, because values past the exact range of float64 were rounded before any later transform ran.

The full release notes and downloads are on the GitHub release page. The tag ships two changes: pull request 720 from rastagan-git, and pull request 721 from akashmalbari. Both authors are new contributors. The full changelog is the diff against v1.7.26.

fromdataframe exposes a pandas DataFrame as a petl table. Iteration goes through DataFrameView.__iter__. On v1.7.26 that path called iterrows(). Pandas builds a Series for the row and selects one dtype for every cell in it. An integer column next to a float column is upcast to float. Float64 represents every integer only through 2^53, which is 9007199254740992. One step past that boundary, the integer rounds.

The report in pull request 720 uses a one row frame. Column id holds 9007199254740993. Column score holds 0.5. The DataFrame still stores the original integer.

import pandas as pd
import petl as etl

df = pd.DataFrame({"id": [9007199254740993], "score": [0.5]},
                  columns=["id", "score"])
print(list(etl.fromdataframe(df)))
# v1.7.26: [("id", "score"), (9007199254740992.0, 0.5)]
# v1.7.27: [("id", "score"), (9007199254740993, 0.5)]

The identifier lost one and changed from int to float. The same rounding shows up when include_index=True.

v1.7.27 iterates with itertuples(index=True, name=None). name=None returns a plain tuple, so duplicate column names and names that are not Python identifiers survive without a namedtuple rename. Header labels stay as they are on the frame. A MultiIndex entry that is itself a tuple stays a tuple.

The index is always read, then removed when the caller did not ask for it. itertuples(index=False) drops rows when the frame has zero columns, so the adapter keeps the index and slices it off. include_index=True still prepends the index. include_index=False still omits it.

Scope stays on petl.io.pandas, with the same public functions and the same dependencies. Twelve parameterized cases cover small integers, large positive and large negative integers, both index modes, a second pass over the same view, a MultiIndex with duplicate columns and names that are not identifiers, a frame with no rows, and a frame with no columns. Assertions use petl.compat.integer_types, so a Python 2 long still counts as an integer. The pull request records no benchmark and claims no speed change.

The tradeoff is the Python type. Identifier columns past 2^53 stay exact, which is what warehouse keys need. A formatter or schema check that required those cells to be float now sees an int beside the float. Local runs cited in the pull request used Python 3.11.9 with pandas 3.0.6 and Python 3.10.0 with pandas 2.3.3.

Pull request 721 extends setheader() and the fluent Table.setheader() method. A callable is applied to each field in the header already on the table. Data rows pass through unchanged. The branch lives in itersetheader, the iterator behind SetHeaderView. The view stays lazy, and the tests check that it can be iterated more than once.

A list, a tuple, or any other iterable still replaces the header as a static sequence. The first revision tested callable before iteration. An object that implements both __iter__ and __call__ would then be invoked as a field function, or would raise from a __call__ that was never meant to see field names. The merged code tests the iterable protocol first. There is no keyword to flip that order. Pass a plain function for a map. Pass a sequence for a replacement.

The change closes issue 431. It rewrites every header field without building the new names before the call. rename stays the selective field tool. A callable passed to setheader maps every field. A function can lowercase each name, or add one prefix to each name.

Tests in petl/test/transform/test_headers.py cover the function, the Table method, a casing change, unchanged rows, a chained view, and an object that is both iterable and callable. The author cites 61 passed on that module plus petl/test/test_method_shadow.py, five doctests in petl/transform/headers.py, and a full pytest petl run of 658 passed with 18 skipped.

Move to v1.7.27 when fromdataframe reads integers beside floats, especially identifiers past 2^53. Those cells change from float to int. A key that arrived as 9007199254740992.0 now arrives as 9007199254740993. Recheck JSON encoding, SQL parameter binding, and any cast that assumed a float. Do that before a backfill.

Call sites of setheader that pass a list, a tuple, or another iterable do not need an edit. An object that is both iterable and callable is still a static header. New call sites can pass a function of one argument, the current field name.

Both pull requests also lengthened a short underline on a heading in docs/changes.rst. The docs build treats warnings as errors, and that short underline was already failing on the previous tag. The underline change does not affect runtime. Merge commits on the release are 6b1b9c2 for the pandas fix and 1725e4a for the header change.