kOps Master Raises Terraform Provider Floors and Validates Registry Mirrors


Thirteen commits landed on kOps master in this window. The diff is 122 files, 1341 insertions, and 923 deletions, and most of that bulk is regenerated Terraform fixtures. The changes that affect an apply are a stricter Terraform target, higher provider floors, and containerd mirror checks that fail before a node boots.

kops update cluster --target=terraform no longer accepts --instance-group or --instance-group-roles. The guard is a PreRunE hook on NewCmdUpdateCluster. Terraform treats a resource missing from the configuration as a deletion, so a render of one instance group would plan destruction of every other resource in that state.

The error, and docs/terraform.md, say to render the full configuration and use terraform apply -target for a staged upgrade. A script that rendered one group into a scratch directory has to switch to that.

Three edits to target_hcl2.go raise required_providers, because the generated HCL now sets attributes the old floors do not have.

AWS and Google both cross a major version. A lock file stuck on a 5.x provider will fail terraform init on the next render from this master. Scaleway stays on 2.x, but 2.2.1 is no longer enough. The fixture refresh rewrites that constraint across tests/integration/update_cluster, including shared_vpc_ipv6/kubernetes.tf.

The same day, the shared VPC IPv6 output dropped try(). The replacement in upup/pkg/fi/cloudup/awstasks/vpc.go uses element and concat. The %s is the ipv6_cidr_block_associations attribute:

element(concat([for a in %s : a.ipv6_cidr_block if a.state == "associated"], [""]), 0)

An empty association list still becomes an empty string. Other lookup failures, which try() used to hide, now fail the plan. Seven fixtures take the new expression.

Since kOps 1.36, nodeup writes spec.containerd.registryMirrors into certs.d/<registry>/hosts.toml instead of the inline registry.mirrors list. containerd appends /v2 unless the table sets override_path = true. The legacy list used a path as written. An endpoint that already contains /v2/<prefix>, such as an ECR pull through cache, was requested as /v2/<prefix>/v2/<repo> and the pull failed.

The hosts.toml writer in nodeup/pkg/model/containerd.go now emits override_path = true when endpointHasPath reports a path. Endpoints with no path stay as they were.

The first path check called url.Parse on the raw string. An endpoint without a scheme parsed wrong, and // or /v2/.. counted as a path. The next commit prefixes https:// when the string does not start with http, matching containerd parseHostConfig, then compares path.Clean of the path to /.

Validation then rejects endpoints containerd would skip, or that would make it ignore every mirror of that registry and pull upstream after a log line. validateContainerdRegistryMirrors rejects:

  • a scheme other than lowercase http or https, so HTTPS:// does not become a host named HTTPS:
  • a value that does not parse, or that has no host, including a host name that starts with http and has no scheme
  • a repeated endpoint, which would be a duplicate TOML table
  • a loopback endpoint with no scheme, because hosts.toml uses https where the legacy config used http

mirror.example.com is still accepted. Loopback must be http://localhost:5000 or https://127.0.0.1:5000. The 1.38 notes document the rule, and the containerd custom test drops the placeholder http://HostIP2:Port2. A small expected file update refreshes nodeup config for that fixture. A bad mirror that used to be a node log now fails kops update, including specs that have been wrong since 1.36.

GetCloudGroups on GCE used to call Instances().Get for every managed instance. instancegroups.go now calls Instances().List once per zone, the first time a MIG in that zone has members, and reuses the map for later MIGs. The list helper already walks pages, so a large zone is not cut off at the first page.

A missing name still logs that the instance may not have been created, and the loop continues. A list error fails the call. Previously a Get error other than not found failed immediately. Large groups spend one paged list instead of one Get per VM. A zone full of unrelated VMs still lists all of them, once. kops rolling-update, kops validate cluster, kops get instances, and instance deletion all call GetCloudGroups on GCE.

kops edit instancegroup --set and --unset learned v1alpha2 paths in a patch written in June and landed on 30 September. InternalPathForInstanceGroupField moves seven flat rootVolume fields under spec.rootVolume, and it fixes the casing of spec.associatePublicIp and spec.kubelet.authenticationTokenWebhookCacheTtl. spec.kubelet.clientCaFile stays unmapped because nodeup overwrites it. Copied v1alpha2 paths used to return “field not found”. v1alpha3 paths pass through.

The scalability scenario change is CI only. Periodics skip make test-e2e-install, CI downloads move to dl.k8s.io, and that scenario defaults to kube-proxy mode nftables. It does not change a running cluster.

Upgrade providers before the next render from this master. AWS >= 6.57.1 and Google >= 6.8.0 will not init against a 5.x lock file. Scaleway needs at least 2.54.0. kOps raised the floor so those attributes exist in the provider schema. Read the provider upgrade notes separately. This tree does not rewrite existing state.

Check spec.containerd.registryMirrors before the next nodeup. Loopback mirrors need an explicit http:// or https://. Endpoints that already carry a path gain override_path = true in hosts.toml. Confirm an ECR pull through cache still serves the repository you expect.

Scripts that pass --instance-group with --target=terraform now fail in PreRunE. Generate the full tree and pass -target to Terraform. On GCE, kops rolling-update and kops validate cluster issue fewer instance Gets, and the list is still the whole zone.