How to Test CI Build Cache Invalidation with a Cached-vs-Clean Check

Use a deterministic Python fixture to expose a missing configuration input in a build cache key, compare cached and clean artifacts, and fail CI when either output is stale.

OpenThunder Editorial · 2026-10-06 · AI-assisted article

Two patterned clay tiles on separate tracks face a gate with a matching template, beside filled and empty cubbies.

AI-generated illustration. A conceptual gate for cache reuse: cached and isolated clean artifacts must each meet independent expectations and match byte-for-byte. That supports only the tested transition.

Test CI build cache invalidation by seeding a cache under configuration A, switching to configuration B, and comparing the cached build with an isolated clean build. Both artifacts must contain the expected configuration B and match byte-for-byte. A successful build exit is not enough.

This creates a decision gate for changes to caching or build scripts, including AI-generated edits: reuse is acceptable for the tested transition only if the result matches independent expectations and a clean build.

The regression: configuration changes, but the cache key does not

Consider a small build that produces a JSON artifact from two inputs:

The artifact embeds both values, but the broken cache key hashes only the source file. Changing the endpoint leaves the key unchanged, allowing the cache to return an artifact with the old endpoint.

Create that history deliberately:

  1. Build with endpoint A and an empty cache.
  2. Keep the source unchanged and switch to endpoint B.
  3. Build again using the seeded cache.
  4. Build with endpoint B in a separate workspace with a separate, empty cache.
  5. Check the expected contents and compare the artifact bytes.

Keep the source unchanged for this test. Changing the source and endpoint together could invalidate the cache through the source change, concealing the missing endpoint input.

A self-contained stale-artifact fixture

Save this as verify_cache.py. It uses Python's standard library and creates all workspaces, outputs, and cache entries in a temporary directory. The .invalid endpoints are placeholders. No network requests are made.

This is a minimal configuration-artifact builder, not a simulation of a CI provider. The explicit env mapping represents the build environment without inheriting unrelated runner settings. For a real build, pass the corresponding variable to the build subprocess.

The default mode uses the broken key. The --include-endpoint flag adds the missing input.

import argparse
import hashlib
import json
import sys
import tempfile
from pathlib import Path


ENDPOINT_A = "https://api-a.invalid"
ENDPOINT_B = "https://api-b.invalid"


def encode(value):
    return (
        json.dumps(value, sort_keys=True, separators=(",", ":")) + "\n"
    ).encode("utf-8")


def build(workspace, cache_dir, env, include_endpoint):
    source = (workspace / "message.txt").read_bytes()
    endpoint = env["API_ENDPOINT"]

    key_inputs = {
        "source_sha256": hashlib.sha256(source).hexdigest()
    }
    if include_endpoint:
        key_inputs["API_ENDPOINT"] = endpoint

    key = hashlib.sha256(encode(key_inputs)).hexdigest()
    cache_dir.mkdir(parents=True, exist_ok=True)
    entry = cache_dir / (key + ".json")
    cache_hit = entry.exists()

    if cache_hit:
        artifact = entry.read_bytes()
    else:
        artifact = encode({
            "api_endpoint": endpoint,
            "message": source.decode("utf-8"),
        })
        entry.write_bytes(artifact)

    output = workspace / "dist" / "config.json"
    output.parent.mkdir(parents=True, exist_ok=True)
    output.write_bytes(artifact)
    return output, cache_hit


def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("--include-endpoint", action="store_true")
    args = parser.parse_args()

    with tempfile.TemporaryDirectory(prefix="cache-check-") as tmp:
        root = Path(tmp)
        warm_workspace = root / "warm-workspace"
        clean_workspace = root / "clean-workspace"
        warm_cache = root / "warm-cache"
        clean_cache = root / "clean-cache"

        for workspace in (warm_workspace, clean_workspace):
            workspace.mkdir()
            (workspace / "message.txt").write_bytes(b"hello")

        seed_path, seed_hit = build(
            warm_workspace, warm_cache,
            {"API_ENDPOINT": ENDPOINT_A}, args.include_endpoint,
        )
        seed = json.loads(seed_path.read_bytes())
        if seed_hit or seed != {
            "api_endpoint": ENDPOINT_A, "message": "hello"
        }:
            print("FAIL: seed build did not establish configuration A")
            return 1

        warm_path, warm_hit = build(
            warm_workspace, warm_cache,
            {"API_ENDPOINT": ENDPOINT_B}, args.include_endpoint,
        )
        clean_path, clean_hit = build(
            clean_workspace, clean_cache,
            {"API_ENDPOINT": ENDPOINT_B}, args.include_endpoint,
        )

        warm_bytes = warm_path.read_bytes()
        clean_bytes = clean_path.read_bytes()
        expected = {"api_endpoint": ENDPOINT_B, "message": "hello"}
        errors = []

        for label, artifact in (
            ("warm", warm_bytes), ("clean", clean_bytes)
        ):
            if json.loads(artifact) != expected:
                errors.append(label + " artifact has unexpected contents")

        if warm_bytes != clean_bytes:
            errors.append("warm and clean artifacts differ byte-for-byte")
        if clean_hit:
            errors.append("clean build unexpectedly used a cached entry")

        print("warm cache:", "HIT" if warm_hit else "MISS")
        print("clean cache:", "HIT" if clean_hit else "MISS")
        print("warm artifact:", repr(warm_bytes))
        print("clean artifact:", repr(clean_bytes))

        if errors:
            for error in errors:
                print("FAIL:", error)
            return 1

        print("PASS: both artifacts contain B and match byte-for-byte")
        return 0


if __name__ == "__main__":
    sys.exit(main())

The serializer fixes key ordering, separators, encoding, and the trailing newline. The artifact contains no timestamps, random identifiers, or workspace paths, so those cannot explain a difference.

The temporary directory is removed when the check finishes. Keep this fixture isolated; do not point it at production caches or customer data.

Reproduce the failure, then add the missing input

Run the broken version first:

python verify_cache.py

The expected result is a nonzero exit status:

These are expected outcomes derived from the fixture, not results from an executed test run.

Now include the endpoint in the key:

python verify_cache.py --include-endpoint

The expected result is exit status zero. Changing the endpoint changes the key, so the build using the seeded cache misses the old entry and produces configuration B. The clean build also produces B, and the bytes match.

The repair is one added input:

key_inputs["API_ENDPOINT"] = endpoint

Here, “warm” means the cache has been populated, not that the next build must hit it. For this transition, a miss is correct.

A striped tile connects to a key and filled pocket; separate paths end in circle-marked and triangle-marked artifact tiles.

AI-generated illustration. The fixture’s intended failure mechanism: unchanged source preserves a source-only cache key despite an endpoint change. The seeded cache can return the old configuration while an isolated clean build produces the new one.

Why comparison alone is not enough

Two matching outputs can both be wrong.

Suppose a build-script regression ignores API_ENDPOINT and always emits a default endpoint. Cached and clean builds could agree while violating the intended configuration. The explicit expected object catches that mistake: each artifact must contain endpoint B and the expected source-derived message.

Checking only the endpoint has the opposite weakness: it could overlook other stale content. Comparing the full artifacts catches differences outside that field.

The verification contract is:

Cached artifact has expected contents. Clean artifact has expected contents. Their bytes are identical.

Put the transition into CI, not just the demonstration

With Python available in the job, the corrected fixture can run as a failure-propagating CI step:

python verify_cache.py --include-endpoint

Do not append || true or allow the step to fail without failing the job. The broken command demonstrates detection; the corrected command is the expected passing version.

A passing toy fixture does not verify your repository's cache. To make this a real regression gate, replace build with an adapter that invokes your build and its actual cache-key logic. Preserve these controls:

A new checkout is not necessarily a clean build. It can still read shared local or remote cache entries, so isolation must cover cache state as well as the working directory.

Expand coverage one input at a time

Test one changed input at a time. Each row below starts with a fresh seed build; update the expected artifact to reflect the target inputs.

Change after seeding Expected target contents Expected behavior for this fixture
No change Original message and endpoint A Existing entry is reusable; warm and clean outputs match
Source only Updated message and endpoint A Source digest changes; old entry is not reused
Endpoint only Original message and endpoint B Endpoint-aware key changes; source-only key fails the check
Source and endpoint Updated message and endpoint B Key changes, but this case alone cannot reveal the missing endpoint input

For a real build, consider other artifact-affecting inputs: templates, build scripts, dependency locks, compiler options, and toolchain identity. Determine which affect the artifact before adding them to a key.

Static analysis can help identify configuration reads missing from key construction. This integration-style check provides different evidence: it exercises the history-dependent sequence that can produce a stale artifact, including cache lookup and restoration.

What a passing result establishes

A pass supports a narrow conclusion: for these inputs and this isolated cache history, the cached build produced the expected artifact and matched the clean build.

It does not prove the cache is universally correct. This fixture does not key on its builder implementation, for example. A later formatting or build-logic change could require another key input or a cache-version boundary. Adding the endpoint does not solve that separate problem.

Nor does the fixture cover concurrent writers, remote-cache behavior, platform differences, or nondeterministic tool output. For a larger output tree, compare both relative file lists and file bytes to catch extra stale files. Check metadata separately if it matters to the output contract, such as executable permissions.

If a real build embeds timestamps or absolute paths, investigate and control those inputs before relaxing the comparison. Keep any normalization narrow and explicit. Removing configuration fields would hide the regression this check is meant to catch.

The practical rule: when an input can change the artifact, test that input independently against a seeded cache. Keep the transition as a regression test alongside the build logic it protects.

For a next step, explore OpenThunder as you plan your pipeline's verification checks.