GCS tag-write stress test

GCS allows about one write per second to any one object, and every server process shares cache/tags/tags.json. Before SITE-6194, each request had its own write buffer, so a burst of cache writes became a burst of writes to that one object: 429 rateLimitExceeded, and concurrent read-modify-writes that dropped each other’s tags.

The button fires a burst of cache misses with unique tags at this instance, revalidates half of them, and watches the GCS generation of tags.json and use-cache/_tags.json. Each write gets a new generation, so the log below counts real writes.

What to check

All green on Pantheon means the fix holds. A burst of 60 writes should produce only a handful of tags.json writes, spaced at least one flush interval apart. Every entry should be in the mapping. revalidateTag() should not write it at all, and nothing should log a 429.

Locally without CACHE_BUCKET, the GCS checks are skipped and only the burst and revalidation run. Requests go to this instance on loopback, and only this instance’s console is checked for warnings, so look at the site logs for other instances too.

This writes to shared cache state. It creates gcsstress:* cache entries and revalidates their tags, then removes them from tags.json. It does not touch any other tag. A run takes about 15 to 30s at the default 5s flush interval.

From a shell

curl -N -X POST 'https://<site>/api/gcs-stress?confirm=mutate-cache&entries=120&concurrency=30' \
  | jq -c 'select(.type=="result") | {v:.result.verdict, t:.result.title, s:.result.summary}'

cleanup=0 leaves the run’s entries in place. SELFTEST_DISABLED=1 turns this endpoint off along with /api/selftest. Set CACHE_TAGS_FLUSH_INTERVAL_MS to change the pacing the checks expect.