You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Without Cache-Control, RFC 9111 §4.2.2
permits a cache to pick its own freshness lifetime, and the common heuristic is
10% of the time since Last-Modified. A file untouched for ten days is then
treated as fresh for a day. Browsers apply this to fetch(), so a client that
re-reads metadata after a publish can get the previous version with no error and
no way to tell.
We hit this on a STAC catalog served from Source Cooperative. After publishing
updated item JSON, a browser kept returning the pre-publish bodies — different ETags for the same URL — while curl returned current ones. The app dropped
26 of 27 archives from its picker, because each item looked like it advertised
no tileset. We now fetch the collection with a cache-busting query parameter
and key the item URLs off a version string in it, so a publish changes every
item URL. That works, but it is a workaround for a missing header, and the
collection itself is now uncacheable.
S3 stores Cache-Control as object metadata and returns it on GET and HEAD, so
the proxy can pass through whatever the object carries. Objects we have checked
do not set it:
Pass through Cache-Control from the object. Publishers then choose per
file — max-age=31536000, immutable for data that is rewritten under a new
name, no-cache for metadata that is overwritten in place. This costs the
proxy one header copy.
Send a default when the object sets none.no-cache is the safe one: it
permits storing but requires revalidation, and revalidation is already a
304 with an empty body.
Option 2 alone fixes the correctness problem. Option 1 lets publishers get
caching back for immutable data.
What value is this feature adding to Source Cooperative?
Clients stop reading stale metadata after a publish. This is currently silent:
the response is a 200 with a plausible body, so nothing downstream detects it.
Publishers stop working around it with cache-busting query strings, which
force a full transfer every time and make caching impossible for anyone.
Whole-object reads of small metadata — STAC JSON, READMEs, styles, TileJSON —
become correctly cacheable by browsers and downstream caches, on the
publisher's terms instead of a heuristic's.
Note
Scope, and how this relates to #188. This header alone does not make
ranged reads edge-cacheable, and an earlier version of this issue claimed it
did. Cloudflare's Cache API refuses to store a 206
(docs), and
the proxy deliberately sets RequestCache::NoStore on any subrequest carrying
a Range header (ForwardRequest::should_bypass_cache in multistore) so a
partial response can never poison the full-object cache entry. Making ranged
reads of large objects (COG, GeoParquet, PMTiles) fast is #188's job —
chunk-aligned caching inside the worker — and the two are complementary rather
than overlapping: #188 governs what the worker's own cache holds, this issue
governs what the proxy tells downstream caches.
Description of Feature:
Send a
Cache-Controlheader on proxied object responses.Responses today carry
ETagandLast-Modifiedbut noCache-Control:Without
Cache-Control, RFC 9111 §4.2.2permits a cache to pick its own freshness lifetime, and the common heuristic is
10% of the time since
Last-Modified. A file untouched for ten days is thentreated as fresh for a day. Browsers apply this to
fetch(), so a client thatre-reads metadata after a publish can get the previous version with no error and
no way to tell.
We hit this on a STAC catalog served from Source Cooperative. After publishing
updated item JSON, a browser kept returning the pre-publish bodies — different
ETags for the same URL — whilecurlreturned current ones. The app dropped26 of 27 archives from its picker, because each item looked like it advertised
no tileset. We now fetch the collection with a cache-busting query parameter
and key the item URLs off a version string in it, so a publish changes every
item URL. That works, but it is a workaround for a missing header, and the
collection itself is now uncacheable.
Revalidation already works, so the fix is cheap:
S3 stores
Cache-Controlas object metadata and returns it on GET and HEAD, sothe proxy can pass through whatever the object carries. Objects we have checked
do not set it:
Two options, not exclusive:
Cache-Controlfrom the object. Publishers then choose perfile —
max-age=31536000, immutablefor data that is rewritten under a newname,
no-cachefor metadata that is overwritten in place. This costs theproxy one header copy.
no-cacheis the safe one: itpermits storing but requires revalidation, and revalidation is already a
304 with an empty body.
Option 2 alone fixes the correctness problem. Option 1 lets publishers get
caching back for immutable data.
What value is this feature adding to Source Cooperative?
the response is a 200 with a plausible body, so nothing downstream detects it.
force a full transfer every time and make caching impossible for anyone.
become correctly cacheable by browsers and downstream caches, on the
publisher's terms instead of a heuristic's.
Note
Scope, and how this relates to #188. This header alone does not make
ranged reads edge-cacheable, and an earlier version of this issue claimed it
did. Cloudflare's Cache API refuses to store a
206(docs), and
the proxy deliberately sets
RequestCache::NoStoreon any subrequest carryinga
Rangeheader (ForwardRequest::should_bypass_cachein multistore) so apartial response can never poison the full-object cache entry. Making ranged
reads of large objects (COG, GeoParquet, PMTiles) fast is #188's job —
chunk-aligned caching inside the worker — and the two are complementary rather
than overlapping: #188 governs what the worker's own cache holds, this issue
governs what the proxy tells downstream caches.