Skip to content

s3proxy: skip uploading a blob the bucket already has - #923

Open
martinpitt wants to merge 1 commit into
buchgr:masterfrom
martinpitt:s3-skip-existing-uploads
Open

martinpitt wants to merge 1 commit into
buchgr:masterfrom
martinpitt:s3-skip-existing-uploads

Conversation

@martinpitt

Copy link
Copy Markdown

The http proxy backend HEADs an object before sending it and logs "SKIP UPLOAD" when it is already there. Do the same in the S3 backend. If the object already exists, only update its timestamp so that lifecycle rules don't evict it.

This avoids multiple uploads of blobs that several parallel actions produced and queued.


I tested this on a mid-sized project build: without the fix that did 278781 uploads for 60638 distinct objects, and sent 6.9 GB on the wire for 4.1 GB that got actually stored. This affects both S3 costs and also upoad time quite considerably. This is sort of the server-side counterpart to #919 (which is client-side) -- I'll look into that as well soon.

Thanks for considering!

The http proxy backend HEADs an object before sending it and logs "SKIP
UPLOAD" when it is already there. Do the same in the S3 backend. If the
object already exists, only update its timestamp so that lifecycle rules
don't evict it.

This avoids multiple uploads of blobs that several parallel actions
produced and queued.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant