A bucket and a volume,in either direction
The AWS CLI over S3, and over anything that speaks it.Pull the inputs in before a run, push the results out after one.
$ dxflow workflow create --identity s3 hub://s3 --start
Whichever sidecarries the uri
SOURCE and TARGET are the two ends, and one of them is an s3:// uri.
s3:// on SOURCE pulls into the volume; on TARGET it pushes back out.
MinIO, Ceph, R2, Wasabi and B2 — set ENDPOINT and it addresses by path.
Pass them on --override, so nothing is written into the definition.
The directioncomes from the uris
Whichever side carries s3:// decides which way the data moves. Nothing else declares it.
$ dxflow workflow start s3 --override env.job.SOURCE=s3://my-bucket/runs/ --override env.job.TARGET=/volume/input
$ dxflow workflow start s3 --override env.job.ACCESS_KEY=AKIA... --override env.job.SECRET_KEY=...
It transfers,and then it stops
A job, not a session — it ends once the last object has landed.
It mounts the wholevolume, on purpose
That reach is what lets a fetch land in another workflow's directory. It is also the thing to be careful with.
The volume is the root of Artifacts, so /volume/input inside the container is the same directory another workflow already reads.
SOURCE=/volume on a push sends every workflow's folder, workflow.json and the logs to the bucket. Point it at a directory inside.
Pass them with --override, scoped to the run. The hub entry is public, and a definition is not the place for a secret.
Pulled once,then it stays
S3 Sync arrives as one image. This is what comes down the first time, and what the disk should have free for it.
What it wants,and what it needs
The definition asks for 2 cores and 2 GB. The image comes up on less than that, and a start given --fit trims the ask to whatever the machine actually has.
Machines that fit it
S3 Sync asks for 2 cores and 2 GB. Cheapest first.